Cancer diagnoses by age¶
I'm a cancer researcher, and I'm interested in profiling adenocarcinoma occurance as a function of age.
First, decide what column to search. I'm looking for columns that have to do with age:
columns(description="age")
Loading ITables v2.9.1 from the init_notebook_mode cell...
(need help?)
|
| ⓘtable | column | data_type | nullable | description |
|---|---|---|---|---|
| file | file_type | text | True | File data type (like "CT image" or "miRNA expression quantification"). In the case of a DICOM series from IDC, this will be a decoded "SOPClassName" string corresponding to the "SOPClassUID" code assigned to DICOM instances belonging to the series. |
| observation | age_at_observation | integer | True | The approximate age at which this observation was made, computed directly as (year_of_observation - subject.year_of_birth). Resulting values may exhibit error margins based on upstream masking or transformation of subject.year_of_birth by upstream data managers to protect personally identifiable information. |
| observation | year_of_observation | integer | True | The year this observation was made. Values may be masked or transformed by upstream data managers to protect personally identifiable information. |
| subject | subject_id | text | False | A unique identifier for this subject minted by CDA. May change release-to-release. Contains no semantically reliable content. (Project codes and acronyms appearing as substrings in subject IDs will, when present, accurately indicate a project that the subject participates in, but coverage isn't complete, level of project varies arbitrarily, and subject-in-project characterization is generally neither complete nor representative.) |
| subject | year_of_birth | integer | True | The year this subject was born. Values may be masked or transformed by upstream data managers to protect personally identifiable information. |
| subject | year_of_death | integer | True | The year this subject died. Values may be masked or transformed by upstream data managers to protect personally identifiable information. |
age_at_observation is exactly what I need, and the description tells me that age is in years. Just to see what the data looks like, I'm going to ask for subjects who have adenocarcinoma and any observation age. The asterisks on either side of *adenocarcinoma* say that anything can be in front of or after adenocarcinoma in the diagnosis, so it will give me back results that have subtypes specified. And the exclaimation point in front of the equals sign != means NOT, so it will give back only not-null values:
summarize_subjects(match_all=["diagnosis = *adenocarcinoma*", "age_at_observation != NULL"])
╔═══════════════════════════════╗ ║ number_of_matching_subjects ║ ╠═══════════════════════════════╣ ║ 1301 ║ ╚═══════════════════════════════╝ ╔════════════════════════════════════════════════╗ ║ number_of_files_related_to_matching_subjects ║ ╠════════════════════════════════════════════════╣ ║ 146512 ║ ╚════════════════════════════════════════════════╝ ╔════════════╦══════════════════════╗ ║ subjects ║ data_source ║ ╠════════════╬══════════════════════╣ ║ 461 ║ PDC + IDC + GDC ║ ║ 358 ║ PDC + IDC + GC + GDC ║ ║ 166 ║ IDC + GDC ║ ║ 150 ║ PDC + GDC ║ ║ 82 ║ IDC only ║ ║ 80 ║ GDC only ║ ║ 2 ║ PDC + IDC + GC ║ ║ 1 ║ PDC + GC ║ ║ 1 ║ PDC + GC + GDC ║ ╚════════════╩══════════════════════╝ ╔════════════════╦═══════════╗ ║ count_result ║ species ║ ╠════════════════╬═══════════╣ ║ 1219 ║ human ║ ║ 82 ║ mouse ║ ╚════════════════╩═══════════╝ ╔════════════════╦══════════════════════════════════╗ ║ count_result ║ race ║ ╠════════════════╬══════════════════════════════════╣ ║ 864 ║ White ║ ║ 253 ║ <NA> ║ ║ 151 ║ Asian ║ ║ 32 ║ Black or African American ║ ║ 1 ║ American Indian or Alaska Native ║ ╚════════════════╩══════════════════════════════════╝ ╔════════════════╦════════════════════╗ ║ count_result ║ ethnicity ║ ╠════════════════╬════════════════════╣ ║ 846 ║ <NA> ║ ║ 414 ║ Non-Hispanic ║ ║ 41 ║ Hispanic or Latino ║ ╚════════════════╩════════════════════╝ ╔════════════════╦══════════════════════════╗ ║ count_result ║ cause_of_death ║ ╠════════════════╬══════════════════════════╣ ║ 1140 ║ <NA> ║ ║ 127 ║ Cancer-Related Death ║ ║ 20 ║ Non-Cancer Related Death ║ ║ 6 ║ Infection ║ ║ 5 ║ Cardiovascular Disorder ║ ║ 3 ║ Surgical Complication ║ ╚════════════════╩══════════════════════════╝ ╔════════════════╦══════════════════════════════════════════╗ ║ count_result ║ diagnosis ║ ╠════════════════╬══════════════════════════════════════════╣ ║ 730 ║ Adenocarcinoma ║ ║ 458 ║ Neoplasm, malignant ║ ║ 247 ║ Endometrioid adenocarcinoma ║ ║ 242 ║ Clear cell adenocarcinoma ║ ║ 222 ║ Renal cell carcinoma ║ ║ 86 ║ Adenocarcinoma, intestinal type ║ ║ 48 ║ Adenocarcinoma, metastatic ║ ║ 33 ║ Adenoma ║ ║ 23 ║ Epithelial tumor, benign ║ ║ 22 ║ Mucinous adenocarcinoma ║ ║ 21 ║ Papillary adenocarcinoma ║ ║ 15 ║ Adenocarcinoma with mixed subtypes ║ ║ 8 ║ Acinar cell carcinoma ║ ║ 3 ║ Cystic, mucinous and serous neoplasms ║ ║ 3 ║ Lepidic adenocarcinoma ║ ║ 3 ║ Oxyphilic adenoma ║ ║ 2 ║ Solid carcinoma ║ ║ 1 ║ Adenosquamous carcinoma ║ ║ 1 ║ Angioimmunoblastic T-cell lymphoma ║ ║ 1 ║ Complex epithelial neoplasms ║ ║ 1 ║ Endometrioid adenocarcinoma, secretor... ║ ║ 1 ║ Hepatoid adenocarcinoma ║ ║ 1 ║ Infiltrating ductular carcinoma ║ ║ 1 ║ Malignant melanoma ║ ║ 1 ║ Minimally invasive adenocarcinoma, mu... ║ ║ 1 ║ Neoplasm, metastatic ║ ║ 1 ║ Neuroendocrine carcinoma ║ ║ 1 ║ Noninfiltrating intraductal papillary... ║ ║ 1 ║ Renal cell carcinoma, chromophobe type ║ ║ 1 ║ Serous carcinoma ║ ║ 1 ║ Small cell carcinoma ║ ║ 1 ║ Squamous cell carcinoma ║ ║ 1 ║ Squamous cell carcinoma, metastatic ║ ║ 1 ║ T-cell large granular lymphocytic leu... ║ ║ 1 ║ Tubular adenocarcinoma ║ ║ 1 ║ Tubulovillous adenoma ║ ╚════════════════╩══════════════════════════════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_birth ║ ╠════════════════╬═════════════════╣ ║ mean ║ 1956 ║ ║ min ║ 1914 ║ ║ lower quartile ║ 1945 ║ ║ median ║ 1953 ║ ║ upper quartile ║ 1961 ║ ║ max ║ 2021 ║ ╚════════════════╩═════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_death ║ ╠════════════════╬═════════════════╣ ║ mean ║ 2016 ║ ║ min ║ 2002 ║ ║ lower quartile ║ 2015 ║ ║ median ║ 2018 ║ ║ upper quartile ║ 2020 ║ ║ max ║ 2023 ║ ╚════════════════╩═════════════════╝ ╔════════════════╦══════════════════════╗ ║ ║ age_at_observation ║ ╠════════════════╬══════════════════════╣ ║ mean ║ 61 ║ ║ min ║ 0 ║ ║ lower quartile ║ 54 ║ ║ median ║ 63 ║ ║ upper quartile ║ 71 ║ ║ max ║ 90 ║ ╚════════════════╩══════════════════════╝
There are just under 950 subjects that fit those criteria, but it looks like some are mice. I don't want mice, so I'm going to add a species filter:
summarize_subjects(match_all=["diagnosis = *adenocarcinoma*", "age_at_observation != NULL", "species = human"])
╔═══════════════════════════════╗ ║ number_of_matching_subjects ║ ╠═══════════════════════════════╣ ║ 1219 ║ ╚═══════════════════════════════╝ ╔════════════════════════════════════════════════╗ ║ number_of_files_related_to_matching_subjects ║ ╠════════════════════════════════════════════════╣ ║ 145794 ║ ╚════════════════════════════════════════════════╝ ╔════════════╦══════════════════════╗ ║ subjects ║ data_source ║ ╠════════════╬══════════════════════╣ ║ 461 ║ PDC + IDC + GDC ║ ║ 358 ║ PDC + GC + IDC + GDC ║ ║ 166 ║ IDC + GDC ║ ║ 150 ║ PDC + GDC ║ ║ 80 ║ GDC only ║ ║ 2 ║ PDC + GC + IDC ║ ║ 1 ║ PDC + GC ║ ║ 1 ║ PDC + GC + GDC ║ ╚════════════╩══════════════════════╝ ╔════════════════╦══════════════════════════╗ ║ count_result ║ cause_of_death ║ ╠════════════════╬══════════════════════════╣ ║ 1058 ║ <NA> ║ ║ 127 ║ Cancer-Related Death ║ ║ 20 ║ Non-Cancer Related Death ║ ║ 6 ║ Infection ║ ║ 5 ║ Cardiovascular Disorder ║ ║ 3 ║ Surgical Complication ║ ╚════════════════╩══════════════════════════╝ ╔════════════════╦══════════════════════════════════╗ ║ count_result ║ race ║ ╠════════════════╬══════════════════════════════════╣ ║ 864 ║ White ║ ║ 171 ║ <NA> ║ ║ 151 ║ Asian ║ ║ 32 ║ Black or African American ║ ║ 1 ║ American Indian or Alaska Native ║ ╚════════════════╩══════════════════════════════════╝ ╔════════════════╦════════════════════╗ ║ count_result ║ ethnicity ║ ╠════════════════╬════════════════════╣ ║ 764 ║ <NA> ║ ║ 414 ║ Non-Hispanic ║ ║ 41 ║ Hispanic or Latino ║ ╚════════════════╩════════════════════╝ ╔════════════════╦═══════════╗ ║ count_result ║ species ║ ╠════════════════╬═══════════╣ ║ 1219 ║ human ║ ╚════════════════╩═══════════╝ ╔════════════════╦══════════════════════════════════════════╗ ║ count_result ║ diagnosis ║ ╠════════════════╬══════════════════════════════════════════╣ ║ 672 ║ Adenocarcinoma ║ ║ 458 ║ Neoplasm, malignant ║ ║ 247 ║ Endometrioid adenocarcinoma ║ ║ 242 ║ Clear cell adenocarcinoma ║ ║ 222 ║ Renal cell carcinoma ║ ║ 62 ║ Adenocarcinoma, intestinal type ║ ║ 48 ║ Adenocarcinoma, metastatic ║ ║ 33 ║ Adenoma ║ ║ 23 ║ Epithelial tumor, benign ║ ║ 22 ║ Mucinous adenocarcinoma ║ ║ 21 ║ Papillary adenocarcinoma ║ ║ 15 ║ Adenocarcinoma with mixed subtypes ║ ║ 8 ║ Acinar cell carcinoma ║ ║ 3 ║ Cystic, mucinous and serous neoplasms ║ ║ 3 ║ Lepidic adenocarcinoma ║ ║ 3 ║ Oxyphilic adenoma ║ ║ 2 ║ Solid carcinoma ║ ║ 1 ║ Adenosquamous carcinoma ║ ║ 1 ║ Angioimmunoblastic T-cell lymphoma ║ ║ 1 ║ Complex epithelial neoplasms ║ ║ 1 ║ Endometrioid adenocarcinoma, secretor... ║ ║ 1 ║ Hepatoid adenocarcinoma ║ ║ 1 ║ Infiltrating ductular carcinoma ║ ║ 1 ║ Malignant melanoma ║ ║ 1 ║ Minimally invasive adenocarcinoma, mu... ║ ║ 1 ║ Neoplasm, metastatic ║ ║ 1 ║ Neuroendocrine carcinoma ║ ║ 1 ║ Noninfiltrating intraductal papillary... ║ ║ 1 ║ Renal cell carcinoma, chromophobe type ║ ║ 1 ║ Serous carcinoma ║ ║ 1 ║ Small cell carcinoma ║ ║ 1 ║ Squamous cell carcinoma ║ ║ 1 ║ Squamous cell carcinoma, metastatic ║ ║ 1 ║ T-cell large granular lymphocytic leu... ║ ║ 1 ║ Tubular adenocarcinoma ║ ║ 1 ║ Tubulovillous adenoma ║ ╚════════════════╩══════════════════════════════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_death ║ ╠════════════════╬═════════════════╣ ║ mean ║ 2016 ║ ║ min ║ 2002 ║ ║ lower quartile ║ 2015 ║ ║ median ║ 2018 ║ ║ upper quartile ║ 2020 ║ ║ max ║ 2023 ║ ╚════════════════╩═════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_birth ║ ╠════════════════╬═════════════════╣ ║ mean ║ 1952 ║ ║ min ║ 1914 ║ ║ lower quartile ║ 1945 ║ ║ median ║ 1952 ║ ║ upper quartile ║ 1960 ║ ║ max ║ 1992 ║ ╚════════════════╩═════════════════╝ ╔════════════════╦══════════════════════╗ ║ ║ age_at_observation ║ ╠════════════════╬══════════════════════╣ ║ mean ║ 63 ║ ║ min ║ 15 ║ ║ lower quartile ║ 55 ║ ║ median ║ 64 ║ ║ upper quartile ║ 71 ║ ║ max ║ 90 ║ ╚════════════════╩══════════════════════╝
Just over 900 humans meet that criteria, that seems promising. I may also want to look at some age ranges, as summaries of the subject data:
summarize_subjects(match_all=["age_at_observation > 80", "diagnosis = *adenocarcinoma*", "species = human"])
╔═══════════════════════════════╗ ║ number_of_matching_subjects ║ ╠═══════════════════════════════╣ ║ 120 ║ ╚═══════════════════════════════╝ ╔════════════════════════════════════════════════╗ ║ number_of_files_related_to_matching_subjects ║ ╠════════════════════════════════════════════════╣ ║ 19629 ║ ╚════════════════════════════════════════════════╝ ╔════════════╦══════════════════════╗ ║ subjects ║ data_source ║ ╠════════════╬══════════════════════╣ ║ 52 ║ PDC + GDC + IDC ║ ║ 33 ║ GC + PDC + GDC + IDC ║ ║ 20 ║ GDC + IDC ║ ║ 15 ║ PDC + GDC ║ ╚════════════╩══════════════════════╝ ╔════════════════╦══════════════╗ ║ count_result ║ ethnicity ║ ╠════════════════╬══════════════╣ ║ 93 ║ <NA> ║ ║ 27 ║ Non-Hispanic ║ ╚════════════════╩══════════════╝ ╔════════════════╦═══════════╗ ║ count_result ║ species ║ ╠════════════════╬═══════════╣ ║ 120 ║ human ║ ╚════════════════╩═══════════╝ ╔════════════════╦═══════════════════════════╗ ║ count_result ║ race ║ ╠════════════════╬═══════════════════════════╣ ║ 62 ║ <NA> ║ ║ 54 ║ White ║ ║ 3 ║ Black or African American ║ ║ 1 ║ Asian ║ ╚════════════════╩═══════════════════════════╝ ╔════════════════╦══════════════════════════╗ ║ count_result ║ cause_of_death ║ ╠════════════════╬══════════════════════════╣ ║ 106 ║ <NA> ║ ║ 5 ║ Non-Cancer Related Death ║ ║ 4 ║ Cancer-Related Death ║ ║ 2 ║ Cardiovascular Disorder ║ ║ 2 ║ Surgical Complication ║ ║ 1 ║ Infection ║ ╚════════════════╩══════════════════════════╝ ╔════════════════╦═════════════════════════════════╗ ║ count_result ║ diagnosis ║ ╠════════════════╬═════════════════════════════════╣ ║ 88 ║ Adenocarcinoma ║ ║ 40 ║ Adenocarcinoma, intestinal type ║ ║ 25 ║ Neoplasm, malignant ║ ║ 18 ║ Adenoma ║ ║ 16 ║ Endometrioid adenocarcinoma ║ ║ 12 ║ Clear cell adenocarcinoma ║ ║ 10 ║ Renal cell carcinoma ║ ║ 5 ║ Mucinous adenocarcinoma ║ ║ 2 ║ Acinar cell carcinoma ║ ║ 1 ║ Adenocarcinoma, metastatic ║ ║ 1 ║ Adenosquamous carcinoma ║ ║ 1 ║ Complex epithelial neoplasms ║ ║ 1 ║ Lepidic adenocarcinoma ║ ║ 1 ║ Papillary adenocarcinoma ║ ╚════════════════╩═════════════════════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_birth ║ ╠════════════════╬═════════════════╣ ║ mean ║ 1932 ║ ║ min ║ 1914 ║ ║ lower quartile ║ 1929 ║ ║ median ║ 1933 ║ ║ upper quartile ║ 1936 ║ ║ max ║ 1942 ║ ╚════════════════╩═════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_death ║ ╠════════════════╬═════════════════╣ ║ mean ║ 2018 ║ ║ min ║ 2007 ║ ║ lower quartile ║ 2017 ║ ║ median ║ 2020 ║ ║ upper quartile ║ 2020 ║ ║ max ║ 2022 ║ ╚════════════════╩═════════════════╝ ╔════════════════╦══════════════════════╗ ║ ║ age_at_observation ║ ╠════════════════╬══════════════════════╣ ║ mean ║ 80 ║ ║ min ║ 63 ║ ║ lower quartile ║ 75 ║ ║ median ║ 82 ║ ║ upper quartile ║ 86 ║ ║ max ║ 90 ║ ╚════════════════╩══════════════════════╝
summarize_subjects(match_all=[ "70 < age_at_observation <= 80", "diagnosis = *adenocarcinoma*", "species = human"])
╔═══════════════════════════════╗ ║ number_of_matching_subjects ║ ╠═══════════════════════════════╣ ║ 314 ║ ╚═══════════════════════════════╝ ╔════════════════════════════════════════════════╗ ║ number_of_files_related_to_matching_subjects ║ ╠════════════════════════════════════════════════╣ ║ 49403 ║ ╚════════════════════════════════════════════════╝ ╔════════════╦══════════════════════╗ ║ subjects ║ data_source ║ ╠════════════╬══════════════════════╣ ║ 116 ║ GDC + PDC + IDC ║ ║ 83 ║ GDC + GC + PDC + IDC ║ ║ 51 ║ GDC + IDC ║ ║ 48 ║ GDC + PDC ║ ║ 13 ║ GDC only ║ ║ 1 ║ GC + PDC ║ ║ 1 ║ GDC + GC + PDC ║ ║ 1 ║ GC + PDC + IDC ║ ╚════════════╩══════════════════════╝ ╔════════════════╦═══════════════════════════╗ ║ count_result ║ race ║ ╠════════════════╬═══════════════════════════╣ ║ 214 ║ White ║ ║ 71 ║ <NA> ║ ║ 19 ║ Asian ║ ║ 10 ║ Black or African American ║ ╚════════════════╩═══════════════════════════╝ ╔════════════════╦═══════════╗ ║ count_result ║ species ║ ╠════════════════╬═══════════╣ ║ 314 ║ human ║ ╚════════════════╩═══════════╝ ╔════════════════╦════════════════════╗ ║ count_result ║ ethnicity ║ ╠════════════════╬════════════════════╣ ║ 194 ║ <NA> ║ ║ 115 ║ Non-Hispanic ║ ║ 5 ║ Hispanic or Latino ║ ╚════════════════╩════════════════════╝ ╔════════════════╦══════════════════════════╗ ║ count_result ║ cause_of_death ║ ╠════════════════╬══════════════════════════╣ ║ 266 ║ <NA> ║ ║ 32 ║ Cancer-Related Death ║ ║ 9 ║ Non-Cancer Related Death ║ ║ 3 ║ Infection ║ ║ 2 ║ Cardiovascular Disorder ║ ║ 2 ║ Surgical Complication ║ ╚════════════════╩══════════════════════════╝ ╔════════════════╦══════════════════════════════════════════╗ ║ count_result ║ diagnosis ║ ╠════════════════╬══════════════════════════════════════════╣ ║ 209 ║ Adenocarcinoma ║ ║ 96 ║ Neoplasm, malignant ║ ║ 56 ║ Endometrioid adenocarcinoma ║ ║ 35 ║ Adenocarcinoma, intestinal type ║ ║ 34 ║ Clear cell adenocarcinoma ║ ║ 30 ║ Renal cell carcinoma ║ ║ 15 ║ Adenoma ║ ║ 13 ║ Adenocarcinoma, metastatic ║ ║ 9 ║ Mucinous adenocarcinoma ║ ║ 4 ║ Epithelial tumor, benign ║ ║ 3 ║ Acinar cell carcinoma ║ ║ 3 ║ Adenocarcinoma with mixed subtypes ║ ║ 2 ║ Papillary adenocarcinoma ║ ║ 1 ║ Adenosquamous carcinoma ║ ║ 1 ║ Complex epithelial neoplasms ║ ║ 1 ║ Cystic, mucinous and serous neoplasms ║ ║ 1 ║ Lepidic adenocarcinoma ║ ║ 1 ║ Malignant melanoma ║ ║ 1 ║ Minimally invasive adenocarcinoma, mu... ║ ║ 1 ║ Neoplasm, metastatic ║ ║ 1 ║ Small cell carcinoma ║ ║ 1 ║ Squamous cell carcinoma, metastatic ║ ║ 1 ║ Tubulovillous adenoma ║ ╚════════════════╩══════════════════════════════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_death ║ ╠════════════════╬═════════════════╣ ║ mean ║ 2016 ║ ║ min ║ 2004 ║ ║ lower quartile ║ 2016 ║ ║ median ║ 2018 ║ ║ upper quartile ║ 2020 ║ ║ max ║ 2023 ║ ╚════════════════╩═════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_birth ║ ╠════════════════╬═════════════════╣ ║ mean ║ 1941 ║ ║ min ║ 1923 ║ ║ lower quartile ║ 1938 ║ ║ median ║ 1943 ║ ║ upper quartile ║ 1946 ║ ║ max ║ 1952 ║ ╚════════════════╩═════════════════╝ ╔════════════════╦══════════════════════╗ ║ ║ age_at_observation ║ ╠════════════════╬══════════════════════╣ ║ mean ║ 73 ║ ║ min ║ 53 ║ ║ lower quartile ║ 71 ║ ║ median ║ 73 ║ ║ upper quartile ║ 77 ║ ║ max ║ 90 ║ ╚════════════════╩══════════════════════╝
summarize_subjects(match_all=["60 < age_at_observation <= 70", "diagnosis = *adenocarcinoma*", "species = human"])
╔═══════════════════════════════╗ ║ number_of_matching_subjects ║ ╠═══════════════════════════════╣ ║ 491 ║ ╚═══════════════════════════════╝ ╔════════════════════════════════════════════════╗ ║ number_of_files_related_to_matching_subjects ║ ╠════════════════════════════════════════════════╣ ║ 70526 ║ ╚════════════════════════════════════════════════╝ ╔════════════╦══════════════════════╗ ║ subjects ║ data_source ║ ╠════════════╬══════════════════════╣ ║ 205 ║ GDC + IDC + PDC ║ ║ 142 ║ GDC + GC + IDC + PDC ║ ║ 59 ║ GDC + PDC ║ ║ 59 ║ GDC + IDC ║ ║ 26 ║ GDC only ║ ╚════════════╩══════════════════════╝ ╔════════════════╦══════════════════════════╗ ║ count_result ║ cause_of_death ║ ╠════════════════╬══════════════════════════╣ ║ 424 ║ <NA> ║ ║ 53 ║ Cancer-Related Death ║ ║ 10 ║ Non-Cancer Related Death ║ ║ 2 ║ Cardiovascular Disorder ║ ║ 2 ║ Infection ║ ╚════════════════╩══════════════════════════╝ ╔════════════════╦════════════════════╗ ║ count_result ║ ethnicity ║ ╠════════════════╬════════════════════╣ ║ 319 ║ <NA> ║ ║ 159 ║ Non-Hispanic ║ ║ 13 ║ Hispanic or Latino ║ ╚════════════════╩════════════════════╝ ╔════════════════╦═══════════╗ ║ count_result ║ species ║ ╠════════════════╬═══════════╣ ║ 491 ║ human ║ ╚════════════════╩═══════════╝ ╔════════════════╦══════════════════════════════════╗ ║ count_result ║ race ║ ╠════════════════╬══════════════════════════════════╣ ║ 370 ║ White ║ ║ 58 ║ <NA> ║ ║ 53 ║ Asian ║ ║ 9 ║ Black or African American ║ ║ 1 ║ American Indian or Alaska Native ║ ╚════════════════╩══════════════════════════════════╝ ╔════════════════╦══════════════════════════════════════════╗ ║ count_result ║ diagnosis ║ ╠════════════════╬══════════════════════════════════════════╣ ║ 264 ║ Adenocarcinoma ║ ║ 196 ║ Neoplasm, malignant ║ ║ 114 ║ Endometrioid adenocarcinoma ║ ║ 96 ║ Clear cell adenocarcinoma ║ ║ 85 ║ Renal cell carcinoma ║ ║ 21 ║ Adenocarcinoma, intestinal type ║ ║ 17 ║ Adenocarcinoma, metastatic ║ ║ 11 ║ Adenoma ║ ║ 10 ║ Epithelial tumor, benign ║ ║ 9 ║ Adenocarcinoma with mixed subtypes ║ ║ 8 ║ Mucinous adenocarcinoma ║ ║ 5 ║ Acinar cell carcinoma ║ ║ 5 ║ Papillary adenocarcinoma ║ ║ 3 ║ Cystic, mucinous and serous neoplasms ║ ║ 2 ║ Oxyphilic adenoma ║ ║ 1 ║ Angioimmunoblastic T-cell lymphoma ║ ║ 1 ║ Malignant melanoma ║ ║ 1 ║ Neoplasm, metastatic ║ ║ 1 ║ Neuroendocrine carcinoma ║ ║ 1 ║ Noninfiltrating intraductal papillary... ║ ║ 1 ║ Serous carcinoma ║ ║ 1 ║ Solid carcinoma ║ ║ 1 ║ Squamous cell carcinoma ║ ║ 1 ║ T-cell large granular lymphocytic leu... ║ ║ 1 ║ Tubular adenocarcinoma ║ ╚════════════════╩══════════════════════════════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_death ║ ╠════════════════╬═════════════════╣ ║ mean ║ 2017 ║ ║ min ║ 2003 ║ ║ lower quartile ║ 2016 ║ ║ median ║ 2019 ║ ║ upper quartile ║ 2020 ║ ║ max ║ 2023 ║ ╚════════════════╩═════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_birth ║ ╠════════════════╬═════════════════╣ ║ mean ║ 1950 ║ ║ min ║ 1931 ║ ║ lower quartile ║ 1948 ║ ║ median ║ 1951 ║ ║ upper quartile ║ 1954 ║ ║ max ║ 1962 ║ ╚════════════════╩═════════════════╝ ╔════════════════╦══════════════════════╗ ║ ║ age_at_observation ║ ╠════════════════╬══════════════════════╣ ║ mean ║ 65 ║ ║ min ║ 43 ║ ║ lower quartile ║ 62 ║ ║ median ║ 65 ║ ║ upper quartile ║ 68 ║ ║ max ║ 88 ║ ╚════════════════╩══════════════════════╝
I'm going to look more closely at one of my age ranges by running get_subject_data instead of summarize_subjects:
get_subject_data(match_all=["60 < age_at_observation <= 70", "diagnosis = *adenocarcinoma*", "species = human"])
Loading ITables v2.9.1 from the init_notebook_mode cell...
(need help?)
|
| ⓘsubject_id | cause_of_death | ethnicity | race | species | year_of_birth | year_of_death | data_source | age_at_observation | diagnosis |
|---|---|---|---|---|---|---|---|---|---|
| CPTAC.C3L-01277 | <NA> | <NA> | White | human | 1955 | <NA> | [GDC, IDC, PDC] | [54, 62] | [Endometrioid adenocarcinoma, Neoplasm, malignant] |
| CPTAC.C3N-02598 | <NA> | <NA> | White | human | 1957 | <NA> | [GDC, IDC, PDC] | [43, 60, 61] | [Endometrioid adenocarcinoma, Neoplasm, malignant] |
| CPTAC.C3L-00361 | <NA> | Non-Hispanic | White | human | 1952 | <NA> | [GC, GDC, IDC, PDC] | [64] | [Endometrioid adenocarcinoma, Neoplasm, malignant] |
| CPTAC.C3N-02234 | <NA> | <NA> | Asian | human | 1950 | <NA> | [GDC, IDC, PDC] | [67, 68] | [Adenocarcinoma] |
| CPTAC.C3N-02028 | <NA> | <NA> | White | human | 1949 | <NA> | [GDC, IDC, PDC] | [68] | [Endometrioid adenocarcinoma, Neoplasm, malignant] |
| CPTAC.C3L-07081 | <NA> | <NA> | White | human | 1952 | <NA> | [GDC, PDC] | [69] | [Adenocarcinoma] |
| CPTAC.C3N-03415 | Non-Cancer Related Death | <NA> | White | human | 1932 | 2020 | [GDC, IDC, PDC] | [69, 85, 86] | [Endometrioid adenocarcinoma, Neoplasm, malignant] |
| CPTAC.C3N-02146 | <NA> | <NA> | Asian | human | 1948 | <NA> | [GDC, IDC, PDC] | [69, 70] | [Adenocarcinoma] |
| TCGA.TCGA-AA-A010 | <NA> | <NA> | <NA> | human | 1962 | <NA> | [GDC, IDC, PDC] | [46, 48, 62] | [Adenocarcinoma, Adenocarcinoma, intestinal type] |
| CPTAC.C3L-09082 | <NA> | <NA> | White | human | 1953 | <NA> | [GDC, PDC] | [69] | [Adenocarcinoma] |
| (481 more rows not shown) | |||||||||
Several of the subjects here have multiple diagnoses and data from multiple sources, but not multiple age_at_observation values. I wonder why that is, so I'm going to have cdapython collate_results, that is, I'm going to have it match up all these observation data points to one another so I can see what the original data looked like:
sixty2seventy_collated = get_subject_data(match_all=["60 < age_at_observation <= 70", "diagnosis = *adenocarcinoma*", "species = human"], collate_results=True)
sixty2seventy_collated
Loading ITables v2.9.1 from the init_notebook_mode cell...
(need help?)
|
| ⓘsubject_id | cause_of_death | ethnicity | race | species | year_of_birth | year_of_death | data_source | observation_data |
|---|---|---|---|---|---|---|---|---|
| CPTAC.C3L-01277 | <NA> | <NA> | White | human | 1955 | <NA> | [GDC, IDC, PDC] | data_source age_at_observation diagnosis 0 IDC 54 Neoplasm, malignant 1 IDC 54 Neoplasm, malignant 2 IDC 62 Neoplasm, malignant 3 PDC 62 Endometrioid adenocarcinoma 4 GDC 62 Endometrioid adenocarcinoma 5 GDC <NA> <NA> 6 PDC <NA> <NA> 7 IDC 54 Neoplasm, malignant |
| CPTAC.C3N-02598 | <NA> | <NA> | White | human | 1957 | <NA> | [GDC, IDC, PDC] | data_source age_at_observation diagnosis 0 IDC 43 Neoplasm, malignant 1 IDC 43 Neoplasm, malignant 2 IDC 43 Neoplasm, malignant 3 PDC 60 Endometrioid adenocarcinoma 4 PDC <NA> <NA> 5 GDC 60 Endometrioid adenocarcinoma 6 IDC 61 Neoplasm, malignant 7 GDC <NA> <NA> |
| CPTAC.C3L-00361 | <NA> | Non-Hispanic | White | human | 1952 | <NA> | [GC, GDC, IDC, PDC] | data_source age_at_observation diagnosis 0 IDC 64 Neoplasm, malignant 1 GDC <NA> <NA> 2 GDC 64 Endometrioid adenocarcinoma 3 PDC <NA> <NA> 4 PDC 64 Endometrioid adenocarcinoma 5 GC <NA> Neoplasm, malignant 6 GC <NA> <NA> 7 GC <NA> <NA> 8 GC <NA> <NA> |
| CPTAC.C3N-02234 | <NA> | <NA> | Asian | human | 1950 | <NA> | [GDC, IDC, PDC] | data_source age_at_observation diagnosis 0 PDC 67 Adenocarcinoma 1 GDC <NA> <NA> 2 GDC 67 Adenocarcinoma 3 PDC <NA> <NA> 4 IDC 68 Adenocarcinoma |
| CPTAC.C3N-02028 | <NA> | <NA> | White | human | 1949 | <NA> | [GDC, IDC, PDC] | data_source age_at_observation diagnosis 0 GDC 68 Endometrioid adenocarcinoma 1 GDC <NA> <NA> 2 PDC 68 Endometrioid adenocarcinoma 3 IDC 68 Neoplasm, malignant 4 PDC <NA> <NA> |
| CPTAC.C3L-07081 | <NA> | <NA> | White | human | 1952 | <NA> | [GDC, PDC] | data_source age_at_observation diagnosis 0 PDC 69 Adenocarcinoma 1 PDC <NA> <NA> 2 GDC 69 Adenocarcinoma 3 GDC <NA> <NA> |
| CPTAC.C3N-03415 | Non-Cancer Related Death | <NA> | White | human | 1932 | 2020 | [GDC, IDC, PDC] | data_source age_at_observation diagnosis 0 GDC <NA> <NA> 1 PDC 85 Endometrioid adenocarcinoma 2 PDC <NA> <NA> 3 IDC 86 Neoplasm, malignant 4 IDC 69 Neoplasm, malignant 5 IDC 69 Neoplasm, malignant 6 GDC 85 Endometrioid adenocarcinoma |
| CPTAC.C3N-02146 | <NA> | <NA> | Asian | human | 1948 | <NA> | [GDC, IDC, PDC] | data_source age_at_observation diagnosis 0 PDC <NA> <NA> 1 IDC 70 Adenocarcinoma 2 PDC 69 Adenocarcinoma 3 GDC <NA> <NA> 4 GDC 69 Adenocarcinoma |
| TCGA.TCGA-AA-A010 | <NA> | <NA> | <NA> | human | 1962 | <NA> | [GDC, IDC, PDC] | data_source age_at_observation diagnosis 0 GDC 46 Adenocarcinoma 1 PDC 46 Adenocarcinoma 2 IDC 48 Adenocarcinoma, intestinal type 3 IDC 62 Adenocarcinoma, intestinal type 4 IDC 46 Adenocarcinoma, intestinal type 5 GDC <NA> <NA> 6 PDC <NA> <NA> 7 IDC <NA> <NA> |
| CPTAC.C3L-09082 | <NA> | <NA> | White | human | 1953 | <NA> | [GDC, PDC] | data_source age_at_observation diagnosis 0 GDC 69 Adenocarcinoma 1 PDC <NA> <NA> 2 GDC <NA> <NA> 3 PDC 69 Adenocarcinoma |
| (481 more rows not shown) | ||||||||
Now each row has an embedded dataframe of all the relevent observation info, lets look at the first row. Here I'm asking for the 'observation_data' column, and the first (zeroth) row:
sixty2seventy_collated['observation_data'][0]
Loading ITables v2.9.1 from the init_notebook_mode cell...
(need help?)
|
| ⓘdata_source | age_at_observation | diagnosis |
|---|---|---|
| IDC | 54 | Neoplasm, malignant |
| IDC | 54 | Neoplasm, malignant |
| IDC | 62 | Neoplasm, malignant |
| PDC | 62 | Endometrioid adenocarcinoma |
| GDC | 62 | Endometrioid adenocarcinoma |
| GDC | <NA> | <NA> |
| PDC | <NA> | <NA> |
| IDC | 54 | Neoplasm, malignant |
It looks like in the original source data, only some of the records included age information. I can look at even more detailed information if I add more columns to my search results table. Here, I'm adding the entire observation table by using asterisks again observation.* means 'anything that is inside the databases observation table'
all_obs_sixty2seventy_collated = get_subject_data(match_all=["60 < age_at_observation <= 70", "diagnosis = *adenocarcinoma*", "species = human"], add_columns='observation.*', collate_results=True)
all_obs_sixty2seventy_collated['observation_data'][0]
Loading ITables v2.9.1 from the init_notebook_mode cell...
(need help?)
|
| ⓘdata_source | age_at_observation | diagnosis | grade | morphology | observed_anatomic_site | resection_anatomic_site | sex | stage | vital_status | year_of_observation |
|---|---|---|---|---|---|---|---|---|---|---|
| IDC | 54 | Neoplasm, malignant | <NA> | <NA> | uterus | <NA> | <NA> | <NA> | <NA> | 2009 |
| IDC | 54 | Neoplasm, malignant | <NA> | <NA> | uterus | <NA> | female | <NA> | <NA> | 2009 |
| IDC | 62 | Neoplasm, malignant | <NA> | <NA> | uterus | <NA> | female | <NA> | <NA> | 2017 |
| PDC | 62 | Endometrioid adenocarcinoma | High Grade G3 | Endometrioid adenocarcinoma | uterus | body of uterus | <NA> | Stage IV | <NA> | 2017 |
| GDC | 62 | Endometrioid adenocarcinoma | High Grade G3 | Endometrioid adenocarcinoma | uterus | body of uterus | <NA> | Stage IV | <NA> | 2017 |
| GDC | <NA> | <NA> | <NA> | <NA> | <NA> | <NA> | female | <NA> | alive | <NA> |
| PDC | <NA> | <NA> | <NA> | <NA> | <NA> | <NA> | female | <NA> | alive | <NA> |
| IDC | 54 | Neoplasm, malignant | <NA> | <NA> | abdomen | <NA> | <NA> | <NA> | <NA> | 2009 |
This is very helpful output for assessing whether individual subjects that came back in my query have the types of data I need, and shows me where each bit of aggregated data came from, but it wouldn't be very effecient to look at each of these subjects one by one. If I want to look at this more detailed information for all the subjects, I can run another function on my results that expands