Find subjects using a file as search input
I have been working with data from IDC for a cohort of 100 individuals, and I'd like to see if any other data is available about my subjects, and where. I want to submit a file that has all of my individuals IDs so I can search them all at once. My file looks like this:
So, I want to match the subject column in mydatafile.tsv to the cda column for subject_id. And since I want to see where data exists about my subjects, I'm doing add_columns and telling it upstream_identifiers.*. * means "anything", so this will add all of the upstream identifier columns so I can see where each of my subjects have data:
get_subject_data(match_from_file = {'input_column': 'subject', 'input_file': 'mydatafile.tsv', 'cda_column_to_match':'subject_id'}, add_columns = 'upstream_identifiers.*')
Loading ITables v2.9.1 from the init_notebook_mode cell...
(need help?)
|
| ⓘsubject_id | cause_of_death | ethnicity | race | species | year_of_birth | year_of_death | data_source | upstream_source | upstream_field | upstream_id |
|---|---|---|---|---|---|---|---|---|---|---|
| CPTAC.01BR001 | <NA> | Non-Hispanic | Black or African American | human | <NA> | <NA> | [GC, GDC, IDC, PDC] | [GC, GDC, IDC, PDC] | [Case.case_id, Case.case_submitter_id, auxiliary_metadata.submitter_case_id, case.case_id, case.submitter_id, dicom_all.PatientID, dicom_all.idc_case_id, participant.dbGaP_subject_id, participant.participant_id, participant.uuid] | [01BR001, 094977cc-0adb-4005-b649-f6ed17f6ec3d, 1362202, 327f4fe6-0a5d-11eb-bc0e-0aad30af8a83, be37f1f7-2f98-4f74-bc04-6dd2ae2afcad, bf9d0398-edbd-5c34-98d2-b49963a62ed2] |
| CPTAC.06BR003 | <NA> | Hispanic or Latino | <NA> | human | <NA> | <NA> | [GC, GDC, IDC, PDC] | [GC, GDC, IDC, PDC] | [Case.case_id, Case.case_submitter_id, auxiliary_metadata.submitter_case_id, case.case_id, case.submitter_id, dicom_all.PatientID, dicom_all.idc_case_id, participant.dbGaP_subject_id, participant.participant_id, participant.uuid] | [06BR003, 1362235, 27e868aa-dbf2-42b3-889d-72f17787b965, 327f9271-0a5d-11eb-bc0e-0aad30af8a83, 5f12b87e-67f5-445c-9694-d70cdc776ced, f06c2a52-a376-5986-91f6-6424c6cceba7] |
| CPTAC.11BR013 | <NA> | Non-Hispanic | White | human | <NA> | <NA> | [GC, GDC, IDC, PDC] | [GC, GDC, IDC, PDC] | [Case.case_id, Case.case_submitter_id, auxiliary_metadata.submitter_case_id, case.case_id, case.submitter_id, dicom_all.PatientID, dicom_all.idc_case_id, participant.dbGaP_subject_id, participant.participant_id, participant.uuid] | [11BR013, 1607135, 327fae36-0a5d-11eb-bc0e-0aad30af8a83, 3b50aeb8-a8ac-5aa3-bdd6-9e00628fd7f6, 68ec7e12-7f4f-442b-a257-bf61f0819a5a, bd79a471-7f1c-4ae9-9a57-ce0ae36f0d88] |
| CPTAC.11BR036 | <NA> | Non-Hispanic | Asian | human | <NA> | <NA> | [GC, GDC, IDC, PDC] | [GC, GDC, IDC, PDC] | [Case.case_id, Case.case_submitter_id, auxiliary_metadata.submitter_case_id, case.case_id, case.submitter_id, dicom_all.PatientID, dicom_all.idc_case_id, participant.dbGaP_subject_id, participant.participant_id, participant.uuid] | [11BR036, 1826126, 327fd1dc-0a5d-11eb-bc0e-0aad30af8a83, b88a9c6a-0d92-563b-93dd-190d2f4d714a, de77f5f4-ae8f-4517-b80f-fbbe4a5b522c, f82b8134-65bc-47f5-9627-ca8b6ed9fc8b] |
| CPTAC.11BR044 | <NA> | Non-Hispanic | White | human | <NA> | <NA> | [GC, GDC, IDC, PDC] | [GC, GDC, IDC, PDC] | [Case.case_id, Case.case_submitter_id, auxiliary_metadata.submitter_case_id, case.case_id, case.submitter_id, dicom_all.PatientID, dicom_all.idc_case_id, participant.dbGaP_subject_id, participant.participant_id, participant.uuid] | [070d82aa-50f9-4be9-b3d0-a664b344e43e, 11BR044, 1826128, 247fa16a-b355-57b0-8f40-2e1f90e0b943, 327fd641-0a5d-11eb-bc0e-0aad30af8a83, b05e56ca-a092-489f-90bb-64246cae3e7d] |
| CPTAC.18BR009 | <NA> | Non-Hispanic | White | human | <NA> | <NA> | [GC, GDC, IDC, PDC] | [GC, GDC, IDC, PDC] | [Case.case_id, Case.case_submitter_id, auxiliary_metadata.submitter_case_id, case.case_id, case.submitter_id, dicom_all.PatientID, dicom_all.idc_case_id, participant.dbGaP_subject_id, participant.participant_id, participant.uuid] | [1826194, 18BR009, 327ff5b2-0a5d-11eb-bc0e-0aad30af8a83, 9678fc40-8167-47f9-90c0-bb7b559038d6, adc1d8b6-a4a2-4cc1-a662-a587e2708ad5, c5f7c308-b237-5078-a895-2a5f0dcbd191] |
| CPTAC.18BR019 | <NA> | Non-Hispanic | White | human | <NA> | <NA> | [GC, GDC, IDC, PDC] | [GC, GDC, IDC, PDC] | [Case.case_id, Case.case_submitter_id, auxiliary_metadata.submitter_case_id, case.case_id, case.submitter_id, dicom_all.PatientID, dicom_all.idc_case_id, participant.dbGaP_subject_id, participant.participant_id, participant.uuid] | [1826198, 18BR019, 327ff8f1-0a5d-11eb-bc0e-0aad30af8a83, 6f659e46-27ce-5296-b4b7-d5d200d66c55, 9b93240d-af41-463a-a87b-9cbfaaecf1a0, c94ca480-957d-4deb-ae9a-c2134c7c55e9] |
| CPTAC.20BR007 | <NA> | Non-Hispanic | Black or African American | human | <NA> | <NA> | [GC, GDC, IDC, PDC] | [GC, GDC, IDC, PDC] | [Case.case_id, Case.case_submitter_id, auxiliary_metadata.submitter_case_id, case.case_id, case.submitter_id, dicom_all.PatientID, dicom_all.idc_case_id, participant.dbGaP_subject_id, participant.participant_id, participant.uuid] | [1826204, 1eef079e-d5cd-4a0c-a9b5-525865aa7847, 20BR007, 327ffd6b-0a5d-11eb-bc0e-0aad30af8a83, 3633b395-4c00-5cff-a5bb-68df8f0e0928, 9a7a0b27-e79a-46b1-b0e4-319d8ef20155] |
| CPTAC.21BR001 | <NA> | Hispanic or Latino | White | human | <NA> | <NA> | [GC, GDC, IDC, PDC] | [GC, GDC, IDC, PDC] | [Case.case_id, Case.case_submitter_id, auxiliary_metadata.submitter_case_id, case.case_id, case.submitter_id, dicom_all.PatientID, dicom_all.idc_case_id, participant.dbGaP_subject_id, participant.participant_id, participant.uuid] | [21BR001, 2546272, 271ff03f-1367-4654-9d67-4456e1ff4eb3, 30df841d-5633-42ff-a0b1-b17335707564, 327fff0c-0a5d-11eb-bc0e-0aad30af8a83, 5286ded4-9f6c-5f94-b305-9a3d7268a2a4] |
| TCGA.TCGA-A2-A0CK | <NA> | Non-Hispanic | White | human | <NA> | <NA> | [GDC, IDC] | [GDC, IDC] | [auxiliary_metadata.submitter_case_id, case.case_id, case.submitter_id, dicom_all.PatientID, dicom_all.idc_case_id] | [5aca16ca-4516-4e27-8249-6914029a7ebf, TCGA-A2-A0CK, e94df7de-ab6a-4be3-ac9d-9a4e73d19161] |
| (90 more rows not shown) | ||||||||||
92 of my subjects have data in CDA, and it looks like they have data spread across GC, GDC, PDC, and IDC. Let's see a summary of what files and subjects there are:
summarize_files(match_from_file = {'input_column': 'subject', 'input_file': 'mydatafile.tsv', 'cda_column_to_match':'subject_id'})
╔════════════════════════════╗ ║ number_of_matching_files ║ ╠════════════════════════════╣ ║ 10655 ║ ╚════════════════════════════╝ ╔════════════════════════════════════════════════╗ ║ number_of_subjects_related_to_matching_files ║ ╠════════════════════════════════════════════════╣ ║ 100 ║ ╚════════════════════════════════════════════════╝ ╔═════════╦═══════════════╗ ║ files ║ data_source ║ ╠═════════╬═══════════════╣ ║ 7085 ║ GDC only ║ ║ 2696 ║ PDC only ║ ║ 824 ║ IDC only ║ ║ 50 ║ GC only ║ ╚═════════╩═══════════════╝ ╔════════════════╦══════════════════════════════════════════╗ ║ count_result ║ file_type ║ ╠════════════════╬══════════════════════════════════════════╣ ║ 1348 ║ Open Standard ║ ║ 1103 ║ Somatic Mutation Index ║ ║ 817 ║ Annotated Somatic Mutation ║ ║ 715 ║ Aligned Reads ║ ║ 674 ║ Proprietary ║ ║ 674 ║ Text ║ ║ 615 ║ Aligned Reads Index ║ ║ 518 ║ Raw Simple Somatic Mutation ║ ║ 392 ║ Transcript Fusion ║ ║ 317 ║ Segmentation Storage ║ ║ 294 ║ VL Whole Slide Microscopy Image Storage ║ ║ 256 ║ Structural Rearrangement ║ ║ 242 ║ Gene Level Copy Number ║ ║ 231 ║ Copy Number Segment ║ ║ 224 ║ Slide Image ║ ║ 190 ║ Masked Intensities ║ ║ 174 ║ Biospecimen Supplement ║ ║ 166 ║ Raw Intensities ║ ║ 166 ║ Simple Germline Variation ║ ║ 163 ║ Allele-specific Copy Number Segment ║ ║ 163 ║ Masked Copy Number Segment ║ ║ 105 ║ MR Image Storage ║ ║ 98 ║ Clinical Supplement ║ ║ 98 ║ Gene Expression Quantification ║ ║ 98 ║ Splice Junction Quantification ║ ║ 95 ║ Isoform Expression Quantification ║ ║ 95 ║ Methylation Beta Value ║ ║ 95 ║ miRNA Expression Quantification ║ ║ 83 ║ Pathology Report ║ ║ 80 ║ Microscopy Bulk Simple Annotations St... ║ ║ 77 ║ Aggregated Somatic Mutation ║ ║ 77 ║ Masked Somatic Mutation ║ ║ 68 ║ Intermediate Analysis Archive ║ ║ 66 ║ Protein Expression Quantification ║ ║ 50 ║ <NA> ║ ║ 16 ║ Advanced Blending Presentation State ... ║ ║ 11 ║ Comprehensive SR Storage ║ ║ 1 ║ Secondary Capture Image Storage ║ ╚════════════════╩══════════════════════════════════════════╝ ╔════════════════╦══════════════════════════════╗ ║ count_result ║ category ║ ╠════════════════╬══════════════════════════════╣ ║ 2630 ║ Simple Nucleotide Variation ║ ║ 1348 ║ Peptide Spectral Matches ║ ║ 1330 ║ Sequencing Reads ║ ║ 1033 ║ Copy Number Variation ║ ║ 674 ║ Processed Mass Spectra ║ ║ 674 ║ Raw Mass Spectra ║ ║ 448 ║ Structural Variation ║ ║ 398 ║ Biospecimen ║ ║ 386 ║ Transcriptome Profiling ║ ║ 328 ║ Somatic Structural Variation ║ ║ 317 ║ Segmentation ║ ║ 294 ║ Slide Microscopy ║ ║ 285 ║ DNA Methylation ║ ║ 203 ║ <NA> ║ ║ 106 ║ Magnetic Resonance ║ ║ 80 ║ Annotation ║ ║ 66 ║ Proteome Profiling ║ ║ 16 ║ Presentation State ║ ║ 16 ║ WXS ║ ║ 11 ║ Structured Report Document ║ ║ 9 ║ RNA-Seq ║ ║ 2 ║ WGS ║ ║ 1 ║ miRNA-Seq ║ ╚════════════════╩══════════════════════════════╝ ╔════════════════╦═════════════╗ ║ count_result ║ format ║ ╠════════════════╬═════════════╣ ║ 1567 ║ TSV ║ ║ 1108 ║ VCF ║ ║ 1103 ║ TBI ║ ║ 842 ║ TXT ║ ║ 824 ║ DICOM ║ ║ 715 ║ BAM ║ ║ 674 ║ <NA> ║ ║ 674 ║ mzIdentML ║ ║ 674 ║ mzML ║ ║ 615 ║ BAI ║ ║ 523 ║ MAF ║ ║ 324 ║ BEDPE ║ ║ 224 ║ SVS ║ ║ 190 ║ IDAT ║ ║ 166 ║ CEL ║ ║ 164 ║ BCR XML ║ ║ 83 ║ PDF ║ ║ 82 ║ BCR SSF XML ║ ║ 68 ║ TAR ║ ║ 19 ║ BCR Biotab ║ ║ 8 ║ GCT ║ ║ 7 ║ BCR OMF XML ║ ║ 1 ║ CSV ║ ╚════════════════╩═════════════╝ ╔════════════════╦════════════╗ ║ count_result ║ access ║ ╠════════════════╬════════════╣ ║ 5614 ║ open ║ ║ 3323 ║ controlled ║ ║ 1718 ║ <NA> ║ ╚════════════════╩════════════╝ ╔════════════════╦══════════════════╗ ║ count_result ║ anatomic_site ║ ╠════════════════╬══════════════════╣ ║ 9831 ║ <NA> ║ ║ 772 ║ breast ║ ║ 48 ║ head of pancreas ║ ║ 48 ║ pancreas ║ ║ 4 ║ colon ║ ║ 1 ║ ascending colon ║ ║ 1 ║ descending colon ║ ║ 1 ║ rectum ║ ║ 1 ║ transverse colon ║ ╚════════════════╩══════════════════╝ ╔════════════════╦═══════════════════╗ ║ count_result ║ tumor_vs_normal ║ ╠════════════════╬═══════════════════╣ ║ 9193 ║ tumor ║ ║ 4325 ║ normal ║ ║ 452 ║ <NA> ║ ╚════════════════╩═══════════════════╝ ╔════════════════╦══════════════╗ ║ ║ size ║ ╠════════════════╬══════════════╣ ║ mean ║ 4233534066 ║ ║ min ║ 72 ║ ║ lower quartile ║ 66397 ║ ║ median ║ 3432701 ║ ║ upper quartile ║ 85253660 ║ ║ max ║ 480722740993 ║ ╚════════════════╩══════════════╝
summarize_subjects(match_from_file = {'input_column': 'subject', 'input_file': 'mydatafile.tsv', 'cda_column_to_match':'subject_id'})
╔═══════════════════════════════╗ ║ number_of_matching_subjects ║ ╠═══════════════════════════════╣ ║ 100 ║ ╚═══════════════════════════════╝ ╔════════════════════════════════════════════════╗ ║ number_of_files_related_to_matching_subjects ║ ╠════════════════════════════════════════════════╣ ║ 10655 ║ ╚════════════════════════════════════════════════╝ ╔════════════╦══════════════════════╗ ║ subjects ║ data_source ║ ╠════════════╬══════════════════════╣ ║ 72 ║ GDC + IDC ║ ║ 9 ║ PDC + GDC + GC + IDC ║ ║ 8 ║ IDC only ║ ║ 5 ║ PDC + GDC + IDC ║ ║ 5 ║ GDC + GC + IDC ║ ║ 1 ║ PDC + IDC ║ ╚════════════╩══════════════════════╝ ╔════════════════╦════════════════════╗ ║ count_result ║ ethnicity ║ ╠════════════════╬════════════════════╣ ║ 75 ║ Non-Hispanic ║ ║ 20 ║ <NA> ║ ║ 5 ║ Hispanic or Latino ║ ╚════════════════╩════════════════════╝ ╔════════════════╦═══════════╗ ║ count_result ║ species ║ ╠════════════════╬═══════════╣ ║ 92 ║ human ║ ║ 8 ║ <NA> ║ ╚════════════════╩═══════════╝ ╔════════════════╦═══════════════════════════╗ ║ count_result ║ race ║ ╠════════════════╬═══════════════════════════╣ ║ 67 ║ White ║ ║ 16 ║ <NA> ║ ║ 11 ║ Black or African American ║ ║ 6 ║ Asian ║ ╚════════════════╩═══════════════════════════╝ ╔════════════════╦══════════════════╗ ║ count_result ║ cause_of_death ║ ╠════════════════╬══════════════════╣ ║ 100 ║ <NA> ║ ╚════════════════╩══════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_birth ║ ╠════════════════╬═════════════════╣ ║ mean ║ 1950 ║ ║ min ║ 1938 ║ ║ lower quartile ║ 1943 ║ ║ median ║ 1950 ║ ║ upper quartile ║ 1958 ║ ║ max ║ 1963 ║ ╚════════════════╩═════════════════╝ ╔════════════════╦═════════════════╗ ║ ║ year_of_death ║ ╠════════════════╬═════════════════╣ ║ mean ║ 2009 ║ ║ min ║ 2008 ║ ║ lower quartile ║ 2008 ║ ║ median ║ 2008 ║ ║ upper quartile ║ 2009 ║ ║ max ║ 2009 ║ ╚════════════════╩═════════════════╝