Connector gdc — NCI Genomic Data Commons (open aggregated data)
| Source | https://portal.gdc.cancer.gov — National Cancer Institute, Center for Cancer Genomics |
| Tier / category | 0 / genomics |
| Access | REST https://api.gdc.cancer.gov, no authentication, open-access tier only |
| Docs verified | 2026-09-08 — https://docs.gdc.cancer.gov/API/Users_Guide/ (Getting Started, Search and Retrieval, Appendix A) + live API |
| License | Open-access GDC data (NCI), citation of GDC and originating program requested — licenseStatus: approved |
| Code | packages/connectors/src/connectors/gdc/ (manifest.ts, normalize.ts, index.ts, gdc.test.ts, fixtures/) |
Data use (quoted)
GDC Policies — https://gdc.cancer.gov/about-gdc/gdc-policies (Data Sharing / Data Use section):
The GDC will also produce harmonized data (raw and derived) based on the originally submitted data including open access somatic variants. […] After data is released, either by submitter request to GDC, or by approval of the Center for Cancer Genomics, harmonized raw and GDC-generated derived data will be made available to the public via the GDC Data Portal and GDC data access tools.
Controlled-access data requires a dbGaP Data Use Certification and is never requested by this connector; only open aggregates are read (project summaries and the portal's "most frequently mutated genes" analysis). NCI asks that publications cite the GDC (Grossman RL et al., Toward a Shared Vision for Cancer Genomic Data, N Engl J Med 2016, PMID 27653561) and the originating program (TCGA, TARGET, CPTAC, …) — the program is stored on every cohort and repeated in the provenance population.
Verified endpoints
| Endpoint | Verified behaviour (2026-09-08) |
|---|---|
GET /status |
{"data_release":"Data Release 46.0 - August 10, 2026","data_release_version":{"major":46,"minor":0,"release_date":"2026-08-10"},"tag":"8.5.0","status":"OK"} → dataset version Data Release 46.0 |
GET /projects?size=100&from=0&fields=project_id,name,primary_site,disease_type,program.name,summary.case_count,summary.file_count |
93 projects; primary_site / disease_type are arrays |
GET /analysis/mutated_cases_count_by_project?size=0 |
aggregations.projects.buckets[]{key, doc_count, case_summary{case_with_ssm{doc_count}}} — cases with simple somatic mutation data per project (denominator). TCGA-PAAD 185 — cross-checked with GET /cases?size=0&filters={files.data_category in ["Simple Nucleotide Variation"], cases.project.project_id in ["TCGA-PAAD"]} → 185 |
GET /analysis/top_mutated_genes_by_project?filters=<JSON>&size=100&fields=symbol,gene_id |
hits {symbol, gene_id (Ensembl), _score} where _score = number of cases with ≥ 1 SSM in the gene. The project_id= query parameter is ignored by the current API (returns pan-GDC counts, TP53 = 9,070); the project must be passed as filters = {"op":"and","content":[{"op":"in","content":{"field":"cases.project.project_id","value":["TCGA-PAAD"]}}]} (TP53 = 124). |
Health check: /status → healthy when status == "OK".
Ingestion
/status→dataRelease./projects→genomic_cohorts(unique onsource_id + study_id):studyId = project_id,name,program,primarySites[],diseaseTypes[],caseCount = summary.case_count,casesWithSsm(from the mutated-cases aggregation),dataRelease,accessLevel open,url https://portal.gdc.cancer.gov/projects/<id>, run-level provenance (dataset 'GDC projects'). Anomaly guard: < 40 projects → refuse to persist. A source record per project (payload includescasesWithSsm).- For each project with
casesWithSsm > 0:top_mutated_genes_by_project(size 100) →cancer_gene_frequencies(unique oncohort_id + gene_symbol + alteration_type):alterationType 'ssm',casesAffected = _score,casesProfiled = casesWithSsm,frequency = affected / profiledvalidated byvalidateFrequency()(rejected rows counted as validation failures, never stored),rank(1-based order of the response),dataRelease,geneIdvia the sharedGeneCache(minimalgenesrow when the symbol is unknown;ensemblGeneIdfilled when empty),cancerIdof the cohort. One provenance row per project (sourceUrl= the analysis URL,dataset 'GDC top mutated genes by project',datasetVersion 'Data Release 46.0',evidenceType cohort,accessLevel open,cohortSize = casesWithSsm,methodology 'GDC analysis endpoint: cases with ≥1 simple somatic mutation in gene / cases with SSM data in project'). TTN, MUC16, CSMD3… appear as observed — no "passenger gene" filtering; the denominator is always shown.
dry_run fetches status, projects and denominators, then the top-5 genes of TCGA-PAAD without writing.
Cancer reconciliation
CancerResolver.byLabel over ordered candidates built by projectLabelCandidates(name) (normalize.ts): the project name first, then a documented rewrite table (TCGA/registry phrasing → NCIt-style labels), e.g. Glioblastoma Multiforme → Glioblastoma, Kidney Renal Clear Cell Carcinoma → Clear Cell Renal Cell Carcinoma, Kidney Chromophobe → Chromophobe Renal Cell Carcinoma, Lymphoid Neoplasm Diffuse Large B-cell Lymphoma → Diffuse Large B-Cell Lymphoma, Liver Hepatocellular Carcinoma → Hepatocellular Carcinoma, Stomach Adenocarcinoma → Gastric Adenocarcinoma, Multiple Myeloma CoMMpass Study → Multiple Myeloma, Functional Genomic Landscape of Acute Myeloid Leukemia → Acute Myeloid Leukemia… Rewrites that broaden the concept (Brain Lower Grade Glioma → Low Grade Glioma, High-Risk Wilms Tumor → Wilms Tumor, Pheochromocytoma and Paraganglioma → Paraganglioma, Cervical Squamous Cell Carcinoma and Endocervical Adenocarcinoma → Cervical Carcinoma, … Metastatic Breast Cancer … → Breast Cancer) record cancerMatchType = CURATED_BROADER; phrasing-only rewrites keep the resolver's match type (ONTOLOGY_EXACT / ALIAS). Resolved projects also get cancer_codes (system 'gdc_project', code = project_id).
Pan-cancer / basket projects (MATCH-*, FM-AD, CPTAC-*, HCMI-*, CCDI-*, EXCEPTIONAL_RESPONDERS, TRIO-*, or ≥ 8 primary sites with generic disease types) are never mapped to a single cancer (isPanCancerProject); they are still cohorts, and their unresolved-label entry carries panCancer: true so curators can ignore them.
Observed run (cancerindex_c, 2026-09-08, OncoTree-only ontology)
| Duration (full) | 28.9 s, ~180 HTTP requests |
| Cohorts | 93 (86 with casesWithSsm > 0) |
| Resolved to a cancer | 41 / 93 (44 %) — 35 ONTOLOGY_EXACT, 2 ALIAS (TCGA-KIRC → Renal Clear Cell Carcinoma, TCGA-READ → Rectal Adenocarcinoma), 4 CURATED_BROADER (TCGA-LGG, TCGA-PCPG, TARGET-WT, OHSU-CNL) |
| Pan-cancer (not mappable by design) | 24 |
| Unresolved | 28 — TCGA-GBM "Glioblastoma Multiforme", TCGA-ESCA "Esophageal Carcinoma", TCGA-MESO "Mesothelioma", TCGA-THCA "Thyroid Carcinoma", TCGA-OV "Ovarian Serous Cystadenocarcinoma", TCGA-TGCT "Testicular Germ Cell Tumors", TCGA-UCS "Uterine Carcinosarcoma", TCGA-CESC, TARGET-CCSK "Clear Cell Sarcoma of the Kidney", TARGET-RT "Rhabdoid Tumor", TARGET-ALL-P1/2/3, MP2PRT-ALL, MMRF-COMMPASS, CMI-MBC, CMI-MPC, WCDT-MCRPC, APOLLO-OV, APOLLO-BRCA-1, REBC-THYR, CGCI-HTMCP-CC/LC, ALCHEMIST-ALCH, ORGANOID-PANCREATIC, PECGS-COPECC, CCG-CUPP, VAREPOP-APOLLO — all have rewrite candidates whose target labels (Glioblastoma, Multiple Myeloma, Acute Lymphoblastic Leukemia, Mesothelioma…) simply do not exist as aliases in the OncoTree-only development ontology; they resolve once NCIt is loaded. |
| Frequency rows | 6,100 in 61 cohorts (3,754 distinct genes; denominators 20 – 3,033), 0 rejected by validateFrequency |
| Cohorts with SSM but no open top-gene aggregate | 25 (FM-AD, all MATCH-* arms, WCDT-MCRPC, APOLLO-OV, ORGANOID-PANCREATIC, OHSU-CNL, BEATAML1.0-CRENOLANIB, VAREPOP-APOLLO, TARGET-ALL-P1…): their SSM files are controlled-access only (files.access = open → 0 cases), so the analysis endpoint returns no hits — nothing is stored, which is the intended open-only behaviour. |
TCGA-PAAD top 5 (cases_affected / cases_profiled): TP53 124/185 (0.670), KRAS 123/185 (0.665), SMAD4 119/185 (0.643), CDKN2A 117/185 (0.632), ME2 114/185 (0.616).
Idempotent: re-running updates cohorts and frequencies in place (unique indexes), reuses cohort ids and creates no duplicate genes.
Limitations / notes
- Only
alteration_type = 'ssm'; CNV / fusion frequencies are not exposed by an open aggregate endpoint. - The
_scoresemantics ("cases with ≥ 1 SSM in the gene") is inferred from the portal's usage of this endpoint and validated numerically against the/casesdenominator; the GDC API guide does not document/analysis/*endpoints formally. - Projects whose SSMs are controlled-access have cohorts but no frequencies (see table).
casesWithSsmfor a project can exceed the number of cases with open SSM files (e.g. TCGA-DLBC 50 vs 47 open); the denominator used is the GDC's owncase_with_ssmcount, matching the portal.