SPB Git forge

spb/cancerindex

Public
37commits 1branches 0releases
2.9 MBsize
maindefault branch
10 days agolast push
TypeScript 97.2% SQL 1.5% CSS 0.6% JavaScript 0.5%

SDK HttpClient: body-read timeout with retry (stalled ChEMBL responses)

Simon-Pierre Boucher committed 16 days ago (Sep 8, 2026) parent f2c4fe3

6 changed files +109 −12

modified docs/connectors/cbioportal.md +1 −1
@@ -41,7 +41,7 @@ Health check: `GET /api/studies?pageSize=1&projection=SUMMARY` → healthy when
41 41 1. `/studies` (DETAILED, paged) + `/cancer-types`. Anomaly guard: < 100 studies → refuse to persist. `datasetVersion = cbioportal-<latest importDate>`.
42 42 2. Studies are processed in `studyId` order; `ctx.cursor.lastStudyId` is saved after each study (restartable, time budget honoured between studies); a completed pass starts over on the next schedule.
43 43 3. **Cohort** (`genomic_cohorts`, unique on `source_id + study_id`): `studyId`, `name`, `program` = first token of the name's trailing parenthetical (`MSK-IMPACT Clinical Sequencing Cohort (MSK, Nat Med 2017)`**MSK**; `… (TCGA, PanCancer Atlas)`**TCGA**), `primarySites` = OncoTree node just below the `tissue` root (luad → nsclc → lung → **Lung**), `diseaseTypes` = cancer-type name, `caseCount = allSampleCount`, `casesWithSsm = sequencedSampleCount`, `dataRelease = importDate (date)`, `accessLevel open`, `url https://www.cbioportal.org/study/summary?id=<studyId>`, provenance per study (`sourceUrl /api/studies/<id>`, `dataset 'cBioPortal public studies'`, `pmid`, `cohortSize`, `population '<studyId> (<program>)'`, methodology with the study citation). An unchanged study keeps its previous provenance row.
44 4. **Cancer reconciliation**: `resolver.byCode('oncotree', CANCERTYPEID.toUpperCase())``EXACT_IDENTIFIER`; fallback `byLabel(cancer-type name)` (match type recorded, and the code stored in `cancer_codes` with that match type); `cancerTypeId = mixed``cancerId null`, `cancerMatchType 'UNRESOLVED'` (pan-cancer, not queued — by design); other misses → `unresolved_labels` with the OncoTree lineage.
44 +4. **Cancer reconciliation**: `resolver.byCode('oncotree', CANCERTYPEID.toUpperCase())``EXACT_IDENTIFIER`; fallback `byLabel(cancer-type name)` (match type recorded, and the code stored in `cancer_codes` with that match type); studies filed under a **tissue-level node** (direct child of the OncoTree root: "Breast", "Prostate", "Bladder/Urinary Tract", "Soft Tissue", "Bowel"…) are rewritten into site-level malignancy labels (`tissueLevelCandidates`: *Breast → Breast Cancer / Malignant Breast Neoplasm / Breast Carcinoma*, *Soft Tissue → Soft Tissue Sarcoma*, *Bowel → Colorectal Cancer*) and recorded as `CURATED_BROADER`; `cancerTypeId = mixed``cancerId null`, `cancerMatchType 'UNRESOLVED'` (pan-cancer, not queued — by design); other misses → `unresolved_labels` with the OncoTree lineage.
45 45 5. **Publication stub** for `pmid` (`publications`, `publicationTypes ['stub']`, title = citation) + `publication_entity_edges` to the cohort's cancer (`method cbioportal_study`) — enriched later by the PubMed connector.
46 46 6. **Frequencies** (`cancer_gene_frequencies`, unique on `cohort_id + gene_symbol + alteration_type`) for studies with `sequencedSampleCount > 0` **and** a `MUTATION_EXTENDED` profile (`/molecular-profiles` and `/mutated-genes/fetch` are requested concurrently — concurrency 2): genes ranked by `numberOfAlteredCases` (desc, symbol asc), **top 200 per study** (`TOP_GENES_PER_STUDY`; WES studies return 7,000–18,000 genes), genes with `numberOfProfiledCases = 0` dropped. `alterationType 'ssm'`, `casesAffected = numberOfAlteredCases`, `casesProfiled = numberOfProfiledCases` (gene-panel aware; equals the `<studyId>_sequenced` count for exome studies), `frequency` validated by `validateFrequency()` (rejected rows counted, never stored), `rank`, `dataRelease`, `geneId` via the shared `GeneCache` (Entrez id filled when empty), `cancerId` of the cohort. Provenance per study (`sourceUrl /api/mutated-genes/fetch`, `dataset 'cBioPortal mutated genes by study'`, `pmid`, `cohortSize = sequencedSampleCount`, methodology: *distinct samples with ≥ 1 mutation in gene / samples profiled for the gene (numberOfProfiledCases; = samples in `<studyId>_sequenced` for whole-exome studies)*, profile id, top-N note). Source record `mutated_genes` per study with the ranked list.
47 47 7. Studies without a mutation profile are cohorts without frequencies (`noMutationProfile` in the run summary); no per-mutation paging is ever needed, so no sample cap applies.
modified docs/connectors/openfda.md +26 −4
@@ -69,13 +69,35 @@ Built once per run from `cancer_aliases` of active, malignant concepts with `top
69 69 - alias length ≥ 6 characters, or ≥ 4 for curated `abbreviation` aliases (NSCLC, HNSCC, GIST, DLBCL); 3-letter abbreviations (AML, CML, ALL) are excluded;
70 70 - generic words are stop-listed (`cancer`, `tumor`, `carcinoma`, `solid tumor`, `leukemia`, `lymphoma`, `sarcoma`, `adenocarcinoma`, `squamous cell carcinoma`, `carcinoma in situ`, `metastatic disease`…, `GENERIC_ALIAS_STOPLIST`);
71 71 - an alias shared by several concepts is kept only when `CancerResolver.byLabel` disambiguates it (preferred name, curated display name, broadest concept of one lineage — e.g. "breast cancer" → Malignant Breast Neoplasm); otherwise it is dropped (`ambiguousDropped`);
72 - matching is a whole-word n-gram lookup over `normalizeLabel(bullet)`; overlapping mentions resolve to the **longest** span ("metastatic breast cancer" beats "breast cancer"), then the bullet is mapped only if **exactly one** distinct cancer remains. "Ph+ CML in blast crisis" → CML; "HES and/or CEL" (two cancers) → null; "melanoma or breast cancer" → null.
72 +- matching is a whole-word n-gram lookup over `normalizeLabel(bullet)`; overlapping mentions resolve to the **longest** span ("metastatic breast cancer" beats "breast cancer"), then the bullet is mapped only if **exactly one** distinct cancer remains — after the `CancerReconciler` has collapsed (a) **equivalent rows** minted by two terminologies (the canonical name of one is an alias of the other, e.g. OncoTree "Non-Small Cell Lung Cancer" and NCIt "Lung Non-Small Cell Carcinoma"; the NCIt-coded row wins) and (b) candidates on **one lineage** to the **narrowest** concept (an indication sentence names the specific disease and uses broader words as context: "Head and Neck Squamous Cell Cancer (HNSCC) … squamous cell cancer" → HNSCC; "HES and/or CEL" → CEL, a descendant of HES in NCIt). "Melanoma or breast cancer" (different diseases) → null. The reconciler also settles aliases shared by equivalent rows at dictionary build time (282 → 19 ambiguous aliases dropped on the dev DB).
73 73
74 Every mapped row is `PROBABILISTIC` (text mention, not an identifier) and says so in `raw.cancer_match`; curators can review `drug_approvals WHERE cancer_id IS NULL` (`raw.bullets[].distinctCancers`).
74 +Every mapped row is `PROBABILISTIC` (text mention, not an identifier) and says so in `raw.cancer_match` (`via` adds `(equivalent)` / `(narrowest_of_lineage)` when the reconciler intervened, `distinctMentions` counts the raw mentions); curators can review `drug_approvals WHERE cancer_id IS NULL` (`raw.bullets[].distinctCancers`).
75 75
76 ## Observed run (cancerindex_b, 2026-09-08)
76 +## Replay without HTTP (`--mode backfill`)
77 77
78 See the run summary appended below after the first 15-minute run.
78 +`pnpm cix run openfda --mode backfill` re-derives every approval row and edge from the payloads already in the raw lake (`source_records.raw_path``RawLake.read`), applying the current dictionary / reconciler rules: no request, cursor untouched, ~15 s for 1,500 records. One application can serve several drug rows ("Abiraterone" / "Abiraterone Acetate", "Osimertinib" / "Osimertinib Mesylate"); the source record keeps only the last canonical id, so the backfill also consults existing `drug_approvals (drug_id, application_number)` pairs. This keeps RAW → NORMALIZED replayable (CLAUDE.md §2) and is the way to refresh mappings when the daily quota is exhausted.
79 +
80 +## Observed run (cancerindex_b, 2026-09-08 — 661 drugs from CIViC therapies, NCIt + HGNC + CIViC loaded)
81 +
82 +| | |
83 +|---|---|
84 +| Requests | 951 (laptop IP, self-stopped at the no-key daily budget after drug 549/661, 5 min 24 s) + 234 (remaining 112 drugs, 77 s, run from a cluster node) = **1,185 requests for one full pass**; `http_failures` 624 are the 404 "no match" answers the SDK counts before the connector treats them as empty |
85 +| Drugs | 661 seen — **208 matched** (≥ 1 accepted Drugs@FDA application), 137 skipped as class/regimen names, 316 unresolved (`no_drugsfda_match`: investigational agents such as Ganetespib, Seribantumab, RapaLink-1, codes like ARS-1620, trial ids…) |
86 +| Applications | 1,221 accepted (316 reference NDA/BLA, 905 ANDA/biosimilar) → 1,200 `application` source records + 332 `label` records |
87 +| Approval rows | **1,902** (`drug_approvals`, US/FDA): 606 ORIG, 1,296 SUPPL (efficacy supplements); **638 rows carry a `cancerId`** (all PROBABILISTIC via label text; 587 single-mention bullets, 51 collapsed by the narrowest-of-lineage / equivalence rule), 1,263 without (0 or ≥ 2 cancers per bullet, tumor-agnostic, or supplements); 15 tumor-agnostic, 32 `accelerated = true` from label wording |
88 +| Drugs with ≥ 1 cancer-mapped row | 161 of 208 |
89 +| Knowledge edges | 638 `drug APPROVED_FOR cancer` (`regulatory_status`, `FDA ORIG`) |
90 +| Ambiguous bullets (≥ 2 cancers, kept unmapped) | 103 |
91 +| Backfill (`--mode backfill`, lake replay, 0 requests) | 15 s for 1,500 lake records — used twice today after mapping-rule changes |
92 +
93 +Examples:
94 +
95 +- **Osimertinib** — NDA208065 (reference; TAGRISSO; UNII 3C06JJ0Z2O filled): ORIG 2015-11-13 split into 5 bullets, all → *Lung Non-Small Cell Carcinoma*; 8 efficacy supplements 2017-03-30 … 2024-09-25 (dated, `cancerId null`).
96 +- **Pembrolizumab** — BLA125514 (KEYTRUDA; BLA761467 KEYTRUDA QLEX rejected as a fixed combination with berahyaluronidase alfa): ORIG 2014-09-04 → 36 bullets, 35 mapped (Melanoma ×2, NSCLC ×6, HNSCC ×4, cHL, PMBCL, urothelial carcinoma, bladder ×2, MSI-H colorectal, gastric, esophageal, cervical ×3, HCC, biliary tract, Merkel cell, RCC ×2, endometrial ×3, TNBC ×3, mesothelioma); 1 unmapped row (paediatric cHL wording + the MSI-H/dMMR and TMB-H **tumor-agnostic** bullets, `tumorAgnostic = true`, `accelerated = true`); 108 efficacy supplements 2015-10-02 … 2026-07-10.
97 +- **Imatinib** — NDA021588 Gleevec (tablets, ORIG 2003-04-18) and NDA219097 Imkeldi (2024-11-22) as reference; 10 ANDAs as source records only. ORIG bullets → Ph+ ALL ×2, GIST (metastatic + adjuvant), DFSP, systemic mastocytosis, MDS/MPN, CEL; the "HES and/or CEL" bullet resolved to CEL by the narrowest-of-lineage rule; the Ph+ CML bullet stays unmapped because "chronic myeloid leukemia" is only a molecular_subtype alias ("…, BCR-ABL1 Positive") whose ambiguity the resolver cannot settle — a curation item. 19 efficacy supplements 2003-05-20 … 2016-08-25.
98 +- Counters: after `pnpm cix counters`, **407 cancers** have `approved_drug_count > 0`; top-level leaders: Malignant Lung Neoplasm 43 drugs, Malignant Breast Neoplasm 35, Leukemia 22, Malignant Kidney Neoplasm 15, Malignant Prostate Neoplasm 14, Malignant Liver Neoplasm 14.
99 +
100 +Idempotent: re-running updates rows in place (`raw.key`), reuses provenance when neither the application nor the label changed, and only removes rows of an application that no longer derive from its freshly fetched payload.
79 101
80 102 ## Limitations / notes
81 103
modified packages/connectors/src/connectors/cbioportal/cbioportal.test.ts +8 −1
@@ -3,7 +3,7 @@ import path from 'node:path';
3 3 import { describe, expect, it } from 'vitest';
4 4 import { validateFrequency } from '../../sdk/validate.js';
5 5 import { TOP_GENES_PER_STUDY, manifest, molecularProfilesUrl, mutatedGenesUrl, studiesUrl, studyUrl } from './manifest.js';
6 import { CancerType, MolecularProfile, MutatedGene, SampleList, Study, findMutationProfile, importDateToRelease, mutatedGenesBody, oncotreeCode, programFromStudyName, rankMutatedGenes, typeLineage } from './normalize.js';
6 +import { CancerType, MolecularProfile, MutatedGene, SampleList, Study, findMutationProfile, importDateToRelease, mutatedGenesBody, oncotreeCode, programFromStudyName, rankMutatedGenes, tissueLevelCandidates, typeLineage } from './normalize.js';
7 7
8 8 const fx = (name: string) => JSON.parse(readFileSync(path.join(import.meta.dirname, 'fixtures', name), 'utf8')) as unknown;
9 9
@@ -52,6 +52,13 @@ describe('studies and cancer types', () => {
52 52 expect(typeLineage('mixed', types)).toMatchObject({ diseaseType: 'Mixed Cancer Types', primarySite: 'Other' });
53 53 expect(typeLineage('nope', types)).toMatchObject({ diseaseType: null, primarySite: null, chain: ['nope'] });
54 54 });
55 + it('rewrites tissue-level nodes into site-level malignancy labels (broadening)', () => {
56 + expect(tissueLevelCandidates('Breast')).toEqual(['Breast Cancer', 'Malignant Breast Neoplasm', 'Breast Carcinoma']);
57 + expect(tissueLevelCandidates('Bladder/Urinary Tract')[0]).toBe('Bladder Cancer');
58 + expect(tissueLevelCandidates('Soft Tissue')[0]).toBe('Soft Tissue Sarcoma');
59 + expect(tissueLevelCandidates('Bowel')[0]).toBe('Colorectal Cancer');
60 + expect(tissueLevelCandidates('')).toEqual([]);
61 + });
55 62 it('rejects a malformed study and accepts an empty page', () => {
56 63 expect(Study.safeParse({ name: 'x' }).success).toBe(false);
57 64 expect(Study.safeParse({ studyId: 'x', cancerTypeId: 'acc', name: 'x', allSampleCount: -1 }).success).toBe(false);
modified packages/connectors/src/connectors/cbioportal/index.ts +10 −2
@@ -6,7 +6,7 @@ import { Connector, type ConnectorHealth, type RunContext } from '../../sdk/run.
6 6 import { validateFrequency } from '../../sdk/validate.js';
7 7 import { GeneCache } from '../civic/genes.js';
8 8 import { CBIO_API, MIN_EXPECTED_STUDIES, STUDIES_PAGE_SIZE, TOP_GENES_PER_STUDY, cancerTypesUrl, manifest, molecularProfilesUrl, mutatedGenesUrl, studiesUrl, studyApiUrl, studyUrl } from './manifest.js';
9 import { CancerType, MolecularProfile, MutatedGene, Study, findMutationProfile, importDateToRelease, mutatedGenesBody, oncotreeCode, programFromStudyName, rankMutatedGenes, typeLineage } from './normalize.js';
9 +import { CancerType, MolecularProfile, MutatedGene, Study, findMutationProfile, importDateToRelease, mutatedGenesBody, oncotreeCode, programFromStudyName, rankMutatedGenes, tissueLevelCandidates, typeLineage } from './normalize.js';
10 10
11 11 interface CbioCursor {
12 12 pass?: string;
@@ -142,9 +142,17 @@ export class CbioportalConnector extends Connector {
142 142 if (!code) return { cancerId: null, matchType: 'UNRESOLVED', via: `cancerTypeId "${study.cancerTypeId}" (pan-cancer / mixed)` };
143 143 const byCode: CancerMatch | null = resolver.byCode('oncotree', code);
144 144 if (byCode) return { cancerId: byCode.cancerId, matchType: byCode.matchType, via: byCode.via };
145 const name = types.get(study.cancerTypeId)?.name ?? study.cancerType?.name;
145 + const node = types.get(study.cancerTypeId);
146 + const name = node?.name ?? study.cancerType?.name;
146 147 const byLabel = name ? resolver.byLabel(name) : null;
147 148 if (byLabel) return { cancerId: byLabel.cancerId, matchType: byLabel.matchType, via: byLabel.via };
149 + // Tissue-level node (direct child of the root): site → site-level malignancy, always a broadening.
150 + if (name && (node?.parent ?? study.cancerType?.parent) === 'tissue') {
151 + for (const cand of tissueLevelCandidates(name)) {
152 + const hit = resolver.byLabel(cand);
153 + if (hit) return { cancerId: hit.cancerId, matchType: 'CURATED_BROADER', via: `tissue-level "${name}" → "${cand}" (${hit.via})` };
154 + }
155 + }
148 156 return { cancerId: null, matchType: 'UNRESOLVED', via: `oncotree:${code} unknown${name ? `, label "${name}" unmatched` : ''}` };
149 157 }
150 158
modified packages/connectors/src/connectors/cbioportal/normalize.ts +15 −0
@@ -131,6 +131,21 @@ export function typeLineage(cancerTypeId: string, types: Map<string, CancerType>
131 131 return { diseaseType: node?.name ?? null, primarySite: site?.name ?? null, chain };
132 132 }
133 133
134 +/**
135 + * Studies filed under a tissue-level OncoTree node ("Breast", "Prostate", "Bladder/Urinary Tract",
136 + * "Soft Tissue") have no disease code to match: the site name is rewritten into the site-level
137 + * malignancy labels the ontology does know. Any hit is a broadening (CURATED_BROADER).
138 + */
139 +export function tissueLevelCandidates(typeName: string): string[] {
140 + const first = typeName.split('/')[0]!.trim();
141 + if (!first) return [];
142 + if (/^soft tissue$/i.test(first)) return ['Soft Tissue Sarcoma', 'Soft Tissue Neoplasm'];
143 + if (/^bowel$/i.test(first)) return ['Colorectal Cancer', 'Malignant Colorectal Neoplasm'];
144 + if (/^cns\/brain$/i.test(typeName) || /^cns$/i.test(first)) return ['Malignant Brain Neoplasm', 'Brain Cancer'];
145 + if (/^(myeloid|lymphoid|blood)$/i.test(first)) return [`${first} Neoplasm`, `${first} Malignancy`];
146 + return [`${first} Cancer`, `Malignant ${first} Neoplasm`, `${first} Carcinoma`];
147 +}
148 +
134 149 /** Rank genes by altered samples (desc), symbol (asc); drop genes without a positive denominator. */
135 150 export function rankMutatedGenes(genes: MutatedGene[], top: number): Array<MutatedGene & { rank: number }> {
136 151 return genes
modified packages/connectors/src/sdk/http.ts +49 −4
@@ -8,6 +8,16 @@ export interface HttpStats {
8 8 retries: number;
9 9 }
10 10
11 +export class BodyTimeoutError extends Error {
12 + constructor(
13 + public readonly url: string,
14 + public readonly timeoutMs: number,
15 + ) {
16 + super(`body read timed out after ${timeoutMs} ms for ${url}`);
17 + this.name = 'BodyTimeoutError';
18 + }
19 +}
20 +
11 21 export class HttpError extends Error {
12 22 constructor(
13 23 public readonly status: number,
@@ -141,14 +151,49 @@ export class HttpClient {
141 151 }
142 152 }
143 153
154 + /**
155 + * Read a body under the same time budget as the headers. Without this a stalled body (server
156 + * accepted the request, sent headers, then went silent — observed with ChEMBL) hangs forever
157 + * because the AbortController is cleared once headers arrive.
158 + */
159 + private async readBody<T>(res: Response, url: string, read: (r: Response) => Promise<T>): Promise<T> {
160 + let timer: ReturnType<typeof setTimeout> | undefined;
161 + const timeout = new Promise<never>((_, reject) => {
162 + timer = setTimeout(() => {
163 + res.body?.cancel().catch(() => {});
164 + reject(new BodyTimeoutError(url, this.timeoutMs));
165 + }, this.timeoutMs);
166 + });
167 + try {
168 + return await Promise.race([read(res), timeout]);
169 + } finally {
170 + clearTimeout(timer);
171 + }
172 + }
173 +
174 + /** Request + body read with retries on body stalls (same backoff as request()). */
175 + private async withBody<T>(url: string, init: RequestInit | undefined, read: (r: Response) => Promise<T>): Promise<T> {
176 + for (let attempt = 0; ; attempt++) {
177 + const res = await this.request(url, init);
178 + try {
179 + return await this.readBody(res, url, read);
180 + } catch (err) {
181 + if (!(err instanceof BodyTimeoutError) || attempt >= this.retry.maxRetries) {
182 + this.stats.failures++;
183 + throw err;
184 + }
185 + this.stats.retries++;
186 + await sleep(this.backoff(attempt));
187 + }
188 + }
189 + }
190 +
144 191 async json<T = unknown>(url: string, init?: RequestInit): Promise<T> {
145 const res = await this.request(url, init);
146 return (await res.json()) as T;
192 + return this.withBody(url, init, (r) => r.json() as Promise<T>);
147 193 }
148 194
149 195 async text(url: string, init?: RequestInit): Promise<string> {
150 const res = await this.request(url, init);
151 return await res.text();
196 + return this.withBody(url, init, (r) => r.text());
152 197 }
153 198
154 199 async postJson<T = unknown>(url: string, body: unknown, init: RequestInit = {}): Promise<T> {
155 200