SPB Git forge

spb/cancerindex

Public
37commits 1branches 0releases
2.9 MBsize
maindefault branch
10 days agolast push
TypeScript 97.2% SQL 1.5% CSS 0.6% JavaScript 0.5%
10.9 KB

# Methodology — Clinical trial intelligence

Formula version ci-trial-intel-v1 · stop-reason rules ci-stop-reasons-v1 · code packages/ranking/src/trial-intelligence.ts, packages/ranking/src/trial-stop-reasons.ts · table trial_intelligence · recomputed by pnpm cix intel after each ClinicalTrials.gov ingest.

Every value described here is a computed metric (claim category computed_metric). Nothing in this layer is an observation: statuses, phases, dates, sponsors, countries, enrollment counts and "why stopped" texts are taken as posted by the registrant on ClinicalTrials.gov, then counted. Each row stores its formula_version and an inputs JSON (reference day, windows, thresholds, active statuses, denominators, burden observation ids) so any figure can be traced back to its inputs.

# 1. Scope and aggregation

# 1.1 Which studies attach to a cancer

A study is attached to a cancer when one of its reconciled conditions (trial_conditions.cancer_id) points to the cancer or to any of its descendants in the NCIt-derived hierarchy (cancer_hierarchy, recursive traversal, depth ≤ 12). Each study is counted once per cancer (DISTINCT), even when several of its conditions map to different descendants. This is the same traversal used by entity_counters and the ranking engine, so active_trials here equals entity_counters.active_trial_count.

Caveat (descendant aggregation). A study registered against a broad condition ("solid tumors", "lymphoma") is attributed to the broad entity only, never pushed down to subtypes. Conversely a study registered against a subtype ("lung adenocarcinoma") also counts for every ancestor ("non-small cell lung carcinoma", "lung cancer"). Figures for a parent are therefore not the sum of its children, and entities at level all overlap: adding them double counts.

# 1.2 Entity levels

entity_level Rows Purpose
top every active cancer with top_level = true (36 today), including true zeros mutually exclusive set — totals may be summed; rankings
all every active malignant entity with ≥ 1 mapped study of any study type subtype pages, exploration; overlapping

A top-level cancer has one row per level with identical figures.

# 1.3 Study type and statuses

Unless stated otherwise, every count uses study_type = 'INTERVENTIONAL' (observational and expanded-access records are excluded). Active means overall status ∈ {RECRUITING, NOT_YET_RECRUITING, ENROLLING_BY_INVITATION, ACTIVE_NOT_RECRUITING} — the same list as entity_counters. UNKNOWN status is neither active nor terminal.

# 2. Counts

Column Definition
total_trials interventional studies mapped to the entity or a descendant, any status
active_trials status ∈ active set
recruiting_trials status = RECRUITING
phase1_active … phase4_active active studies with PHASEn ∈ phases. A PHASE2|PHASE3 study counts in both phase 2 and phase 3. EARLY_PHASE1 is not counted as phase 1.
phase3_recruiting status = RECRUITING and PHASE3 ∈ phases
completed_trials, terminated_trials, withdrawn_trials, suspended_trials by overall status
with_results has_results = true

# 3. Growth

Registration date = first_posted_date. With asOf = the day of computation (stored in inputs.asOf) and calendar-month arithmetic (day clamped to the month end):

text
new_trials_12m       = studies with first_posted_date ∈ [asOf − 12 months, asOf)
new_trials_prior_12m = studies with first_posted_date ∈ [asOf − 24 months, asOf − 12 months)
trial_growth_yoy     = (new_trials_12m − new_trials_prior_12m) / new_trials_prior_12m

trial_growth_yoy is null when new_trials_prior_12m < 20 (inputs.thresholds.growthMinPriorTrials) so small denominators do not produce spurious growth. The exact window bounds are stored in inputs.windows. Because the reference day moves, two runs on different days are not comparable window for window; the ranking snapshot keeps the run's windows in each row's inputs.

# 4. Enrollment

Over active interventional studies: avg_enrollment = mean of enrollment_count, median_enrollment = percentile_cont(0.5), total_enrollment_active = sum. enrollment_count is the registrant's figure and may be anticipated rather than actual (enrollment_type); studies without a count are ignored by the mean and median.

# 5. Sponsors (active interventional studies)

  • distinct_sponsors — distinct lead_sponsor strings (no normalisation: "NCI" and "National Cancer Institute (NCI)" are two sponsors).
  • industry_share = studies with lead_sponsor_class = 'INDUSTRY' / active studies.
  • sponsor_hhi — Herfindahl–Hirschman index of lead sponsors: HHI = Σ_s (n_s / N)² where n_s = active studies led by sponsor s and N = active studies with a lead sponsor (inputs.denominators.sponsor). Range (0, 1]; 1 = a single sponsor; 1/k for k equal sponsors. Null when active studies < 10 (hhiMinActiveTrials).
  • top_sponsor, top_sponsor_share — the sponsor with the most active studies (ties broken alphabetically) and n_top / N.

# 6. Geography (active interventional studies)

Based on the study-level countries array (distinct countries with at least one site).

  • distinct_countries — countries appearing in at least one active study.
  • us_share = studies with 'United States' ∈ countries / active studies.
  • top_country, top_country_share — most frequent country and its share of active studies.
  • country_hhi — HHI over trial–country pairs: Σ_c (p_c / P)² where p_c = active studies listing country c and P = Σ p_c (inputs.denominators.countryPairs).

Caveat. A multinational study contributes one pair to every country it lists, so country shares can sum above 100 % and country_hhi measures the concentration of site presence, not of studies.

# 7. Failures

# 7.1 Termination share

Over interventional studies with first_posted_date ≥ 2010-01-01 (terminationSince):

text
termination_share = (terminated + withdrawn) / (completed + terminated + withdrawn)

Null when the denominator (terminal studies, inputs.denominators.terminal) is < 30 (terminationMinTerminalTrials). Suspended, active and unknown-status studies are not terminal and are excluded from both numerator and denominator. Statuses are registrant-reported; a TERMINATED status does not imply a negative result (many studies stop for accrual reasons).

# 7.2 Stop-reason classification (why_stopped_breakdown)

For interventional studies with status TERMINATED, WITHDRAWN or SUSPENDED, the free-text why_stopped is mapped to one category using explicit keyword rules only (case-insensitive, whole-word or word-prefix matches). Rules are tested in the order below; the first match wins so specific causes take precedence over broad ones, and sponsor_decision — the broadest — is tested last. The API also returns every matched category (reasonMatches). Nothing is inferred from the design, sponsor or outcome of a study.

Order Category Keywords / patterns
1 covid covid…, pandemic
2 safety safety, toxicit…, adverse
3 efficacy efficacy, futility, lack of [clinical/therapeutic] benefit, interim analysis/analyses
4 drug_supply supply, drug availability, manufactur…
5 investigator PI left, investigator
6 enrollment accru… (accrual), enrol… (enrollment/enrolment), recruit… (recruitment)
7 funding fund/funds/funded/funding, financ…, budget…
8 sponsor_decision business, sponsor decision / sponsor's decision / decision of/by the sponsor / sponsor decided, strateg…, portfolio, company
— other_stated text present, no rule matched
— not_stated why_stopped null or blank

The breakdown is a JSON object { category: count } over the entity's stopped interventional studies; categories with zero studies are omitted. inputs.stopReasonRulesVersion records the rule set. The public pages /trials/terminated and the API route /v1/trials/terminated apply the same rules to any filtered set of studies (all study types there, with studyType as a filter).

# 8. Burden-normalized intensity (top-level cancers only)

text
trials_per_1000_deaths = active_trials / (deaths / 1 000)
trials_per_100k_cases  = active_trials / (incidence / 100 000)

deaths and incidence come from epidemiology_observations for geography USA (burden_geography), sex = 'all', age_group = 'all', metrics mortality_count and incidence_count, from the same source_id and the same year. The latest year where both counts exist is used (burden_year, burden_source_id); ties are broken by source id then observation id so the choice is deterministic. Deaths must be ≥ 100 (burdenMinDeaths), otherwise both ratios are null. The observation ids used are stored in inputs.burden.

Level all rows carry the burden ratios only when the entity is itself top-level; subtypes have no site-level burden counts in the US registries and get null (shown as "—", never zero).

Caveats. Trial counts are worldwide while the burden is US; the ratio therefore measures worldwide research intensity relative to US burden, not US research per US death. Trial counts depend on how registrants phrase conditions; burden counts depend on the registry's site definition (site_definition on the observation).

# 9. Rankings derived from this layer

rankTrialIntelligence (called at the end of computeTrialIntelligence) persists snapshots for scope geo=WORLD | sex=all | age=all | year=latest | level=top|all:

Metric slug Value Eligibility Direction
phase3_recruiting_trials phase3_recruiting > 0 descending
trial_growth_yoy trial_growth_yoy non-null (prior window ≥ 20) descending
trial_termination_share termination_share non-null (terminal ≥ 30) descending — rank 1 = highest share (higher_is_worse is display information)
sponsor_concentration sponsor_hhi non-null (active ≥ 10) descending

Each ranking row's inputs holds the counts and windows behind its value. A metric is skipped when fewer than three entities are eligible.

# 10. Reproducibility

The computation is a single SQL transaction: descendants temp table → cancer × study map → per-cancer aggregates → delete + insert of both levels. Given the same database state and the same asOf day it is deterministic (~14 s on the reference database: 126 195 studies, 187 523 condition mappings, 9 510 cancers → 36 top rows + ~2 200 all rows). Pure formula helpers (growthWindows, growthYoy, hhi, terminationShare, burdenNormalized, classifyStopReason) are unit-tested in packages/ranking/src/*.test.ts.