SPB Git forge

spb/cancerindex

Public
37commits 1branches 0releases
2.9 MBsize
maindefault branch
10 days agolast push
TypeScript 97.2% SQL 1.5% CSS 0.6% JavaScript 0.5%
36.7 KB

CancerIndex — the global index of cancer

CancerIndex.io

Understand cancer through data.
A provenance-first, continuously updated, transparently sourced index of every recognized cancer entity — epidemiology, genomics, biomarkers, drugs, regulatory approvals, clinical trials, research and derived intelligence, cross-linked in one coherent dataset.

www.cancerindex.io · API docs · Methodology · Sources · Pulse


# Table of contents

  1. What CancerIndex is
  2. What it refuses to do
  3. The index in numbers
  4. Product map
  5. Architecture
  6. Repository layout
  7. Data model
  8. Connectors
  9. Derived intelligence layer
  10. Rankings and metrics
  11. Public API
  12. Operator CLI
  13. Local development
  14. Configuration
  15. Quality: tests, QA, accessibility, performance
  16. Deployment and operations
  17. Brand
  18. Design language
  19. Known gotchas
  20. Roadmap
  21. Licensing, attribution and contact

# What CancerIndex is

CancerIndex organizes fragmented oncology data into one structured, navigable system. A reader moves from a cancer to its incidence and mortality, to the genes altered in it, to a variant, to the drugs that target it, to the regulatory approvals of those drugs in each jurisdiction, to the clinical trials that investigate them and to the papers that describe them — every step a link, every number a sourced record.

It is built for researchers, students, journalists, policymakers, biotech professionals and informed members of the public. It feels like a research terminal, not a health-content website: dense tables, sparse charts, one accent colour, a source badge on every figure.

Five things make it different from a cancer information site:

  • A canonical taxonomy first. 9 510 active disease entities anchored on the NCI Thesaurus, with OncoTree, ICD-10, ICD-O, DOID, UMLS, MeSH and MONDO cross-references; 36 mutually exclusive top-level site groups (GLOBOCAN / ICD-10 ranges) for burden rankings without double counting.
  • Provenance on every value. Each observation, edge, approval and count points to a provenance row (source, dataset, version, URL, retrieval date, population, methodology); raw payloads are kept in a gzip JSON lake and can be replayed.
  • Layers that stay separable. RAW → NORMALIZED → CANONICAL → DERIVED → RANKED. Derived values carry a formula_version and the inputs that produced them ("Why this rank?").
  • Claim categories that never merge. Observed · published · curated · regulatory · guideline · computed · AI-generated — labelled everywhere, mixed nowhere.
  • No invented numbers. Missing data is shown as Data not yet available. No estimates, no placeholder statistics, no composite "worst cancer" score.

# What it refuses to do

  • It is not a physician: no diagnosis, no individual prognosis, no dosing, no treatment recommendation. Population statistics never predict an individual outcome and the UI says so.
  • It does not rank cancers by an opaque composite; each ranking is one metric, one scope, one formula.
  • It does not infer facts: no edge is created by a language model, no reason for a terminated trial is guessed, no indication is derived from an ATC class, no cancer is attached to an approval unless the official text names exactly one.
  • It does not redistribute data it may not redistribute: every connector manifest is licence-reviewed before it goes live (IARC / GLOBOCAN stays under review, SEER awaits credentials).

# The index in numbers

Production database on 2026-09-11 (SELECT count(*) — live counts, never estimates):

Layer Entity Count
Taxonomy active malignant cancer entities (of 9 510 active entities) 5 595
Taxonomy top-level site groups (ranking scope) 36
Genomics genes (HGNC) 45 170
Genomics variants (CIViC, ClinVar) 417 000
Genomics curated CIViC evidence items 11 984
Biomarkers canonical biomarkers (NCIt-verified) 54
Drugs canonical drugs (CIViC, ChEMBL, openFDA, Health Canada, EMA) 814
Regulatory approval records — US 1 570 · CA 1 534 · EU 459 3 563
Trials ClinicalTrials.gov oncology studies 126 238
Trials registrant-entered study locations (1.19 M geocoded) 1 212 144
Literature PubMed records linked to entities (plus 19 080 per-cancer count windows) 4 287
Epidemiology US observations (CDC WONDER, U.S. Cancer Statistics), 1999–2024 10 674
Graph active knowledge edges with cancer context 18 822
Derived trial-intelligence rows · drug-pipeline rows · research-gap components 2 240 · 5 498 · 2 966
Derived country × cancer × phase trial-site aggregates 11 805
Rankings current snapshots over 25 metric definitions 875
Provenance source records · provenance rows 629 889 · 106 263
Sources connectors registered (17 active, 1 under licence review, 1 awaiting credentials) 19
Storage PostgreSQL 17 2.5 GB

# Product map

Area Route(s) What it shows
Home / Hero, live ticker (cancers, active trials, recruiting Phase III, approved drugs), US burden, data-explorer chart, trial intelligence, new approvals, rankings preview, fastest-rising incidence, gap ratios, curated evidence; aside: where trials recruit, graph teaser, taxonomy, rare spotlight, sources
Cancer profile /cancer/[slug] + tabs statistics survival genomics evidence drugs trials research rankings sources Header with codes and badges; observations by geography/year/sex with charts; cohort alteration frequencies with denominators; CIViC evidence with native levels; jurisdiction-aware approvals; trials with an intelligence strip; literature; "Why this rank?"; every source behind the page
Cancers, taxonomy /cancers, /taxonomy Filterable entity table (level, type, malignant, hematologic, pediatric, rare); tree browser across hierarchy types
Data explorer /explore, /explore/coverage, /api/export/epidemiology.csv Metric × cancers × geography × sex × age × years → comparable chart groups (metric, unit, geography, source, standard population, age never mixed), observations table, permalink, CSV with attribution, API call; coverage matrix
Countries /countries, /country/[slug] Latest-year summary, top cancers per metric, trends, clinical trial activity (sites by cancer and phase), the jurisdiction's approval records, sources
Compare /compare?ids= 2–4 cancers side by side (registry figures, counters, ranks)
Trials /trials, /trial/[nct], /trials/intelligence, /trials/terminated, /trials/map, /api/export/trial-intelligence.csv Search and filters; study record with mappings, locations, publications; per-cancer trial metrics (growth, enrollment, sponsor and country concentration, termination share, trials per 1 000 deaths); failure tracking with registrant-stated reasons classified by explicit rules; Equal Earth choropleth of sites
Drugs /drugs, /drug/[slug], /pipeline Canonical drugs with brands as aliases; identifiers (ATC, DIN, UNII, ChEMBL…), approvals by jurisdiction, development stage per cancer, evidence, trials; pipeline funnel per cancer
Approvals /approvals Dated feed of FDA, Health Canada and EMA records grouped by month, filters by authority, jurisdiction, cancer, status, year
Genes, variants /genes, /gene/[symbol], /variant/[slug] HGNC genes; evidence by cancer; variants with ClinVar interpretations; cohort frequencies; publications
Biomarkers /biomarkers, /biomarker/[slug] 54 curated biomarkers with NCIt codes verified against EVS; derived cancers, drugs, approvals (tumour-agnostic from real rows), trials, publications
Rankings /rankings, /rankings/[metric] One metric, one scope, one formula version per table; lineage on every row; CSV export
Research gap /research-gap, /api/export/research-gap.csv Death share vs trial and publication shares, log₂ gap ratios, per-1 000-deaths intensities, log-log scatter
Knowledge graph /graph?focus=type:ref Contextual neighbourhood (source-native edges vs derived registry links), cancer → gene → variant → drug → approval → trials paths, accessible edge table
Pulse, history /pulse, /year/[year], /data-updates What changed (approvals, new recruiting Phase III, registration momentum, ranking moves, publications, dataset refreshes); a year in cancer 1999–now; connector state and ingestion log
Sources, method /sources, /source/[slug], /methodology, /methodology/trial-map, /trust, /about, /developers, /data Licence registry and connector health; every formula and threshold; policies; team and hosting; API guide; redistributable downloads
Search ⌘K / Ctrl+K anywhere, /search Cancer, gene, variant, drug, trial, publication, source — exact › alias › prefix › fuzzy
Admin /admin/* (token) Connectors, runs, unresolved labels, rankings, trace

Every page has a light and a dark theme (toggle in the header, prefers-color-scheme by default), a canonical URL, structured metadata and a server-rendered Open Graph / Twitter image.

# Architecture

text
                   ┌──────────────────────── sources (19 connectors) ────────────────────────┐
                   │ NCIt EVS · OncoTree · HGNC · CIViC · ClinVar · GDC · cBioPortal · MeSH   │
                   │ ClinicalTrials.gov · PubMed · CDC WONDER · CDC USCS · ChEMBL · openFDA   │
                   │ Health Canada DPD · EMA · (IARC GLOBOCAN: review) · (SEER: credentials)  │
                   └───────────────┬───────────────────────────────────────────────────────────┘
                                   │ HTTP / bulk files, rate-limited, restartable cursors

   workers/main.ts  ──►  packages/connectors (SDK: manifest · HttpClient · RawLake · RunContext)
   pg-boss scheduler          │ RAW  → data/raw/{source}/{date}/{entity}/*.jsonl.gz + source_records
   cron per manifest          │ NORMALIZED → validators, schema-drift stats, unresolved_labels
   counters 06:00 UTC         │ CANONICAL → cancers · genes · variants · drugs · trials · publications
   intel    06:15 UTC         │             observations · approvals · knowledge_edges · provenance
   rank     06:30 UTC         ▼
                        PostgreSQL 17 (+ pg_trgm, unaccent, pgvector) — Drizzle schema, snake_case

                    packages/ranking: counters → intelligence → rankings (DERIVED / RANKED)

              ┌────────────────────┴────────────────────┐
              ▼                                         ▼
   apps/api  Fastify 5 · /v1 · zod · OpenAPI      apps/web  Next.js 16 (webpack) · server components
   envelope { data, sources, dataRelease }         Tailwind v4 tokens · pure SVG charts · dark mode
   rate limits · API keys · x-request-id           /api/v1/* proxied to the API · ISR caching
              └────────────────────┬────────────────────┘

             MacLustr node M4M64b · PM2 (web :8250, api :8251, worker, backup)
             MacLustr Tunnel (WireGuard + Caddy on BHS64) → https://www.cancerindex.io

Principles baked into the code

  • Reconciliation before ingestion. Labels map to entities by shared identifier, then curated alias, then normalized string; every mapping stores a match_type (EXACT_IDENTIFIER, CURATED_EXACT, ONTOLOGY_EXACT, CURATED_BROADER, ALIAS, PROBABILISTIC, UNRESOLVED). Unknown labels go to unresolved_labels, never to /dev/null.
  • Time-aware observations. A value for year X never overwrites year Y; multi-year aggregates keep both bounds; rates keep their standard population.
  • Idempotent, restartable connectors. Payload hashes, cursor checkpoints every 2 000 records or 60 s, SIGTERM-safe aborts, anomaly guard refusing destructive updates when a source shrinks by more than half.
  • Deterministic derived layers. Counters, intelligence tables and rankings are full rebuilds in one transaction; snapshots are immutable and keep previous_rank for change explanation.

# Repository layout

text
apps/
  web/                 Next.js 16 app (React 19, server components, Tailwind v4, --webpack)
    src/app/           routes (see Product map), opengraph-image.tsx per entity, icon.svg, manifest.ts
    src/components/    ui · charts (SVG) · cancer tabs · graph · explorer · country · home modules · layout
    src/lib/           db · format · queries/* (server-only SQL) · graph-model · explorer-* · og.tsx · site.ts
    src/assets/fonts/  WOFF copies of Newsreader / Inter / IBM Plex Mono for server-side images
    public/brand/      logo.svg · logo-mark.svg · logo-mark-dark.svg · PNG icons
    qa/smoke.mjs       read-only smoke suite (44 routes, weight caps, mobile overflow, console errors)
  api/                 Fastify 5 public API (/v1), zod schemas, OpenAPI at /v1/docs, admin routes
packages/
  shared/              CI-XXX-00000001 ids, provenance types, normalization, logger, env
  database/            Drizzle schema (12 files, ~50 tables), migrations 0000–0002, seeds (metrics,
                       geographies, biomarkers), alerts, ids
  ontology/            qualifier rules, CancerResolver (alias/code reconciliation), TOP_LEVEL_CANCERS
  connectors/          SDK (manifest · HttpClient · RawLake · RunContext · validators · doctor) and
                       connectors/<id>/{manifest.ts,index.ts,normalize.ts,fixtures/,*.test.ts}
  ranking/             counters · intelligence (trial-intelligence, trial-sites, drug-pipeline,
                       drug-duplicates, research-gap) · engine (snapshots) · trace · country-codes
workers/               pg-boss scheduler and job handlers
scripts/ci.ts          operator CLI (`pnpm cix …`)
deploy/                mld manifest, first-run bootstrap, backup / restore scripts
docs/                  ARCHITECTURE · DATA-MODEL · METHODOLOGY · API · SECURITY · AI · source-policy ·
                       schema-changes-ops · adr/ (6 ADRs) · connectors/ (18 pages) · methodology/ (7 pages)
CLAUDE.md              operational rules every contributor (human or agent) follows

# Data model

Source of truth: packages/database/src/schema/*.ts (documented in docs/DATA-MODEL.md). Public identifiers are CI-<NS>-00000001 minted per namespace (CAN, GENE, VAR, DRUG, TRIAL, PUB, BIO, STUDY, METRIC, SOURCE, GEO, PROV…) — never database integers.

Group Tables Notes
Registry & operations sources, ingest_runs, connector_cursors, connector_field_stats, source_records, provenance, unresolved_labels, change_events, audit_log, entity_merges, system_alerts, api_keys, id_sequences licence status per source; one row per run with counters and log; raw-lake index with payload hashes; merge queue (proposed, never automatic)
Cancer ontology cancers, cancer_aliases, cancer_hierarchy (ncit · oncotree · anatomical · …), cancer_codes, anatomical_sites, cancer_anatomy, geographies, cohort_definitions one row per disease concept, aliases and codes searchable, several trees coexist
Genomics genes, gene_aliases, variants, variant_aliases, variant_clinical_significance, biomarkers, genomic_cohorts, cancer_gene_frequencies, entity_embeddings coordinates carry their assembly; frequencies carry their denominator
Drugs & regulatory drugs, drug_aliases, drug_codes, treatment_regimens, drug_approvals brands are aliases; approvals are per jurisdiction, authority, application/DIN/EMA number, dated, with status and verbatim source status
Trials clinical_trials, trial_conditions, trial_interventions, trial_locations, trial_pulse conditions and interventions reconciled with match_type; 1.2 M locations
Literature publications, publication_entity_edges, literature_counts the exact PubMed query is stored with every count
Evidence & graph knowledge_edges, civic_evidence_items, risk_factors edges carry cancer context, direction, native evidence level, category, provenance ids, status (active · superseded)
Epidemiology epidemiology_observations, survival_observations metric, unit, sex, age group, year(s), CI, standard population, estimate type, site definition
Derived entity_counters, trial_intelligence, trial_site_country_counts, drug_pipeline, research_gap_components, metric_definitions, ranking_snapshots, rankings, ai_answers every row has a formula_version and inputs

Schema changes are migrations generated with pnpm db:generate (one per integration) and applied by pnpm db:migrate, which also creates the extensions and the trigram / GIN performance indexes.

# Connectors

Each connector is a class with a manifest (licence, terms review date, documentation verification date, rate limits, schedule, anomaly guard), a health check and a sync(); it is tested against sanitized fixtures (normal, empty, pagination, rate-limit, server error, malformed) and documented in docs/connectors/<id>.md. pnpm cix sources:sync seeds manifests into sources; /sources shows them.

Connector Organization Category Licence status What it feeds
ncit-evs NCI EVS terminology approved (CC BY 4.0) cancers, aliases, codes, hierarchy — the ontology backbone
oncotree MSKCC terminology approved (CC BY 4.0) OncoTree codes and tree mapped onto NCIt concepts
mesh NLM terminology approved MeSH headings as aliases (PubMed queries, EMA therapeutic areas)
hgnc HGNC genes approved (CC0) 45 k gene records and aliases
civic CIViC evidence approved (CC0) evidence items, variants, therapies (mints drugs), edges
clinvar NCBI variants approved (public domain) 417 k variants with interpretations
gdc NCI GDC genomics approved (open tier) TCGA cohorts and gene alteration frequencies
cbioportal cBioPortal genomics approved additional cohorts and frequencies
clinicaltrials ClinicalTrials.gov trials approved (public domain) 126 k studies, conditions, interventions, locations, trial_pulse; intervention → drug reconciliation
pubmed NLM literature approved per-cancer literature counts (query stored) and linked records
cdc-wonder CDC epidemiology approved US mortality 1999–2024
cdc-uscs CDC epidemiology approved US incidence and mortality (USCS), sex-specific handling
chembl EMBL-EBI drugs approved (CC BY-SA) drug enrichment (ids, mechanism, targets)
openfda US FDA regulatory approved (CC0) Drugs@FDA applications and labels → US approval rows, indication-text cancer mapping, edges
health-canada-dpd Health Canada regulatory approved (OGL Canada) Drug Product Database (ATC L01/L02/L03/V10) → DIN-level Canadian records, drug minting, codes
ema European Medicines Agency regulatory approved (attribution) medicines data xlsx → EU authorisations with dates and statuses, indication mapping, edges
seer-explorer NCI SEER epidemiology approved explorer datasets (no schedule)
seer NCI SEER API epidemiology awaiting credentials survival and rates once SEER_API_KEY is set
iarc-globocan IARC epidemiology review global burden — gated by IARC_TERMS_ACCEPTED_BY; nothing is ingested without a human decision

Run order for a fresh database: deploy/first-run.sh (terminology → genes → evidence/genomics/variants → trials → literature → epidemiology), then the regulatory connectors, pnpm cix reconcile-drugs, pnpm cix counters, pnpm cix intel, pnpm cix rank.

# Derived intelligence layer

Recomputed daily by the worker (maintenance.intel, between counters and rankings) or with pnpm cix intel (≈ 50 s on the production database). Every row stores formula_version and inputs. Methods: docs/methodology/*.md and /methodology.

Module Table Formula version Highlights
Clinical trial intelligence trial_intelligence ci-trial-intel-v1, stop reasons ci-stop-reasons-v1 per cancer (top and all levels, descendants included): total / active / recruiting / Phase I–IV counts, growth of first-posted studies (12 m vs prior 12 m, ≥ 20), enrollment mean/median, sponsor and country HHI, industry and US shares, termination share (≥ 2010, ≥ 30 terminal), stop-reason breakdown by explicit keyword rules, trials per 1 000 deaths / per 100 k cases (US, latest year, deaths ≥ 100)
Trial map trial_site_country_counts ci-trial-sites-v1 sites and studies per country × (all / each top-level cancer) × (any / each phase) × (all / recruiting only); ISO 3166-1 alpha-3 mapping with explicit unmapped historical names
Drug pipeline drug_pipeline, entity_merges ci-drug-pipeline-v1 stage per drug and per drug × top-level cancer (approved › withdrawn › highest registry phase › phase not stated); salt-form duplicates proposed to the merge queue
Research Gap Index research_gap_components + 4 ranking metrics ci-research-gap-components-v1 for each burden scope: death share vs active-trial and publication shares over the eligible set (deaths ≥ 100), log₂ gap ratios, per-1 000-deaths intensities
Trial → drug reconciliation trial_interventions.drug_id (in the ClinicalTrials connector) exact alias › salt/dose/label-stripped alias › probabilistic head; shared aliases resolved by rule (own generic name › brand › base molecule) or left unresolved — 78 956 rows linked

# Rankings and metrics

25 metric definitions live in metric_definitions (seeded from packages/database/src/seed-data/metrics.ts) and are rendered live on /methodology#metrics. A ranking snapshot is one metric × one scope (geo=USA|sex=all|age=all|year=2024|level=top) × one formula version; rows keep rank, percentile, confidence, previous rank and the exact inputs.

Category Metrics
Burden incidence_count, mortality_count, as_incidence_rate, as_mortality_rate
Lethality mortality_incidence_ratio, five_year_survival (awaits survival observations)
Clinical research active_trials, recruiting_trials, phase3_trials, phase3_recruiting_trials, trial_termination_share, sponsor_concentration
Research activity & trends publications_5y, publications_12m, publication_growth, trial_growth_yoy
Molecular knowledge curated_evidence_items, associated_genes, genomic_cohorts
Unmet need trial_gap, research_gap (percentile-based), trial_gap_ratio, research_gap_ratio (share-based, log₂), trials_per_1000_deaths, publications_per_1000_deaths

Burden, lethality and gap metrics exist only for scopes with licensed observations (currently the United States, per year and sex, top level). No composite score is published (ADR-006).

# Public API

Base URL https://www.cancerindex.io/api/v1 (proxied to Fastify). Read-only JSON. Every response is an envelope { data, sources, dataRelease, generatedAt, total?, limit?, offset?, hasMore? } where sources lists the upstream sources, licences and attributions behind the returned data. Swagger UI at /api/v1/docs, OpenAPI 3.1 at /api/v1/openapi.json. Rate limits per IP; an optional bearer API key raises them.

Endpoints Purpose
/cancers, /cancers/:id, /cancers/:id/{statistics,survival,genes,variants,drugs,trials,publications} taxonomy, observations, evidence, approvals, trials per cancer (descendants included)
/genes, /genes/:symbol, /variants/:id, /drugs, /drugs/:id, /biomarkers, /biomarkers/:slug entity details with derived links
/trials, /trials/:nct, /trials/intelligence[/:cancer], /trials/terminated, /trials/sites search, study record, per-cancer intelligence, failures with classified reasons, country/city site aggregates
/epidemiology, /epidemiology/coverage, /epidemiology/metrics time-aware observations with provenance (Data explorer)
/approvals, /approvals/recent, /pipeline, /pipeline/summary jurisdiction-aware regulatory records and development stages
/research-gap, /research-gap/scopes gap components per burden scope
/graph/:type/:id, /graph/cancer/:id/paths knowledge-graph neighbourhoods and chains
/rankings/metrics, /rankings, /rankings/:metric/:cancerId/explain catalogue, snapshots, "Why this rank?"
/search, /sources, /sources/:slug, /stats, /changes, /healthz search, registry and health, live counts, change events
/admin/* (x-admin-token) connectors, runs, unresolved labels, jobs, trace, audit

Details and examples: docs/API.md.

# Operator CLI

text
pnpm cix connectors                          list connectors, licence status, health, last success
pnpm cix run <id> [--mode full|incremental|backfill|dry_run] [--max-records N] [--max-minutes M] [--reset-cursor]
pnpm cix run-all [--max-minutes M]           every active connector in registry order
pnpm cix health <id>                         source liveness probe
pnpm cix sources:sync                        manifests → sources table
pnpm cix reconcile-drugs [--remap]           trial interventions → canonical drugs
pnpm cix counters                            rebuild entity_counters
pnpm cix intel                               trial intelligence · trial sites · drug pipeline · research gap
pnpm cix rank                                recompute every ranking snapshot
pnpm cix stats | trace <type> <id> | doctor [--no-disk] | alerts [ack|resolve <id>]

--mode backfill replays a connector's raw lake without HTTP (used after a mapping-rule change); dry_run never writes and never moves a cursor.

# Local development

Requirements: Node ≥ 22, pnpm 11, PostgreSQL 17 with pg_trgm, unaccent and vector; rsvg-convert and Python Pillow only if you regenerate brand rasters.

bash
createdb cancerindex && cp .env.example .env          # set NCBI_EMAIL, ADMIN_TOKEN
pnpm install
pnpm db:migrate && pnpm db:seed && pnpm cix sources:sync
pnpm cix run oncotree --mode dry_run                  # smoke against the live API, no writes
bash deploy/first-run.sh                              # ordered first ingestion (hours; restartable)
pnpm cix run health-canada-dpd && pnpm cix run ema && pnpm cix run openfda
pnpm cix reconcile-drugs && pnpm cix counters && pnpm cix intel && pnpm cix rank
pnpm dev:api &                                        # http://127.0.0.1:8251/v1/docs
pnpm worker &                                         # schedules from connector manifests
pnpm dev:web                                          # http://localhost:8250

A faster path for UI work is to restore a production dump (bash deploy/restore.sh <dump> --target cancerindex_prodcopy, then rename databases) and run pnpm db:migrate && pnpm db:seed. Note that Next.js allows a single next dev per app directory (.next/dev/lock).

# Configuration

.env at the repository root (read by the API, the worker and the web app).

Variable Purpose
DATABASE_URL, DB_POOL_MAX PostgreSQL connection
WEB_PORT (8250), API_PORT (8251), API_HOST, CI_API_URL ports and the API URL the web app proxies to
NEXT_PUBLIC_SITE_URL canonical site URL (metadata, sitemaps, share images)
CI_DATA_DIR raw data lake root (data/raw, data/cache)
ADMIN_TOKEN /admin console and /v1/admin/*
NCBI_TOOL, NCBI_EMAIL, NCBI_API_KEY E-utilities (PubMed, ClinVar); the key raises the limit to 10 req/s
SEER_API_KEY unlocks the SEER connector
OPENFDA_API_KEY raises openFDA from 1 000 to 120 000 requests/day
IARC_TERMS_ACCEPTED_BY human gate for GLOBOCAN (stays unset until terms are settled)
WORKER_CONCURRENCY, CI_MAX_RUN_MINUTES, LOG_LEVEL worker tuning
ANTHROPIC_API_KEY, OPENAI_API_KEY, OPENAI_BASE_URL reserved for the future /ask layer (not enabled)

# Quality: tests, QA, accessibility, performance

  • pnpm typecheck — strict TypeScript across the 8 workspaces.
  • pnpm test — 528 vitest tests: connector fixtures (every connector), SDK (lake, run, checkpoints), formulas (shares, log ratios, HHI, growth, termination share, stop-reason rules, pipeline stages, country codes, map scales, projections, explorer comparability, CSV escaping, graph layout), API smoke tests against the local database (skipped when unreachable).
  • pnpm --filter @cancerindex/web qa / qa:prod — read-only smoke suite over 44 routes: HTTP 200, expected text, per-route weight caps, no "Data not yet available" where data must exist, no horizontal overflow at 390 px, console-error scan (Playwright when available).
  • Accessibility: keyboard navigation, focus rings, colour never the only carrier (badges carry text, charts have legends and data tables, maps have equivalent tables), aria labels on SVG figures.
  • Performance: ISR caching (revalidate per page, Cache-Control on public routes), materialized derived tables, GIN / trigram indexes, sql.param for array parameters, request budgets on heavy pages (graph groups, city layers).

# Deployment and operations

Production runs on the MacLustr cluster (node M4M64b: Postgres 17 + pgvector, PM2) behind the MacLustr Tunnel (WireGuard + Caddy on the BHS64 gateway) at https://www.cancerindex.io.

bash
# from the laptop — gateway M1M32 orchestrates (mld)
rsync -a --exclude '/data/' --exclude node_modules --exclude .next --exclude .git --exclude '.env*' ./ /tmp/cancerindex-deploy/
mld stage /tmp/cancerindex-deploy cancerindex
mld deploy cancerindex --node M4M64b          # install, migrate, seed, sources:sync, build, PM2 restart, health checks
Process (PM2) Role
cancerindex-web (:8250) next start
cancerindex-api (:8251) Fastify API
cancerindex-worker pg-boss: connector schedules, maintenance.counters 06:00 UTC → maintenance.intel 06:15 → maintenance.rank 06:30, hourly health probes and alerts
cancerindex-backup nightly pg_dump 05:20, 14 daily + 8 weekly retained (deploy/backup.sh, restore.sh)

Operational visibility: /data-updates (connector state, ingestion log), /admin (runs, unresolved labels, alerts), pnpm cix doctor, system_alerts. Deploy manifest: deploy/mld-manifest.cancerindex.json (note the excludes are anchored — /coverage/, /data/ — so route folders with those names are kept).

# Brand

CancerIndex mark The mark reads CI: an open ring — a cell with its nucleus, the "C" — and a graduated index scale in the accent teal, the "I". It is drawn once in SVG and used everywhere: the header and footer (BrandMark, theme-aware through currentColor and the accent token), the favicon (app/icon.svg, favicon.ico, apple-icon.png), the web manifest and the share images.


Asset Path
Logo with wordmark apps/web/public/brand/logo.svg
Mark (light / dark backgrounds) apps/web/public/brand/logo-mark.svg, logo-mark-dark.svg (+ PNG 512)
Favicon and app icons apps/web/src/app/icon.svg, favicon.ico, apple-icon.png, public/brand/icon-{192,512}.png
Share images apps/web/src/lib/og.tsx (next/og, 1200 × 630, site fonts); app/opengraph-image.tsx (live counts) and opengraph-image.tsx under cancer/[slug], gene/[symbol], drug/[slug], biomarker/[slug], trial/[nct]

Colours: paper #fafaf7, ink #1c1c1a, accent teal #0f5f63 (dark theme #6cc3c6). Type: Newsreader (display), Inter (UI), IBM Plex Mono (identifiers).

# Design language

Scientific, editorial, institutional. Off-white paper, charcoal ink, one restrained teal accent, thin rules, tabular numerals, dense tables, sparklines and small multiples rather than decorative charts. All charts are server-rendered SVG with a legend, units, the dataset year and a source line; every value shows a claim badge (observed · published · curated · regulatory · computed) and a source badge whose popover gives dataset, version, retrieval date and URL. Mobile is first-class (tables scroll inside their wrap; no horizontal overflow at 390 px). Light and dark themes share one token set (--color-*, --color-series-*).

# Known gotchas

  • A JS array inside a Drizzle template — sql`x = ANY(${arr}::text[])` — becomes a ($1,$2,…) tuple and fails; use sql.param(arr) or sql.raw for constant lists.
  • Columns declared with the updatedAt() helper are physically updated_at even when named computedAt.
  • db.execute<T> needs a type alias, not an interface, for T.
  • A prop named ref cannot be passed to a server component (React reserves it) — the page fails in production with minified error #441; use another name (entityRef).
  • SVG <title> must be a single text node or hydration fails; array_agg over empty arrays errors — use unnest / string_agg.
  • opengraph-image.tsx cannot live inside an optional catch-all segment ([[...tab]]); place it one level up.
  • ranking_snapshots.scope_key has no source dimension: two sources for the same geography/year/sex overwrite each other's current snapshot, so a preferred source is chosen per scope.
  • A shared alias between a concept and its descendant ("breast cancer") must resolve to the broadest concept when building a dictionary; the narrowest-of-lineage rule is only for sentence-level mentions.
  • OncoTree's WAF rejects any User-Agent containing a URL; NCBI throttles at 3 req/s without a key; openFDA allows 1 000 requests/day without a key; the EMA xlsx has 39 columns and header row 9.
  • mld sync_excludes patterns must be anchored (/data/, /coverage/), otherwise route folders with the same name are dropped from the deploy.

# Roadmap

  • Global burden once IARC / GLOBOCAN terms are settled (human decision) and SEER credentials are set; survival observations and the five_year_survival metric follow.
  • MHRA and TGA regulatory connectors; populating drug_approvals.biomarker_ids from label text.
  • Risk factors and attributable burden, screening and guideline registries (metadata and links only, with change detection), hereditary syndromes, pediatric and rare-cancer views.
  • Researcher, institution and funding entities (OpenAlex, NIH RePORTER, CIHR) for the funding-gap view.
  • Watchlists, alerts and a grounded, cited /ask layer that summarizes indexed sources only.
  • Lighter default views for the heaviest pages (/graph, biomarker pages with hundreds of approvals).

# Licensing, attribution and contact

  • Code: private repository (spbgit cancerindex.git).
  • Derived data (rankings, intelligence tables, CSV exports): © CancerIndex, CC BY 4.0, with attribution rows naming the underlying providers.
  • Source data remains under the licence of each provider — see /sources and the sources array of every API response (NCIt CC BY 4.0, OncoTree CC BY 4.0, HGNC CC0, CIViC CC0, openFDA CC0, Health Canada OGL, EMA with acknowledgement, ClinVar / PubMed / ClinicalTrials.gov / CDC public domain, GDC open tier, ChEMBL CC BY-SA).
  • Not medical advice. CancerIndex provides research and educational information only.

Built and maintained by Simon-Pierre Boucher — contact contact@spboucher.ai — hosted on MacLustr (https://www.maclustr.io).