CancerIndex.io
Understand cancer through data.
A provenance-first, continuously updated, transparently sourced index of every recognized cancer entity —
epidemiology, genomics, biomarkers, drugs, regulatory approvals, clinical trials, research and derived
intelligence, cross-linked in one coherent dataset.
www.cancerindex.io · API docs · Methodology · Sources · Pulse
Table of contents
- What CancerIndex is
- What it refuses to do
- The index in numbers
- Product map
- Architecture
- Repository layout
- Data model
- Connectors
- Derived intelligence layer
- Rankings and metrics
- Public API
- Operator CLI
- Local development
- Configuration
- Quality: tests, QA, accessibility, performance
- Deployment and operations
- Brand
- Design language
- Known gotchas
- Roadmap
- Licensing, attribution and contact
What CancerIndex is
CancerIndex organizes fragmented oncology data into one structured, navigable system. A reader moves from a cancer to its incidence and mortality, to the genes altered in it, to a variant, to the drugs that target it, to the regulatory approvals of those drugs in each jurisdiction, to the clinical trials that investigate them and to the papers that describe them — every step a link, every number a sourced record.
It is built for researchers, students, journalists, policymakers, biotech professionals and informed members of the public. It feels like a research terminal, not a health-content website: dense tables, sparse charts, one accent colour, a source badge on every figure.
Five things make it different from a cancer information site:
- A canonical taxonomy first. 9 510 active disease entities anchored on the NCI Thesaurus, with OncoTree, ICD-10, ICD-O, DOID, UMLS, MeSH and MONDO cross-references; 36 mutually exclusive top-level site groups (GLOBOCAN / ICD-10 ranges) for burden rankings without double counting.
- Provenance on every value. Each observation, edge, approval and count points to a
provenancerow (source, dataset, version, URL, retrieval date, population, methodology); raw payloads are kept in a gzip JSON lake and can be replayed. - Layers that stay separable. RAW → NORMALIZED → CANONICAL → DERIVED → RANKED. Derived values
carry a
formula_versionand the inputs that produced them ("Why this rank?"). - Claim categories that never merge. Observed · published · curated · regulatory · guideline · computed · AI-generated — labelled everywhere, mixed nowhere.
- No invented numbers. Missing data is shown as Data not yet available. No estimates, no placeholder statistics, no composite "worst cancer" score.
What it refuses to do
- It is not a physician: no diagnosis, no individual prognosis, no dosing, no treatment recommendation. Population statistics never predict an individual outcome and the UI says so.
- It does not rank cancers by an opaque composite; each ranking is one metric, one scope, one formula.
- It does not infer facts: no edge is created by a language model, no reason for a terminated trial is guessed, no indication is derived from an ATC class, no cancer is attached to an approval unless the official text names exactly one.
- It does not redistribute data it may not redistribute: every connector manifest is licence-reviewed before it goes live (IARC / GLOBOCAN stays under review, SEER awaits credentials).
The index in numbers
Production database on 2026-09-11 (SELECT count(*) — live counts, never estimates):
| Layer | Entity | Count |
|---|---|---|
| Taxonomy | active malignant cancer entities (of 9 510 active entities) | 5 595 |
| Taxonomy | top-level site groups (ranking scope) | 36 |
| Genomics | genes (HGNC) | 45 170 |
| Genomics | variants (CIViC, ClinVar) | 417 000 |
| Genomics | curated CIViC evidence items | 11 984 |
| Biomarkers | canonical biomarkers (NCIt-verified) | 54 |
| Drugs | canonical drugs (CIViC, ChEMBL, openFDA, Health Canada, EMA) | 814 |
| Regulatory | approval records — US 1 570 · CA 1 534 · EU 459 | 3 563 |
| Trials | ClinicalTrials.gov oncology studies | 126 238 |
| Trials | registrant-entered study locations (1.19 M geocoded) | 1 212 144 |
| Literature | PubMed records linked to entities (plus 19 080 per-cancer count windows) | 4 287 |
| Epidemiology | US observations (CDC WONDER, U.S. Cancer Statistics), 1999–2024 | 10 674 |
| Graph | active knowledge edges with cancer context | 18 822 |
| Derived | trial-intelligence rows · drug-pipeline rows · research-gap components | 2 240 · 5 498 · 2 966 |
| Derived | country × cancer × phase trial-site aggregates | 11 805 |
| Rankings | current snapshots over 25 metric definitions | 875 |
| Provenance | source records · provenance rows | 629 889 · 106 263 |
| Sources | connectors registered (17 active, 1 under licence review, 1 awaiting credentials) | 19 |
| Storage | PostgreSQL 17 | 2.5 GB |
Product map
| Area | Route(s) | What it shows |
|---|---|---|
| Home | / |
Hero, live ticker (cancers, active trials, recruiting Phase III, approved drugs), US burden, data-explorer chart, trial intelligence, new approvals, rankings preview, fastest-rising incidence, gap ratios, curated evidence; aside: where trials recruit, graph teaser, taxonomy, rare spotlight, sources |
| Cancer profile | /cancer/[slug] + tabs statistics survival genomics evidence drugs trials research rankings sources |
Header with codes and badges; observations by geography/year/sex with charts; cohort alteration frequencies with denominators; CIViC evidence with native levels; jurisdiction-aware approvals; trials with an intelligence strip; literature; "Why this rank?"; every source behind the page |
| Cancers, taxonomy | /cancers, /taxonomy |
Filterable entity table (level, type, malignant, hematologic, pediatric, rare); tree browser across hierarchy types |
| Data explorer | /explore, /explore/coverage, /api/export/epidemiology.csv |
Metric × cancers × geography × sex × age × years → comparable chart groups (metric, unit, geography, source, standard population, age never mixed), observations table, permalink, CSV with attribution, API call; coverage matrix |
| Countries | /countries, /country/[slug] |
Latest-year summary, top cancers per metric, trends, clinical trial activity (sites by cancer and phase), the jurisdiction's approval records, sources |
| Compare | /compare?ids= |
2–4 cancers side by side (registry figures, counters, ranks) |
| Trials | /trials, /trial/[nct], /trials/intelligence, /trials/terminated, /trials/map, /api/export/trial-intelligence.csv |
Search and filters; study record with mappings, locations, publications; per-cancer trial metrics (growth, enrollment, sponsor and country concentration, termination share, trials per 1 000 deaths); failure tracking with registrant-stated reasons classified by explicit rules; Equal Earth choropleth of sites |
| Drugs | /drugs, /drug/[slug], /pipeline |
Canonical drugs with brands as aliases; identifiers (ATC, DIN, UNII, ChEMBL…), approvals by jurisdiction, development stage per cancer, evidence, trials; pipeline funnel per cancer |
| Approvals | /approvals |
Dated feed of FDA, Health Canada and EMA records grouped by month, filters by authority, jurisdiction, cancer, status, year |
| Genes, variants | /genes, /gene/[symbol], /variant/[slug] |
HGNC genes; evidence by cancer; variants with ClinVar interpretations; cohort frequencies; publications |
| Biomarkers | /biomarkers, /biomarker/[slug] |
54 curated biomarkers with NCIt codes verified against EVS; derived cancers, drugs, approvals (tumour-agnostic from real rows), trials, publications |
| Rankings | /rankings, /rankings/[metric] |
One metric, one scope, one formula version per table; lineage on every row; CSV export |
| Research gap | /research-gap, /api/export/research-gap.csv |
Death share vs trial and publication shares, log₂ gap ratios, per-1 000-deaths intensities, log-log scatter |
| Knowledge graph | /graph?focus=type:ref |
Contextual neighbourhood (source-native edges vs derived registry links), cancer → gene → variant → drug → approval → trials paths, accessible edge table |
| Pulse, history | /pulse, /year/[year], /data-updates |
What changed (approvals, new recruiting Phase III, registration momentum, ranking moves, publications, dataset refreshes); a year in cancer 1999–now; connector state and ingestion log |
| Sources, method | /sources, /source/[slug], /methodology, /methodology/trial-map, /trust, /about, /developers, /data |
Licence registry and connector health; every formula and threshold; policies; team and hosting; API guide; redistributable downloads |
| Search | ⌘K / Ctrl+K anywhere, /search |
Cancer, gene, variant, drug, trial, publication, source — exact › alias › prefix › fuzzy |
| Admin | /admin/* (token) |
Connectors, runs, unresolved labels, rankings, trace |
Every page has a light and a dark theme (toggle in the header, prefers-color-scheme by default),
a canonical URL, structured metadata and a server-rendered Open Graph / Twitter image.
Architecture
┌──────────────────────── sources (19 connectors) ────────────────────────┐
│ NCIt EVS · OncoTree · HGNC · CIViC · ClinVar · GDC · cBioPortal · MeSH │
│ ClinicalTrials.gov · PubMed · CDC WONDER · CDC USCS · ChEMBL · openFDA │
│ Health Canada DPD · EMA · (IARC GLOBOCAN: review) · (SEER: credentials) │
└───────────────┬───────────────────────────────────────────────────────────┘
│ HTTP / bulk files, rate-limited, restartable cursors
▼
workers/main.ts ──► packages/connectors (SDK: manifest · HttpClient · RawLake · RunContext)
pg-boss scheduler │ RAW → data/raw/{source}/{date}/{entity}/*.jsonl.gz + source_records
cron per manifest │ NORMALIZED → validators, schema-drift stats, unresolved_labels
counters 06:00 UTC │ CANONICAL → cancers · genes · variants · drugs · trials · publications
intel 06:15 UTC │ observations · approvals · knowledge_edges · provenance
rank 06:30 UTC ▼
PostgreSQL 17 (+ pg_trgm, unaccent, pgvector) — Drizzle schema, snake_case
│
packages/ranking: counters → intelligence → rankings (DERIVED / RANKED)
│
┌────────────────────┴────────────────────┐
▼ ▼
apps/api Fastify 5 · /v1 · zod · OpenAPI apps/web Next.js 16 (webpack) · server components
envelope { data, sources, dataRelease } Tailwind v4 tokens · pure SVG charts · dark mode
rate limits · API keys · x-request-id /api/v1/* proxied to the API · ISR caching
└────────────────────┬────────────────────┘
▼
MacLustr node M4M64b · PM2 (web :8250, api :8251, worker, backup)
MacLustr Tunnel (WireGuard + Caddy on BHS64) → https://www.cancerindex.ioPrinciples baked into the code
- Reconciliation before ingestion. Labels map to entities by shared identifier, then curated alias,
then normalized string; every mapping stores a
match_type(EXACT_IDENTIFIER,CURATED_EXACT,ONTOLOGY_EXACT,CURATED_BROADER,ALIAS,PROBABILISTIC,UNRESOLVED). Unknown labels go tounresolved_labels, never to/dev/null. - Time-aware observations. A value for year X never overwrites year Y; multi-year aggregates keep both bounds; rates keep their standard population.
- Idempotent, restartable connectors. Payload hashes, cursor checkpoints every 2 000 records or 60 s, SIGTERM-safe aborts, anomaly guard refusing destructive updates when a source shrinks by more than half.
- Deterministic derived layers. Counters, intelligence tables and rankings are full rebuilds in one
transaction; snapshots are immutable and keep
previous_rankfor change explanation.
Repository layout
apps/
web/ Next.js 16 app (React 19, server components, Tailwind v4, --webpack)
src/app/ routes (see Product map), opengraph-image.tsx per entity, icon.svg, manifest.ts
src/components/ ui · charts (SVG) · cancer tabs · graph · explorer · country · home modules · layout
src/lib/ db · format · queries/* (server-only SQL) · graph-model · explorer-* · og.tsx · site.ts
src/assets/fonts/ WOFF copies of Newsreader / Inter / IBM Plex Mono for server-side images
public/brand/ logo.svg · logo-mark.svg · logo-mark-dark.svg · PNG icons
qa/smoke.mjs read-only smoke suite (44 routes, weight caps, mobile overflow, console errors)
api/ Fastify 5 public API (/v1), zod schemas, OpenAPI at /v1/docs, admin routes
packages/
shared/ CI-XXX-00000001 ids, provenance types, normalization, logger, env
database/ Drizzle schema (12 files, ~50 tables), migrations 0000–0002, seeds (metrics,
geographies, biomarkers), alerts, ids
ontology/ qualifier rules, CancerResolver (alias/code reconciliation), TOP_LEVEL_CANCERS
connectors/ SDK (manifest · HttpClient · RawLake · RunContext · validators · doctor) and
connectors/<id>/{manifest.ts,index.ts,normalize.ts,fixtures/,*.test.ts}
ranking/ counters · intelligence (trial-intelligence, trial-sites, drug-pipeline,
drug-duplicates, research-gap) · engine (snapshots) · trace · country-codes
workers/ pg-boss scheduler and job handlers
scripts/ci.ts operator CLI (`pnpm cix …`)
deploy/ mld manifest, first-run bootstrap, backup / restore scripts
docs/ ARCHITECTURE · DATA-MODEL · METHODOLOGY · API · SECURITY · AI · source-policy ·
schema-changes-ops · adr/ (6 ADRs) · connectors/ (18 pages) · methodology/ (7 pages)
CLAUDE.md operational rules every contributor (human or agent) followsData model
Source of truth: packages/database/src/schema/*.ts (documented in docs/DATA-MODEL.md). Public
identifiers are CI-<NS>-00000001 minted per namespace (CAN, GENE, VAR, DRUG, TRIAL, PUB,
BIO, STUDY, METRIC, SOURCE, GEO, PROV…) — never database integers.
| Group | Tables | Notes |
|---|---|---|
| Registry & operations | sources, ingest_runs, connector_cursors, connector_field_stats, source_records, provenance, unresolved_labels, change_events, audit_log, entity_merges, system_alerts, api_keys, id_sequences |
licence status per source; one row per run with counters and log; raw-lake index with payload hashes; merge queue (proposed, never automatic) |
| Cancer ontology | cancers, cancer_aliases, cancer_hierarchy (ncit · oncotree · anatomical · …), cancer_codes, anatomical_sites, cancer_anatomy, geographies, cohort_definitions |
one row per disease concept, aliases and codes searchable, several trees coexist |
| Genomics | genes, gene_aliases, variants, variant_aliases, variant_clinical_significance, biomarkers, genomic_cohorts, cancer_gene_frequencies, entity_embeddings |
coordinates carry their assembly; frequencies carry their denominator |
| Drugs & regulatory | drugs, drug_aliases, drug_codes, treatment_regimens, drug_approvals |
brands are aliases; approvals are per jurisdiction, authority, application/DIN/EMA number, dated, with status and verbatim source status |
| Trials | clinical_trials, trial_conditions, trial_interventions, trial_locations, trial_pulse |
conditions and interventions reconciled with match_type; 1.2 M locations |
| Literature | publications, publication_entity_edges, literature_counts |
the exact PubMed query is stored with every count |
| Evidence & graph | knowledge_edges, civic_evidence_items, risk_factors |
edges carry cancer context, direction, native evidence level, category, provenance ids, status (active · superseded) |
| Epidemiology | epidemiology_observations, survival_observations |
metric, unit, sex, age group, year(s), CI, standard population, estimate type, site definition |
| Derived | entity_counters, trial_intelligence, trial_site_country_counts, drug_pipeline, research_gap_components, metric_definitions, ranking_snapshots, rankings, ai_answers |
every row has a formula_version and inputs |
Schema changes are migrations generated with pnpm db:generate (one per integration) and applied by
pnpm db:migrate, which also creates the extensions and the trigram / GIN performance indexes.
Connectors
Each connector is a class with a manifest (licence, terms review date, documentation verification
date, rate limits, schedule, anomaly guard), a health check and a sync(); it is tested against
sanitized fixtures (normal, empty, pagination, rate-limit, server error, malformed) and documented in
docs/connectors/<id>.md. pnpm cix sources:sync seeds manifests into sources; /sources shows them.
| Connector | Organization | Category | Licence status | What it feeds |
|---|---|---|---|---|
ncit-evs |
NCI EVS | terminology | approved (CC BY 4.0) | cancers, aliases, codes, hierarchy — the ontology backbone |
oncotree |
MSKCC | terminology | approved (CC BY 4.0) | OncoTree codes and tree mapped onto NCIt concepts |
mesh |
NLM | terminology | approved | MeSH headings as aliases (PubMed queries, EMA therapeutic areas) |
hgnc |
HGNC | genes | approved (CC0) | 45 k gene records and aliases |
civic |
CIViC | evidence | approved (CC0) | evidence items, variants, therapies (mints drugs), edges |
clinvar |
NCBI | variants | approved (public domain) | 417 k variants with interpretations |
gdc |
NCI GDC | genomics | approved (open tier) | TCGA cohorts and gene alteration frequencies |
cbioportal |
cBioPortal | genomics | approved | additional cohorts and frequencies |
clinicaltrials |
ClinicalTrials.gov | trials | approved (public domain) | 126 k studies, conditions, interventions, locations, trial_pulse; intervention → drug reconciliation |
pubmed |
NLM | literature | approved | per-cancer literature counts (query stored) and linked records |
cdc-wonder |
CDC | epidemiology | approved | US mortality 1999–2024 |
cdc-uscs |
CDC | epidemiology | approved | US incidence and mortality (USCS), sex-specific handling |
chembl |
EMBL-EBI | drugs | approved (CC BY-SA) | drug enrichment (ids, mechanism, targets) |
openfda |
US FDA | regulatory | approved (CC0) | Drugs@FDA applications and labels → US approval rows, indication-text cancer mapping, edges |
health-canada-dpd |
Health Canada | regulatory | approved (OGL Canada) | Drug Product Database (ATC L01/L02/L03/V10) → DIN-level Canadian records, drug minting, codes |
ema |
European Medicines Agency | regulatory | approved (attribution) | medicines data xlsx → EU authorisations with dates and statuses, indication mapping, edges |
seer-explorer |
NCI SEER | epidemiology | approved | explorer datasets (no schedule) |
seer |
NCI SEER API | epidemiology | awaiting credentials | survival and rates once SEER_API_KEY is set |
iarc-globocan |
IARC | epidemiology | review | global burden — gated by IARC_TERMS_ACCEPTED_BY; nothing is ingested without a human decision |
Run order for a fresh database: deploy/first-run.sh (terminology → genes → evidence/genomics/variants
→ trials → literature → epidemiology), then the regulatory connectors, pnpm cix reconcile-drugs,
pnpm cix counters, pnpm cix intel, pnpm cix rank.
Derived intelligence layer
Recomputed daily by the worker (maintenance.intel, between counters and rankings) or with
pnpm cix intel (≈ 50 s on the production database). Every row stores formula_version and inputs.
Methods: docs/methodology/*.md and /methodology.
| Module | Table | Formula version | Highlights |
|---|---|---|---|
| Clinical trial intelligence | trial_intelligence |
ci-trial-intel-v1, stop reasons ci-stop-reasons-v1 |
per cancer (top and all levels, descendants included): total / active / recruiting / Phase I–IV counts, growth of first-posted studies (12 m vs prior 12 m, ≥ 20), enrollment mean/median, sponsor and country HHI, industry and US shares, termination share (≥ 2010, ≥ 30 terminal), stop-reason breakdown by explicit keyword rules, trials per 1 000 deaths / per 100 k cases (US, latest year, deaths ≥ 100) |
| Trial map | trial_site_country_counts |
ci-trial-sites-v1 |
sites and studies per country × (all / each top-level cancer) × (any / each phase) × (all / recruiting only); ISO 3166-1 alpha-3 mapping with explicit unmapped historical names |
| Drug pipeline | drug_pipeline, entity_merges |
ci-drug-pipeline-v1 |
stage per drug and per drug × top-level cancer (approved › withdrawn › highest registry phase › phase not stated); salt-form duplicates proposed to the merge queue |
| Research Gap Index | research_gap_components + 4 ranking metrics |
ci-research-gap-components-v1 |
for each burden scope: death share vs active-trial and publication shares over the eligible set (deaths ≥ 100), log₂ gap ratios, per-1 000-deaths intensities |
| Trial → drug reconciliation | trial_interventions.drug_id |
(in the ClinicalTrials connector) | exact alias › salt/dose/label-stripped alias › probabilistic head; shared aliases resolved by rule (own generic name › brand › base molecule) or left unresolved — 78 956 rows linked |
Rankings and metrics
25 metric definitions live in metric_definitions (seeded from packages/database/src/seed-data/metrics.ts)
and are rendered live on /methodology#metrics. A ranking snapshot is one metric × one scope
(geo=USA|sex=all|age=all|year=2024|level=top) × one formula version; rows keep rank, percentile,
confidence, previous rank and the exact inputs.
| Category | Metrics |
|---|---|
| Burden | incidence_count, mortality_count, as_incidence_rate, as_mortality_rate |
| Lethality | mortality_incidence_ratio, five_year_survival (awaits survival observations) |
| Clinical research | active_trials, recruiting_trials, phase3_trials, phase3_recruiting_trials, trial_termination_share, sponsor_concentration |
| Research activity & trends | publications_5y, publications_12m, publication_growth, trial_growth_yoy |
| Molecular knowledge | curated_evidence_items, associated_genes, genomic_cohorts |
| Unmet need | trial_gap, research_gap (percentile-based), trial_gap_ratio, research_gap_ratio (share-based, log₂), trials_per_1000_deaths, publications_per_1000_deaths |
Burden, lethality and gap metrics exist only for scopes with licensed observations (currently the United States, per year and sex, top level). No composite score is published (ADR-006).
Public API
Base URL https://www.cancerindex.io/api/v1 (proxied to Fastify). Read-only JSON. Every response is an
envelope { data, sources, dataRelease, generatedAt, total?, limit?, offset?, hasMore? } where sources
lists the upstream sources, licences and attributions behind the returned data. Swagger UI at /api/v1/docs,
OpenAPI 3.1 at /api/v1/openapi.json. Rate limits per IP; an optional bearer API key raises them.
| Endpoints | Purpose |
|---|---|
/cancers, /cancers/:id, /cancers/:id/{statistics,survival,genes,variants,drugs,trials,publications} |
taxonomy, observations, evidence, approvals, trials per cancer (descendants included) |
/genes, /genes/:symbol, /variants/:id, /drugs, /drugs/:id, /biomarkers, /biomarkers/:slug |
entity details with derived links |
/trials, /trials/:nct, /trials/intelligence[/:cancer], /trials/terminated, /trials/sites |
search, study record, per-cancer intelligence, failures with classified reasons, country/city site aggregates |
/epidemiology, /epidemiology/coverage, /epidemiology/metrics |
time-aware observations with provenance (Data explorer) |
/approvals, /approvals/recent, /pipeline, /pipeline/summary |
jurisdiction-aware regulatory records and development stages |
/research-gap, /research-gap/scopes |
gap components per burden scope |
/graph/:type/:id, /graph/cancer/:id/paths |
knowledge-graph neighbourhoods and chains |
/rankings/metrics, /rankings, /rankings/:metric/:cancerId/explain |
catalogue, snapshots, "Why this rank?" |
/search, /sources, /sources/:slug, /stats, /changes, /healthz |
search, registry and health, live counts, change events |
/admin/* (x-admin-token) |
connectors, runs, unresolved labels, jobs, trace, audit |
Details and examples: docs/API.md.
Operator CLI
pnpm cix connectors list connectors, licence status, health, last success
pnpm cix run <id> [--mode full|incremental|backfill|dry_run] [--max-records N] [--max-minutes M] [--reset-cursor]
pnpm cix run-all [--max-minutes M] every active connector in registry order
pnpm cix health <id> source liveness probe
pnpm cix sources:sync manifests → sources table
pnpm cix reconcile-drugs [--remap] trial interventions → canonical drugs
pnpm cix counters rebuild entity_counters
pnpm cix intel trial intelligence · trial sites · drug pipeline · research gap
pnpm cix rank recompute every ranking snapshot
pnpm cix stats | trace <type> <id> | doctor [--no-disk] | alerts [ack|resolve <id>]--mode backfill replays a connector's raw lake without HTTP (used after a mapping-rule change); dry_run
never writes and never moves a cursor.
Local development
Requirements: Node ≥ 22, pnpm 11, PostgreSQL 17 with pg_trgm, unaccent and vector; rsvg-convert
and Python Pillow only if you regenerate brand rasters.
createdb cancerindex && cp .env.example .env # set NCBI_EMAIL, ADMIN_TOKEN
pnpm install
pnpm db:migrate && pnpm db:seed && pnpm cix sources:sync
pnpm cix run oncotree --mode dry_run # smoke against the live API, no writes
bash deploy/first-run.sh # ordered first ingestion (hours; restartable)
pnpm cix run health-canada-dpd && pnpm cix run ema && pnpm cix run openfda
pnpm cix reconcile-drugs && pnpm cix counters && pnpm cix intel && pnpm cix rank
pnpm dev:api & # http://127.0.0.1:8251/v1/docs
pnpm worker & # schedules from connector manifests
pnpm dev:web # http://localhost:8250A faster path for UI work is to restore a production dump (bash deploy/restore.sh <dump> --target cancerindex_prodcopy, then rename databases) and run pnpm db:migrate && pnpm db:seed.
Note that Next.js allows a single next dev per app directory (.next/dev/lock).
Configuration
.env at the repository root (read by the API, the worker and the web app).
| Variable | Purpose |
|---|---|
DATABASE_URL, DB_POOL_MAX |
PostgreSQL connection |
WEB_PORT (8250), API_PORT (8251), API_HOST, CI_API_URL |
ports and the API URL the web app proxies to |
NEXT_PUBLIC_SITE_URL |
canonical site URL (metadata, sitemaps, share images) |
CI_DATA_DIR |
raw data lake root (data/raw, data/cache) |
ADMIN_TOKEN |
/admin console and /v1/admin/* |
NCBI_TOOL, NCBI_EMAIL, NCBI_API_KEY |
E-utilities (PubMed, ClinVar); the key raises the limit to 10 req/s |
SEER_API_KEY |
unlocks the SEER connector |
OPENFDA_API_KEY |
raises openFDA from 1 000 to 120 000 requests/day |
IARC_TERMS_ACCEPTED_BY |
human gate for GLOBOCAN (stays unset until terms are settled) |
WORKER_CONCURRENCY, CI_MAX_RUN_MINUTES, LOG_LEVEL |
worker tuning |
ANTHROPIC_API_KEY, OPENAI_API_KEY, OPENAI_BASE_URL |
reserved for the future /ask layer (not enabled) |
Quality: tests, QA, accessibility, performance
pnpm typecheck— strict TypeScript across the 8 workspaces.pnpm test— 528 vitest tests: connector fixtures (every connector), SDK (lake, run, checkpoints), formulas (shares, log ratios, HHI, growth, termination share, stop-reason rules, pipeline stages, country codes, map scales, projections, explorer comparability, CSV escaping, graph layout), API smoke tests against the local database (skipped when unreachable).pnpm --filter @cancerindex/web qa/qa:prod— read-only smoke suite over 44 routes: HTTP 200, expected text, per-route weight caps, no "Data not yet available" where data must exist, no horizontal overflow at 390 px, console-error scan (Playwright when available).- Accessibility: keyboard navigation, focus rings, colour never the only carrier (badges carry text,
charts have legends and data tables, maps have equivalent tables),
arialabels on SVG figures. - Performance: ISR caching (
revalidateper page,Cache-Controlon public routes), materialized derived tables, GIN / trigram indexes,sql.paramfor array parameters, request budgets on heavy pages (graph groups, city layers).
Deployment and operations
Production runs on the MacLustr cluster (node M4M64b: Postgres 17 + pgvector, PM2) behind the MacLustr Tunnel (WireGuard + Caddy on the BHS64 gateway) at https://www.cancerindex.io.
# from the laptop — gateway M1M32 orchestrates (mld)
rsync -a --exclude '/data/' --exclude node_modules --exclude .next --exclude .git --exclude '.env*' ./ /tmp/cancerindex-deploy/
mld stage /tmp/cancerindex-deploy cancerindex
mld deploy cancerindex --node M4M64b # install, migrate, seed, sources:sync, build, PM2 restart, health checks| Process (PM2) | Role |
|---|---|
cancerindex-web (:8250) |
next start |
cancerindex-api (:8251) |
Fastify API |
cancerindex-worker |
pg-boss: connector schedules, maintenance.counters 06:00 UTC → maintenance.intel 06:15 → maintenance.rank 06:30, hourly health probes and alerts |
cancerindex-backup |
nightly pg_dump 05:20, 14 daily + 8 weekly retained (deploy/backup.sh, restore.sh) |
Operational visibility: /data-updates (connector state, ingestion log), /admin (runs, unresolved
labels, alerts), pnpm cix doctor, system_alerts. Deploy manifest: deploy/mld-manifest.cancerindex.json
(note the excludes are anchored — /coverage/, /data/ — so route folders with those names are kept).
Brand
The mark reads CI: an open ring — a cell with its nucleus, the "C" — and a graduated
index scale in the accent teal, the "I". It is drawn once in SVG and used everywhere: the header and
footer (
BrandMark, theme-aware through currentColor and the accent token),
the favicon (app/icon.svg, favicon.ico, apple-icon.png), the web
manifest and the share images.
| Asset | Path |
|---|---|
| Logo with wordmark | apps/web/public/brand/logo.svg |
| Mark (light / dark backgrounds) | apps/web/public/brand/logo-mark.svg, logo-mark-dark.svg (+ PNG 512) |
| Favicon and app icons | apps/web/src/app/icon.svg, favicon.ico, apple-icon.png, public/brand/icon-{192,512}.png |
| Share images | apps/web/src/lib/og.tsx (next/og, 1200 × 630, site fonts); app/opengraph-image.tsx (live counts) and opengraph-image.tsx under cancer/[slug], gene/[symbol], drug/[slug], biomarker/[slug], trial/[nct] |
Colours: paper #fafaf7, ink #1c1c1a, accent teal #0f5f63 (dark theme #6cc3c6). Type: Newsreader
(display), Inter (UI), IBM Plex Mono (identifiers).
Design language
Scientific, editorial, institutional. Off-white paper, charcoal ink, one restrained teal accent, thin
rules, tabular numerals, dense tables, sparklines and small multiples rather than decorative charts. All
charts are server-rendered SVG with a legend, units, the dataset year and a source line; every value shows
a claim badge (observed · published · curated · regulatory · computed) and a source badge whose popover
gives dataset, version, retrieval date and URL. Mobile is first-class (tables scroll inside their wrap; no
horizontal overflow at 390 px). Light and dark themes share one token set (--color-*, --color-series-*).
Known gotchas
- A JS array inside a Drizzle template —
sql`x = ANY(${arr}::text[])`— becomes a($1,$2,…)tuple and fails; usesql.param(arr)orsql.rawfor constant lists. - Columns declared with the
updatedAt()helper are physicallyupdated_ateven when namedcomputedAt. db.execute<T>needs atypealias, not aninterface, forT.- A prop named
refcannot be passed to a server component (React reserves it) — the page fails in production with minified error #441; use another name (entityRef). - SVG
<title>must be a single text node or hydration fails;array_aggover empty arrays errors — useunnest/string_agg. opengraph-image.tsxcannot live inside an optional catch-all segment ([[...tab]]); place it one level up.ranking_snapshots.scope_keyhas no source dimension: two sources for the same geography/year/sex overwrite each other's current snapshot, so a preferred source is chosen per scope.- A shared alias between a concept and its descendant ("breast cancer") must resolve to the broadest concept when building a dictionary; the narrowest-of-lineage rule is only for sentence-level mentions.
- OncoTree's WAF rejects any User-Agent containing a URL; NCBI throttles at 3 req/s without a key; openFDA allows 1 000 requests/day without a key; the EMA xlsx has 39 columns and header row 9.
- mld
sync_excludespatterns must be anchored (/data/,/coverage/), otherwise route folders with the same name are dropped from the deploy.
Roadmap
- Global burden once IARC / GLOBOCAN terms are settled (human decision) and SEER credentials are set;
survival observations and the
five_year_survivalmetric follow. - MHRA and TGA regulatory connectors; populating
drug_approvals.biomarker_idsfrom label text. - Risk factors and attributable burden, screening and guideline registries (metadata and links only, with change detection), hereditary syndromes, pediatric and rare-cancer views.
- Researcher, institution and funding entities (OpenAlex, NIH RePORTER, CIHR) for the funding-gap view.
- Watchlists, alerts and a grounded, cited
/asklayer that summarizes indexed sources only. - Lighter default views for the heaviest pages (
/graph, biomarker pages with hundreds of approvals).
Licensing, attribution and contact
- Code: private repository (spbgit
cancerindex.git). - Derived data (rankings, intelligence tables, CSV exports): © CancerIndex, CC BY 4.0, with attribution rows naming the underlying providers.
- Source data remains under the licence of each provider — see
/sourcesand thesourcesarray of every API response (NCIt CC BY 4.0, OncoTree CC BY 4.0, HGNC CC0, CIViC CC0, openFDA CC0, Health Canada OGL, EMA with acknowledgement, ClinVar / PubMed / ClinicalTrials.gov / CDC public domain, GDC open tier, ChEMBL CC BY-SA). - Not medical advice. CancerIndex provides research and educational information only.
Built and maintained by Simon-Pierre Boucher — contact contact@spboucher.ai — hosted on MacLustr (https://www.maclustr.io).