CancerIndex.io
Understand cancer through data.
A provenance-first, continuously updated, transparently sourced index of every recognized cancer entity —
epidemiology, genomics, biomarkers, drugs, regulatory approvals, clinical trials, research and derived
intelligence, cross-linked in one coherent dataset.
www.cancerindex.io ·
API docs ·
Methodology ·
Sources ·
Pulse
---
## Table of contents
1. [What CancerIndex is](#what-cancerindex-is)
2. [What it refuses to do](#what-it-refuses-to-do)
3. [The index in numbers](#the-index-in-numbers)
4. [Product map](#product-map)
5. [Architecture](#architecture)
6. [Repository layout](#repository-layout)
7. [Data model](#data-model)
8. [Connectors](#connectors)
9. [Derived intelligence layer](#derived-intelligence-layer)
10. [Rankings and metrics](#rankings-and-metrics)
11. [Public API](#public-api)
12. [Operator CLI](#operator-cli)
13. [Local development](#local-development)
14. [Configuration](#configuration)
15. [Quality: tests, QA, accessibility, performance](#quality-tests-qa-accessibility-performance)
16. [Deployment and operations](#deployment-and-operations)
17. [Brand](#brand)
18. [Design language](#design-language)
19. [Known gotchas](#known-gotchas)
20. [Roadmap](#roadmap)
21. [Licensing, attribution and contact](#licensing-attribution-and-contact)
---
## What CancerIndex is
CancerIndex organizes fragmented oncology data into one structured, navigable system. A reader moves
from a **cancer** to its **incidence and mortality**, to the **genes** altered in it, to a **variant**, to the
**drugs** that target it, to the **regulatory approvals** of those drugs in each jurisdiction, to the
**clinical trials** that investigate them and to the **papers** that describe them — every step a link, every
number a sourced record.
It is built for researchers, students, journalists, policymakers, biotech professionals and informed
members of the public. It feels like a research terminal, not a health-content website: dense tables,
sparse charts, one accent colour, a source badge on every figure.
Five things make it different from a cancer information site:
- **A canonical taxonomy first.** 9 510 active disease entities anchored on the NCI Thesaurus, with
OncoTree, ICD-10, ICD-O, DOID, UMLS, MeSH and MONDO cross-references; 36 mutually exclusive
top-level site groups (GLOBOCAN / ICD-10 ranges) for burden rankings without double counting.
- **Provenance on every value.** Each observation, edge, approval and count points to a `provenance`
row (source, dataset, version, URL, retrieval date, population, methodology); raw payloads are kept
in a gzip JSON lake and can be replayed.
- **Layers that stay separable.** RAW → NORMALIZED → CANONICAL → DERIVED → RANKED. Derived values
carry a `formula_version` and the inputs that produced them ("Why this rank?").
- **Claim categories that never merge.** Observed · published · curated · regulatory · guideline ·
computed · AI-generated — labelled everywhere, mixed nowhere.
- **No invented numbers.** Missing data is shown as *Data not yet available*. No estimates, no
placeholder statistics, no composite "worst cancer" score.
## What it refuses to do
- It is **not a physician**: no diagnosis, no individual prognosis, no dosing, no treatment
recommendation. Population statistics never predict an individual outcome and the UI says so.
- It does not rank cancers by an opaque composite; each ranking is one metric, one scope, one formula.
- It does not infer facts: no edge is created by a language model, no reason for a terminated trial is
guessed, no indication is derived from an ATC class, no cancer is attached to an approval unless the
official text names exactly one.
- It does not redistribute data it may not redistribute: every connector manifest is licence-reviewed
before it goes live (IARC / GLOBOCAN stays under review, SEER awaits credentials).
## The index in numbers
Production database on 2026-09-11 (`SELECT count(*)` — live counts, never estimates):
| Layer | Entity | Count |
|---|---|---:|
| Taxonomy | active malignant cancer entities (of 9 510 active entities) | 5 595 |
| Taxonomy | top-level site groups (ranking scope) | 36 |
| Genomics | genes (HGNC) | 45 170 |
| Genomics | variants (CIViC, ClinVar) | 417 000 |
| Genomics | curated CIViC evidence items | 11 984 |
| Biomarkers | canonical biomarkers (NCIt-verified) | 54 |
| Drugs | canonical drugs (CIViC, ChEMBL, openFDA, Health Canada, EMA) | 814 |
| Regulatory | approval records — US 1 570 · CA 1 534 · EU 459 | 3 563 |
| Trials | ClinicalTrials.gov oncology studies | 126 238 |
| Trials | registrant-entered study locations (1.19 M geocoded) | 1 212 144 |
| Literature | PubMed records linked to entities (plus 19 080 per-cancer count windows) | 4 287 |
| Epidemiology | US observations (CDC WONDER, U.S. Cancer Statistics), 1999–2024 | 10 674 |
| Graph | active knowledge edges with cancer context | 18 822 |
| Derived | trial-intelligence rows · drug-pipeline rows · research-gap components | 2 240 · 5 498 · 2 966 |
| Derived | country × cancer × phase trial-site aggregates | 11 805 |
| Rankings | current snapshots over 25 metric definitions | 875 |
| Provenance | source records · provenance rows | 629 889 · 106 263 |
| Sources | connectors registered (17 active, 1 under licence review, 1 awaiting credentials) | 19 |
| Storage | PostgreSQL 17 | 2.5 GB |
## Product map
| Area | Route(s) | What it shows |
|---|---|---|
| Home | `/` | Hero, live ticker (cancers, active trials, recruiting Phase III, approved drugs), US burden, data-explorer chart, trial intelligence, new approvals, rankings preview, fastest-rising incidence, gap ratios, curated evidence; aside: where trials recruit, graph teaser, taxonomy, rare spotlight, sources |
| Cancer profile | `/cancer/[slug]` + tabs `statistics` `survival` `genomics` `evidence` `drugs` `trials` `research` `rankings` `sources` | Header with codes and badges; observations by geography/year/sex with charts; cohort alteration frequencies with denominators; CIViC evidence with native levels; jurisdiction-aware approvals; trials with an intelligence strip; literature; "Why this rank?"; every source behind the page |
| Cancers, taxonomy | `/cancers`, `/taxonomy` | Filterable entity table (level, type, malignant, hematologic, pediatric, rare); tree browser across hierarchy types |
| Data explorer | `/explore`, `/explore/coverage`, `/api/export/epidemiology.csv` | Metric × cancers × geography × sex × age × years → comparable chart groups (metric, unit, geography, source, standard population, age never mixed), observations table, permalink, CSV with attribution, API call; coverage matrix |
| Countries | `/countries`, `/country/[slug]` | Latest-year summary, top cancers per metric, trends, clinical trial activity (sites by cancer and phase), the jurisdiction's approval records, sources |
| Compare | `/compare?ids=` | 2–4 cancers side by side (registry figures, counters, ranks) |
| Trials | `/trials`, `/trial/[nct]`, `/trials/intelligence`, `/trials/terminated`, `/trials/map`, `/api/export/trial-intelligence.csv` | Search and filters; study record with mappings, locations, publications; per-cancer trial metrics (growth, enrollment, sponsor and country concentration, termination share, trials per 1 000 deaths); failure tracking with registrant-stated reasons classified by explicit rules; Equal Earth choropleth of sites |
| Drugs | `/drugs`, `/drug/[slug]`, `/pipeline` | Canonical drugs with brands as aliases; identifiers (ATC, DIN, UNII, ChEMBL…), approvals by jurisdiction, development stage per cancer, evidence, trials; pipeline funnel per cancer |
| Approvals | `/approvals` | Dated feed of FDA, Health Canada and EMA records grouped by month, filters by authority, jurisdiction, cancer, status, year |
| Genes, variants | `/genes`, `/gene/[symbol]`, `/variant/[slug]` | HGNC genes; evidence by cancer; variants with ClinVar interpretations; cohort frequencies; publications |
| Biomarkers | `/biomarkers`, `/biomarker/[slug]` | 54 curated biomarkers with NCIt codes verified against EVS; derived cancers, drugs, approvals (tumour-agnostic from real rows), trials, publications |
| Rankings | `/rankings`, `/rankings/[metric]` | One metric, one scope, one formula version per table; lineage on every row; CSV export |
| Research gap | `/research-gap`, `/api/export/research-gap.csv` | Death share vs trial and publication shares, log₂ gap ratios, per-1 000-deaths intensities, log-log scatter |
| Knowledge graph | `/graph?focus=type:ref` | Contextual neighbourhood (source-native edges vs derived registry links), cancer → gene → variant → drug → approval → trials paths, accessible edge table |
| Pulse, history | `/pulse`, `/year/[year]`, `/data-updates` | What changed (approvals, new recruiting Phase III, registration momentum, ranking moves, publications, dataset refreshes); a year in cancer 1999–now; connector state and ingestion log |
| Sources, method | `/sources`, `/source/[slug]`, `/methodology`, `/methodology/trial-map`, `/trust`, `/about`, `/developers`, `/data` | Licence registry and connector health; every formula and threshold; policies; team and hosting; API guide; redistributable downloads |
| Search | ⌘K / Ctrl+K anywhere, `/search` | Cancer, gene, variant, drug, trial, publication, source — exact › alias › prefix › fuzzy |
| Admin | `/admin/*` (token) | Connectors, runs, unresolved labels, rankings, trace |
Every page has a light and a dark theme (toggle in the header, `prefers-color-scheme` by default),
a canonical URL, structured metadata and a server-rendered Open Graph / Twitter image.
## Architecture
```
┌──────────────────────── sources (19 connectors) ────────────────────────┐
│ NCIt EVS · OncoTree · HGNC · CIViC · ClinVar · GDC · cBioPortal · MeSH │
│ ClinicalTrials.gov · PubMed · CDC WONDER · CDC USCS · ChEMBL · openFDA │
│ Health Canada DPD · EMA · (IARC GLOBOCAN: review) · (SEER: credentials) │
└───────────────┬───────────────────────────────────────────────────────────┘
│ HTTP / bulk files, rate-limited, restartable cursors
▼
workers/main.ts ──► packages/connectors (SDK: manifest · HttpClient · RawLake · RunContext)
pg-boss scheduler │ RAW → data/raw/{source}/{date}/{entity}/*.jsonl.gz + source_records
cron per manifest │ NORMALIZED → validators, schema-drift stats, unresolved_labels
counters 06:00 UTC │ CANONICAL → cancers · genes · variants · drugs · trials · publications
intel 06:15 UTC │ observations · approvals · knowledge_edges · provenance
rank 06:30 UTC ▼
PostgreSQL 17 (+ pg_trgm, unaccent, pgvector) — Drizzle schema, snake_case
│
packages/ranking: counters → intelligence → rankings (DERIVED / RANKED)
│
┌────────────────────┴────────────────────┐
▼ ▼
apps/api Fastify 5 · /v1 · zod · OpenAPI apps/web Next.js 16 (webpack) · server components
envelope { data, sources, dataRelease } Tailwind v4 tokens · pure SVG charts · dark mode
rate limits · API keys · x-request-id /api/v1/* proxied to the API · ISR caching
└────────────────────┬────────────────────┘
▼
MacLustr node M4M64b · PM2 (web :8250, api :8251, worker, backup)
MacLustr Tunnel (WireGuard + Caddy on BHS64) → https://www.cancerindex.io
```
**Principles baked into the code**
- *Reconciliation before ingestion.* Labels map to entities by shared identifier, then curated alias,
then normalized string; every mapping stores a `match_type` (`EXACT_IDENTIFIER`, `CURATED_EXACT`,
`ONTOLOGY_EXACT`, `CURATED_BROADER`, `ALIAS`, `PROBABILISTIC`, `UNRESOLVED`). Unknown labels go to
`unresolved_labels`, never to `/dev/null`.
- *Time-aware observations.* A value for year X never overwrites year Y; multi-year aggregates keep
both bounds; rates keep their standard population.
- *Idempotent, restartable connectors.* Payload hashes, cursor checkpoints every 2 000 records or 60 s,
SIGTERM-safe aborts, anomaly guard refusing destructive updates when a source shrinks by more than half.
- *Deterministic derived layers.* Counters, intelligence tables and rankings are full rebuilds in one
transaction; snapshots are immutable and keep `previous_rank` for change explanation.
## Repository layout
```
apps/
web/ Next.js 16 app (React 19, server components, Tailwind v4, --webpack)
src/app/ routes (see Product map), opengraph-image.tsx per entity, icon.svg, manifest.ts
src/components/ ui · charts (SVG) · cancer tabs · graph · explorer · country · home modules · layout
src/lib/ db · format · queries/* (server-only SQL) · graph-model · explorer-* · og.tsx · site.ts
src/assets/fonts/ WOFF copies of Newsreader / Inter / IBM Plex Mono for server-side images
public/brand/ logo.svg · logo-mark.svg · logo-mark-dark.svg · PNG icons
qa/smoke.mjs read-only smoke suite (44 routes, weight caps, mobile overflow, console errors)
api/ Fastify 5 public API (/v1), zod schemas, OpenAPI at /v1/docs, admin routes
packages/
shared/ CI-XXX-00000001 ids, provenance types, normalization, logger, env
database/ Drizzle schema (12 files, ~50 tables), migrations 0000–0002, seeds (metrics,
geographies, biomarkers), alerts, ids
ontology/ qualifier rules, CancerResolver (alias/code reconciliation), TOP_LEVEL_CANCERS
connectors/ SDK (manifest · HttpClient · RawLake · RunContext · validators · doctor) and
connectors//{manifest.ts,index.ts,normalize.ts,fixtures/,*.test.ts}
ranking/ counters · intelligence (trial-intelligence, trial-sites, drug-pipeline,
drug-duplicates, research-gap) · engine (snapshots) · trace · country-codes
workers/ pg-boss scheduler and job handlers
scripts/ci.ts operator CLI (`pnpm cix …`)
deploy/ mld manifest, first-run bootstrap, backup / restore scripts
docs/ ARCHITECTURE · DATA-MODEL · METHODOLOGY · API · SECURITY · AI · source-policy ·
schema-changes-ops · adr/ (6 ADRs) · connectors/ (18 pages) · methodology/ (7 pages)
CLAUDE.md operational rules every contributor (human or agent) follows
```
## Data model
Source of truth: `packages/database/src/schema/*.ts` (documented in `docs/DATA-MODEL.md`). Public
identifiers are `CI--00000001` minted per namespace (`CAN`, `GENE`, `VAR`, `DRUG`, `TRIAL`, `PUB`,
`BIO`, `STUDY`, `METRIC`, `SOURCE`, `GEO`, `PROV`…) — never database integers.
| Group | Tables | Notes |
|---|---|---|
| Registry & operations | `sources`, `ingest_runs`, `connector_cursors`, `connector_field_stats`, `source_records`, `provenance`, `unresolved_labels`, `change_events`, `audit_log`, `entity_merges`, `system_alerts`, `api_keys`, `id_sequences` | licence status per source; one row per run with counters and log; raw-lake index with payload hashes; merge queue (proposed, never automatic) |
| Cancer ontology | `cancers`, `cancer_aliases`, `cancer_hierarchy` (ncit · oncotree · anatomical · …), `cancer_codes`, `anatomical_sites`, `cancer_anatomy`, `geographies`, `cohort_definitions` | one row per disease concept, aliases and codes searchable, several trees coexist |
| Genomics | `genes`, `gene_aliases`, `variants`, `variant_aliases`, `variant_clinical_significance`, `biomarkers`, `genomic_cohorts`, `cancer_gene_frequencies`, `entity_embeddings` | coordinates carry their assembly; frequencies carry their denominator |
| Drugs & regulatory | `drugs`, `drug_aliases`, `drug_codes`, `treatment_regimens`, `drug_approvals` | brands are aliases; approvals are per jurisdiction, authority, application/DIN/EMA number, dated, with status and verbatim source status |
| Trials | `clinical_trials`, `trial_conditions`, `trial_interventions`, `trial_locations`, `trial_pulse` | conditions and interventions reconciled with `match_type`; 1.2 M locations |
| Literature | `publications`, `publication_entity_edges`, `literature_counts` | the exact PubMed query is stored with every count |
| Evidence & graph | `knowledge_edges`, `civic_evidence_items`, `risk_factors` | edges carry cancer context, direction, native evidence level, category, provenance ids, status (`active` · `superseded`) |
| Epidemiology | `epidemiology_observations`, `survival_observations` | metric, unit, sex, age group, year(s), CI, standard population, estimate type, site definition |
| Derived | `entity_counters`, `trial_intelligence`, `trial_site_country_counts`, `drug_pipeline`, `research_gap_components`, `metric_definitions`, `ranking_snapshots`, `rankings`, `ai_answers` | every row has a `formula_version` and `inputs` |
Schema changes are migrations generated with `pnpm db:generate` (one per integration) and applied by
`pnpm db:migrate`, which also creates the extensions and the trigram / GIN performance indexes.
## Connectors
Each connector is a class with a manifest (licence, terms review date, documentation verification
date, rate limits, schedule, anomaly guard), a health check and a `sync()`; it is tested against
sanitized fixtures (normal, empty, pagination, rate-limit, server error, malformed) and documented in
`docs/connectors/.md`. `pnpm cix sources:sync` seeds manifests into `sources`; `/sources` shows them.
| Connector | Organization | Category | Licence status | What it feeds |
|---|---|---|---|---|
| `ncit-evs` | NCI EVS | terminology | approved (CC BY 4.0) | cancers, aliases, codes, hierarchy — the ontology backbone |
| `oncotree` | MSKCC | terminology | approved (CC BY 4.0) | OncoTree codes and tree mapped onto NCIt concepts |
| `mesh` | NLM | terminology | approved | MeSH headings as aliases (PubMed queries, EMA therapeutic areas) |
| `hgnc` | HGNC | genes | approved (CC0) | 45 k gene records and aliases |
| `civic` | CIViC | evidence | approved (CC0) | evidence items, variants, therapies (mints drugs), edges |
| `clinvar` | NCBI | variants | approved (public domain) | 417 k variants with interpretations |
| `gdc` | NCI GDC | genomics | approved (open tier) | TCGA cohorts and gene alteration frequencies |
| `cbioportal` | cBioPortal | genomics | approved | additional cohorts and frequencies |
| `clinicaltrials` | ClinicalTrials.gov | trials | approved (public domain) | 126 k studies, conditions, interventions, locations, `trial_pulse`; intervention → drug reconciliation |
| `pubmed` | NLM | literature | approved | per-cancer literature counts (query stored) and linked records |
| `cdc-wonder` | CDC | epidemiology | approved | US mortality 1999–2024 |
| `cdc-uscs` | CDC | epidemiology | approved | US incidence and mortality (USCS), sex-specific handling |
| `chembl` | EMBL-EBI | drugs | approved (CC BY-SA) | drug enrichment (ids, mechanism, targets) |
| `openfda` | US FDA | regulatory | approved (CC0) | Drugs@FDA applications and labels → US approval rows, indication-text cancer mapping, edges |
| `health-canada-dpd` | Health Canada | regulatory | approved (OGL Canada) | Drug Product Database (ATC L01/L02/L03/V10) → DIN-level Canadian records, drug minting, codes |
| `ema` | European Medicines Agency | regulatory | approved (attribution) | medicines data xlsx → EU authorisations with dates and statuses, indication mapping, edges |
| `seer-explorer` | NCI SEER | epidemiology | approved | explorer datasets (no schedule) |
| `seer` | NCI SEER API | epidemiology | awaiting credentials | survival and rates once `SEER_API_KEY` is set |
| `iarc-globocan` | IARC | epidemiology | **review** | global burden — gated by `IARC_TERMS_ACCEPTED_BY`; nothing is ingested without a human decision |
Run order for a fresh database: `deploy/first-run.sh` (terminology → genes → evidence/genomics/variants
→ trials → literature → epidemiology), then the regulatory connectors, `pnpm cix reconcile-drugs`,
`pnpm cix counters`, `pnpm cix intel`, `pnpm cix rank`.
## Derived intelligence layer
Recomputed daily by the worker (`maintenance.intel`, between counters and rankings) or with
`pnpm cix intel` (≈ 50 s on the production database). Every row stores `formula_version` and `inputs`.
Methods: `docs/methodology/*.md` and `/methodology`.
| Module | Table | Formula version | Highlights |
|---|---|---|---|
| Clinical trial intelligence | `trial_intelligence` | `ci-trial-intel-v1`, stop reasons `ci-stop-reasons-v1` | per cancer (top and all levels, descendants included): total / active / recruiting / Phase I–IV counts, growth of first-posted studies (12 m vs prior 12 m, ≥ 20), enrollment mean/median, sponsor and country HHI, industry and US shares, termination share (≥ 2010, ≥ 30 terminal), stop-reason breakdown by explicit keyword rules, trials per 1 000 deaths / per 100 k cases (US, latest year, deaths ≥ 100) |
| Trial map | `trial_site_country_counts` | `ci-trial-sites-v1` | sites and studies per country × (all / each top-level cancer) × (any / each phase) × (all / recruiting only); ISO 3166-1 alpha-3 mapping with explicit unmapped historical names |
| Drug pipeline | `drug_pipeline`, `entity_merges` | `ci-drug-pipeline-v1` | stage per drug and per drug × top-level cancer (approved › withdrawn › highest registry phase › phase not stated); salt-form duplicates proposed to the merge queue |
| Research Gap Index | `research_gap_components` + 4 ranking metrics | `ci-research-gap-components-v1` | for each burden scope: death share vs active-trial and publication shares over the eligible set (deaths ≥ 100), `log₂` gap ratios, per-1 000-deaths intensities |
| Trial → drug reconciliation | `trial_interventions.drug_id` | (in the ClinicalTrials connector) | exact alias › salt/dose/label-stripped alias › probabilistic head; shared aliases resolved by rule (own generic name › brand › base molecule) or left unresolved — 78 956 rows linked |
## Rankings and metrics
25 metric definitions live in `metric_definitions` (seeded from `packages/database/src/seed-data/metrics.ts`)
and are rendered live on `/methodology#metrics`. A ranking snapshot is one metric × one scope
(`geo=USA|sex=all|age=all|year=2024|level=top`) × one formula version; rows keep rank, percentile,
confidence, previous rank and the exact inputs.
| Category | Metrics |
|---|---|
| Burden | `incidence_count`, `mortality_count`, `as_incidence_rate`, `as_mortality_rate` |
| Lethality | `mortality_incidence_ratio`, `five_year_survival` (awaits survival observations) |
| Clinical research | `active_trials`, `recruiting_trials`, `phase3_trials`, `phase3_recruiting_trials`, `trial_termination_share`, `sponsor_concentration` |
| Research activity & trends | `publications_5y`, `publications_12m`, `publication_growth`, `trial_growth_yoy` |
| Molecular knowledge | `curated_evidence_items`, `associated_genes`, `genomic_cohorts` |
| Unmet need | `trial_gap`, `research_gap` (percentile-based), `trial_gap_ratio`, `research_gap_ratio` (share-based, log₂), `trials_per_1000_deaths`, `publications_per_1000_deaths` |
Burden, lethality and gap metrics exist only for scopes with licensed observations (currently the
United States, per year and sex, top level). No composite score is published (ADR-006).
## Public API
Base URL `https://www.cancerindex.io/api/v1` (proxied to Fastify). Read-only JSON. Every response is an
envelope `{ data, sources, dataRelease, generatedAt, total?, limit?, offset?, hasMore? }` where `sources`
lists the upstream sources, licences and attributions behind the returned data. Swagger UI at `/api/v1/docs`,
OpenAPI 3.1 at `/api/v1/openapi.json`. Rate limits per IP; an optional bearer API key raises them.
| Endpoints | Purpose |
|---|---|
| `/cancers`, `/cancers/:id`, `/cancers/:id/{statistics,survival,genes,variants,drugs,trials,publications}` | taxonomy, observations, evidence, approvals, trials per cancer (descendants included) |
| `/genes`, `/genes/:symbol`, `/variants/:id`, `/drugs`, `/drugs/:id`, `/biomarkers`, `/biomarkers/:slug` | entity details with derived links |
| `/trials`, `/trials/:nct`, `/trials/intelligence[/:cancer]`, `/trials/terminated`, `/trials/sites` | search, study record, per-cancer intelligence, failures with classified reasons, country/city site aggregates |
| `/epidemiology`, `/epidemiology/coverage`, `/epidemiology/metrics` | time-aware observations with provenance (Data explorer) |
| `/approvals`, `/approvals/recent`, `/pipeline`, `/pipeline/summary` | jurisdiction-aware regulatory records and development stages |
| `/research-gap`, `/research-gap/scopes` | gap components per burden scope |
| `/graph/:type/:id`, `/graph/cancer/:id/paths` | knowledge-graph neighbourhoods and chains |
| `/rankings/metrics`, `/rankings`, `/rankings/:metric/:cancerId/explain` | catalogue, snapshots, "Why this rank?" |
| `/search`, `/sources`, `/sources/:slug`, `/stats`, `/changes`, `/healthz` | search, registry and health, live counts, change events |
| `/admin/*` (`x-admin-token`) | connectors, runs, unresolved labels, jobs, trace, audit |
Details and examples: `docs/API.md`.
## Operator CLI
```
pnpm cix connectors list connectors, licence status, health, last success
pnpm cix run [--mode full|incremental|backfill|dry_run] [--max-records N] [--max-minutes M] [--reset-cursor]
pnpm cix run-all [--max-minutes M] every active connector in registry order
pnpm cix health source liveness probe
pnpm cix sources:sync manifests → sources table
pnpm cix reconcile-drugs [--remap] trial interventions → canonical drugs
pnpm cix counters rebuild entity_counters
pnpm cix intel trial intelligence · trial sites · drug pipeline · research gap
pnpm cix rank recompute every ranking snapshot
pnpm cix stats | trace | doctor [--no-disk] | alerts [ack|resolve ]
```
`--mode backfill` replays a connector's raw lake without HTTP (used after a mapping-rule change); `dry_run`
never writes and never moves a cursor.
## Local development
Requirements: Node ≥ 22, pnpm 11, PostgreSQL 17 with `pg_trgm`, `unaccent` and `vector`; `rsvg-convert`
and Python Pillow only if you regenerate brand rasters.
```bash
createdb cancerindex && cp .env.example .env # set NCBI_EMAIL, ADMIN_TOKEN
pnpm install
pnpm db:migrate && pnpm db:seed && pnpm cix sources:sync
pnpm cix run oncotree --mode dry_run # smoke against the live API, no writes
bash deploy/first-run.sh # ordered first ingestion (hours; restartable)
pnpm cix run health-canada-dpd && pnpm cix run ema && pnpm cix run openfda
pnpm cix reconcile-drugs && pnpm cix counters && pnpm cix intel && pnpm cix rank
pnpm dev:api & # http://127.0.0.1:8251/v1/docs
pnpm worker & # schedules from connector manifests
pnpm dev:web # http://localhost:8250
```
A faster path for UI work is to restore a production dump (`bash deploy/restore.sh --target
cancerindex_prodcopy`, then rename databases) and run `pnpm db:migrate && pnpm db:seed`.
Note that Next.js allows a single `next dev` per app directory (`.next/dev/lock`).
## Configuration
`.env` at the repository root (read by the API, the worker and the web app).
| Variable | Purpose |
|---|---|
| `DATABASE_URL`, `DB_POOL_MAX` | PostgreSQL connection |
| `WEB_PORT` (8250), `API_PORT` (8251), `API_HOST`, `CI_API_URL` | ports and the API URL the web app proxies to |
| `NEXT_PUBLIC_SITE_URL` | canonical site URL (metadata, sitemaps, share images) |
| `CI_DATA_DIR` | raw data lake root (`data/raw`, `data/cache`) |
| `ADMIN_TOKEN` | `/admin` console and `/v1/admin/*` |
| `NCBI_TOOL`, `NCBI_EMAIL`, `NCBI_API_KEY` | E-utilities (PubMed, ClinVar); the key raises the limit to 10 req/s |
| `SEER_API_KEY` | unlocks the SEER connector |
| `OPENFDA_API_KEY` | raises openFDA from 1 000 to 120 000 requests/day |
| `IARC_TERMS_ACCEPTED_BY` | human gate for GLOBOCAN (stays unset until terms are settled) |
| `WORKER_CONCURRENCY`, `CI_MAX_RUN_MINUTES`, `LOG_LEVEL` | worker tuning |
| `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `OPENAI_BASE_URL` | reserved for the future `/ask` layer (not enabled) |
## Quality: tests, QA, accessibility, performance
- `pnpm typecheck` — strict TypeScript across the 8 workspaces.
- `pnpm test` — 528 vitest tests: connector fixtures (every connector), SDK (lake, run, checkpoints),
formulas (shares, log ratios, HHI, growth, termination share, stop-reason rules, pipeline stages,
country codes, map scales, projections, explorer comparability, CSV escaping, graph layout), API
smoke tests against the local database (skipped when unreachable).
- `pnpm --filter @cancerindex/web qa` / `qa:prod` — read-only smoke suite over 44 routes: HTTP 200,
expected text, per-route weight caps, no "Data not yet available" where data must exist, no horizontal
overflow at 390 px, console-error scan (Playwright when available).
- Accessibility: keyboard navigation, focus rings, colour never the only carrier (badges carry text,
charts have legends and data tables, maps have equivalent tables), `aria` labels on SVG figures.
- Performance: ISR caching (`revalidate` per page, `Cache-Control` on public routes), materialized
derived tables, GIN / trigram indexes, `sql.param` for array parameters, request budgets on heavy
pages (graph groups, city layers).
## Deployment and operations
Production runs on the MacLustr cluster (node **M4M64b**: Postgres 17 + pgvector, PM2) behind the
MacLustr Tunnel (WireGuard + Caddy on the BHS64 gateway) at https://www.cancerindex.io.
```bash
# from the laptop — gateway M1M32 orchestrates (mld)
rsync -a --exclude '/data/' --exclude node_modules --exclude .next --exclude .git --exclude '.env*' ./ /tmp/cancerindex-deploy/
mld stage /tmp/cancerindex-deploy cancerindex
mld deploy cancerindex --node M4M64b # install, migrate, seed, sources:sync, build, PM2 restart, health checks
```
| Process (PM2) | Role |
|---|---|
| `cancerindex-web` (:8250) | `next start` |
| `cancerindex-api` (:8251) | Fastify API |
| `cancerindex-worker` | pg-boss: connector schedules, `maintenance.counters` 06:00 UTC → `maintenance.intel` 06:15 → `maintenance.rank` 06:30, hourly health probes and alerts |
| `cancerindex-backup` | nightly `pg_dump` 05:20, 14 daily + 8 weekly retained (`deploy/backup.sh`, `restore.sh`) |
Operational visibility: `/data-updates` (connector state, ingestion log), `/admin` (runs, unresolved
labels, alerts), `pnpm cix doctor`, `system_alerts`. Deploy manifest: `deploy/mld-manifest.cancerindex.json`
(note the excludes are anchored — `/coverage/`, `/data/` — so route folders with those names are kept).
## Brand
The mark reads CI: an open ring — a cell with its nucleus, the "C" — and a graduated
index scale in the accent teal, the "I". It is drawn once in SVG and used everywhere: the header and
footer (BrandMark, theme-aware through currentColor and the accent token),
the favicon (app/icon.svg, favicon.ico, apple-icon.png), the web
manifest and the share images.
| Asset | Path |
|---|---|
| Logo with wordmark | `apps/web/public/brand/logo.svg` |
| Mark (light / dark backgrounds) | `apps/web/public/brand/logo-mark.svg`, `logo-mark-dark.svg` (+ PNG 512) |
| Favicon and app icons | `apps/web/src/app/icon.svg`, `favicon.ico`, `apple-icon.png`, `public/brand/icon-{192,512}.png` |
| Share images | `apps/web/src/lib/og.tsx` (next/og, 1200 × 630, site fonts); `app/opengraph-image.tsx` (live counts) and `opengraph-image.tsx` under `cancer/[slug]`, `gene/[symbol]`, `drug/[slug]`, `biomarker/[slug]`, `trial/[nct]` |
Colours: paper `#fafaf7`, ink `#1c1c1a`, accent teal `#0f5f63` (dark theme `#6cc3c6`). Type: Newsreader
(display), Inter (UI), IBM Plex Mono (identifiers).
## Design language
Scientific, editorial, institutional. Off-white paper, charcoal ink, one restrained teal accent, thin
rules, tabular numerals, dense tables, sparklines and small multiples rather than decorative charts. All
charts are server-rendered SVG with a legend, units, the dataset year and a source line; every value shows
a claim badge (observed · published · curated · regulatory · computed) and a source badge whose popover
gives dataset, version, retrieval date and URL. Mobile is first-class (tables scroll inside their wrap; no
horizontal overflow at 390 px). Light and dark themes share one token set (`--color-*`, `--color-series-*`).
## Known gotchas
- A JS array inside a Drizzle template — `` sql`x = ANY(${arr}::text[])` `` — becomes a `($1,$2,…)` tuple
and fails; use `sql.param(arr)` or `sql.raw` for constant lists.
- Columns declared with the `updatedAt()` helper are physically `updated_at` even when named `computedAt`.
- `db.execute` needs a `type` alias, not an `interface`, for `T`.
- A prop named `ref` cannot be passed to a server component (React reserves it) — the page fails in
production with minified error #441; use another name (`entityRef`).
- SVG `` must be a single text node or hydration fails; `array_agg` over empty arrays errors — use
`unnest` / `string_agg`.
- `opengraph-image.tsx` cannot live inside an optional catch-all segment (`[[...tab]]`); place it one
level up.
- `ranking_snapshots.scope_key` has no source dimension: two sources for the same geography/year/sex
overwrite each other's current snapshot, so a preferred source is chosen per scope.
- A shared alias between a concept and its descendant ("breast cancer") must resolve to the **broadest**
concept when building a dictionary; the narrowest-of-lineage rule is only for sentence-level mentions.
- OncoTree's WAF rejects any User-Agent containing a URL; NCBI throttles at 3 req/s without a key;
openFDA allows 1 000 requests/day without a key; the EMA xlsx has 39 columns and header row 9.
- mld `sync_excludes` patterns must be anchored (`/data/`, `/coverage/`), otherwise route folders with
the same name are dropped from the deploy.
## Roadmap
- Global burden once IARC / GLOBOCAN terms are settled (human decision) and SEER credentials are set;
survival observations and the `five_year_survival` metric follow.
- MHRA and TGA regulatory connectors; populating `drug_approvals.biomarker_ids` from label text.
- Risk factors and attributable burden, screening and guideline registries (metadata and links only,
with change detection), hereditary syndromes, pediatric and rare-cancer views.
- Researcher, institution and funding entities (OpenAlex, NIH RePORTER, CIHR) for the funding-gap view.
- Watchlists, alerts and a grounded, cited `/ask` layer that summarizes indexed sources only.
- Lighter default views for the heaviest pages (`/graph`, biomarker pages with hundreds of approvals).
## Licensing, attribution and contact
- **Code**: private repository (spbgit `cancerindex.git`).
- **Derived data** (rankings, intelligence tables, CSV exports): © CancerIndex, CC BY 4.0, with
attribution rows naming the underlying providers.
- **Source data** remains under the licence of each provider — see `/sources` and the `sources` array of
every API response (NCIt CC BY 4.0, OncoTree CC BY 4.0, HGNC CC0, CIViC CC0, openFDA CC0, Health Canada
OGL, EMA with acknowledgement, ClinVar / PubMed / ClinicalTrials.gov / CDC public domain, GDC open tier,
ChEMBL CC BY-SA).
- **Not medical advice.** CancerIndex provides research and educational information only.
Built and maintained by **Simon-Pierre Boucher** — contact **contact@spboucher.ai** — hosted on
**MacLustr** ().