AI Risk Index — Methodology v1.0.0 (DRAFT)
Status: draft, unreleased. This document is the source of truth for the scoring methodology (see CLAUDE.md §9). Code in
packages/scoringlinks to the section anchors below; if code and this document disagree, stop and flag it.Related work and evidence base:
docs/research/01-existing-indices.md(indices),02-data-sources.md(data),04-landscape-and-evidence.md(adoption & barriers evidence).
1. Principles
- Task-based, not occupation-based. Occupations are bundles of tasks with very different AI exposure; whole-occupation scoring (Frey & Osborne 2013) has a poor empirical record. Occupation scores are always derived from task scores (Arntz, Gregory & Zierahn 2016; Eloundou et al. 2024).
- Three concepts, never collapsed (see §5):
- Exposure — AI is technically capable of performing the task.
- Substitution — AI actually replaces the human performing it. This is the headline composite.
- Augmentation — AI assists the human, increasing productivity. Realized labor-market effects concentrate in automation-classified usage, not augmentation (Brynjolfsson, Chandar & Chen 2025), so conflating the three is not just imprecise — it is empirically wrong.
- Uncertainty is part of the score. Every published score carries
score_low / score / score_high. Interval width is driven primarily by disagreement between independent LLM raters (single-model ratings show up to a 19× spread in headline statistics across frontier models — Yin, Vu & Persico 2026) plus human-calibration error on the expert anchor set. - Fully reproducible. Weights, formulas, prompt versions, rater model IDs and
per-model ratings are all published. Every score traces to a stored
score_runsrow; historical runs are immutable. - Versioned. Any change to weights, formulas, or rater prompts bumps
INDEX_VERSION(semver) with a changelog entry and regenerated sensitivity outputs.
2. Data
| Input | Source | Role |
|---|---|---|
| Occupations & task statements | ONET 30.x (ONET-SOC 2019 taxonomy), CC BY 4.0 | Unit of analysis (~18k tasks, ~900 data-level occupations) |
| Task importance weights | O*NET Task Ratings (IM scale 1–5) | Aggregation weights (§6) |
| EU/France occupations | ESCO v1.2.x + official ESCO↔O*NET crosswalk; ROME 4.0 | Crosswalked scores (crosswalk loss documented per occupation) |
| Wages | BLS OEWS (latest May release); Eurostat SES; INSEE | cost_ratio denominator |
| Adoption data | Census BTOS AI supplement, Ramp AI Index, Anthropic Economic Index (Hugging Face), Challenger reports | adoption_velocity inputs (§4.5) |
| Human anchor ratings | Expert Delphi panel (data/derived/expert_overrides/) |
LLM-rater calibration (§3) |
Raw dumps are immutable (data/raw/, gitignored); all transformations go through
apps/etl; derived artifacts commit only manifests (hashes + row counts).
3. Task rating (LLM-as-evaluator)
Each O*NET task statement is rated on a 5-point scale per dimension (§4) by a
panel of ≥2 (target 3) frontier LLMs (RATER_MODELS), using versioned rubric
prompts (apps/worker/src/raters/prompts/). Per task × dimension:
rating_mid= mean of panel ratings;rating_low/rating_high= min / max of panel ratings (rater-disagreement band).
Rubric anchors follow the citable standards: automatability uses the Eloundou et al. "≥50% time saving at equal quality" threshold, decomposed into named criteria (SML-style multi-criterion rubric); ratings require structured justifications and are stored with model ID, prompt version and raw response (full audit trail).
Calibration: 5% of ratings are sampled for human review; panel-vs-human disagreement > 1 point routes the task to the expert queue. Expert overrides replace the LLM band for that task and are flagged in the API output.
4. Dimensions
Composite weights live in packages/scoring/src/weights.ts and are published at
/api/v1/methodology. Weights sum to 1.0.
| Dimension | Key | Weight | Orientation | Rubric anchor (rating of 5 means…) |
|---|---|---|---|---|
| Task automatability | automatability |
0.35 | direct | Current AI (incl. tooling/agents) can do the task with ≥50% time saving at equal quality |
| Current technical feasibility | feasibility |
0.20 | direct | Deployed products demonstrably perform this task reliably today (not merely conceivable) |
| Cost of substitution vs. wage | cost_ratio |
0.15 | direct | AI cost per task-equivalent ≪ loaded human wage for the same output |
| Adoption barriers | barriers |
0.20 | inverted | Strong barriers: licensing/liability/regulation/human-contact requirements block substitution |
| Sector adoption velocity | adoption_velocity |
0.10 | direct | Occupation's dominant sectors adopt AI fast and deep (agents in production, measured displacement) |
Orientation: a direct dimension's higher rating increases substitution pressure; an
inverted dimension's higher rating decreases it. barriers is the only inverted
dimension in v1. Orientation is encoded once, in packages/scoring/src/weights.ts,
next to the weights.
Additionally, augmentation is rated per task on the same 5-point scale (does AI
assist the human on this task, raising productivity without replacing them?). It is
not part of the substitution composite; it feeds the augmentation sub-score (§5).
4.4 barriers components
Rated against five named criteria (see research doc 04): licensing/authorization requirement (0.30), liability & error-cost asymmetry (0.25), regulatory-process coverage (0.20, jurisdiction-specific), human-contact requirement (0.15), organizational friction (0.10). Pre-2022 null results (OECD 2021; Acemoglu et al. 2022) show barriers dominate short-run outcomes — hence the 0.20 weight.
4.5 adoption_velocity components
Grounded in measured adoption, not forecasts: sector AI-use rate employment-weighted (0.35), sector adoption momentum (0.20), agentic deployment depth (0.20), realized displacement intensity (0.15), occupation-level usage intensity from the Anthropic Economic Index (0.10). Refreshed each index release; sources are public and dated.
5. Scoring formulas
All formulas are implemented, pure and deterministic, in packages/scoring.
Ratings r ∈ [1,5] normalize to pressure p ∈ [0,1]:
p = (r − 1) / 4 (direct dimensions)
p = 1 − (r − 1) / 4 (inverted dimensions)Per task:
substitution_task = 100 · Σ_d w_d · p_d (all five dimensions)
exposure_task = 100 · (w_auto·p_auto + w_feas·p_feas) / (w_auto + w_feas)
augmentation_task = 100 · p_augmentationConfidence bounds: score_low is computed with each dimension's
pressure-minimizing rating bound (for direct dimensions the low rating; for inverted
dimensions the high rating), score_high symmetrically. Bounds are therefore
worst/best-case envelopes over rater disagreement, and low ≤ score ≤ high always
holds.
6. Aggregation to occupations
Occupation scores are the importance-weighted mean of task scores, weights from
O*NET Task Ratings importance (IM, 1–5), normalized to sum to 1 within the
occupation. Applied identically to low, score and high. Tasks lacking
importance ratings receive the occupation-mean importance.
The API additionally reports the share of tasks with substitution_task ≥ 70
("highly exposed task share") — this, not the composite alone, is the preferred
headline in UI copy ("X% of tasks in this occupation are highly exposed").
7. Versioning & runs
INDEX_VERSION(semver) inpackages/scoring/src/version.ts.- MAJOR: formula/weight changes. MINOR: data-source version bumps (new O*NET release, new adoption data). PATCH: recomputation with refreshed adoption inputs, prompt clarifications that don't change the rubric semantics.
- Every computation writes a
score_runsrow (index version, prompt version, rater models). Old runs stay queryable forever; the API serves the latest by default and any run on request.
8. Validation & sensitivity
Published with every MAJOR/MINOR release under docs/methodology/sensitivity/:
- Convergent validity: Spearman correlation of our occupation scores against Felten AIOE, Eloundou β, and ILO WP140 gradients (ρ ≈ 0.84 between independent modern methodologies is the reference bar).
- Rater stability: cross-model agreement distribution; occupations with the widest bands flagged in-product.
- Outcome tracking: correlation against the Stanford "Canaries" dashboard (entry-level employment in exposed occupations) and AEI usage shares — exposure indices individually explain <11% of realized unemployment risk (Frank, Ahn & Moro 2025), so we report outcome tracking honestly rather than claiming prediction.
- Weight sensitivity: composite rank stability under ±25% perturbation of each weight.
9. Known limitations
- LLM raters co-evolve with the technology they measure (the "ruler" problem); multi-model panels bound but do not eliminate this.
- Sector-level adoption inputs are priors, corrected by occupation-level usage data.
- ESCO/ROME crosswalked scores inherit crosswalk loss; flagged per occupation.
- Scores describe tasks as currently constituted; occupations reorganize.
- This index measures exposure and substitution pressure, not certainty of job loss. Product copy follows the adaptation-not-doom tone guide accordingly.