SPB Git

spb/airiskindex Public

The most methodologically rigorous, fully transparent AI job-exposure index.

TypeScript 88% Python 6.1% SQL 2.7% CSS 1.2% JavaScript 0.9% Shell 0.8%
9.7 KB

# AI Risk Index — Methodology v1.0.0 (DRAFT)

Status: draft, unreleased. This document is the source of truth for the scoring methodology (see CLAUDE.md §9). Code in packages/scoring links to the section anchors below; if code and this document disagree, stop and flag it.

Related work and evidence base: docs/research/01-existing-indices.md (indices), 02-data-sources.md (data), 04-landscape-and-evidence.md (adoption & barriers evidence).


# 1. Principles

  1. Task-based, not occupation-based. Occupations are bundles of tasks with very different AI exposure; whole-occupation scoring (Frey & Osborne 2013) has a poor empirical record. Occupation scores are always derived from task scores (Arntz, Gregory & Zierahn 2016; Eloundou et al. 2024).
  2. Three concepts, never collapsed (see §5):
    • Exposure — AI is technically capable of performing the task.
    • Substitution — AI actually replaces the human performing it. This is the headline composite.
    • Augmentation — AI assists the human, increasing productivity. Realized labor-market effects concentrate in automation-classified usage, not augmentation (Brynjolfsson, Chandar & Chen 2025), so conflating the three is not just imprecise — it is empirically wrong.
  3. Uncertainty is part of the score. Every published score carries score_low / score / score_high. Interval width is driven primarily by disagreement between independent LLM raters (single-model ratings show up to a 19× spread in headline statistics across frontier models — Yin, Vu & Persico 2026) plus human-calibration error on the expert anchor set.
  4. Fully reproducible. Weights, formulas, prompt versions, rater model IDs and per-model ratings are all published. Every score traces to a stored score_runs row; historical runs are immutable.
  5. Versioned. Any change to weights, formulas, or rater prompts bumps INDEX_VERSION (semver) with a changelog entry and regenerated sensitivity outputs.

# 2. Data

Input Source Role
Occupations & task statements ONET 30.x (ONET-SOC 2019 taxonomy), CC BY 4.0 Unit of analysis (~18k tasks, ~900 data-level occupations)
Task importance weights O*NET Task Ratings (IM scale 1–5) Aggregation weights (§6)
EU/France occupations ESCO v1.2.x + official ESCO↔O*NET crosswalk; ROME 4.0 Crosswalked scores (crosswalk loss documented per occupation)
Wages BLS OEWS (latest May release); Eurostat SES; INSEE cost_ratio denominator
Adoption data Census BTOS AI supplement, Ramp AI Index, Anthropic Economic Index (Hugging Face), Challenger reports adoption_velocity inputs (§4.5)
Human anchor ratings Expert Delphi panel (data/derived/expert_overrides/) LLM-rater calibration (§3)

Raw dumps are immutable (data/raw/, gitignored); all transformations go through apps/etl; derived artifacts commit only manifests (hashes + row counts).

# 3. Task rating (LLM-as-evaluator)

Each O*NET task statement is rated on a 5-point scale per dimension (§4) by a panel of ≥2 (target 3) frontier LLMs (RATER_MODELS), using versioned rubric prompts (apps/worker/src/raters/prompts/). Per task × dimension:

  • rating_mid = mean of panel ratings;
  • rating_low / rating_high = min / max of panel ratings (rater-disagreement band).

Rubric anchors follow the citable standards: automatability uses the Eloundou et al. "≥50% time saving at equal quality" threshold, decomposed into named criteria (SML-style multi-criterion rubric); ratings require structured justifications and are stored with model ID, prompt version and raw response (full audit trail).

Calibration: 5% of ratings are sampled for human review; panel-vs-human disagreement > 1 point routes the task to the expert queue. Expert overrides replace the LLM band for that task and are flagged in the API output.

# 4. Dimensions

Composite weights live in packages/scoring/src/weights.ts and are published at /api/v1/methodology. Weights sum to 1.0.

Dimension Key Weight Orientation Rubric anchor (rating of 5 means…)
Task automatability automatability 0.35 direct Current AI (incl. tooling/agents) can do the task with ≥50% time saving at equal quality
Current technical feasibility feasibility 0.20 direct Deployed products demonstrably perform this task reliably today (not merely conceivable)
Cost of substitution vs. wage cost_ratio 0.15 direct AI cost per task-equivalent ≪ loaded human wage for the same output
Adoption barriers barriers 0.20 inverted Strong barriers: licensing/liability/regulation/human-contact requirements block substitution
Sector adoption velocity adoption_velocity 0.10 direct Occupation's dominant sectors adopt AI fast and deep (agents in production, measured displacement)

Orientation: a direct dimension's higher rating increases substitution pressure; an inverted dimension's higher rating decreases it. barriers is the only inverted dimension in v1. Orientation is encoded once, in packages/scoring/src/weights.ts, next to the weights.

Additionally, augmentation is rated per task on the same 5-point scale (does AI assist the human on this task, raising productivity without replacing them?). It is not part of the substitution composite; it feeds the augmentation sub-score (§5).

# 4.4 barriers components

Rated against five named criteria (see research doc 04): licensing/authorization requirement (0.30), liability & error-cost asymmetry (0.25), regulatory-process coverage (0.20, jurisdiction-specific), human-contact requirement (0.15), organizational friction (0.10). Pre-2022 null results (OECD 2021; Acemoglu et al. 2022) show barriers dominate short-run outcomes — hence the 0.20 weight.

# 4.5 adoption_velocity components

Grounded in measured adoption, not forecasts: sector AI-use rate employment-weighted (0.35), sector adoption momentum (0.20), agentic deployment depth (0.20), realized displacement intensity (0.15), occupation-level usage intensity from the Anthropic Economic Index (0.10). Refreshed each index release; sources are public and dated.

# 5. Scoring formulas

All formulas are implemented, pure and deterministic, in packages/scoring. Ratings r ∈ [1,5] normalize to pressure p ∈ [0,1]:

text
p = (r − 1) / 4            (direct dimensions)
p = 1 − (r − 1) / 4        (inverted dimensions)

Per task:

text
substitution_task = 100 · Σ_d  w_d · p_d                     (all five dimensions)
exposure_task     = 100 · (w_auto·p_auto + w_feas·p_feas) / (w_auto + w_feas)
augmentation_task = 100 · p_augmentation

Confidence bounds: score_low is computed with each dimension's pressure-minimizing rating bound (for direct dimensions the low rating; for inverted dimensions the high rating), score_high symmetrically. Bounds are therefore worst/best-case envelopes over rater disagreement, and low ≤ score ≤ high always holds.

# 6. Aggregation to occupations

Occupation scores are the importance-weighted mean of task scores, weights from O*NET Task Ratings importance (IM, 1–5), normalized to sum to 1 within the occupation. Applied identically to low, score and high. Tasks lacking importance ratings receive the occupation-mean importance.

The API additionally reports the share of tasks with substitution_task ≥ 70 ("highly exposed task share") — this, not the composite alone, is the preferred headline in UI copy ("X% of tasks in this occupation are highly exposed").

# 7. Versioning & runs

  • INDEX_VERSION (semver) in packages/scoring/src/version.ts.
  • MAJOR: formula/weight changes. MINOR: data-source version bumps (new O*NET release, new adoption data). PATCH: recomputation with refreshed adoption inputs, prompt clarifications that don't change the rubric semantics.
  • Every computation writes a score_runs row (index version, prompt version, rater models). Old runs stay queryable forever; the API serves the latest by default and any run on request.

# 8. Validation & sensitivity

Published with every MAJOR/MINOR release under docs/methodology/sensitivity/:

  1. Convergent validity: Spearman correlation of our occupation scores against Felten AIOE, Eloundou β, and ILO WP140 gradients (ρ ≈ 0.84 between independent modern methodologies is the reference bar).
  2. Rater stability: cross-model agreement distribution; occupations with the widest bands flagged in-product.
  3. Outcome tracking: correlation against the Stanford "Canaries" dashboard (entry-level employment in exposed occupations) and AEI usage shares — exposure indices individually explain <11% of realized unemployment risk (Frank, Ahn & Moro 2025), so we report outcome tracking honestly rather than claiming prediction.
  4. Weight sensitivity: composite rank stability under ±25% perturbation of each weight.

# 9. Known limitations

  • LLM raters co-evolve with the technology they measure (the "ruler" problem); multi-model panels bound but do not eliminate this.
  • Sector-level adoption inputs are priors, corrected by occupation-level usage data.
  • ESCO/ROME crosswalked scores inherit crosswalk loss; flagged per occupation.
  • Scores describe tasks as currently constituted; occupations reorganize.
  • This index measures exposure and substitution pressure, not certainty of job loss. Product copy follows the adaptation-not-doom tone guide accordingly.