SPB Git

spb/airiskindex Public

The most methodologically rigorous, fully transparent AI job-exposure index.

TypeScript 88% Python 6.1% SQL 2.7% CSS 1.2% JavaScript 0.9% Shell 0.8%
53.6 KB

# Existing AI Job-Exposure / Automation-Risk Indices and Methodologies

Research memo — AI Risk Index (airiskindex.io) Date compiled: 2026-08-05 (web research current to August 2026) Scope: All major academic and industry indices measuring occupational exposure to AI/automation, their methodologies, criticisms, empirical validation, and lessons for the airiskindex.io v1 scoring model.


# Table of Contents

  1. First wave: pre-generative-AI automation risk (2013–2019)
  2. Second wave: AI-specific exposure measures (2018–2021)
  3. Third wave: LLM/generative-AI exposure (2023–2024)
  4. Institutional indices: ILO, OECD, IMF (2023–2026)
  5. Industry & consultancy estimates
  6. Usage-based measures: Anthropic Economic Index & OpenAI (2025–2026)
  7. Fourth wave: 2025–2026 indices and meta-critiques
  8. Empirical validation: do exposure scores predict real outcomes?
  9. Master comparison table
  10. Implications for airiskindex.io v1 methodology

# 1. First wave: pre-generative-AI automation risk (2013–2019)

# 1.1 Frey & Osborne — "The Future of Employment" (2013 working paper; 2017 published)

  • Authors/year: Carl Benedikt Frey & Michael A. Osborne (Oxford Martin School). Working paper Sept 2013; published in Technological Forecasting and Social Change 114 (2017): 254–280.
  • Unit of analysis: Whole occupations (702 SOC occupations).
  • Methodology:
    • ML experts hand-labeled ~70 occupations as automatable (1) or not (0) at a workshop.
    • Identified three "engineering bottlenecks" to computerisation: perception & manipulation, creative intelligence, social intelligence, operationalized via 9 O*NET variables (e.g., finger dexterity, originality, social perceptiveness, persuasion, negotiation, assisting/caring for others, cramped work spaces).
    • A Gaussian process classifier trained on the 70 labels extrapolated a "probability of computerisation" (0–1) to all 702 occupations.
    • Occupations bucketed: high risk (p > 0.7), medium (0.3–0.7), low (p < 0.3).
  • Key numbers: 47% of US employment at "high risk" of computerisation "over the next decade or two" (i.e., roughly by 2030). Transportation, logistics, office/administrative support, and production occupations most at risk.
  • Criticisms (extensive):
    • Occupation-level, not task-level: treats occupations as monolithic. Arntz et al. (2016) showed within-occupation task heterogeneity slashes the estimate to ~9%.
    • Technical capability ≠ adoption: no economics (cost, wages, regulation, preferences) in the model.
    • Subjective training labels from a small expert workshop; only ~70 seed labels drive all 702 predictions.
    • Model-selection sensitivity: "The Future of Employment Revisited" (arXiv:2104.13747) shows automation forecasts swing heavily with classifier choice on the same labels.
    • Poor ex-post predictive record: occupations flagged high-risk did not experience differential employment declines through the late 2010s (see §8; Frank, Ahn & Moro 2025 find F&O scores explain ❤️% of unemployment-risk variation individually).
  • URLs:

# 1.2 Arntz, Gregory & Zierahn — OECD task-based critique (2016)

  • Authors/year: Melanie Arntz, Terry Gregory, Ulrich Zierahn (ZEW/OECD). "The Risk of Automation for Jobs in OECD Countries", OECD Social, Employment and Migration Working Paper No. 189 (2016); follow-up "Revisiting the risk of automation" in Economics Letters 159 (2017).
  • Unit of analysis: Individual workers' task bundles (PIAAC Survey of Adult Skills microdata), not occupation averages.
  • Methodology: Transferred Frey–Osborne occupation-level risk to individual workers, then re-estimated risk as a function of each worker's actual reported tasks (PIAAC), allowing within-occupation heterogeneity. High risk = automatability > 70%.
  • Key numbers: Only ~9% of jobs across 21 OECD countries at high risk (US 9%, Germany 12%, Korea 6%) — versus 47% under F&O.
  • Significance: Founded the task-based paradigm that every serious index since has adopted — including airiskindex.io. Key insight: workers in the "same" occupation do different task mixes; scoring must start at the task level.
  • Criticisms: Still anchored on F&O's original subjective labels; PIAAC task self-reports are coarse; may understate risk if task bundles themselves adjust post-automation.
  • URLs:

# 1.3 Brynjolfsson, Mitchell & Rock — Suitability for Machine Learning (SML) (2017–2018)

  • Authors/year: Erik Brynjolfsson, Tom Mitchell, Daniel Rock. "What Can Machine Learning Do? Workforce Implications" (Science, 2017); "What Can Machines Learn, and What Does It Mean for Occupations and the Economy?" (AEA Papers & Proceedings 108, 2018: 43–47).
  • Unit of analysis: O*NET tasks (18,156 tasks; ~2,069 Detailed Work Activities), aggregated to ~950 occupations.
  • Methodology:
    • A 23-question rubric capturing what (then-current, supervised) ML can do — e.g., mapping well-defined inputs to outputs, tolerance for error, no long chains of reasoning, digital data availability, no need for detailed physical manipulation.
    • Each task scored 1–5 per question; rubric validated by ML experts, then scaled via CrowdFlower crowd workers; aggregated to task SML then occupation SML (task-importance weighted).
  • Key findings: (1) ML affects different occupations than earlier automation waves; (2) most occupations have at least some high-SML tasks; (3) almost no occupation is fully automatable; (4) capturing value requires task re-bundling / job redesign.
  • Criticisms: Rubric was tuned to pre-LLM supervised ML (weak on generation, reasoning, dialogue); crowd ratings noisy; scores never strongly validated against outcomes (explains ~0–3% of unemployment-risk variation individually per Frank et al. 2025).
  • Relevance to us: Direct methodological ancestor of a multi-criterion task rubric — our automatability + feasibility split echoes SML's separation of "could ML do it" from "is it practical".
  • URLs:

# 2. Second wave: AI-specific exposure measures (2018–2021)

# 2.1 Webb (2020) — Patent-based exposure

  • Author/year: Michael Webb (Stanford). "The Impact of Artificial Intelligence on the Labor Market" (SSRN 3482150, Nov 2019/2020; still unpublished but heavily cited).
  • Unit of analysis: O*NET task descriptions × patent text; aggregated to occupations.
  • Methodology:
    • Selected AI patents (~16,400) by keyword; dependency-parsed titles to extract verb–object pairs (~8,000 pairs, e.g., "diagnose disease", "detect fraud").
    • Extracted verb–object pairs from O*NET task statements; scored each task by the frequency with which its verb-object pairs appear in AI patents.
    • Occupation score = task-importance-weighted average; reported as exposure percentiles.
    • Built-in validation strategy: applied the same method to robots and software patents and showed those historical exposure measures predicted realized employment/wage declines in exposed occupations — then applied it to AI.
  • Key findings: AI (unlike robots/software) exposes high-skilled, high-wage, older workers most: e.g., clinical lab technicians, chemical engineers, optometrists, radiologic technicians. Robots hit low-skill physical work; software hit mid-skill routine work.
  • Criticisms: Patents lag and imperfectly reflect deployable capability (pre-LLM corpus, so misses generative AI entirely); verb-object matching is crude semantics; no distinction between substitution and augmentation.
  • URLs:

# 2.2 Felten, Raj & Seamans — AI Occupational Exposure (AIOE) (2018, 2021, 2023)

  • Authors/year: Edward Felten (Princeton), Manav Raj (Wharton), Robert Seamans (NYU). "Occupational, Industry, and Geographic Exposure to Artificial Intelligence: A Novel Dataset and Its Potential Uses", Strategic Management Journal 42(12), 2021: 2195–2217. Generative-AI update: "Occupational Heterogeneity in Exposure to Generative AI" (SSRN 4414065, April 2023).
  • Unit of analysis: 52 O*NET abilities (not tasks) × 10 AI application areas; aggregated to occupations (also industry AIIE and county-level geography).
  • Methodology:
    • 10 AI applications from the EFF AI Progress Measurement project (image recognition, language modeling, translation, speech recognition, abstract strategy games, etc.).
    • Amazon Mechanical Turk crowd workers rated relatedness of each application to each of 52 O*NET abilities → ability-level exposure = sum of relatedness scores.
    • Occupation AIOE = weighted sum of ability exposures using O*NET ability importance and prevalence weights.
    • 2023 generative-AI variant re-weights toward language modeling and image generation: top exposed = telemarketers, then post-secondary teachers (languages, history, law), sociologists, judges.
  • Key properties: Continuous z-scored index; famously positively correlated with wages and education (white-collar exposure) — opposite sign to Frey–Osborne.
  • Criticisms: Ability-level (even further from tasks than occupations); MTurk relatedness judgments are lay opinions; "exposure" deliberately neutral between substitution and augmentation (authors are explicit about this); static.
  • Data: Public GitHub (occupation-, industry-, geography-level scores): https://github.com/AIOE-Data/AIOE
  • Adoption: The IMF (Cazzaniga et al. 2024) and Pew (2023) analyses build directly on AIOE; the ECB European work (Albanesi et al.) uses AIOE + Webb.
  • URLs:

# 2.3 Other second-wave measures (brief)

  • Georgieff & Hyee (OECD 2021), "Artificial intelligence and employment: New cross-country evidence" — applied Felten-style AIOE to PIAAC across 23 OECD countries; most-exposed: business professionals, managers, chief executives, science/engineering professionals. Found no negative employment relationship 2012–2019; in high-computer-use occupations, higher AI exposure correlated with higher employment growth. URL: https://doi.org/10.1787/c2c1d276-en
  • Lassébie & Quintini (OECD 2022), "What skills and abilities can automation technologies replicate?" — expert survey on automatability of ~100 O*NET skills/abilities; basis for OECD Employment Outlook 2023 statement that occupations at highest risk of automation account for ~27% of OECD employment. URL: https://doi.org/10.1787/646aad77-en ; Employment Outlook 2023 AI chapter: https://www.oecd.org/en/publications/oecd-employment-outlook-2023_08785bba-en.html
  • Tolan et al. (2021, JRC/EC) — mapped AI research benchmarks to cognitive abilities to tasks; early "capability→ability→occupation" chain, precursor of the 2026 OECD capability-indicator approach.
  • Startup-based exposure ("Follow the money", arXiv 2024) — measures exposure via commercial AI startup activity mapped to occupations; argues patent/ability measures miss commercialization. URL: https://arxiv.org/abs/2412.04924

# 3. Third wave: LLM/generative-AI exposure (2023–2024)

# 3.1 Eloundou, Manning, Mishkin & Rock — "GPTs are GPTs" (OpenAI/Wharton, 2023; Science 2024)

The single most influential template for airiskindex.io's LLM-as-evaluator pipeline.

  • Authors/year: Tyna Eloundou, Sam Manning, Pamela Mishkin (OpenAI), Daniel Rock (Wharton). arXiv:2303.10130 (Mar 2023); published as "GPTs are GPTs: Labor market impact potential of LLMs", Science 384(6702), June 2024: 1306–1308.
  • Unit of analysis: O*NET task/DWA level (19,265 tasks; 2,087 DWAs), aggregated to 1,016 occupations with task weights; combined with BLS employment/wage data.
  • Exposure rubric (the core innovation): Exposure = "whether access to an LLM or LLM-powered system would reduce the time required for a human to perform a specific task by at least 50% while maintaining quality". Three levels:
    • E0 — no exposure.
    • E1 — direct exposure: LLM alone (via chat/API) achieves the 50% time reduction.
    • E2 — LLM+ exposure: achievable only with additional software/tooling built on the LLM (image input, retrieval, agents, etc.).
    • Aggregates: α = E1 (lower bound), β = E1 + 0.5·E2 (expected), ζ = E1 + E2 (upper bound).
  • Raters: Both human annotators (OpenAI staff, trained on rubric) and GPT-4 itself as a rater with the rubric as prompt; human–GPT-4 agreement was high (occupation-level correlations ≈ 0.80), pioneering the LLM-as-evaluator design we plan to use.
  • Key numbers:
    • ~80% of US workers have ≥10% of tasks exposed (β); ~19% of workers have ≥50% of tasks exposed.
    • ~1.8% of jobs have >half their tasks E1-exposed; rises to ~46% of jobs under ζ (with LLM-powered software).
    • Exposure increases with wage and education (up to a point); science and critical-thinking-intensive skills correlate negatively; programming and writing positively.
  • Criticisms:
    • Measures potential time savings, not substitution vs augmentation, adoption, or net employment effect (authors are explicit).
    • Rater instability: the flagship statistic is wildly model-dependent (see Yin et al. 2026, §7.2: 2.7%–51.5% across frontier raters).
    • Static snapshot of March-2023 GPT-4 capability; the "E2 software will exist" counterfactual is speculative.
    • 50%-time-saving threshold is arbitrary; binary-ish levels lose information.
  • URLs:

# 3.2 Goldman Sachs — Briggs & Kodnani (March 2023)

See §5.1. Methodologically an O*NET task-importance exercise inspired by Eloundou-style exposure; headline "300 million FTE jobs exposed" globally.


# 4. Institutional indices: ILO, OECD, IMF (2023–2026)

# 4.1 ILO — Gmyrek, Berg & Bescond (2023): "Generative AI and Jobs: A Global Analysis"

  • Authors/year: Paweł Gmyrek, Janine Berg, David Bescond. ILO Working Paper 96, August 2023.
  • Methodology: Scored ISCO-08 occupation task lists (not O*NET) with GPT-4 as rater (multiple prompts, averaged), producing task-level automation-potential scores; distinguished automation potential vs augmentation potential at occupation level; mapped to global employment via ILO harmonized microdata for 100+ countries, by income group and sex.
  • Key numbers:
    • Only clerical support work is highly exposed as a group: 24% of clerical tasks highly exposed, +58% medium exposure. Other occupational groups: 1–4% of tasks highly exposed.
    • Globally, ~2.3% of employment (~75M jobs) in the top automation-potential bucket; 13.4% (~427M) in augmentation potential.
    • Exposure concentrated in high/upper-middle-income countries (more clerical employment) and strongly gendered (clerical work is female-dominated: in high-income countries, several times more female than male employment in the highest-exposure category).
  • Framing: "Augmentation, not automation, is the most likely impact" — the origin of the transformation-over-replacement institutional narrative.
  • URLs:

# 4.2 ILO 2024 interim work

  • "Mind the AI Divide: Shaping a Global Perspective on the Future of Work" (ILO & World Bank, Aug 2024) — applies the 2023 index with a digital-infrastructure overlay: poor countries are less exposed but also less able to capture augmentation gains (the "AI divide").
  • Gmyrek, Winkler & Garganta (2024, ILO/World Bank) — "Buffer or bottleneck? Employment exposure to generative AI and the digital divide in Latin America": 26–38% of LAC jobs exposed; digital access gates both risk and benefit.
  • URL hub: https://www.ilo.org/publications (search "generative AI"); LAC paper: https://openknowledge.worldbank.org/handle/10986/41808

# 4.3 ILO–NASK (2025): "Generative AI and Jobs: A Refined Global Index of Occupational Exposure" — current institutional state of the art

# 4.4 OECD (2023–2026)

# 4.5 IMF — Cazzaniga et al. (2024) and the AI Preparedness Index


# 5. Industry & consultancy estimates

# 5.1 Goldman Sachs (Briggs & Kodnani, March 2023)

# 5.2 McKinsey Global Institute (June 2023, updated)

  • "The economic potential of generative AI: the next productivity frontier". Proprietary work-activity/capability model (~2,100 work activities, 850 occupations).
  • Key numbers: GenAI + existing tech could automate activities absorbing 60–70% of employees' time; genAI value $2.6–4.4T/yr; half of today's work activities automated between 2030 and 2060 (midpoint ~2045) — pulled forward ~a decade vs pre-genAI estimate.
  • Criticism: proprietary/black-box capability ratings; "time automatable" ≠ jobs; adoption scenarios highly assumption-driven.
  • URL: https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier

# 5.3 Pew Research Center (Kochhar, July 2023)

# 5.4 PwC Global AI Jobs Barometer (2024, 2025, 2026)


# 6. Usage-based measures: Anthropic Economic Index & OpenAI (2025–2026)

The decisive innovation of 2025–26: replacing predicted exposure with observed AI usage mapped to the same O*NET task taxonomy. This is the empirical anchor airiskindex.io should exploit for adoption_velocity and to validate automatability.

# 6.1 Anthropic Economic Index (AEI) — all releases to August 2026

Methodology (constant across releases): Clio, a privacy-preserving analysis pipeline in which Claude classifies large samples of real Claude.ai/API conversations against the O*NET task taxonomy (~20,000 tasks) and SOC occupations, plus interaction-mode classification (automation = full delegation/directive; augmentation = iterative collaboration, learning, validation). All aggregated data released openly on Hugging Face.

Release Report Data & model Headline findings
Feb 10, 2025 (paper arXiv:2503.04761, Handa et al., "Which Economic Tasks are Performed with AI?") Launch report 4M+ Claude.ai conversations 36% of occupations used AI for ≥25% of their tasks; only ~4% for ≥75%. Usage concentrated in software development & writing (computer/math ≈ 37% of conversations); peaks in mid-to-high-wage occupations, low at both wage extremes. 57% augmentation / 43% automation.
Mar 27, 2025 v2 (Claude 3.7 Sonnet) New conversations + cluster-level data Usage patterns stable; extended-thinking usage concentrated in technical tasks; released bottom-up task clusters.
Sep 15, 2025 (arXiv:2511.15080) "Uneven geographic and enterprise adoption" 1P API + geographic breakdowns; Anthropic AI Usage Index (AUI) = country share of usage ÷ share of working-age population US 21.6% of usage; per-capita leaders Israel, Singapore, Australia, NZ, S. Korea. +1% GDP/capita ↔ +0.7% AUI (US states: 1.8% elasticity). DC highest state AUI (3.82). Automation rose 27%→39% of conversations since Dec 2024, surpassing augmentation for the first time; API usage even more automation-heavy.
Jan 15, 2026 "New building blocks" (economic primitives) 1M Claude.ai + 1M 1P API transcripts (Sonnet 4.5); Nov 2025 data Five primitives: task complexity, human/AI skill level, use case, AI autonomy, task success. College-level tasks: 12x estimated speedup but 66% success rate vs 70% for simpler tasks; revised aggregate productivity estimate +1.2 pp/yr (down from 1.8 after reliability adjustment). Augmentation back above automation on Claude.ai (52% vs 45%).
Mar 24, 2026 "Learning curves" Feb 2026 data (Opus 4.5/4.6) Claude.ai task mix de-concentrating (top-10 tasks 24%→19%) while API concentrates (28%→33%). 49% of jobs in sample now see Claude used for ≥25% of tasks (up from 36% in Jan 2025). 6-month+ tenure users: +10% conversation success; usage value ≈ $48–49/hr wage-equivalent tasks.
Apr 2026 AEI Survey launched 9,700 Claude users, linked usage+perceptions See below.
Jun 26, 2026 "Cadences" Apr–Jun 2026, hourly sampling; artifact classifier 93% of conversations produce artifacts (explanations 17%, documents/reports 15%). Higher-wage occupations' conversations consume 2.07x tokens. Survey: >⅓ of users expect AI to handle most of their work tasks within 12 months; only 10% rate own job loss likely; heavier automation users are more optimistic. Women use Claude less in automated modes (−0.33 SD).

# 6.2 OpenAI — "How People Use ChatGPT" (Chatterji et al., NBER w34255, Sept 2025)

  • Aaron Chatterji, Tom Cunningham, David Deming, Zoë Hitzig, Christopher Ong, Carl Shan, Kevin Wadman. Privacy-preserving classification of a representative sample of ChatGPT consumer conversations, Nov 2022–Jul 2025 (~10% of world adult population using ChatGPT).
  • Findings: non-work usage grew from 53%→>70% of messages; work usage concentrated in decision support (advice, writing, information) rather than task execution; work usage highest among educated, high-paid professionals. Three-quarters of work messages: writing, information seeking, decision support.
  • Relevance: independent replication that realized usage is augmentation-tilted and knowledge-work-concentrated — cross-provider triangulation for adoption_velocity.
  • URLs: https://www.nber.org/papers/w34255 (PDF: https://www.nber.org/system/files/working_papers/w34255.pdf)

# 7. Fourth wave: 2025–2026 indices and meta-critiques

# 7.1 Stanford Digital Economy Lab — "Canaries in the Coal Mine?" (Brynjolfsson, Chandar & Chen, Aug/Nov 2025)

# 7.2 Yin, Vu & Persico (2026) — multi-model instability of LLM-rated exposure ("When the ruler is made of the thing it measures")

  • NBER WP 35110, "How (un)stable are LLM occupational exposure scores? Evidence from multi-model replication"; VoxEU column May 2026.
  • Replicated the Eloundou rubric with four frontier models on identical O*NET data: share of US occupations with >50% of tasks at high direct exposure = 2.7% (Gemini 2.5) … 3.8% (GPT-4) … 20.3% (GPT-5) … 51.5% (Claude 4.5) — a 19x spread. Management occupations: >80% high-exposure under Claude, <20% under Gemini. Downstream diff-in-diff employment estimates flip sign across raters. Bias is systematic per model, doesn't wash out with sample size, and co-evolves with the technology being measured (feedback channel).
  • Recommendation (directly applicable to us): any LLM-rated exposure analysis must report results from ≥2–3 different frontier models; convergence ⇒ robust, divergence ⇒ model artifact.
  • URL: https://cepr.org/voxeu/columns/when-ruler-made-thing-it-measures-multi-model-evidence-ai-occupational-exposure

# 7.3 "AI Exposure Scores: what they measure, what they miss, and what comes next" (Lund, Euyang, Munyikwa & Fadaee, arXiv June 2026)

  • Field review. Diagnoses a structural gap (static scores can't answer dynamic who/when/where policy questions) and a coordination gap (policy still cites static 2023 GPTs-are-GPTs numbers despite methodological advances). Surveys five successor families: dynamic/benchmark-based measures, ensembles, task-framework extensions, worker-centered metrics, adoption/usage data. Recommends moving "from prediction to preparedness".
  • URL: https://arxiv.org/abs/2606.23633

# 7.4 Other notable 2025–2026 entries

  • Iceberg Index (Chopra et al., MIT + Oak Ridge National Laboratory, arXiv 2510.25137, late 2025): skills-centered simulation — 151M US workers, 923 occupations, 32,000+ skills, ~3,000 counties; catalogued 13,000+ AI tools; agent-based simulation (AgentTorch on Frontier supercomputer). Visible "surface" tech-sector exposure = 2.2% of wage bill (~$211B); full skill-overlap exposure = 11.7% of US wage bill (~$1.2T). Explicitly technical exposure, not displacement. URLs: https://arxiv.org/abs/2510.25137 ; https://iceberg.mit.edu/report.pdf
  • Yale Budget Lab — "Evaluating the Impact of AI on the Labor Market: Current State of Affairs" (Gimbel, Kinder, Kendall & Lee, Oct 2025, updated): occupational-mix dissimilarity analysis; finds no broad acceleration in labor-market compositional change attributable to AI 33 months post-ChatGPT — important null-result counterweight to Canaries. URL: https://budgetlab.yale.edu/research/evaluating-impact-ai-labor-market-current-state-affairs
  • UK task-based GenAI exposure index (arXiv 2507.22748, 2025): novel LLM-scored task index applied to UK SOC codes — example of the national-adaptation pattern relevant to our ESCO/ROME crosswalk. URL: https://arxiv.org/abs/2507.22748
  • OAIES / capability-staged exposure (2025–2026): scores O*NET task automatable share at discrete AI capability stages (pre-LLM ML → early LLMs → multimodal → reasoning → agentic), multi-model rated (GPT-4o + Claude 3.5); cross-methodology Spearman ρ = 0.84 against independent scores. Overview: https://www.emergentmind.com/topics/ai-exposed-occupations ; theory-based variant (Moravec-paradox index): https://arxiv.org/abs/2510.13369
  • "AI and jobs: A review of theory, estimates, and evidence" (arXiv 2509.15265, 2025): comprehensive literature review; useful bibliography. URL: https://arxiv.org/abs/2509.15265
  • Agentic-AI exposure analyses (2026): e.g., "Agentic AI and Occupational Displacement" (arXiv 2604.00186) extends task exposure to autonomous multi-step agents across regions. URL: https://arxiv.org/abs/2604.00186
  • "The Jagged Global Economy" (arXiv 2607.05404, 2026): frontier-AI benchmark performance mapped to national economies — capability-grounded, benchmark-updated exposure. URL: https://arxiv.org/abs/2607.05404

# 8. Empirical validation: do exposure scores predict real outcomes?

Bottom line: individually, classic exposure scores are weak predictors; ensembles + adoption data + post-2022 windows perform much better. Realized effects so far are concentrated (entry-level, automation-tilted tasks, online freelancing), not economy-wide.

  1. Frank, Ahn & Moro (PNAS Nexus, April 2025), "AI exposure predicts unemployment risk" — built occupation-level unemployment risk from US unemployment-insurance claims (2010–2020); tested 10 exposure scores. Every individual score performs poorly (best single: Arntz automation probability, R² = 0.107; most < 3%; Frey–Osborne, SML, Felten, Webb all weak alone). An ensemble of all scores explains 29.8% (75.5% with education/skill/region controls) — +18 pp over baseline. Lesson: no single score suffices; combine dimensions. URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC11983276/ (arXiv:2308.02624)
  2. Acemoglu, Autor, Hazell & Restrepo (JOLE 2022), "AI and Jobs: Evidence from Online Vacancies" — AI-exposed establishments (Burning Glass) post more AI vacancies and reduce non-AI hiring, but no detectable aggregate occupation-level employment effects through 2018. URL: https://jadhazell.github.io/website/AI_And_Jobs.pdf
  3. Brynjolfsson, Chandar & Chen (2025) "Canaries" — first large-scale realized-effect finding: −13–16% relative employment for early-career workers in most-exposed occupations; effects load on automation-classified (AEI) usage, not augmentation. (§7.1)
  4. Hui, Reshef & Zhou (2024), "The Short-Term Effects of Generative AI on Online Labor Markets" — after ChatGPT, exposed freelancers (writing-heavy) on a large platform saw ~2% fewer jobs and ~5% lower earnings; top performers not spared. VoxEU: https://cepr.org/voxeu/columns/artificial-intelligence-and-its-short-term-effects-employment
  5. Hampole, Papanikolaou, Schmidt & Seegmiller (NBER w33509, 2025), "Artificial Intelligence and the Labor Market" — vacancy-based measure of firm AI adoption; AI adoption predicts declining demand for exposed occupations within adopting firms, with reallocation toward AI-complementary roles. URL: https://www.nber.org/papers/w33509
  6. Humlum & Vestergaard (2025), "Large Language Models, Small Labor Market Effects" (Denmark, NBER w33777) — despite rapid ChatGPT adoption among exposed workers, no detectable effects on earnings or hours in 2023–24 administrative data; average time savings ~3%. Counterweight showing adoption ≠ displacement in the short run. URL: https://www.nber.org/papers/w33777
  7. Wage-growth cross-section (2019 vs 2023, arXiv 2312.04714 & follow-ups): one-unit higher AI exposure ↔ −6.5 pp wage growth, explaining ~34% of cross-sectional variation post-ChatGPT; late-2022→early-2025 CPS analyses find high-exposure occupations losing 5.6–8.5 pp employment per 10-point exposure. URL: https://arxiv.org/abs/2312.04714
  8. Georgieff & Hyee (OECD 2021) / Employment Outlook 2023: 2012–2019 — no negative employment relationship; exposure without adoption predicts nothing pre-2022. (§2.3, §4.4)
  9. Yale Budget Lab (2025–2026): aggregate occupational mix shifting no faster than historical benchmarks (§7.4) — realized effects are cohort- and task-specific, not (yet) aggregate.

Synthesis for validation design: (a) validate at task/cohort level, not aggregate; (b) test against unemployment/UI-claim risk and entry-level hiring, not just employment stocks; (c) use ensembles; (d) treat pre-2022 null results as evidence about adoption gating, which is exactly what our barriers and adoption_velocity dimensions model.


# 9. Master comparison table

Index / Study Year Approach Unit of analysis Scale / output Substitution vs augmentation split? Validation status Data availability
Frey & Osborne 2013/2017 Expert labels + Gaussian process classifier on O*NET bottleneck variables Occupation (702 SOC) P(computerisation) 0–1 No Poor ex-post; explains ❤️% of unemployment risk alone Scores in paper appendix (public)
Arntz, Gregory & Zierahn (OECD) 2016/2017 F&O risk re-estimated on individual PIAAC task bundles Worker/task bundle P(automation) 0–1 No Best single predictor in Frank et al. ensemble (R²=0.107) PIAAC public; scores replicable
Brynjolfsson, Mitchell & Rock (SML) 2017/2018 23-question rubric, crowd-rated O*NET task (18k) → occupation SML 1–5 Implicit (redesign framing) Weak alone openICPSR replication archive
Webb 2020 Patent–task text overlap (verb–object pairs) Task → occupation Exposure percentile No Historical validation on robots/software; AI portion pre-LLM Author site / SSRN
Felten, Raj & Seamans (AIOE) 2021 (genAI 2023) 10 AI apps × 52 abilities, MTurk relatedness, importance-weighted Ability → occupation (+industry, county) Continuous z-score No (explicitly neutral) Weak alone; base of IMF/Pew analyses GitHub (AIOE-Data/AIOE)
Eloundou et al. "GPTs are GPTs" 2023/2024 Rubric (≥50% time saving), human + GPT-4 raters O*NET task/DWA → occupation E0/E1/E2; α, β, ζ shares 0–1 No (time-savings only) Predicts Canaries cohort effects; rater-unstable (19x across models) arXiv appendix; rubric public
Goldman Sachs (Briggs & Kodnani) 2023 Task-importance share judged automatable Occupation % tasks automatable Partial (25% substitution assumption) n/a Report only (proprietary)
McKinsey MGI 2023 Proprietary activity–capability model Work activity (~2,100) % of work time automatable; adoption scenarios Partial n/a Report only (proprietary)
Pew Research (Kochhar) 2023 Felten-style ability importance Occupation High/medium/low exposure No n/a Report + appendix
ILO Gmyrek et al. WP96 2023 GPT-4-rated ISCO task scores ISCO task → occupation → global employment Automation vs augmentation potential, 0–1 Yes n/a Scores in WP annexes
IMF Cazzaniga et al. (AIOE + C-AIOE) 2024 AIOE + complementarity shielding index Occupation → country employment Exposure × complementarity quadrants Yes (complementarity) n/a AIPI dashboard; WP data
ILO–NASK refined index (WP140) 2025 Hybrid: 52,558 human ratings calibrating GPT-4o, 4-gradient scale Task → ISCO occupation → 100+ countries 4 exposure gradients Yes (gradient) n/a WP + interactive tool
OECD AI Exposure Measure 2025/2026 Occupation capability requirements vs OECD AI Capability Indicators (9 domains, expert-set levels) Ability-domain → occupation Capability-overlap exposure; versioned as AI levels advance No New OECD publication + indicators
PwC AI Jobs Barometer 2024–2026 ~1B job ads; demand-side Vacancy/occupation/industry Growth, wage premium, skill-change rates No Is itself outcome data Annual reports
Anthropic Economic Index 2025–2026 (6 releases) Observed Claude usage classified to O*NET tasks (Clio); AUI; primitives Conversation → task → occupation, geo Usage shares; automation vs augmentation %; complexity/success Yes (measured, not predicted) Is itself adoption data; used in Canaries Hugging Face (open)
OpenAI / Chatterji et al. 2025 Observed ChatGPT usage classification Message → task category Usage shares by intent Partial (Asking/Doing/Expressing) Is itself adoption data NBER paper (aggregates)
Canaries in the Coal Mine (Stanford DEL) 2025 ADP payroll × exposure scores (realized effects) Worker-level panel Employment effects by age × exposure Uses AEI automation/augmentation Is the validation Dashboard public; ADP restricted
Iceberg Index (MIT/ORNL) 2025 Skill-level tool coverage + agent-based simulation 32k skills → 923 occupations → counties % of wage bill exposed ($) No New iceberg.mit.edu; arXiv
Yin, Vu & Persico (multi-model replication) 2026 Meta: Eloundou rubric × 4 frontier raters Task → occupation Rater-dispersion bounds n/a Meta-validation NBER WP 35110
Frank, Ahn & Moro (ensemble) 2025 Ensemble of 10 exposure scores vs UI-claims risk Occupation × state × month Unemployment-risk R² No Is the validation PNAS Nexus (open access)

# 10. Implications for airiskindex.io v1 methodology

Mapping the literature onto our five dimensions (weights from packages/scoring/src/weights.ts) and our exposure/substitution/augmentation triad.

# Cross-cutting lessons (apply to the whole index)

  1. Task-based is settled science (Arntz 2016 → everyone since). Our O*NET task-level scoring with importance/frequency weights is the correct v1 backbone. Keep occupation scores as derived, never primary.
  2. Never collapse to one number without sub-scores. Felten's "neutral exposure", ILO's automation/augmentation split, IMF's complementarity, and AEI's measured automation:augmentation ratio all show the field converging on our exposure/substitution/augmentation triad. This is a genuine differentiator — most indices still publish one headline number and get misquoted (Goldman's "300M jobs lost" problem). Our tone rule ("adaptation, not doom") is empirically supported: realized effects so far are cohort-specific task reallocation, not mass unemployment (Yale Budget Lab; Humlum & Vestergaard) — with real, measurable pain at the entry level (Canaries).
  3. LLM-as-evaluator is standard but fragile — multi-model rating is now table stakes. Yin et al. (2026): 19x spread in headline statistics across frontier raters on identical data; Claude-family raters score exposure highest of all models. Concrete requirements for apps/worker/src/raters/:
    • Rate every task with ≥2 (ideally 3) different frontier models (RATER_MODEL must become a list or we add RATER_MODEL_SECONDARY); store per-model scores; publish cross-model agreement per occupation.
    • Report confidence intervals derived from rater disagreement — this slots directly into our existing score_low/score/score_high schema.
    • Calibrate LLM ratings against a human-rated anchor set (ILO–NASK's 52,558-judgment survey design is the gold standard; our expert Delphi overrides in data/derived/expert_overrides/ serve this role — sample deliberately across the exposure spectrum, not just flagged disagreements).
    • Prompt-version everything (we already do) and re-rate on model change — because the instrument co-evolves with the phenomenon (the "ruler" problem).
  4. Anchor thresholds in explicit, documented rubrics. Eloundou's "≥50% time saving at equal quality" is the citable standard; SML's 23 questions show multi-criterion rubrics beat single judgments. Our prompts should decompose ratings into named criteria and require structured justifications (auditable per §6 of CLAUDE.md).
  5. Validate against outcomes, and say so publicly. Frank et al. (2025): single scores explain <11% of unemployment risk; ensembles ~30–75%. We should (a) benchmark our composite against the public AEI usage data, PwC vacancy signals, and the Canaries dashboard; (b) publish a docs/methodology/sensitivity/ correlation report against AIOE, GPTs-are-GPTs β, and ILO WP140 scores each release. Spearman ρ ≈ 0.84 between independent modern methodologies is the bar for "capturing the same signal".

# Dimension-by-dimension

Dimension (weight) Lessons from the literature
automatability (0.35) This is Eloundou's E1/E2 construct + SML's rubric. Use a multi-criterion, multi-model LLM rubric with a time-savings-at-quality threshold; keep 5-point scale (matches ILO gradient practice and our expert-review disagreement rule). Distinguish "LLM alone" (E1) from "LLM + tooling/agents" (E2) — capability-staged variants (OAIES; OECD capability levels) show staging by AI generation makes scores updateable rather than obsolete. Critically: rate automation vs augmentation potential separately per task (ILO 2023/2025) so the composite's sub-scores are computed, not asserted.
feasibility (0.20) Separating "conceivable" from "deployable now" is what F&O failed to do and what killed their forecast. Ground feasibility in observed evidence: AEI task-success rates (66–70% by complexity — Jan 2026 primitives), benchmark-linked measures (OECD AI Capability Indicators' 9 domains with attained levels; "Jagged Global Economy"), and Iceberg's tool-catalogue approach (is there an actual product performing this skill?). Feasibility should decay-adjust automatability: high automatability + low current success rate ⇒ wide CI, lower composite.
cost_ratio (0.15) Least developed dimension in the literature — a genuine gap we can own. Only Webb (implicitly, via wages), Goldman (25% substitution assumption), and AEI's $/hr wage-equivalent task values touch it. Use O*NET-linked BLS wages (as AEI does: $48–49/hr average task value) vs API/inference cost per task-equivalent; store as integer cents per our conventions. Note Acemoglu's caution ("so-so automation"): low cost ratio can drive adoption even with mediocre quality — interact with feasibility.
barriers (0.20) Directly validated by IMF's C-AIOE complementarity (physical presence, human contact, legal responsibility) — the reason judges are exposed but safe and telemarketers are not. Pre-2022 null results (OECD 2021; Acemoglu et al. 2022) prove barriers dominate short-run outcomes. Operationalize the IMF/Pizzinelli shielding factors at task level: regulation/licensing, liability, required physical presence, human-contact preference, data confidentiality. Humlum & Vestergaard (Denmark) show even adoption without workflow redesign yields ~3% time savings — organizational barriers belong here too.
adoption_velocity (0.10) The 2025–26 revolution: use measured adoption, not guesses. Sources: AEI Hugging Face releases (task-level usage shares, automation:augmentation ratio, AUI by geography — open data, quarterly cadence), PwC Jobs Barometer (vacancy-side skill change 66% faster in exposed occupations), OpenAI usage paper. Sector velocity is empirically uneven (API vs consumer concentration diverging; GDP-elasticity of adoption 0.7) — justify per-sector velocity scores with these citations. Design for time-series updates: AEI shows adoption shares move 10+ pp in a year (automation 27%→39%→45–52% oscillation), so this dimension must re-score every INDEX_VERSION.

# Positioning / product implications

  • Transparency is our moat and the field's known weakness: McKinsey/Goldman are black boxes; even academic scores rarely ship rater-level data. We publish weights (/api/v1/methodology), prompts (versioned), per-model ratings, and CIs — no major index does all four.
  • Versioning is becoming an explicit norm (OECD exposure measure designed to be updateable; AEI releases dated datasets). Our INDEX_VERSION + immutable score_runs architecture matches best practice; cite OECD/AEI precedent in METHODOLOGY.md.
  • CI bounds have empirical semantics now: rater disagreement (Yin et al.) + human-LLM calibration error (ILO–NASK) + feasibility uncertainty (AEI success rates) are the three quantifiable components of score_low/score_high.
  • EU/France (ESCO/ROME) crosswalk: ILO WP140 (ISCO-based) and the UK index (arXiv 2507.22748) are the reference patterns for adapting O*NET-trained scores to other taxonomies; document crosswalk loss explicitly.
  • Watch list for future versions: agentic-AI exposure extensions (arXiv 2604.00186), benchmark-grounded dynamic scores (OECD capability levels; "Jagged Global Economy"), worker-centered metrics (Lund et al. taxonomy), and the Canaries dashboard as a rolling validation target.

Compiled via WebSearch, Tavily, WebFetch, and OpenAlex queries, 2026-08-05. All URLs verified live at compile time unless noted. This document feeds docs/methodology/METHODOLOGY.md §Related Work.