# Existing AI Job-Exposure / Automation-Risk Indices and Methodologies **Research memo — AI Risk Index (airiskindex.io)** **Date compiled:** 2026-08-05 (web research current to August 2026) **Scope:** All major academic and industry indices measuring occupational exposure to AI/automation, their methodologies, criticisms, empirical validation, and lessons for the airiskindex.io v1 scoring model. --- ## Table of Contents 1. [First wave: pre-generative-AI automation risk (2013–2019)](#1-first-wave) 2. [Second wave: AI-specific exposure measures (2018–2021)](#2-second-wave) 3. [Third wave: LLM/generative-AI exposure (2023–2024)](#3-third-wave) 4. [Institutional indices: ILO, OECD, IMF (2023–2026)](#4-institutional-indices) 5. [Industry & consultancy estimates](#5-industry-estimates) 6. [Usage-based measures: Anthropic Economic Index & OpenAI (2025–2026)](#6-usage-based-measures) 7. [Fourth wave: 2025–2026 indices and meta-critiques](#7-fourth-wave) 8. [Empirical validation: do exposure scores predict real outcomes?](#8-empirical-validation) 9. [Master comparison table](#9-comparison-table) 10. [Implications for airiskindex.io v1 methodology](#10-implications) --- ## 1. First wave: pre-generative-AI automation risk (2013–2019) ### 1.1 Frey & Osborne — "The Future of Employment" (2013 working paper; 2017 published) - **Authors/year:** Carl Benedikt Frey & Michael A. Osborne (Oxford Martin School). Working paper Sept 2013; published in *Technological Forecasting and Social Change* 114 (2017): 254–280. - **Unit of analysis:** Whole occupations (702 SOC occupations). - **Methodology:** - ML experts hand-labeled ~70 occupations as automatable (1) or not (0) at a workshop. - Identified three "engineering bottlenecks" to computerisation: **perception & manipulation**, **creative intelligence**, **social intelligence**, operationalized via 9 O*NET variables (e.g., finger dexterity, originality, social perceptiveness, persuasion, negotiation, assisting/caring for others, cramped work spaces). - A **Gaussian process classifier** trained on the 70 labels extrapolated a "probability of computerisation" (0–1) to all 702 occupations. - Occupations bucketed: high risk (p > 0.7), medium (0.3–0.7), low (p < 0.3). - **Key numbers:** **47% of US employment** at "high risk" of computerisation "over the next decade or two" (i.e., roughly by 2030). Transportation, logistics, office/administrative support, and production occupations most at risk. - **Criticisms (extensive):** - **Occupation-level, not task-level:** treats occupations as monolithic. Arntz et al. (2016) showed within-occupation task heterogeneity slashes the estimate to ~9%. - **Technical capability ≠ adoption:** no economics (cost, wages, regulation, preferences) in the model. - **Subjective training labels** from a small expert workshop; only ~70 seed labels drive all 702 predictions. - **Model-selection sensitivity:** "The Future of Employment Revisited" (arXiv:2104.13747) shows automation forecasts swing heavily with classifier choice on the same labels. - **Poor ex-post predictive record:** occupations flagged high-risk did not experience differential employment declines through the late 2010s (see §8; Frank, Ahn & Moro 2025 find F&O scores explain <3% of unemployment-risk variation individually). - **URLs:** - Published paper: https://doi.org/10.1016/j.techfore.2016.08.019 (PDF mirror: http://reparti.free.fr/freyosborne17.pdf) - Oxford Martin 2013 working paper: https://www.oxfordmartin.ox.ac.uk/downloads/academic/The_Future_of_Employment.pdf - Critique (Melbourne Institute): https://melbourneinstitute.unimelb.edu.au/__data/assets/pdf_file/0005/3197111/wp2019n10.pdf - Critique (model selection): https://arxiv.org/abs/2104.13747 ### 1.2 Arntz, Gregory & Zierahn — OECD task-based critique (2016) - **Authors/year:** Melanie Arntz, Terry Gregory, Ulrich Zierahn (ZEW/OECD). *"The Risk of Automation for Jobs in OECD Countries"*, OECD Social, Employment and Migration Working Paper No. 189 (2016); follow-up "Revisiting the risk of automation" in *Economics Letters* 159 (2017). - **Unit of analysis:** Individual workers' **task bundles** (PIAAC Survey of Adult Skills microdata), not occupation averages. - **Methodology:** Transferred Frey–Osborne occupation-level risk to individual workers, then re-estimated risk as a function of each worker's *actual reported tasks* (PIAAC), allowing within-occupation heterogeneity. High risk = automatability > 70%. - **Key numbers:** Only **~9% of jobs across 21 OECD countries** at high risk (US 9%, Germany 12%, Korea 6%) — versus 47% under F&O. - **Significance:** Founded the **task-based paradigm** that every serious index since has adopted — including airiskindex.io. Key insight: workers in the "same" occupation do different task mixes; scoring must start at the task level. - **Criticisms:** Still anchored on F&O's original subjective labels; PIAAC task self-reports are coarse; may *understate* risk if task bundles themselves adjust post-automation. - **URLs:** - OECD WP 189: https://doi.org/10.1787/5jlz9h56dvq7-en - Economics Letters 2017: https://doi.org/10.1016/j.econlet.2017.07.001 ### 1.3 Brynjolfsson, Mitchell & Rock — Suitability for Machine Learning (SML) (2017–2018) - **Authors/year:** Erik Brynjolfsson, Tom Mitchell, Daniel Rock. "What Can Machine Learning Do? Workforce Implications" (*Science*, 2017); "What Can Machines Learn, and What Does It Mean for Occupations and the Economy?" (*AEA Papers & Proceedings* 108, 2018: 43–47). - **Unit of analysis:** O*NET tasks (18,156 tasks; ~2,069 Detailed Work Activities), aggregated to ~950 occupations. - **Methodology:** - A **23-question rubric** capturing what (then-current, supervised) ML can do — e.g., mapping well-defined inputs to outputs, tolerance for error, no long chains of reasoning, digital data availability, no need for detailed physical manipulation. - Each task scored 1–5 per question; rubric validated by ML experts, then scaled via CrowdFlower crowd workers; aggregated to task SML then occupation SML (task-importance weighted). - **Key findings:** (1) ML affects *different* occupations than earlier automation waves; (2) most occupations have at least some high-SML tasks; (3) **almost no occupation is fully automatable**; (4) capturing value requires **task re-bundling / job redesign**. - **Criticisms:** Rubric was tuned to pre-LLM supervised ML (weak on generation, reasoning, dialogue); crowd ratings noisy; scores never strongly validated against outcomes (explains ~0–3% of unemployment-risk variation individually per Frank et al. 2025). - **Relevance to us:** Direct methodological ancestor of a **multi-criterion task rubric** — our `automatability` + `feasibility` split echoes SML's separation of "could ML do it" from "is it practical". - **URLs:** - AEA P&P: https://www.aeaweb.org/articles?id=10.1257/pandp.20181019 - SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3224100 - Replication data: https://www.openicpsr.org/openicpsr/project/114436 - WorldSML rubric (Stanford Digital Economy Lab, ongoing): https://digitaleconomy.stanford.edu/research/suitability-for-machine-learning-rubric-worldsml/ --- ## 2. Second wave: AI-specific exposure measures (2018–2021) ### 2.1 Webb (2020) — Patent-based exposure - **Author/year:** Michael Webb (Stanford). "The Impact of Artificial Intelligence on the Labor Market" (SSRN 3482150, Nov 2019/2020; still unpublished but heavily cited). - **Unit of analysis:** O*NET task descriptions × patent text; aggregated to occupations. - **Methodology:** - Selected AI patents (~16,400) by keyword; dependency-parsed titles to extract **verb–object pairs** (~8,000 pairs, e.g., "diagnose disease", "detect fraud"). - Extracted verb–object pairs from O*NET task statements; scored each task by the frequency with which its verb-object pairs appear in AI patents. - Occupation score = task-importance-weighted average; reported as **exposure percentiles**. - **Built-in validation strategy:** applied the same method to *robots* and *software* patents and showed those historical exposure measures predicted realized employment/wage declines in exposed occupations — then applied it to AI. - **Key findings:** AI (unlike robots/software) exposes **high-skilled, high-wage, older** workers most: e.g., clinical lab technicians, chemical engineers, optometrists, radiologic technicians. Robots hit low-skill physical work; software hit mid-skill routine work. - **Criticisms:** Patents lag and imperfectly reflect deployable capability (pre-LLM corpus, so misses generative AI entirely); verb-object matching is crude semantics; no distinction between substitution and augmentation. - **URLs:** - Paper: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3482150 / https://www.michaelwebb.co/webb_ai.pdf - Brookings explainer: https://www.brookings.edu/articles/how-patents-can-tell-us-what-jobs-ai-is-poised-to-disrupt/ ### 2.2 Felten, Raj & Seamans — AI Occupational Exposure (AIOE) (2018, 2021, 2023) - **Authors/year:** Edward Felten (Princeton), Manav Raj (Wharton), Robert Seamans (NYU). "Occupational, Industry, and Geographic Exposure to Artificial Intelligence: A Novel Dataset and Its Potential Uses", *Strategic Management Journal* 42(12), 2021: 2195–2217. Generative-AI update: "Occupational Heterogeneity in Exposure to Generative AI" (SSRN 4414065, April 2023). - **Unit of analysis:** 52 O*NET **abilities** (not tasks) × 10 AI application areas; aggregated to occupations (also industry AIIE and county-level geography). - **Methodology:** - 10 AI applications from the EFF AI Progress Measurement project (image recognition, language modeling, translation, speech recognition, abstract strategy games, etc.). - Amazon Mechanical Turk crowd workers rated **relatedness** of each application to each of 52 O*NET abilities → ability-level exposure = sum of relatedness scores. - Occupation AIOE = weighted sum of ability exposures using O*NET ability **importance and prevalence** weights. - 2023 generative-AI variant re-weights toward language modeling and image generation: top exposed = telemarketers, then post-secondary teachers (languages, history, law), sociologists, judges. - **Key properties:** Continuous z-scored index; famously **positively correlated with wages and education** (white-collar exposure) — opposite sign to Frey–Osborne. - **Criticisms:** Ability-level (even further from tasks than occupations); MTurk relatedness judgments are lay opinions; "exposure" deliberately **neutral between substitution and augmentation** (authors are explicit about this); static. - **Data:** Public GitHub (occupation-, industry-, geography-level scores): https://github.com/AIOE-Data/AIOE - **Adoption:** The **IMF** (Cazzaniga et al. 2024) and Pew (2023) analyses build directly on AIOE; the ECB European work (Albanesi et al.) uses AIOE + Webb. - **URLs:** - SMJ paper: https://doi.org/10.1002/smj.3286 - SSRN 2021: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3822412 - GenAI variant: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4414065 ### 2.3 Other second-wave measures (brief) - **Georgieff & Hyee (OECD 2021), "Artificial intelligence and employment: New cross-country evidence"** — applied Felten-style AIOE to PIAAC across 23 OECD countries; most-exposed: business professionals, managers, chief executives, science/engineering professionals. Found **no negative employment relationship 2012–2019**; in high-computer-use occupations, higher AI exposure correlated with *higher* employment growth. URL: https://doi.org/10.1787/c2c1d276-en - **Lassébie & Quintini (OECD 2022), "What skills and abilities can automation technologies replicate?"** — expert survey on automatability of ~100 O*NET skills/abilities; basis for OECD Employment Outlook 2023 statement that occupations at highest risk of automation account for **~27% of OECD employment**. URL: https://doi.org/10.1787/646aad77-en ; Employment Outlook 2023 AI chapter: https://www.oecd.org/en/publications/oecd-employment-outlook-2023_08785bba-en.html - **Tolan et al. (2021, JRC/EC)** — mapped AI research benchmarks to cognitive abilities to tasks; early "capability→ability→occupation" chain, precursor of the 2026 OECD capability-indicator approach. - **Startup-based exposure ("Follow the money", arXiv 2024)** — measures exposure via commercial AI startup activity mapped to occupations; argues patent/ability measures miss commercialization. URL: https://arxiv.org/abs/2412.04924 --- ## 3. Third wave: LLM/generative-AI exposure (2023–2024) ### 3.1 Eloundou, Manning, Mishkin & Rock — "GPTs are GPTs" (OpenAI/Wharton, 2023; *Science* 2024) **The single most influential template for airiskindex.io's LLM-as-evaluator pipeline.** - **Authors/year:** Tyna Eloundou, Sam Manning, Pamela Mishkin (OpenAI), Daniel Rock (Wharton). arXiv:2303.10130 (Mar 2023); published as "GPTs are GPTs: Labor market impact potential of LLMs", *Science* 384(6702), June 2024: 1306–1308. - **Unit of analysis:** O*NET task/DWA level (19,265 tasks; 2,087 DWAs), aggregated to 1,016 occupations with task weights; combined with BLS employment/wage data. - **Exposure rubric (the core innovation):** Exposure = "whether access to an LLM or LLM-powered system would reduce the time required for a human to perform a specific task **by at least 50% while maintaining quality**". Three levels: - **E0** — no exposure. - **E1** — direct exposure: LLM alone (via chat/API) achieves the 50% time reduction. - **E2** — LLM+ exposure: achievable only with additional software/tooling built on the LLM (image input, retrieval, agents, etc.). - Aggregates: **α = E1** (lower bound), **β = E1 + 0.5·E2** (expected), **ζ = E1 + E2** (upper bound). - **Raters:** Both human annotators (OpenAI staff, trained on rubric) and **GPT-4 itself as a rater** with the rubric as prompt; human–GPT-4 agreement was high (occupation-level correlations ≈ 0.80), pioneering the LLM-as-evaluator design we plan to use. - **Key numbers:** - ~**80% of US workers** have ≥10% of tasks exposed (β); **~19% of workers** have ≥50% of tasks exposed. - ~1.8% of jobs have >half their tasks E1-exposed; rises to **~46% of jobs** under ζ (with LLM-powered software). - Exposure **increases with wage and education** (up to a point); science and critical-thinking-intensive skills correlate negatively; programming and writing positively. - **Criticisms:** - Measures *potential time savings*, not substitution vs augmentation, adoption, or net employment effect (authors are explicit). - Rater instability: the flagship statistic is wildly model-dependent (see Yin et al. 2026, §7.2: 2.7%–51.5% across frontier raters). - Static snapshot of March-2023 GPT-4 capability; the "E2 software will exist" counterfactual is speculative. - 50%-time-saving threshold is arbitrary; binary-ish levels lose information. - **URLs:** - arXiv: https://arxiv.org/abs/2303.10130 - Science: https://www.science.org/doi/10.1126/science.adj0998 - OpenAI page: https://openai.com/index/gpts-are-gpts/ - Follow-up "Extending GPTs Are GPTs to Firms" (AEA P&P 2025): https://www.aeaweb.org/articles?id=10.1257/pandp.20251045 ### 3.2 Goldman Sachs — Briggs & Kodnani (March 2023) See §5.1. Methodologically an O*NET task-importance exercise inspired by Eloundou-style exposure; headline "300 million FTE jobs exposed" globally. --- ## 4. Institutional indices: ILO, OECD, IMF (2023–2026) ### 4.1 ILO — Gmyrek, Berg & Bescond (2023): "Generative AI and Jobs: A Global Analysis" - **Authors/year:** Paweł Gmyrek, Janine Berg, David Bescond. ILO Working Paper 96, August 2023. - **Methodology:** Scored **ISCO-08** occupation task lists (not O*NET) with **GPT-4 as rater** (multiple prompts, averaged), producing task-level automation-potential scores; distinguished **automation potential** vs **augmentation potential** at occupation level; mapped to global employment via ILO harmonized microdata for 100+ countries, by income group and sex. - **Key numbers:** - Only **clerical support work** is highly exposed as a group: **24% of clerical tasks highly exposed**, +58% medium exposure. Other occupational groups: 1–4% of tasks highly exposed. - Globally, ~**2.3% of employment (~75M jobs)** in the top automation-potential bucket; **13.4% (~427M)** in augmentation potential. - Exposure concentrated in **high/upper-middle-income countries** (more clerical employment) and **strongly gendered** (clerical work is female-dominated: in high-income countries, several times more female than male employment in the highest-exposure category). - **Framing:** "Augmentation, not automation, is the most likely impact" — the origin of the transformation-over-replacement institutional narrative. - **URLs:** - WP96 PDF: https://www.ilo.org/sites/default/files/2024-07/WP96_web.pdf - SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4584219 - Policy companion "Generative AI and Jobs: Policies to Manage the Transition": https://www.ilo.org/publications/generative-ai-and-jobs-policies-manage-transition ### 4.2 ILO 2024 interim work - **"Mind the AI Divide: Shaping a Global Perspective on the Future of Work"** (ILO & World Bank, Aug 2024) — applies the 2023 index with a digital-infrastructure overlay: poor countries are less *exposed* but also less able to *capture augmentation gains* (the "AI divide"). - **Gmyrek, Winkler & Garganta (2024, ILO/World Bank)** — "Buffer or bottleneck? Employment exposure to generative AI and the digital divide in Latin America": 26–38% of LAC jobs exposed; digital access gates both risk and benefit. - URL hub: https://www.ilo.org/publications (search "generative AI"); LAC paper: https://openknowledge.worldbank.org/handle/10986/41808 ### 4.3 ILO–NASK (2025): "Generative AI and Jobs: A Refined Global Index of Occupational Exposure" — current institutional state of the art - **Authors/year:** Paweł Gmyrek, Janine Berg, K. Kamiński, F. Konopczyński, A. Ładna, B. Nafradi, K. Rosłaniec, M. Troszyński (ILO + Poland's NASK). ILO Working Paper 140, May 2025 + Research Brief "Generative AI and jobs: a 2025 update". - **Methodology (major upgrade over 2023):** - **Hybrid human+LLM pipeline:** 52,558 human judgments on automation potential of 2,861 tasks (representative sample of 29,753 tasks in the Polish occupational classification), from a survey of 1,640 people (workers/experts), used to calibrate and validate GPT-4o task scoring; then scaled to the full ISCO task universe. - Replaced the binary automation/augmentation split with a **4-gradient exposure spectrum** (from marginal exposure to highest exposure), acknowledging most jobs are partially transformed. - Occupation scores mapped to global employment microdata by country income group, sex, and region. - **Key numbers:** - **1 in 4 jobs worldwide (25% of global employment)** has measurable GenAI exposure; **34% in high-income countries**. - Highest-gradient (transformation most likely) ≈ 3–4% of global employment, still concentrated in clerical work. - Gender gap persists: in high-income countries, ~**9.6% of female employment** vs ~3.5% of male employment in the top exposure gradient. - Headline framing: "**transformation, not replacement**". - **Criticisms:** LLM-rater dependence (flagged by Yin et al. 2026); Polish task-survey generalizability; exposure ≠ adoption in low-connectivity countries (self-acknowledged). - **URLs:** - WP140 PDF: https://www.ilo.org/sites/default/files/2025-05/WP140_web.pdf - WP140 interactive: https://webapps.ilo.org/static/english/intserv/working-papers/wp140/index.html - 2025 update brief: https://www.ilo.org/publications/generative-ai-and-jobs-2025-update (PDF: https://www.ilo.org/sites/default/files/2025-05/Research%20brief_GenAI%202025%20Update.pdf) - Press release: https://www.ilo.org/resource/news/one-four-jobs-risk-being-transformed-genai-new-ilo–nask-global-index-shows ### 4.4 OECD (2023–2026) - **Employment Outlook 2023 (AI chapters):** using Lassébie–Quintini expert-based measure, occupations at highest automation risk = **~27% of employment** across OECD; "no signs of slowing labour demand (yet)" in AI-exposed occupations. URLs: https://www.oecd.org/en/publications/oecd-employment-outlook-2023_08785bba-en/full-report/artificial-intelligence-and-jobs-no-signs-of-slowing-labour-demand-yet_5aebe670.html - **AI case studies & job quality (2023–2024):** firm case studies in finance/manufacturing across 8 countries: 23% of firms reported AI reduced employment in affected roles; wages mostly unchanged. Georgieff (2024), "Artificial intelligence and wage inequality": https://www.oecd.org/en/publications/artificial-intelligence-and-wage-inequality_bf98a45c-en.html - **OECD AI Capability Indicators (2025):** 5-year effort, 50+ experts; 9 capability domains (Language; Social interaction; Problem solving; Creativity; Metacognition & critical thinking; Knowledge/learning/memory; Vision; Manipulation; Robotic intelligence) each on an ordinal capability scale. URL: https://www.oecd.org/en/publications/introducing-the-oecd-ai-capability-indicators_be745f04-en.html - **The OECD AI Exposure Measure (2025/2026):** maps occupations' required capability *levels* in each of the 9 domains against AI's *current attained level* per the Capability Indicators → exposure = overlap. Explicitly designed to be **forward-looking, transparent, and updateable** as AI capability levels advance — the first institutional index architected for versioned re-scoring (same philosophy as our `INDEX_VERSION`). URL: https://www.oecd.org/en/publications/the-oecd-ai-exposure-measure_f3da0f0a-en.html - **Skills in the AI age (OECD AI Papers No. 60, July 2026):** applies the exposure measure to skills demand. URL: https://www.oecd.org/content/dam/oecd/en/publications/reports/2026/07/skills-in-the-ai-age_e8d8c1e6/972bd15e-en.pdf ### 4.5 IMF — Cazzaniga et al. (2024) and the AI Preparedness Index - **Authors/year:** Mauro Cazzaniga, Florence Jaumotte, Longji Li, Giovanni Melina, Augustus Panton, Carlo Pizzinelli, Emma Rockall, Marina M. Tavares. "Gen-AI: Artificial Intelligence and the Future of Work", IMF Staff Discussion Note SDN/2024/001, January 2024. - **Methodology:** Takes Felten's **AIOE** and adds a **potential complementarity index (C-AIOE)** (from Pizzinelli et al. 2023, IMF WP/23/216): occupations scored on shielding factors — required physical presence, human interaction, legal/social responsibility (judges are exposed *and* complemented; telemarketers exposed and *not*). Splits employment into: high exposure + high complementarity (augmentation likely) vs high exposure + low complementarity (displacement risk) vs low exposure. - **Key numbers:** **~40% of global employment exposed** to AI (**60% advanced economies, 40% emerging, 26% low-income**). In AEs, roughly half of exposed jobs are high-complementarity. Women and college-educated more exposed but better positioned for gains. - **AI Preparedness Index (AIPI):** country-level (174 economies) readiness across digital infrastructure, human capital & labor policies, innovation & integration, regulation & ethics — the macro complement to occupational exposure. Dashboard: https://www.imf.org/external/datamapper/AIPI@AIPI - **Criticisms:** Inherits all AIOE limitations; complementarity ratings are judgment calls; country mapping via ISCO crosswalks is coarse. - **URLs:** - SDN PDF: https://www.imf.org/-/media/files/publications/sdn/2024/english/sdnea2024001.pdf - eLibrary: https://www.elibrary.imf.org/view/journals/006/2024/001/006.2024.issue-001-en.xml - Pizzinelli et al. WP/23/216 (C-AIOE): https://www.imf.org/en/Publications/WP/Issues/2023/10/04/Labor-Market-Exposure-to-AI-Cross-country-Differences-and-Distributional-Implications-539656 - Follow-up: "Exposure to Artificial Intelligence and Occupational Mobility" (WP/24/116): https://www.imf.org/-/media/files/publications/wp/2024/english/wpiea2024116-print-pdf.pdf --- ## 5. Industry & consultancy estimates ### 5.1 Goldman Sachs (Briggs & Kodnani, March 2023) - "The Potentially Large Effects of Artificial Intelligence on Economic Growth". O*NET task-level judgment of automatable share per occupation (26 US, 24 European task categories importance/complexity weighted). - **Key numbers:** ~**2/3 of US/European occupations partially exposed**; generative AI could substitute up to **25% of current work** = **300M FTE jobs** globally exposed; +7% global GDP over 10 years. Most exposed: office/admin support (46% of tasks automatable), legal (44%), architecture/engineering (37%). - **Criticism:** binary "automatable share" judgments, no adoption model; the 300M number is routinely misquoted as "job losses". - URLs: https://www.goldmansachs.com/insights/articles/generative-ai-could-raise-global-gdp-by-7-percent ; follow-up US labor analysis: https://www.goldmansachs.com/insights/articles/how-will-ai-affect-the-us-labor-market ### 5.2 McKinsey Global Institute (June 2023, updated) - "The economic potential of generative AI: the next productivity frontier". Proprietary work-activity/capability model (~2,100 work activities, 850 occupations). - **Key numbers:** GenAI + existing tech could automate activities absorbing **60–70% of employees' time**; genAI value $2.6–4.4T/yr; **half of today's work activities automated between 2030 and 2060 (midpoint ~2045)** — pulled forward ~a decade vs pre-genAI estimate. - **Criticism:** proprietary/black-box capability ratings; "time automatable" ≠ jobs; adoption scenarios highly assumption-driven. - URL: https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier ### 5.3 Pew Research Center (Kochhar, July 2023) - "Which U.S. Workers Are More Exposed to AI on Their Jobs?" — Felten-style ability-importance approach on O*NET. - **Key numbers:** **19% of US workers in most-exposed jobs** vs 23% in least-exposed (2022). Most-exposed jobs pay *more* ($33/hr vs $20/hr); exposure higher for women, Asian, college-educated workers. Notably, workers in exposed industries did **not** feel their jobs at risk. - URL: https://www.pewresearch.org/social-trends/2023/07/26/which-u-s-workers-are-more-exposed-to-ai-on-their-jobs/ ### 5.4 PwC Global AI Jobs Barometer (2024, 2025, 2026) - Analyzes ~**1 billion job ads** worldwide + firm financials. 2025 edition ("The Fearless Future"): industries most exposed to AI saw productivity growth nearly **4x** (7%→27%); **56% wage premium** for AI-skilled workers (up from 25%); skills in AI-exposed occupations changing **66% faster**; employment *still growing* even in highly automatable roles. 2026 edition: labor market splitting into "two distinct paths", rewarding human skills. - **Value to us:** the best large-scale *demand-side* signal (vacancies), useful to calibrate `adoption_velocity`. - URLs: https://www.pwc.com/gx/en/issues/artificial-intelligence/job-barometer/2025/report.pdf ; 2026 PR: https://www.pwc.com/gx/en/news-room/press-releases/2026/pwc-2026-ai-jobs-barometer.html --- ## 6. Usage-based measures: Anthropic Economic Index & OpenAI (2025–2026) The decisive innovation of 2025–26: replacing *predicted* exposure with **observed AI usage** mapped to the same O*NET task taxonomy. This is the empirical anchor airiskindex.io should exploit for `adoption_velocity` and to validate `automatability`. ### 6.1 Anthropic Economic Index (AEI) — all releases to August 2026 **Methodology (constant across releases):** Clio, a privacy-preserving analysis pipeline in which Claude classifies large samples of real Claude.ai/API conversations against the **O*NET task taxonomy (~20,000 tasks)** and SOC occupations, plus interaction-mode classification (**automation** = full delegation/directive; **augmentation** = iterative collaboration, learning, validation). All aggregated data released openly on Hugging Face. | Release | Report | Data & model | Headline findings | |---|---|---|---| | **Feb 10, 2025** (paper arXiv:2503.04761, Handa et al., "Which Economic Tasks are Performed with AI?") | Launch report | 4M+ Claude.ai conversations | **36% of occupations** used AI for ≥25% of their tasks; only ~4% for ≥75%. Usage concentrated in **software development & writing** (computer/math ≈ 37% of conversations); peaks in **mid-to-high-wage** occupations, low at both wage extremes. **57% augmentation / 43% automation**. | | **Mar 27, 2025** | v2 (Claude 3.7 Sonnet) | New conversations + cluster-level data | Usage patterns stable; extended-thinking usage concentrated in technical tasks; released bottom-up task clusters. | | **Sep 15, 2025** (arXiv:2511.15080) | "Uneven geographic and enterprise adoption" | 1P API + geographic breakdowns; **Anthropic AI Usage Index (AUI)** = country share of usage ÷ share of working-age population | US 21.6% of usage; per-capita leaders Israel, Singapore, Australia, NZ, S. Korea. **+1% GDP/capita ↔ +0.7% AUI** (US states: 1.8% elasticity). DC highest state AUI (3.82). **Automation rose 27%→39%** of conversations since Dec 2024, surpassing augmentation for the first time; API usage even more automation-heavy. | | **Jan 15, 2026** | "New building blocks" (economic primitives) | 1M Claude.ai + 1M 1P API transcripts (Sonnet 4.5); Nov 2025 data | Five **primitives**: task complexity, human/AI skill level, use case, AI autonomy, task success. College-level tasks: **12x estimated speedup** but 66% success rate vs 70% for simpler tasks; revised aggregate productivity estimate **+1.2 pp/yr** (down from 1.8 after reliability adjustment). Augmentation back above automation on Claude.ai (52% vs 45%). | | **Mar 24, 2026** | "Learning curves" | Feb 2026 data (Opus 4.5/4.6) | Claude.ai task mix **de-concentrating** (top-10 tasks 24%→19%) while API concentrates (28%→33%). **49% of jobs in sample** now see Claude used for ≥25% of tasks (up from 36% in Jan 2025). 6-month+ tenure users: +10% conversation success; usage value ≈ $48–49/hr wage-equivalent tasks. | | **Apr 2026** | AEI **Survey** launched | 9,700 Claude users, linked usage+perceptions | See below. | | **Jun 26, 2026** | "Cadences" | Apr–Jun 2026, hourly sampling; artifact classifier | 93% of conversations produce artifacts (explanations 17%, documents/reports 15%). Higher-wage occupations' conversations consume 2.07x tokens. Survey: >⅓ of users expect AI to handle most of their work tasks within 12 months; only 10% rate own job loss likely; heavier automation users are *more* optimistic. Women use Claude less in automated modes (−0.33 SD). | - **Data availability (all releases):** https://huggingface.co/datasets/Anthropic/EconomicIndex (per-release folders `release_2025_03_27/`, `release_2025_09_15/`, etc., with documentation + replication notebooks; R package `aieconindex` on CRAN). - **Index hub:** https://www.anthropic.com/economic-index — reports: https://www.anthropic.com/news/the-anthropic-economic-index ; https://www.anthropic.com/research/economic-index-geography ; https://www.anthropic.com/research/economic-index-primitives ; https://www.anthropic.com/research/economic-index-march-2026-report ; https://www.anthropic.com/research/economic-index-june-2026-report - **Limitations (self-acknowledged):** Claude users ≠ workforce (selection bias toward developers/knowledge workers); conversation ≠ completed work; O*NET classification by LLM inherits classifier error; per-provider view only. ### 6.2 OpenAI — "How People Use ChatGPT" (Chatterji et al., NBER w34255, Sept 2025) - Aaron Chatterji, Tom Cunningham, David Deming, Zoë Hitzig, Christopher Ong, Carl Shan, Kevin Wadman. Privacy-preserving classification of a representative sample of ChatGPT consumer conversations, Nov 2022–Jul 2025 (~10% of world adult population using ChatGPT). - **Findings:** non-work usage grew from 53%→>70% of messages; work usage concentrated in **decision support** (advice, writing, information) rather than task execution; work usage highest among educated, high-paid professionals. Three-quarters of work messages: writing, information seeking, decision support. - **Relevance:** independent replication that *realized* usage is augmentation-tilted and knowledge-work-concentrated — cross-provider triangulation for `adoption_velocity`. - URLs: https://www.nber.org/papers/w34255 (PDF: https://www.nber.org/system/files/working_papers/w34255.pdf) --- ## 7. Fourth wave: 2025–2026 indices and meta-critiques ### 7.1 Stanford Digital Economy Lab — "Canaries in the Coal Mine?" (Brynjolfsson, Chandar & Chen, Aug/Nov 2025) - **The most important realized-effects paper to date.** Uses **ADP payroll microdata** (millions of workers, monthly) linked to occupational AI-exposure measures (Eloundou/GPTs-are-GPTs based, cross-checked with Anthropic Economic Index automation/augmentation shares). - **Six facts**, headline: since late 2022, **early-career workers (22–25) in the most AI-exposed occupations saw a ~13–16% relative employment decline** (16% in the Nov 2025 revision, controlling for firm-level shocks), while older workers in the same occupations and less-exposed young workers kept growing. Adjustment happens via **employment, not wages**. Declines concentrated where AEI data says AI **automates** rather than augments. Entry-level hiring is the "canary". - **Live monitoring:** "Canaries Dashboard" — https://digitaleconomy.stanford.edu/project/indicators/canaries-dashboard/ - **URLs:** paper page https://digitaleconomy.stanford.edu/publications/canaries-in-the-coal-mine ; PDF (Nov 2025) https://digitaleconomy.stanford.edu/app/uploads/2025/11/CanariesintheCoalMine_Nov25.pdf ; SIEPR WP: https://siepr.stanford.edu/publications/working-paper/canaries-coal-mine-six-facts-about-recent-employment-effects-artificial ### 7.2 Yin, Vu & Persico (2026) — multi-model instability of LLM-rated exposure ("When the ruler is made of the thing it measures") - NBER WP 35110, "How (un)stable are LLM occupational exposure scores? Evidence from multi-model replication"; VoxEU column May 2026. - Replicated the Eloundou rubric with **four frontier models on identical O*NET data**: share of US occupations with >50% of tasks at high direct exposure = **2.7% (Gemini 2.5) … 3.8% (GPT-4) … 20.3% (GPT-5) … 51.5% (Claude 4.5)** — a **19x spread**. Management occupations: >80% high-exposure under Claude, <20% under Gemini. Downstream diff-in-diff employment estimates **flip sign** across raters. Bias is systematic per model, doesn't wash out with sample size, and co-evolves with the technology being measured (feedback channel). - **Recommendation (directly applicable to us):** any LLM-rated exposure analysis must report results from **≥2–3 different frontier models**; convergence ⇒ robust, divergence ⇒ model artifact. - URL: https://cepr.org/voxeu/columns/when-ruler-made-thing-it-measures-multi-model-evidence-ai-occupational-exposure ### 7.3 "AI Exposure Scores: what they measure, what they miss, and what comes next" (Lund, Euyang, Munyikwa & Fadaee, arXiv June 2026) - Field review. Diagnoses a **structural gap** (static scores can't answer dynamic who/when/where policy questions) and a **coordination gap** (policy still cites static 2023 GPTs-are-GPTs numbers despite methodological advances). Surveys five successor families: **dynamic/benchmark-based measures, ensembles, task-framework extensions, worker-centered metrics, adoption/usage data**. Recommends moving "from prediction to preparedness". - URL: https://arxiv.org/abs/2606.23633 ### 7.4 Other notable 2025–2026 entries - **Iceberg Index (Chopra et al., MIT + Oak Ridge National Laboratory, arXiv 2510.25137, late 2025):** skills-centered simulation — 151M US workers, 923 occupations, 32,000+ skills, ~3,000 counties; catalogued 13,000+ AI tools; agent-based simulation (AgentTorch on Frontier supercomputer). Visible "surface" tech-sector exposure = 2.2% of wage bill (~$211B); full skill-overlap exposure = **11.7% of US wage bill (~$1.2T)**. Explicitly technical exposure, not displacement. URLs: https://arxiv.org/abs/2510.25137 ; https://iceberg.mit.edu/report.pdf - **Yale Budget Lab — "Evaluating the Impact of AI on the Labor Market: Current State of Affairs" (Gimbel, Kinder, Kendall & Lee, Oct 2025, updated):** occupational-mix dissimilarity analysis; finds **no broad acceleration** in labor-market compositional change attributable to AI 33 months post-ChatGPT — important null-result counterweight to Canaries. URL: https://budgetlab.yale.edu/research/evaluating-impact-ai-labor-market-current-state-affairs - **UK task-based GenAI exposure index (arXiv 2507.22748, 2025):** novel LLM-scored task index applied to UK SOC codes — example of the national-adaptation pattern relevant to our ESCO/ROME crosswalk. URL: https://arxiv.org/abs/2507.22748 - **OAIES / capability-staged exposure (2025–2026):** scores O*NET task automatable share at discrete **AI capability stages** (pre-LLM ML → early LLMs → multimodal → reasoning → agentic), multi-model rated (GPT-4o + Claude 3.5); cross-methodology Spearman ρ = 0.84 against independent scores. Overview: https://www.emergentmind.com/topics/ai-exposed-occupations ; theory-based variant (Moravec-paradox index): https://arxiv.org/abs/2510.13369 - **"AI and jobs: A review of theory, estimates, and evidence" (arXiv 2509.15265, 2025):** comprehensive literature review; useful bibliography. URL: https://arxiv.org/abs/2509.15265 - **Agentic-AI exposure analyses (2026):** e.g., "Agentic AI and Occupational Displacement" (arXiv 2604.00186) extends task exposure to autonomous multi-step agents across regions. URL: https://arxiv.org/abs/2604.00186 - **"The Jagged Global Economy" (arXiv 2607.05404, 2026):** frontier-AI benchmark performance mapped to national economies — capability-grounded, benchmark-updated exposure. URL: https://arxiv.org/abs/2607.05404 --- ## 8. Empirical validation: do exposure scores predict real outcomes? **Bottom line: individually, classic exposure scores are weak predictors; ensembles + adoption data + post-2022 windows perform much better. Realized effects so far are concentrated (entry-level, automation-tilted tasks, online freelancing), not economy-wide.** 1. **Frank, Ahn & Moro (PNAS Nexus, April 2025), "AI exposure predicts unemployment risk"** — built occupation-level *unemployment risk* from US unemployment-insurance claims (2010–2020); tested 10 exposure scores. **Every individual score performs poorly** (best single: Arntz automation probability, R² = 0.107; most < 3%; Frey–Osborne, SML, Felten, Webb all weak alone). An **ensemble of all scores** explains 29.8% (75.5% with education/skill/region controls) — +18 pp over baseline. Lesson: **no single score suffices; combine dimensions**. URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC11983276/ (arXiv:2308.02624) 2. **Acemoglu, Autor, Hazell & Restrepo (JOLE 2022), "AI and Jobs: Evidence from Online Vacancies"** — AI-exposed establishments (Burning Glass) post more AI vacancies and *reduce* non-AI hiring, but **no detectable aggregate occupation-level employment effects** through 2018. URL: https://jadhazell.github.io/website/AI_And_Jobs.pdf 3. **Brynjolfsson, Chandar & Chen (2025) "Canaries"** — first large-scale realized-effect finding: −13–16% relative employment for early-career workers in most-exposed occupations; effects load on **automation-classified** (AEI) usage, not augmentation. (§7.1) 4. **Hui, Reshef & Zhou (2024), "The Short-Term Effects of Generative AI on Online Labor Markets"** — after ChatGPT, exposed freelancers (writing-heavy) on a large platform saw ~2% fewer jobs and ~5% lower earnings; top performers not spared. VoxEU: https://cepr.org/voxeu/columns/artificial-intelligence-and-its-short-term-effects-employment 5. **Hampole, Papanikolaou, Schmidt & Seegmiller (NBER w33509, 2025), "Artificial Intelligence and the Labor Market"** — vacancy-based measure of firm AI adoption; AI adoption predicts declining demand for exposed occupations within adopting firms, with reallocation toward AI-complementary roles. URL: https://www.nber.org/papers/w33509 6. **Humlum & Vestergaard (2025), "Large Language Models, Small Labor Market Effects" (Denmark, NBER w33777)** — despite rapid ChatGPT adoption among exposed workers, **no detectable effects on earnings or hours** in 2023–24 administrative data; average time savings ~3%. Counterweight showing adoption ≠ displacement in the short run. URL: https://www.nber.org/papers/w33777 7. **Wage-growth cross-section (2019 vs 2023, arXiv 2312.04714 & follow-ups):** one-unit higher AI exposure ↔ **−6.5 pp wage growth**, explaining ~34% of cross-sectional variation post-ChatGPT; late-2022→early-2025 CPS analyses find high-exposure occupations losing 5.6–8.5 pp employment per 10-point exposure. URL: https://arxiv.org/abs/2312.04714 8. **Georgieff & Hyee (OECD 2021) / Employment Outlook 2023:** 2012–2019 — no negative employment relationship; **exposure without adoption predicts nothing** pre-2022. (§2.3, §4.4) 9. **Yale Budget Lab (2025–2026):** aggregate occupational mix shifting no faster than historical benchmarks (§7.4) — realized effects are **cohort- and task-specific, not (yet) aggregate**. **Synthesis for validation design:** (a) validate at task/cohort level, not aggregate; (b) test against unemployment/UI-claim risk and entry-level hiring, not just employment stocks; (c) use ensembles; (d) treat pre-2022 null results as evidence about *adoption gating*, which is exactly what our `barriers` and `adoption_velocity` dimensions model. --- ## 9. Master comparison table | Index / Study | Year | Approach | Unit of analysis | Scale / output | Substitution vs augmentation split? | Validation status | Data availability | |---|---|---|---|---|---|---|---| | **Frey & Osborne** | 2013/2017 | Expert labels + Gaussian process classifier on O*NET bottleneck variables | Occupation (702 SOC) | P(computerisation) 0–1 | No | Poor ex-post; explains <3% of unemployment risk alone | Scores in paper appendix (public) | | **Arntz, Gregory & Zierahn (OECD)** | 2016/2017 | F&O risk re-estimated on individual PIAAC task bundles | Worker/task bundle | P(automation) 0–1 | No | Best single predictor in Frank et al. ensemble (R²=0.107) | PIAAC public; scores replicable | | **Brynjolfsson, Mitchell & Rock (SML)** | 2017/2018 | 23-question rubric, crowd-rated | O*NET task (18k) → occupation | SML 1–5 | Implicit (redesign framing) | Weak alone | openICPSR replication archive | | **Webb** | 2020 | Patent–task text overlap (verb–object pairs) | Task → occupation | Exposure percentile | No | Historical validation on robots/software; AI portion pre-LLM | Author site / SSRN | | **Felten, Raj & Seamans (AIOE)** | 2021 (genAI 2023) | 10 AI apps × 52 abilities, MTurk relatedness, importance-weighted | Ability → occupation (+industry, county) | Continuous z-score | No (explicitly neutral) | Weak alone; base of IMF/Pew analyses | GitHub (AIOE-Data/AIOE) | | **Eloundou et al. "GPTs are GPTs"** | 2023/2024 | Rubric (≥50% time saving), human + GPT-4 raters | O*NET task/DWA → occupation | E0/E1/E2; α, β, ζ shares 0–1 | No (time-savings only) | Predicts Canaries cohort effects; rater-unstable (19x across models) | arXiv appendix; rubric public | | **Goldman Sachs (Briggs & Kodnani)** | 2023 | Task-importance share judged automatable | Occupation | % tasks automatable | Partial (25% substitution assumption) | n/a | Report only (proprietary) | | **McKinsey MGI** | 2023 | Proprietary activity–capability model | Work activity (~2,100) | % of work time automatable; adoption scenarios | Partial | n/a | Report only (proprietary) | | **Pew Research (Kochhar)** | 2023 | Felten-style ability importance | Occupation | High/medium/low exposure | No | n/a | Report + appendix | | **ILO Gmyrek et al. WP96** | 2023 | GPT-4-rated ISCO task scores | ISCO task → occupation → global employment | Automation vs augmentation potential, 0–1 | **Yes** | n/a | Scores in WP annexes | | **IMF Cazzaniga et al. (AIOE + C-AIOE)** | 2024 | AIOE + complementarity shielding index | Occupation → country employment | Exposure × complementarity quadrants | **Yes** (complementarity) | n/a | AIPI dashboard; WP data | | **ILO–NASK refined index (WP140)** | 2025 | Hybrid: 52,558 human ratings calibrating GPT-4o, 4-gradient scale | Task → ISCO occupation → 100+ countries | 4 exposure gradients | **Yes** (gradient) | n/a | WP + interactive tool | | **OECD AI Exposure Measure** | 2025/2026 | Occupation capability requirements vs OECD AI Capability Indicators (9 domains, expert-set levels) | Ability-domain → occupation | Capability-overlap exposure; versioned as AI levels advance | No | New | OECD publication + indicators | | **PwC AI Jobs Barometer** | 2024–2026 | ~1B job ads; demand-side | Vacancy/occupation/industry | Growth, wage premium, skill-change rates | No | Is itself outcome data | Annual reports | | **Anthropic Economic Index** | 2025–2026 (6 releases) | Observed Claude usage classified to O*NET tasks (Clio); AUI; primitives | Conversation → task → occupation, geo | Usage shares; automation vs augmentation %; complexity/success | **Yes** (measured, not predicted) | Is itself adoption data; used in Canaries | **Hugging Face (open)** | | **OpenAI / Chatterji et al.** | 2025 | Observed ChatGPT usage classification | Message → task category | Usage shares by intent | Partial (Asking/Doing/Expressing) | Is itself adoption data | NBER paper (aggregates) | | **Canaries in the Coal Mine (Stanford DEL)** | 2025 | ADP payroll × exposure scores (realized effects) | Worker-level panel | Employment effects by age × exposure | Uses AEI automation/augmentation | **Is the validation** | Dashboard public; ADP restricted | | **Iceberg Index (MIT/ORNL)** | 2025 | Skill-level tool coverage + agent-based simulation | 32k skills → 923 occupations → counties | % of wage bill exposed ($) | No | New | iceberg.mit.edu; arXiv | | **Yin, Vu & Persico (multi-model replication)** | 2026 | Meta: Eloundou rubric × 4 frontier raters | Task → occupation | Rater-dispersion bounds | n/a | Meta-validation | NBER WP 35110 | | **Frank, Ahn & Moro (ensemble)** | 2025 | Ensemble of 10 exposure scores vs UI-claims risk | Occupation × state × month | Unemployment-risk R² | No | **Is the validation** | PNAS Nexus (open access) | --- ## 10. Implications for airiskindex.io v1 methodology Mapping the literature onto our five dimensions (weights from `packages/scoring/src/weights.ts`) and our exposure/substitution/augmentation triad. ### Cross-cutting lessons (apply to the whole index) 1. **Task-based is settled science** (Arntz 2016 → everyone since). Our O*NET task-level scoring with importance/frequency weights is the correct v1 backbone. Keep occupation scores as *derived*, never primary. 2. **Never collapse to one number without sub-scores.** Felten's "neutral exposure", ILO's automation/augmentation split, IMF's complementarity, and AEI's measured automation:augmentation ratio all show the field converging on our exposure/substitution/augmentation triad. This is a genuine differentiator — most indices still publish one headline number and get misquoted (Goldman's "300M jobs lost" problem). Our tone rule ("adaptation, not doom") is empirically supported: realized effects so far are cohort-specific task reallocation, not mass unemployment (Yale Budget Lab; Humlum & Vestergaard) — with real, measurable pain at the entry level (Canaries). 3. **LLM-as-evaluator is standard but fragile — multi-model rating is now table stakes.** Yin et al. (2026): 19x spread in headline statistics across frontier raters on identical data; Claude-family raters score exposure *highest* of all models. Concrete requirements for `apps/worker/src/raters/`: - Rate every task with **≥2 (ideally 3) different frontier models** (`RATER_MODEL` must become a list or we add `RATER_MODEL_SECONDARY`); store per-model scores; publish cross-model agreement per occupation. - Report **confidence intervals derived from rater disagreement** — this slots directly into our existing `score_low`/`score`/`score_high` schema. - Calibrate LLM ratings against a **human-rated anchor set** (ILO–NASK's 52,558-judgment survey design is the gold standard; our expert Delphi overrides in `data/derived/expert_overrides/` serve this role — sample deliberately across the exposure spectrum, not just flagged disagreements). - Prompt-version everything (we already do) and re-rate on model change — because the instrument co-evolves with the phenomenon (the "ruler" problem). 4. **Anchor thresholds in explicit, documented rubrics.** Eloundou's "≥50% time saving at equal quality" is the citable standard; SML's 23 questions show multi-criterion rubrics beat single judgments. Our prompts should decompose ratings into named criteria and require structured justifications (auditable per §6 of CLAUDE.md). 5. **Validate against outcomes, and say so publicly.** Frank et al. (2025): single scores explain <11% of unemployment risk; ensembles ~30–75%. We should (a) benchmark our composite against the public AEI usage data, PwC vacancy signals, and the Canaries dashboard; (b) publish a `docs/methodology/sensitivity/` correlation report against AIOE, GPTs-are-GPTs β, and ILO WP140 scores each release. Spearman ρ ≈ 0.84 between independent modern methodologies is the bar for "capturing the same signal". ### Dimension-by-dimension | Dimension (weight) | Lessons from the literature | |---|---| | **`automatability` (0.35)** | This is Eloundou's E1/E2 construct + SML's rubric. Use a multi-criterion, multi-model LLM rubric with a time-savings-at-quality threshold; keep 5-point scale (matches ILO gradient practice and our expert-review disagreement rule). Distinguish "LLM alone" (E1) from "LLM + tooling/agents" (E2) — capability-staged variants (OAIES; OECD capability levels) show staging by AI generation makes scores updateable rather than obsolete. Critically: rate **automation vs augmentation potential separately per task** (ILO 2023/2025) so the composite's sub-scores are computed, not asserted. | | **`feasibility` (0.20)** | Separating "conceivable" from "deployable now" is what F&O failed to do and what killed their forecast. Ground feasibility in *observed evidence*: AEI task-success rates (66–70% by complexity — Jan 2026 primitives), benchmark-linked measures (OECD AI Capability Indicators' 9 domains with attained levels; "Jagged Global Economy"), and Iceberg's tool-catalogue approach (is there an actual product performing this skill?). Feasibility should decay-adjust automatability: high automatability + low current success rate ⇒ wide CI, lower composite. | | **`cost_ratio` (0.15)** | Least developed dimension in the literature — a genuine gap we can own. Only Webb (implicitly, via wages), Goldman (25% substitution assumption), and AEI's $/hr wage-equivalent task values touch it. Use O*NET-linked BLS wages (as AEI does: $48–49/hr average task value) vs API/inference cost per task-equivalent; store as integer cents per our conventions. Note Acemoglu's caution ("so-so automation"): low cost ratio can drive adoption even with mediocre quality — interact with feasibility. | | **`barriers` (0.20)** | Directly validated by IMF's C-AIOE complementarity (physical presence, human contact, legal responsibility) — the reason judges are exposed but safe and telemarketers are not. Pre-2022 null results (OECD 2021; Acemoglu et al. 2022) prove barriers dominate short-run outcomes. Operationalize the IMF/Pizzinelli shielding factors at task level: regulation/licensing, liability, required physical presence, human-contact preference, data confidentiality. Humlum & Vestergaard (Denmark) show even *adoption* without workflow redesign yields ~3% time savings — organizational barriers belong here too. | | **`adoption_velocity` (0.10)** | The 2025–26 revolution: use **measured** adoption, not guesses. Sources: AEI Hugging Face releases (task-level usage shares, automation:augmentation ratio, AUI by geography — open data, quarterly cadence), PwC Jobs Barometer (vacancy-side skill change 66% faster in exposed occupations), OpenAI usage paper. Sector velocity is empirically uneven (API vs consumer concentration diverging; GDP-elasticity of adoption 0.7) — justify per-sector velocity scores with these citations. Design for **time-series updates**: AEI shows adoption shares move 10+ pp in a year (automation 27%→39%→45–52% oscillation), so this dimension must re-score every `INDEX_VERSION`. | ### Positioning / product implications - **Transparency is our moat and the field's known weakness:** McKinsey/Goldman are black boxes; even academic scores rarely ship rater-level data. We publish weights (`/api/v1/methodology`), prompts (versioned), per-model ratings, and CIs — no major index does all four. - **Versioning is becoming an explicit norm** (OECD exposure measure designed to be updateable; AEI releases dated datasets). Our `INDEX_VERSION` + immutable `score_runs` architecture matches best practice; cite OECD/AEI precedent in METHODOLOGY.md. - **CI bounds have empirical semantics now:** rater disagreement (Yin et al.) + human-LLM calibration error (ILO–NASK) + feasibility uncertainty (AEI success rates) are the three quantifiable components of `score_low`/`score_high`. - **EU/France (ESCO/ROME) crosswalk:** ILO WP140 (ISCO-based) and the UK index (arXiv 2507.22748) are the reference patterns for adapting O*NET-trained scores to other taxonomies; document crosswalk loss explicitly. - **Watch list for future versions:** agentic-AI exposure extensions (arXiv 2604.00186), benchmark-grounded dynamic scores (OECD capability levels; "Jagged Global Economy"), worker-centered metrics (Lund et al. taxonomy), and the Canaries dashboard as a rolling validation target. --- *Compiled via WebSearch, Tavily, WebFetch, and OpenAlex queries, 2026-08-05. All URLs verified live at compile time unless noted. This document feeds `docs/methodology/METHODOLOGY.md` §Related Work.*