CLAUDE.md — AI Risk Index Platform (www.airiskindex.io)
This file provides guidance to Claude Code when working in this repository.
1. Project Overview
AI Risk Index (airiskindex.io) is a public platform that scores occupations on their exposure to AI-driven automation, using a task-based methodology (not whole-occupation scoring). Every score must be transparent, versioned, and reproducible.
Core product promise: "The most methodologically rigorous, fully transparent AI job-exposure index."
The three concepts (never confuse them)
- Exposure — AI is technically capable of performing the task
- Substitution — AI actually replaces the human performing it
- Augmentation — AI assists the human, increasing productivity
The composite index reports all three as separate sub-scores. Never collapse them into a single undifferentiated "risk" number in code, UI copy, or API responses without also exposing the sub-scores.
Scoring model (v1)
Composite score 0–100 per occupation, computed from task-level scores weighted by task importance/frequency (O*NET weights). Dimensions:
| Dimension | Key | Weight (v1) |
|---|---|---|
| Task automatability | automatability |
0.35 |
| Current technical feasibility | feasibility |
0.20 |
| Cost of substitution vs. wage | cost_ratio |
0.15 |
| Adoption barriers (regulation, liability, human-contact requirement) | barriers |
0.20 |
| Sector adoption velocity | adoption_velocity |
0.10 |
- Weights live in
packages/scoring/src/weights.ts— never hardcode weights anywhere else. - Every scoring release gets a semver version (
INDEX_VERSIONinpackages/scoring/src/version.ts) and a changelog entry indocs/methodology/CHANGELOG.md. - Scores are stored with confidence intervals (
score_low,score,score_high). UI must always be able to display uncertainty.
Data sources
- O*NET 30.x database dumps (occupations + tasks + importance ratings) →
data/raw/onet/— 30.3 as of May 2026, 31.0 expected late Aug 2026. ⚠️ 30.x renamed Technology Skills → Software Skills and split Skills into Essential/Transferable; ETL loaders must target the 30.x schema. License CC BY 4.0 (attribution required). - ESCO crosswalk for EU/France occupations (ROME codes) →
data/raw/esco/ - LLM-as-evaluator task ratings (see §6) →
data/derived/ratings/ - Expert panel (Delphi) validation overrides →
data/derived/expert_overrides/
Raw data is never edited in place. All transformations go through the ETL pipeline (apps/etl).
2. Tech Stack
- Monorepo: pnpm workspaces + Turborepo
- Language: TypeScript everywhere (strict mode). Python 3.12 only inside
apps/etlfor data science steps. - Frontend: Next.js 14 (App Router), React 18, Tailwind CSS, shadcn/ui, Recharts for visualizations
- API: Next.js route handlers for public API (
/api/v1/*) + tRPC for internal app calls - Database: PostgreSQL 16 (Prisma ORM). Read-heavy → materialized views for score lookups.
- Cache: Redis (score lookups, rate limiting)
- Jobs: BullMQ workers in
apps/worker(score recomputation, LLM rating batches) - Auth: Auth.js (email magic link + OAuth). Public browsing requires no auth.
- Testing: Vitest (unit), Playwright (e2e), pytest (ETL)
- Lint/format: ESLint + Prettier, ruff for Python. CI fails on warnings.
3. Repository Layout
airiskindex/
├── apps/
│ ├── web/ # Next.js app (site + public API routes)
│ ├── worker/ # BullMQ background workers
│ └── etl/ # Python data pipeline (raw → derived → DB)
├── packages/
│ ├── scoring/ # Pure TS scoring engine — NO I/O, fully deterministic
│ ├── db/ # Prisma schema + client + seeds
│ ├── ui/ # Shared React components
│ └── config/ # Shared eslint/ts/tailwind configs
├── data/
│ ├── raw/ # Immutable source dumps (gitignored, fetched by script)
│ └── derived/ # Pipeline outputs (versioned manifests committed)
├── docs/
│ └── methodology/ # Public methodology doc, changelog, sensitivity analyses
├── infra/ # Deployment scripts, ngrok config, systemd units
└── CLAUDE.md4. Commands
pnpm install # install all workspaces
pnpm dev # run web app on :3000 (+ worker with --filter)
pnpm build # turbo build all
pnpm test # vitest across packages
pnpm test:e2e # playwright (requires pnpm dev running)
pnpm lint && pnpm typecheck # must pass before any commit
pnpm db:migrate # prisma migrate dev
pnpm db:seed # seed occupations + demo scores
pnpm score:recompute # full index recomputation (writes new INDEX_VERSION run)
cd apps/etl && make pipeline # full ETL: raw → derived → DB loadSingle test file: pnpm vitest run packages/scoring/src/composite.test.ts
5. Coding Conventions
packages/scoringmust stay pure and deterministic: no network, no DB, noDate.now(), no randomness. It takes typed inputs and returns typed scores. This is what makes the methodology auditable.- Every scoring function has property-based tests (fast-check) + snapshot tests against the published methodology examples in
docs/methodology/examples/. - All user-facing risk language follows the tone guide: the product frames results as adaptation guidance, not doom. Avoid copy like "your job will disappear"; prefer "X% of tasks in this occupation are highly exposed".
- API responses are versioned (
/api/v1/...) and includeindex_versionin every payload. - Money/wages: store as integer cents + ISO currency. Percentages: store as 0–1 floats, format only at the UI layer.
- Never commit anything under
data/raw/. Derived data commits only the manifest JSON (hashes + row counts), not the payloads. - Migrations: additive-first. Destructive migrations require a
-- DESTRUCTIVEcomment and a manual approval in PR review. - Accessibility: all charts need a data-table fallback (
<VisuallyHidden>table) — this is a public-interest tool.
6. LLM-as-Evaluator (task rating pipeline)
Task automatability ratings are produced by an LLM rater in apps/worker/src/raters/, then validated by human experts.
- Use the Anthropic API via the official SDK. Model names come from the
RATER_MODELSenv var (comma-separated list — multi-model rating is required: single-model LLM exposure ratings show up to 19× spread across frontier raters; seedocs/research/01-existing-indices.md§7.2). Never hardcode model IDs in source. - Newer models (Sonnet 5 / Opus 5) reject the
temperatureparameter — rating variance comes from the multi-model panel, not sampling temperature. - Prompts live in
apps/worker/src/raters/prompts/*.mdand are versioned; a prompt change bumpsRATER_PROMPT_VERSIONand invalidates cached ratings. - Every rating stores: model, prompt version, raw response, parsed score, timestamp. Full audit trail, always.
- Ratings are sampled (5%) for human review; disagreement > 1 point on the 5-point scale flags the task for the expert panel queue.
- Batch jobs must be idempotent and resumable (BullMQ job IDs = deterministic hash of task_id + prompt version).
- For current API details (batch endpoints, rate limits, model names), check https://platform.claude.com/docs (docs.claude.com redirects there) rather than relying on memory. A vetted snapshot lives in
docs/research/03-llm-rater-api.md.
7. Environments & Deployment
Environments
local— developer machine, SQLite optional shortcut is not allowed; always Postgres via Docker (infra/docker-compose.dev.yml).staging— node m3u96b, exposed via ngrok (see below).production— airiskindex.io (target: VPS/managed later; staging setup is the current deployment).
Deploying to node m3u96b (staging/current prod)
The app currently runs on the self-hosted node m3u96b and is exposed publicly through an ngrok tunnel mapped to www.airiskindex.io.
Process layout on m3u96b (macOS node — PM2, not systemd; deployed 2026-08-05):
- App dir:
~/apps/airiskindex(deployed by rsync from the laptop — no git remote yet) airiskindex-web(PM2) →./run-web.sh→next start -p 3000(cwdapps/web)airiskindex-ngrok(PM2) →./run-ngrok.sh→ngrok http 3000 --url=https://www.airiskindex.iopostgresql@16+redisviabrew services(NOT Docker on this node)- Worker (
airiskindex-worker) not started yet — starting it submits a live Anthropic rating batch .envat repo root (chmod 600), symlinked intoapps/web/.env,packages/db/.env,apps/worker/.env(Next.js and Prisma each load env from their own directory in a monorepo)- Registered in
~/Desktop/cluster-skill/cluster-deployments.json
Deploy procedure:
# from the laptop
rsync -az --delete --exclude node_modules --exclude .next --exclude .turbo \
--exclude .env --exclude "data/raw/*" ~/Desktop/airiskindex/ M3U96b:apps/airiskindex/
ssh M3U96b
cd ~/apps/airiskindex
pnpm install --no-frozen-lockfile
pnpm build
set -a && . ./.env && set +a && pnpm db:migrate:deploy
pm2 restart airiskindex-web && pm2 save
curl -fsS http://127.0.0.1:3000/api/v1/health # must return {"ok":true,...}
curl -fsS https://www.airiskindex.io/api/v1/health # ALWAYS check the public URL toongrok configuration (infra/ngrok.yml, copied to /etc/ngrok/ngrok.yml on the node):
version: 3
agent:
authtoken: ${NGROK_AUTHTOKEN} # from env, never committed
endpoints:
- name: airiskindex
url: https://www.airiskindex.io # requires custom domain configured in ngrok dashboard + CNAME
upstream:
url: 3000- The custom domain must be added in the ngrok dashboard and DNS
CNAMEforwww.airiskindex.iopointed at the ngrok edge target they provide. Verify withdig CNAME www.airiskindex.io. - ngrok runs as
ngrok.service(systemd,Restart=always). Logs:journalctl -u ngrok -f. - Health rule: after every deploy, hit the public URL, not just localhost — tunnel failures are the most common outage cause.
- Because the app sits behind ngrok, trust
X-Forwarded-*headers: Next.js config already sets trusted proxy handling; do not remove it. Rate limiting keys offx-forwarded-forfirst IP.
Environment variables (/srv/airiskindex/.env, template in infra/.env.example):
DATABASE_URL=postgresql://...
REDIS_URL=redis://...
ANTHROPIC_API_KEY=... # rater pipeline
RATER_MODELS=... # comma-separated model IDs (multi-model rating panel)
NGROK_AUTHTOKEN=...
NEXTAUTH_URL=https://www.airiskindex.io
NEXTAUTH_SECRET=...
PUBLIC_BASE_URL=https://www.airiskindex.ioNever print, log, or commit secrets. infra/.env.example lists keys with empty values only.
Rollback
ssh m3u96b
cd /srv/airiskindex
git checkout <previous-tag>
pnpm install --frozen-lockfile && pnpm build
sudo systemctl restart airiskindex-web airiskindex-workerDB rollbacks: only via forward-fix migrations. Nightly pg_dump to /srv/backups (retained 14 days) — verify the cron is alive when touching infra.
8. Public API rules
/api/v1/occupations— list/search (paginated, cached 1h)/api/v1/occupations/:code— full score breakdown incl. sub-scores, CI bounds, task list,index_version/api/v1/methodology— machine-readable weights + version metadata- Rate limit: 60 req/min unauthenticated, 600 with API key. Return
429withRetry-After. - Breaking changes require a new
/api/v2— never mutate v1 response shapes.
9. Methodology Integrity Rules (non-negotiable)
- Any change to weights, formulas, or rater prompts ⇒ bump
INDEX_VERSION, add changelog entry, regeneratedocs/methodology/sensitivity/outputs. - Scores shown anywhere must be traceable to a stored computation run (
score_runstable) — no ad-hoc numbers. - Historical scores are immutable; recomputations create new runs, old runs remain queryable.
- The public methodology doc (
docs/methodology/METHODOLOGY.md) is the source of truth; code comments link to its section anchors. If code and doc disagree, stop and flag it — do not silently pick one.
10. When Working in This Repo, Claude Should
- Run
pnpm typecheck && pnpm lint && pnpm testbefore declaring any task done. - Touch
packages/scoringonly with accompanying tests and a methodology changelog note. - Ask before: destructive migrations, changing index weights, editing ngrok/systemd config, or anything that alters public API shapes.
- Prefer small, reviewable commits with conventional-commit messages (
feat(scoring): ...,fix(api): ...,infra(deploy): ...). - Keep UI copy aligned with the "adaptation, not doom" tone guide.
- When uncertain about Anthropic API specifics (models, batch API, limits), consult the official docs instead of guessing.