LLM Index logo # LLM Index **The discriminative, contamination-resistant, fully transparent LLM ranking — updated live.** [**www.llmindex.io**](https://www.llmindex.io) · [Methodology](https://www.llmindex.io/methodology) · [White Paper (PDF)](docs/whitepaper/llmindex-whitepaper.pdf) · [Public API](https://www.llmindex.io/api/v1/leaderboard) [![Website](https://img.shields.io/website?url=https%3A%2F%2Fwww.llmindex.io&label=llmindex.io&up_color=10b981)](https://www.llmindex.io) ![Index](https://img.shields.io/badge/index-v0.2.0-10b981?logo=vercel&logoColor=white) ![Live](https://img.shields.io/badge/benchmark-LIVE%20%E2%80%A2%20updating-06b6d4) ![Models](https://img.shields.io/badge/models-135-8b5cf6) ![Domains](https://img.shields.io/badge/domains-12-f59e0b) ![Judges](https://img.shields.io/badge/judge%20panel-3%20cross--provider-ec4899) [![White paper](https://img.shields.io/badge/white%20paper-19%20pages%20·%20LaTeX-b91c1c?logo=latex&logoColor=white)](docs/whitepaper/llmindex-whitepaper.pdf) ![IRT](https://img.shields.io/badge/scoring-IRT%202PL%20%CE%B8-10b981) ![BT](https://img.shields.io/badge/duels-Bradley--Terry-0ea5e9) ![CI](https://img.shields.io/badge/every%20score-95%25%20CI-64748b) ![accuracy](https://img.shields.io/badge/accuracy__irt-0.60-047857) ![consistency](https://img.shields.io/badge/consistency-0.15-059669) ![contamination](https://img.shields.io/badge/contamination__resistance-0.15-0d9488) ![calibration](https://img.shields.io/badge/calibration-0.10-0891b2) ![pareto](https://img.shields.io/badge/cost%20%26%20latency-Pareto%20frontier%2C%20never%20blended-475569) ![Next.js](https://img.shields.io/badge/Next.js%2014-000000?logo=nextdotjs&logoColor=white) ![TypeScript](https://img.shields.io/badge/TypeScript%20strict-3178C6?logo=typescript&logoColor=white) ![Tailwind](https://img.shields.io/badge/Tailwind-06B6D4?logo=tailwindcss&logoColor=white) ![Prisma](https://img.shields.io/badge/Prisma-2D3748?logo=prisma&logoColor=white) ![PostgreSQL](https://img.shields.io/badge/PostgreSQL%2016-4169E1?logo=postgresql&logoColor=white) ![Redis](https://img.shields.io/badge/Redis-DC382D?logo=redis&logoColor=white) ![Python](https://img.shields.io/badge/Python%20%2B%20NumPy-3776AB?logo=python&logoColor=white) ![pnpm](https://img.shields.io/badge/pnpm%20%2B%20Turborepo-F69220?logo=pnpm&logoColor=white) ![OpenRouter](https://img.shields.io/badge/models%20via-OpenRouter-6467F2)
--- ## Why another index? Because leaderboards are broken. Classic leaderboards fail twice: **top models cluster at 95%+ on saturated benchmarks** (zero discrimination), and **fixed test sets leak into training data** (contamination). LLM Index is engineered against both, from the psychometrics up: | | | |---|---| | 🎯 **IRT 2PL scoring** | Every item has a fitted difficulty *b* and discrimination *a*; ability θ is a MAP estimate with Fisher-information standard errors. Items that don't separate models are auto-retired. Raw accuracy is never the score. | | 🎲 **Dynamic item generation** | Every scored batch is freshly generated from seeded template generators (values, paraphrases, structures). There is no fixed test set to memorize — the fixed-vs-fresh gap is *published* per model as `contamination_delta`. | | 🤖 **Home-made agentic bench** | Simulated tool-calling environments (triage-under-policy, treasury ledger, deployment DAGs) with distractor tools; a deterministic simulator computes the unique correct call sequence — graded by canonical-JSON equality, no judges. | | 🧠 **Agentic under context load** | 120–300-row generated ledgers packed with near-miss decoys: the model must find, order, and act on the few matching records. Prompt length is the difficulty knob. | | 💻 **Home-made terminal bench** | No shell executes: a closed, unambiguous POSIX subset is simulated in TypeScript. Models predict exact pipeline stdout, file trees after `mv/cp/rm/cd`, and `&&`/`\|\|` exit-code traces. | | 🎨 **SVG logo duels** | Models reproduce real-world logos in raw SVG from memory; a 3-judge cross-provider panel (position-swapped, never self-judging) feeds a Bradley-Terry fit. | | 👁 **Vision OCR under clutter** | Generated scenes (rotated codes, noise, low-contrast, decoys) rasterized to PNG; ambiguous glyphs excluded by design. Text-only models skip; weights renormalize. | | 🛡 **Extraction that never cheats models** | Lenient cascade (`ANSWER:` in any markdown, `FINAL ANSWER`, `\boxed{}`, fenced blocks), number normalization, 16k completion budgets, truncated ≠ wrong. Formatting is measured in its own domain — never silently everywhere. | | 📡 **Live, one model at a time** | Parallel evaluation lanes stream results; after every completed model the IRT refit re-runs and the public leaderboard re-ranks in real time. | | 🔍 **Total transparency** | Every model has a page showing **every answer on every test**, judge verdicts, confidence, latency and cost. Every number traces to an immutable score run. Answer keys never leave the server. | ## Screenshots
**Live leaderboard — desktop** LLM Index home, live leaderboard
**Mobile** Mobile view **Model transparency page** Model profile with every answer
**Methodology — public, detailed, with live sample items and full references** Methodology page
## The 12 domains ![code](https://img.shields.io/badge/code-trace%20%2B%20nested%20control%20flow-0ea5e9) ![math](https://img.shields.io/badge/math-chains%20%C2%B7%20counterfactual%20bases%20%C2%B7%20distractors-10b981) ![reasoning](https://img.shields.io/badge/reasoning-7--entity%20deduction%20%2B%20decoys-8b5cf6) ![agentic](https://img.shields.io/badge/agentic-simulated%20tool%20calling-f59e0b) ![terminal](https://img.shields.io/badge/terminal-simulated%20POSIX%20subset-64748b) ![knowledge](https://img.shields.io/badge/knowledge-free%20response%2C%20no%20guessing%20floor-06b6d4) ![multilingual](https://img.shields.io/badge/multilingual-0--999%20number%20words%20FR%2FES-ec4899) ![instruction](https://img.shields.io/badge/instruction__following-5%20stacked%20constraints-84cc16) ![writing](https://img.shields.io/badge/writing-judged%20duels%20%2B%20BT-a855f7) ![safety](https://img.shields.io/badge/safety__refusal__quality-gray--zone%20duels-ef4444) ![svg](https://img.shields.io/badge/svg__design-logo%20reproduction%20duels-f97316) ![vision](https://img.shields.io/badge/vision__ocr-clutter%20%2B%20grounded%20arithmetic-14b8a6) Domain weights are **equal by design** — the maximum-entropy prior; any other weighting is an editorial value judgment. Per-domain scores are always published so you can re-weight. Cost and latency are **never** blended into quality: they live on a separate Pareto frontier. ## Architecture ``` llmindex/ ├── apps/ │ ├── web/ # Next.js 14 — live leaderboard, transparency pages, /api/v1/* │ ├── worker/ # eval runner · duel runner · parallel benchmark orchestrator · refit │ └── psychometrics/ # Python: 2PL IRT (MAP + Fisher SE), Bradley-Terry (MM + SE), calibration ├── packages/ │ ├── scoring/ # domains, weights, INDEX_VERSION, pure score aggregation (no I/O) │ ├── items/ # 25 template generators + simulators + robust grading cascade │ ├── openrouter/ # typed client: retries, timeouts, cost tracking, multimodal │ └── db/ # Prisma: models, eval_items, model_responses, pairwise_duels, score_runs ├── docs/methodology/ # METHODOLOGY, CHANGELOG (semver), IRT_SPEC, JUDGE_PROTOCOL └── data/item-bank/ # public template manifest (instantiated items stay server-side) ``` **Pipeline:** OpenRouter catalog sync → seeded batch generation → parallel evaluation lanes (temperature 0 scored + k consistency samples, full audit trail: raw responses, tokens, latency, cost) → lenient extraction → 2PL fit per domain + Bradley-Terry for duels → scores with 95% CIs → leaderboard updates live → discrimination hygiene auto-flags dead items. ## Public API ```bash curl https://www.llmindex.io/api/v1/leaderboard # global ranking + CIs curl https://www.llmindex.io/api/v1/leaderboard/agentic # per-domain + sub-metrics curl https://www.llmindex.io/api/v1/models/anthropic/claude-sonnet-5 curl https://www.llmindex.io/api/v1/methodology # machine-readable weights & hyperparams curl https://www.llmindex.io/api/v1/benchmark/progress # live run status ``` Rate limits: 60 req/min anonymous, 600 with an API key. Breaking changes ship as `/api/v2` — v1 shapes are frozen. ## Run it ```bash pnpm install pnpm db:migrate && pnpm db:seed # sync models + pricing from OpenRouter pnpm dev # web on :3000 pnpm eval:run --model --domain math --n 30 --k 2 pnpm duel:run --domain svg_design --pairs 100 pnpm benchmark:run --parallel 6 # full live benchmark, refit after every model pnpm index:refit # IRT + BT refit → new immutable score run pnpm typecheck && pnpm lint && pnpm lint:headers && pnpm test # all gates ``` ## Integrity rules (non-negotiable) 1. Answer keys and rubrics never ship to the client or public API. 2. Scored batches use freshly perturbed items; the frozen anchor subset is ≤20% of any run. 3. Every displayed number traces to an immutable `score_runs` row (item-set hash, model set, index version, fit diagnostics). Batches with >2% failed calls are excluded until re-run. 4. Weights and IRT hyperparameters live in exactly one file and are served machine-readable. 5. Every methodology change bumps the semver `INDEX_VERSION` with a public changelog entry. The design draws on 60+ published sources (metabench, GSM-Symbolic, ZebraLogic, R-Horizon, τ-bench, BFCL, Terminal-Bench audits, Math-Verify, ReasonIF, …) — full list with links on the [methodology page](https://www.llmindex.io/methodology). Everything here is original: our own environments, items, simulators and graders. ---
**Author & maintainer:** Simon-Pierre Boucher · [contact@spboucher.ai](mailto:contact@spboucher.ai) License: Proprietary — © Simon-Pierre Boucher, all rights reserved. Source visible for transparency and auditability.