# Benchmarks reported by the leaderboard connectors (SWE-bench boards, Artificial Analysis component evaluations). Fragment of registry/benchmarks.yaml # (same canonical fields: family / variant / version / family_head / metric / metric_min / metric_max / higher_is_better / harness / comparability_note). benchmarks: - key: swe-bench-full name: SWE-bench (full test split) aliases: [SWE-bench Test, SWE-bench Full, SWE-bench (full), swebench full] family: swe-bench variant: full family_head: true category: coding task: resolve real GitHub issues (2,294 instances) metric: resolved metric_label: "% resolved" unit: "%" metric_min: 0 metric_max: 100 higher_is_better: true harness: submitter's agent scaffold (config.system); evaluation by the SWE-bench harness comparability_note: Scaffold-dependent; compare within one scaffold and attempt regime only. website: https://www.swebench.com paper: https://arxiv.org/abs/2310.06770 source_url: https://www.swebench.com/ - key: swe-bench-lite name: SWE-bench Lite aliases: [SWE-Bench Lite, swebench lite] family: swe-bench variant: Lite category: coding task: resolve real GitHub issues (300-instance subset) metric: resolved metric_label: "% resolved" unit: "%" metric_min: 0 metric_max: 100 higher_is_better: true harness: submitter's agent scaffold (config.system); evaluation by the SWE-bench harness comparability_note: Scaffold-dependent; Lite is an easier subset — never compare with Verified or full scores. website: https://www.swebench.com paper: https://arxiv.org/abs/2310.06770 known_limitations: Scaffold/agent dependent; results are not comparable across harnesses. source_url: https://www.swebench.com/ - key: swe-bench-multimodal name: SWE-bench Multimodal aliases: [SWE-Bench Multimodal, SWE-bench MM] family: swe-bench variant: Multimodal category: coding task: resolve visual JavaScript issues (screenshots + code) metric: resolved metric_label: "% resolved" unit: "%" metric_min: 0 metric_max: 100 higher_is_better: true harness: submitter's agent scaffold (config.system); evaluation by the SWE-bench harness comparability_note: Requires image input; scaffold-dependent. website: https://www.swebench.com/multimodal source_url: https://www.swebench.com/ - key: swe-bench-multilingual name: SWE-bench Multilingual aliases: [SWE-Bench Multilingual] family: swe-bench variant: Multilingual category: coding task: resolve GitHub issues across 9 programming languages metric: resolved metric_label: "% resolved" unit: "%" metric_min: 0 metric_max: 100 higher_is_better: true harness: submitter's agent scaffold (config.system); evaluation by the SWE-bench harness comparability_note: Scaffold-dependent; different instance set from Verified. website: https://www.swebench.com/multilingual source_url: https://www.swebench.com/ - key: swe-bench-pro name: SWE-Bench Pro aliases: [SWE-bench Pro, SWE Bench Pro, swebench pro] family: swe-bench variant: Pro category: coding task: long-horizon, enterprise-grade software engineering tasks (public set) curated by Scale AI metric: resolved metric_label: "% resolved" unit: "%" metric_min: 0 metric_max: 100 higher_is_better: true harness: Scale AI evaluation with a fixed scaffold per leaderboard column comparability_note: Public vs commercial (held-out) sets are different leaderboards; not comparable with SWE-bench Verified. website: https://scale.com/leaderboard/swe_bench_pro_public source_url: https://scale.com/leaderboard/swe_bench_pro_public - key: mmmu-pro name: MMMU-Pro aliases: [MMMU Pro] family: mmmu variant: Pro category: multimodal task: robust multimodal understanding (10-option, vision-only variants) metric: accuracy metric_label: accuracy unit: "%" metric_min: 0 metric_max: 100 higher_is_better: true harness: Artificial Analysis runs its own evaluation; labs self-report comparability_note: Standard (10-option) vs vision-only settings are different numbers. paper: https://arxiv.org/abs/2409.02813 source_url: https://arxiv.org/abs/2409.02813 - key: livecodebench name: LiveCodeBench aliases: [LCB, Live Code Bench] family: livecodebench variant: rolling family_head: true category: coding task: contamination-free competitive programming problems metric: pass@1 metric_label: pass@1 unit: "%" metric_min: 0 metric_max: 100 higher_is_better: true harness: Artificial Analysis runs its own evaluation; labs self-report on a chosen date window comparability_note: The problem window (release dates) differs between reporters — scores from different windows are not comparable. website: https://livecodebench.github.io paper: https://arxiv.org/abs/2403.07974 source_url: https://arxiv.org/abs/2403.07974 - key: scicode name: SciCode aliases: [SciCode benchmark] family: scicode variant: main family_head: true category: coding task: research-level scientific coding problems metric: accuracy metric_label: accuracy unit: "%" metric_min: 0 metric_max: 100 higher_is_better: true harness: Artificial Analysis runs its own evaluation comparability_note: Sub-problem vs main-problem accuracy are different numbers; with/without background prompts. website: https://scicode-bench.github.io paper: https://arxiv.org/abs/2407.13168 source_url: https://arxiv.org/abs/2407.13168 - key: ifbench name: IFBench aliases: [IF-Bench] family: ifeval variant: IFBench category: instruction-following task: precise instruction following with novel constraints metric: accuracy metric_label: accuracy unit: "%" metric_min: 0 metric_max: 100 higher_is_better: true harness: Artificial Analysis runs its own evaluation comparability_note: Different constraint set from IFEval; not comparable with IFEval scores. paper: https://arxiv.org/abs/2507.02833 source_url: https://arxiv.org/abs/2507.02833 - key: tau2-bench name: τ²-bench aliases: [tau2-bench, TAU2-bench, tau-squared bench, τ²-Bench, τ2-bench, τ²-Bench Telecom, tau2 bench telecom] family: tau-bench variant: τ² category: agentic task: dual-control tool-agent-user interaction (telecom, retail, airline) metric: pass^1 metric_label: pass^1 unit: "%" metric_min: 0 metric_max: 100 higher_is_better: true harness: simulated user (LLM) + dual-control tool environment; Artificial Analysis reports the Telecom domain comparability_note: Domain (config.variant = Telecom / Retail / Airline) and the user-simulator model must match. paper: https://arxiv.org/abs/2506.07982 source_url: https://arxiv.org/abs/2506.07982