SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
63.6 KB

# Models — OpenAI ↔ Anthropic ↔ xAI ↔ Gemini equivalence and use-case table

Status: every value is copied from generated/models.json (statuses, context, output, cutoffs, capability flags, pricing blocks; live-verified on 2026-09-18/19 where marked), generated/pricing.json and generated/deprecations.json. Equivalence rows are use-case pairings by price/tier and capability, not benchmark claims. xAI probes ran on a Tier 0 team, Gemini probes on a free-tier key (paid-only Gemini models are ACCOUNT_RESTRICTED, not absent). Sources: https://developers.openai.com/api/docs/models · https://developers.openai.com/api/docs/pricing · https://developers.openai.com/api/docs/deprecations · https://platform.claude.com/docs/en/models/overview · https://platform.claude.com/docs/en/about-claude/pricing · https://platform.claude.com/docs/en/about-claude/model-deprecations · https://docs.x.ai/developers/models · https://docs.x.ai/developers/pricing · https://docs.x.ai/developers/migration/may-15-retirement · https://ai.google.dev/gemini-api/docs/models · https://ai.google.dev/gemini-api/docs/pricing · https://ai.google.dev/gemini-api/docs/deprecations · docs/models/openai-models.md · docs/models/anthropic-models.md · docs/models/xai-models.md · docs/models/gemini-models.md Last verified: 2026-09-18

# 1. Use-case pairing (current models, four providers)

Prices USD per 1M tokens, standard tier: in / cached-in / out. "Cached-in" is a cache read. Context = documented window; Out = documented max output. xAI prices double for requests whose prompt is ≥200k tokens (all tokens of that request); Gemini Pro prices rise for prompts >200k; Gemini 3.6–3.8 Flash prices are introductory (double on 2027-01-01).

Use case OpenAI Anthropic xAI Gemini Why they pair
Frontier reasoning, hardest tasks gpt-6-astra — 1,050k ctx / 128k out · text+image → text · effort low…max (no none) · $10 / $1 / $50 · cutoff 2026-04-30 · rel. 2026-09-03 · LIVE_VERIFIED claude-fable-5-1 — 1M / 128k · text+image+PDF → text · adaptive thinking always on, effort low…max · $10 / $0.25 / $50 · cutoff Jun 2026 · rel. 2026-09-01 · LIVE_VERIFIED grok-4.6 — 500k / no published limit · text+image → text · reasoning always on, effort low…xhigh (default high) · $2 / $0.50 / $6 (≥200k: $4 / $1 / $12) · cutoff 2026-02-01 · rel. 2026-08 · LIVE_VERIFIED gemini-3.1-pro-preview — 1,048,576 / 65,536 · text+image+audio+video+PDF → text · thinkingLevel low/medium/high (default high, no minimal) · $2 / $0.20 / $12 (>200k: $4 / $0.40 / $18) · cutoff Jan 2025 (model card) · rel. 2026-02-19 · PREVIEW · ACCOUNT_RESTRICTED (paid tier only) Each provider's top reasoning line. Price spread is 5× ($2 vs $10 in). Astra and Fable are always-on reasoners; Grok 4.6 too; Gemini 3.1 Pro cannot go below low. Only Gemini's is still a preview.
Flagship general / agentic gpt-5.6-sol (alias gpt-5.6) — 1,050k / 128k · effort none…max, reasoning.mode: pro · $4 / $0.40 / $20 (promo ≥ 2026-11-21) · cutoff 2026-02-16 · rel. 2026-07-09 claude-opus-5 — 1M / 128k · adaptive (on by default, can disable except xhigh/max) · $5 / $0.50 / $25 · cutoff May 2026 · rel. 2026-07-24 · fast mode $10 / $50 grok-4.5 — 500k · reasoning always on, effort low…high (xhigh treated as high) · $2 / $0.30 / $6 · rel. 2026-07 · priority 2× · no batch · aliases grok-4.5-latest, grok-build-latest gemini-3.8-flash — 1,048,576 / 65,536 · thinking low/medium/high (default medium; minimal → 400) · $0.75 / $0.075 / $3.75 (→ $1.50 / $0.15 / $7.50 in 2027) · rel. 2026-09-02 · GA · LIVE_VERIFIED · free tier Newest general-purpose line per provider. Gemini's Flash line is the current flagship for the Developer API (Pro is preview); Grok 4.5/4.6 sit between OpenAI's mini and flagship prices.
Balanced mid-tier gpt-5.6-terra — 1,050k / 128k · $2 / $0.20 / $12; also gpt-5.5 $5 / $0.50 / $30, gpt-5.4 $2.5 / $0.25 / $15 claude-sonnet-5 — 1M / 128k · adaptive on by default · $2 / $0.20 / $10 · cutoff Jan 2026 · rel. 2026-06-30 grok-4.3 — 1M · reasoning optional (reasoning_effort: none LIVE_DISCOVERED → 0 reasoning tokens; default low) · $1.25 / $0.20 / $2.50 (≥200k ×2) · batch −20 % · rel. 2026-04 · redirect target of grok-3, grok-4-0709, grok-4-fast-* gemini-3.5-flash — 1,048,576 / 65,536 · thinking minimal…high (default medium; budget 0 disables) · $1.50 / $0.15 / $9 · rel. 2026-05-19 · GA · LIVE_VERIFIED Identical $2 input on Terra / Sonnet 5 / 3.1 Pro; grok-4.3 and gemini-3.5-flash undercut them. All support the full tool set of their provider (§3–§5).
Small / high-volume gpt-5.6-luna — $0.20 / $0.02 / $1.20; gpt-5.4-mini — 400k / 128k · $0.75 / $0.075 / $4.50; gpt-5.4-nano — $0.20 / $0.02 / $1.25 claude-haiku-4-5-20251001 (alias claude-haiku-4-5) — 200k / 64k · manual extended thinking only, no effort · $1 / $0.10 / $5 · cutoff Feb 2025 · rel. 2025-10-15 grok-build-0.1 — 256k · $1 / $0.20 / $2 · PREVIEW ('early access'); grok-4.20-0309-non-reasoning — 1M · no reasoning · $1.25 / $0.20 / $2.50 · batch −20 % gemini-3.5-flash-lite — 1,048,576 / 65,536 · thinking default minimal · $0.30 / $0.03 / $2.50 · rel. 2026-07-21 · GA · LIVE_VERIFIED · free tier; gemini-3.6-flash / -3.7-flash $0.75 / $0.075 / $3.75 (introductory) Cheapest current text model per provider. OpenAI nano/Luna and Gemini Flash-Lite are ~5× cheaper than Haiku 4.5; xAI has no sub-$1 text model.
Deep-compute "pro" gpt-5.5-pro / gpt-5.4-pro — 1,050k / 128k · effort medium…xhigh · no prompt caching, no fast · $30 / — / $180; gpt-5.6-sol with reasoning.mode: pro (standard rates) no pro SKU — output_config.effort: max on Opus 5 / Fable 5.1 grok-4.20-multi-agent-0309 — 1M · one Responses call fans out to 4 or 16 agents (reasoning.effort low/medium vs high/xhigh) · $1.25 / $0.20 / $2.50 per agent-token · BETA · Responses-only (Chat → 400) Deep Research agents deep-research-preview-04-2026 / -max- (Interactions, background: true, ≤60 min, est. $1–3 / $3–7 per task, list-rate tokens + tools) · PREVIEW; gemini-3.8-live-extended-thinking for voice OpenAI sells extra compute as a model/mode, xAI as a multi-agent model, Gemini as an agent, Anthropic as the top effort value.
Coding agents gpt-5.3-codex — 400k / 128k · $1.75 / $0.175 / $14 · standard + fast only; GPT-5.6 family lists hosted shell / apply_patch / skills claude-opus-5, claude-sonnet-5 (bash, text_editor, code_execution, skills, computer/browser toolsets, programmatic tool calling); claude-opus-4-8 legacy grok-build-0.1 (grok-code-fast-1 redirect; 256k; $1 / $0.20 / $2; PREVIEW) + Grok Build CLI (grok, BETA: TUI, headless, ACP, MCP, hooks, skills, sandbox) antigravity-preview-09-2026 (Interactions coding agent with Linux sandbox, default model gemini-3.8-flash, max_total_tokens) · PREVIEW; gemini-3.1-pro-preview-customtools (bash-style custom tools, 3.1 Pro price) Only xAI and Gemini ship a coding-specific model id; OpenAI ships codex; Anthropic ships coding tools on its general models.
Previous-generation flagships still served gpt-5.4 ($2.5 / $0.25 / $15), gpt-5.2 ($1.75 / $0.175 / $14), gpt-5.1 & gpt-5 ($1.25 / $0.125 / $10), o3 ($2 / $0.50 / $8), o3-pro ($20 / — / $80) claude-opus-4-8, -4-7, -4-6 ($5 / $0.50 / $25, LEGACY), claude-opus-4-5-20251101 ($5, 200k, LEGACY), claude-sonnet-4-6 ($3 / $0.30 / $15, LEGACY), claude-sonnet-4-5-20250929 ($3, 200k, LEGACY) grok-4.20-0309-reasoning — 1M · reasoning always on (reasoning_effort → 400) · $1.25 / $0.20 / $2.50 · batch −20 % · 15 aliases incl. grok-4.20, grok-4.20-beta*, grok-4.20-reasoning-gv2 gemini-3-flash-preview ($0.50 / $0.05 / $3, PREVIEW, replacement 3.6 Flash, no date), gemini-3.1-flash-lite ($0.25 / $0.025 / $1.50, DEPRECATED → 2027-05-07), gemini-2.5-pro / -flash / -flash-lite (priced and listed but 404 'no longer available to new users') OpenAI keeps old flagships "Active" until a dated deprecation; Anthropic marks "Active (legacy)"; xAI keeps 4.20 on the price list; Gemini keeps 2.5 listed while blocking new users.
Deprecated but callable o4-mini, o3-mini, o1, o1-pro, gpt-4.1-nano (2026-10-23), gpt-5-*-2025-08-07, o3-2025-04-16, o3-pro-2025-06-10 (2026-12-11) claude-mythos-preview (invite, replacement claude-mythos-5, date TBA) grok-imagine-image-quality (→ 2026-11-02, redirects to grok-imagine-image-2.0 low); /v1/messages compat surface gemini-3.1-flash-lite (→ 2027-05-07), gemini-2.5-flash-image (→ 2026-10-02), gemini-omni-flash-preview (→ 2026-09-30), antigravity-preview-05-2026 (→ 2026-10-05), gemini-embedding-001 (→ 2028-05-14) see §7
Non-reasoning models gpt-4.1 (1,047k / 32k, $2 / $0.50 / $8, fine-tunable), gpt-4.1-mini, gpt-4o (128k / 16k), gpt-4o-mini, chat-latest (400k, $5 / $0.50 / $30) none current — every active Claude model reasons (Haiku/Sonnet 4.5 optionally) grok-4.20-0309-non-reasoning (no reasoning_content, reasoning_tokens: 0); grok-4.3 with reasoning_effort: none none — every Gemini 3.x text model thinks (minimal on Flash-Lite / 3.6 / 3.5 / 3 Flash; thinkingBudget: 0 on 3.6 / 3.5 Flash only); Gemma 4 live returned thoughtsTokenCount: 5 OpenAI and xAI keep explicit non-reasoning ids; Anthropic and Gemini only expose reasoning-off switches on some models.
Gated / invite / paid-only gpt-5.6-cyber, gpt-5.5-cyber ($12.5 / $1.25 / $75), gpt-daybreak-blue-latest, gpt-daybreak-red-latest, gpt-rosalind-research (ACCOUNT_RESTRICTED) claude-mythos-5-1, claude-mythos-5 ($10 / $0.25–$1 / $50, Project Glasswing, ACCOUNT_RESTRICTED, 404 for our key) grok-embedding-small (404, no price), Skills API (404), tool_search (403 alpha), custom-voice creation (Enterprise), Management API (needs a Management key) — all ACCOUNT_RESTRICTED Pro models and every image / video / music model are paid-tier only (limit: 0 on the free tier): gemini-3.1-pro-preview, gemini-pro-latest, gemini-3.1-flash-image, gemini-3-pro-image, Veo 3.1, Lyria 3.x; cachedContents and Batch also 429/400 on free Different gates: invite programs (OpenAI, Anthropic), ACLs/alpha (xAI), billing tier (Gemini).
Free of charge none none none (prepaid credits) free-tier rows on 3.8/3.7/3.6/3.5 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite, 3 Flash Preview, Live 3.8 / 3.1, TTS 3.1 / 2.5 Flash, gemini-3.5-transcribe, gemini-embedding-2; Gemma 4 (gemma-4-31b-it, gemma-4-26b-a4b-it, 262,144 / 32,768, text only) is free-only; content may be used to improve Google products Gemini is the only provider with a documented $0 tier.

# 2. Capability table — current OpenAI text models

From capabilities in generated/models.json. ✔ true · ✘ false · ? "unknown". Effort = accepted reasoning.effort values (default in bold when recorded). Cache = prompt caching / 24 h retention. Tiers = service tiers listed as available (S standard, B batch, X flex, F fast).

Model Status Ctx Out Cutoff Modalities Effort SO FC Cache / 24h Web Code MCP Comp. Shell Patch Skills ToolSearch FT Tiers
gpt-6-astra DOCUMENTED · LIVE_VERIFIED 1,050,000 128,000 2026-04-30 text, image → text low, medium, high, xhigh, max ✔ ✔ ✔ / ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✘ S B X F
gpt-5.6-sol (= gpt-5.6) DOCUMENTED · LIVE_VERIFIED 1,050,000 128,000 2026-02-16 text, image → text none, low, medium, high, xhigh, max · mode pro ✔ ✔ ✔ / ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✘ S B X F
gpt-5.6-terra DOCUMENTED · LIVE_VERIFIED 1,050,000 128,000 2026-02-16 text, image → text none…max · pro ✔ ✔ ✔ / ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✘ S B X F
gpt-5.6-luna DOCUMENTED · LIVE_VERIFIED 1,050,000 128,000 2026-02-16 text, image → text none…max · pro ✔ ✔ ✔ / ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✘ S B X F
gpt-5.5 DOCUMENTED · LIVE_VERIFIED 1,050,000 128,000 2025-12-01 text, image → text none, low, medium, high, xhigh ✔ ✔ ✔ / ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✘ S B X F
gpt-5.5-pro DOCUMENTED · LIVE_VERIFIED 1,050,000 128,000 2025-12-01 text, image → text medium, high, xhigh ✔ ✔ ✘ / ✔ ✔ ✔ ✔ ✘ ✔ ✘ ✘ ✘ ✘ S B X
gpt-5.4 DOCUMENTED · LIVE_VERIFIED 1,050,000 128,000 2025-08-31 text, image → text none, low, medium, high, xhigh ✔ ✔ ✔ / ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✘ S B X F
gpt-5.4-pro DOCUMENTED · LIVE_VERIFIED 1,050,000 128,000 2025-08-31 text, image → text medium, high, xhigh ✘ ✔ ✘ / ✘ ✔ ✘ ✔ ✔ ✘ ✔ ✘ ✔ ✘ S B X
gpt-5.4-mini DOCUMENTED · LIVE_VERIFIED 400,000 128,000 2025-08-31 text, image → text none, low, medium, high, xhigh ✔ ✔ ✔ / ? ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✘ S B X F
gpt-5.4-nano DOCUMENTED · LIVE_VERIFIED 400,000 128,000 2025-08-31 text, image → text none, low, medium, high, xhigh ✔ ✔ ✔ / ? ✔ ✔ ✔ ✘ ✔ ✔ ✔ ✘ ✘ S B X
gpt-5.3-codex DOCUMENTED · LIVE_VERIFIED 400,000 128,000 2025-08-31 text, image → text low, medium, high, xhigh ✔ ✔ ✔ / ? ✔ ✘ ✘ ✘ ✔ ✘ ✔ ✘ ✘ S F
gpt-5.2 DOCUMENTED · LIVE_VERIFIED 400,000 128,000 2025-08-31 text, image → text none, low, medium, high, xhigh ✔ ✔ ✔ / ✔ ✔ ✔ ✔ ✘ ✔ ✔ ✔ ✘ ✘ S B X F
gpt-5.2-pro DOCUMENTED · LIVE_VERIFIED 400,000 128,000 2025-08-31 text, image → text (model default) ✘ ✔ ✘ / ✘ ✔ ✘ ✔ ✘ ✘ ✘ ✘ ✘ ✘ S B
gpt-5.1 DOCUMENTED · LIVE_VERIFIED 400,000 128,000 2024-09-30 text, image → text none, low, medium, high ✔ ✔ ✔ / ✔ ✔ ✔ ✔ ✘ ✘ ✔ ✘ ✘ ✘ S B X F
gpt-5 DOCUMENTED · LIVE_VERIFIED 400,000 128,000 2024-09-30 text, image → text minimal, low, medium, high ✔ ✔ ✔ / ✔ ✔ ✔ ✔ ✘ ✘ ✘ ✘ ✘ ✘ S B X F
gpt-5-mini DOCUMENTED · LIVE_VERIFIED 400,000 128,000 2024-05-31 text, image → text (default) ✔ ✔ ✘ / ✘ ✔ ✔ ✔ ✘ ✘ ✘ ✘ ✘ ✘ S B X F
gpt-5-nano DOCUMENTED · LIVE_VERIFIED 400,000 128,000 2024-05-31 text, image → text (default) ✔ ✔ ✔ / ? ✔ ✔ ✔ ✘ ✘ ✘ ✘ ✘ ✘ S B X
gpt-5-pro DOCUMENTED · LIVE_VERIFIED 400,000 272,000 2024-09-30 text, image → text high ✔ ✔ ✘ / ✘ ✔ ✘ ✔ ✘ ✘ ✘ ✘ ✘ ✘ S B
o3 DOCUMENTED · LIVE_VERIFIED 200,000 100,000 2024-06-01 text, image → text (default) ✔ ✔ ✔ / ? ✔ ✔ ✔ ✘ ✘ ✘ ✘ ✘ ✘ S B X F
o3-pro DOCUMENTED · LIVE_VERIFIED 200,000 100,000 2024-06-01 text, image → text (default) ✔ ✔ ✘ / ✘ ✔ ✘ ✔ ✘ ✘ ✘ ✘ ✘ ✘ S B
o4-mini DOCUMENTED · LIVE_VERIFIED · DEPRECATED (2026-10-23) 200,000 100,000 2024-06-01 text, image → text (default) ✔ ✔ ✔ / ? ✔ ✔ ✔ ✘ ✘ ✘ ✘ ✘ ✔ S B X F
gpt-4.1 DOCUMENTED · LIVE_VERIFIED 1,047,576 32,768 2024-06-01 text, image → text non-reasoning ✔ ✔ ✔ / ✔ ✔ ✔ ✔ ✘ ✘ ✘ ✘ ✘ ✔ S B F
gpt-4.1-mini DOCUMENTED · LIVE_VERIFIED 1,047,576 32,768 2024-06-01 text, image → text non-reasoning ✔ ✔ ✘ / ✘ ✔ ✔ ✔ ✘ ✘ ✘ ✘ ✘ ✔ S B F
gpt-4o / gpt-4o-mini DOCUMENTED · LIVE_VERIFIED 128,000 16,384 2023-10-01 text, image → text non-reasoning ✔ ✔ ✘ / ✘ ✔ ✔ ✔ ✘ ✘ ✘ ✘ ✘ ✔ S B F
chat-latest DOCUMENTED · LIVE_VERIFIED 400,000 128,000 2025-08-31 text, image → text non-reasoning ✔ ✔ ✘ / ✘ ✔ ✔ ✔ ✘ ✘ ✘ ✘ ✘ ✘ S

Data note: for the GPT-5.6 family, GPT-6 Astra and the Daybreak/cyber ids the records have function_calling: true and tool_function_calling: false — the second flag was read from a model-page tool list that omits "function calling" for those pages; treat function_calling as authoritative (function tools were used live on gpt-5.6-luna in the Agents API run).

# 3. Capability table — current Anthropic models

From capabilities and pricing in generated/models.json. Cache min = minimum cacheable prefix tokens. Prefill = assistant prefill accepted. Sampling = temperature/top_p/top_k behaviour. ZDR = zero-data-retention eligible.

Model Status Lifecycle Ctx Out Cutoff (reliable / training) Thinking Effort (default high) SO / strict Cache min · 1h Fast tool_choice any/tool Prefill Sampling Web / Fetch / Code / MCP / PTC Computer toolset GA / Browser ZDR Retirement not before
claude-fable-5-1 DOCUMENTED · LIVE_VERIFIED Active (latest) 1,000,000 128,000 Jun 2026 / Jun 2026 adaptive, always on low, medium, high, xhigh, max · per-message (beta) ✔ / ✔ 512 · ✔ ✘ ✘ (400) ✘ 400 if non-default ✔ ✔ ✔ ✔ ✔ ✔ / ✔ ✘ 2027-09-01
claude-mythos-5-1 DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW Active (invite) 1,000,000 128,000 Jun 2026 adaptive, always on low…max · per-message ✔ / ✔ 512 · ✔ ✘ ✘ ✘ 400 ✔ ✔ ✔ ✔ ✔ ✔ / ✔ ✘ 2027-09-01
claude-fable-5 DOCUMENTED · LIVE_VERIFIED · LEGACY Active (legacy) 1,000,000 128,000 Jan 2026 adaptive, always on low…max ✔ / ✔ 512 · ✔ ✘ ✔ ✘ 400 ✔ ✔ ✔ ✔ ✔ ✔ / ✔ ✘ 2027-06-09
claude-mythos-5 DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW Active (invite) 1,000,000 128,000 Jan 2026 adaptive, always on low…max ✔ / ✔ 512 · ✔ ✘ ✔ ✘ 400 ✔ ✔ ✔ ✔ ✔ ✔ / ✔ ✘ 2027-06-09
claude-opus-5 DOCUMENTED · LIVE_VERIFIED Active (latest) 1,000,000 128,000 May 2026 adaptive, on by default (disable except xhigh/max) low…max · per-message ✔ / ✔ 512 · ✔ ✔ $10/$50 ✔ ✘ 400 ✔ ✔ ✔ ✔ ✔ ✔ / ✔ ✔ 2027-07-24
claude-opus-4-8 DOCUMENTED · LIVE_VERIFIED · LEGACY Active (legacy) 1,000,000 128,000 Jan 2026 adaptive (off by default) low…max ✔ / ✔ 1,024 · ✔ ✔ $10/$50 ✔ ? 400 ✔ ✔ ✔ ✔ ✔ ✔ / ✔ ✔ 2027-05-28
claude-opus-4-7 DOCUMENTED · LIVE_VERIFIED · LEGACY Active (legacy) 1,000,000 128,000 Jan 2026 adaptive low…max ✔ / ✔ 2,048 · ✔ ✘ (removed) ✔ ? 400 ✔ ✔ ✔ ✔ ✔ ✘ (beta computer_20251124) / ✘ ✔ 2027-04-16
claude-opus-4-6 DOCUMENTED · LIVE_VERIFIED · LEGACY Active (legacy) 1,000,000 128,000 May 2025 / Aug 2025 adaptive; manual enabled deprecated low, medium, high, max ✔ / ✔ 4,096 · ✔ ✘ (runs standard) ✔ ✘ accepted (not with thinking) ✔ ✔ ✔ ✔ ✔ ✘ / ✘ ✔ 2027-02-05
claude-opus-4-5-20251101 (alias claude-opus-4-5) DOCUMENTED · LIVE_VERIFIED · LEGACY Active (legacy) 200,000 64,000 May 2025 / Aug 2025 extended (manual budget_tokens) low, medium, high ✔ / ✔ 4,096 · ✔ ✘ ✔ ? accepted ✔ ✔ ✔ ✔ ✔ ✘ / ✘ ✔ 2026-11-24
claude-sonnet-5 DOCUMENTED · LIVE_VERIFIED Active (latest) 1,000,000 128,000 Jan 2026 adaptive, on by default low…max ✔ / ✔ 1,024 · ✔ ✘ ✔ ✘ (400 live) 400 ✔ ✔ ✔ ✔ ✔ ✔ / ✔ ✔ 2027-06-30
claude-sonnet-4-6 DOCUMENTED · LIVE_VERIFIED · LEGACY Active (legacy) 1,000,000 128,000 Aug 2025 / Jan 2026 adaptive; manual deprecated low, medium, high, max ✔ / ✔ 1,024 · ✔ ✘ ✔ ? accepted ✔ ✔ ✔ ✔ ✔ ✘ / ✘ ✔ 2027-02-17
claude-sonnet-4-5-20250929 (alias claude-sonnet-4-5) DOCUMENTED · LIVE_VERIFIED · LEGACY Active (legacy) 200,000 64,000 Jan 2025 / Jul 2025 extended (manual) ✘ (400) ✔ / ✔ 1,024 · ✔ ✘ ✔ ? accepted ✔ ✔ ✔ ✔ ✔ ✘ / ✘ ✔ 2026-09-29
claude-haiku-4-5-20251001 (alias claude-haiku-4-5) DOCUMENTED · LIVE_VERIFIED Active (latest) 200,000 64,000 Feb 2025 / Jul 2025 extended (manual) ✘ (400) ✔ / ✔ 4,096 · ✔ ✘ ✔ ✔ (live) accepted ✔ ✔ ✔ ✔ ✘ (PTC 400) ✘ / ✘ ✔ 2026-10-15

Modalities for every Claude model above: text + image + PDF in → text out. Tool support per exact tool type (e.g. web_search_20260318 vs 20250305) is in generated/compatibility/model-tool-matrix.json.

# 4. Capability table — current xAI Grok models

From capabilities, pricing and rate_limits in generated/models.json (kind: model). Reasoning = always-on / optional / none and accepted reasoning_effort values (default in bold). Cache = automatic prefix cache (no markers). Tools = all server tools (web_search, x_search, code_interpreter, file_search, mcp, image_generation, attachment search) — tools.json lists every Grok text model as compatible. Prices standard (< 200k prompt) in / cached / out; ≥200k = 2× on all tokens. Batch = 20 % discount where supported. Priority = service_tier: priority 2×. RPS/TPM = documented Tier 0 → Tier 4.

Model Status Ctx Out Cutoff Release Modalities Reasoning · effort SO FC Cache Tools Compact / WS Responses Batch Priority US 1.1× In / cached / out RPS · TPM (T0→T4) Aliases
grok-4.6 DOCUMENTED · LIVE_VERIFIED 500,000 none published 2026-02-01 2026-08 text, image → text always on · low, medium, high, xhigh ✔ ✔ ✔ (0.25×) ✔ all ✔ / ✔ ✘ 2× ✔ (only model) 2.00 / 0.50 / 6.00 150→500 · 50M→100M —
grok-4.5 DOCUMENTED · LIVE_VERIFIED 500,000 null null 2026-07 text, image → text always on · low, medium, high (xhigh = high) ✔ ✔ ✔ (0.15×) ✔ all ✔ / ✔ ✘ 2× ✘ 2.00 / 0.30 / 6.00 150→500 · 50M→100M grok-4.5-latest, grok-build-latest
grok-4.3 DOCUMENTED · LIVE_VERIFIED 1,000,000 null null 2026-04 text, image → text optional · low, medium, high, xhigh + none (LIVE_DISCOVERED) ✔ ✔ ✔ (0.16×) ✔ all ✔ / ✔ −20 % 2× ✘ (eu-west-1 host) 1.25 / 0.20 / 2.50 37→208 · 10M→85M grok-4.3-latest; redirects grok-3, grok-4-0709, grok-4-fast-*, grok-4-1-fast-*
grok-4.20-0309-reasoning DOCUMENTED · LIVE_VERIFIED 1,000,000 null null 2026-03-09 text, image → text always on · reasoning_effort → 400 ✔ ✔ ✔ ✔ all ✔ / ✔ −20 % 2× ✘ 1.25 / 0.20 / 2.50 37→208 · 10M→85M grok-4.20, grok-4.20-reasoning(-latest), grok-4.20-0309, grok-4.20-beta*, grok-4.20-experimental-beta*, grok-4.20-reasoning-gv2 (15)
grok-4.20-0309-non-reasoning DOCUMENTED · LIVE_VERIFIED 1,000,000 null null 2026-03-09 text, image → text none (0 reasoning tokens) ✔ ✔ ✔ ✔ all ✔ / ✔ −20 % 2× ✘ 1.25 / 0.20 / 2.50 37→208 · 10M→85M grok-4.20-non-reasoning(-latest), -beta-*, -gv2 (8)
grok-4.20-multi-agent-0309 DOCUMENTED · BETA · LIVE_VERIFIED 1,000,000 null null 2026-03 text, image → text multi-agent · effort = agent count (low/medium 4, high/xhigh 16; default medium) ✔ ✔ ✔ ✔ all Responses only (Chat → 400) −20 % not documented ✘ 1.25 / 0.20 / 2.50 (per agent token) 9→56 · 2.5M→21M grok-4.20-multi-agent(-latest), -beta-*, -experimental-beta-* (6)
grok-build-0.1 DOCUMENTED · PREVIEW · LIVE_VERIFIED 256,000 null null 2026-05 ('early access') text, image → text always on · reasoning_effort → 400 ✔ ✔ ✔ (0.2×) ✔ all ✔ / ✔ ✘ 2× ✘ 1.00 / 0.20 / 2.00 37→208 · 10M→85M grok-code-fast-1, grok-code-fast, grok-code-fast-1-0825

Every Grok text model also accepts the Anthropic-compatible /v1/messages (DEPRECATED), POST /v1/responses/compact, wss://api.x.ai/v1/responses, deferred Chat Completions and /v1/tokenize-text. usage.cost_in_usd_ticks (1 USD = 10^10 ticks) is returned on every call. Fine-tuning: none offered. logprobs: ignored on 4.20+.

# 4b. xAI media, voice and embedding models

Model Status Modalities Price Endpoints Notes
grok-imagine-image DOCUMENTED · LIVE_VERIFIED text+image → image $0.02 / image POST /v1/images/generations, /edits (JSON only), Batch, gRPC prompt limit 16,000 tokens; alias grok-imagine-image-2026-03-02; JPEG only
grok-imagine-image-2.0 DOCUMENTED · LIVE_VERIFIED text+image → image $0.04–$0.08 (quality low/medium × resolution 1k/1.5k/2k; catalogue default $0.06) same 5 reference images, 21:9 / 5:2 aspect ratios (Aug 2026); quality default moved to auto
grok-imagine-image-quality DOCUMENTED · DEPRECATED · LIVE_VERIFIED text+image → image $0.05 same (no Batch) retires 2026-11-02 → grok-imagine-image-2.0 quality=low; aliases -20260403, -latest, grok-imagine-image-pro
grok-imagine-video DOCUMENTED · LIVE_VERIFIED text+image+video → video $0.05 / s POST /v1/videos/generations, /edits, /extensions, GET /v1/videos/{request_id} (202 pending) 1–15 s, 480p–1080p, audio; batch URLs expire 1 h
grok-imagine-video-1.5 DOCUMENTED · LIVE_VERIFIED text+image+audio → video $0.08 / s same native 1080p, reference-to-video, reference_audios, last_frame; aliases -preview, -2026-05-30
grok-voice-think-fast-2.0 (grok-voice-latest) DOCUMENTED audio+text ↔ audio+text $0.08 / min ($4.80 / h) + $0.004 per text item wss://api.x.ai/v1/realtime, POST /v1/realtime/client_secrets, SIP /v2/phone-numbers, /v1/realtime/calls/{id}/refer and …/hangup OpenAI Realtime event vocabulary; tools in-session; 120-min sessions; -1.0 LEGACY; not in GET /v1/models
grok-voice-transcribe-2.0 / -1.0 DOCUMENTED audio → text $0.10 / h REST, $0.20 / h streaming POST /v1/stt, wss://api.x.ai/v1/stt 25 languages, diarization, smart_turn; default model ambiguous in docs
text-to-speech (no model id) DOCUMENTED text → audio $15 / 1M characters POST /v1/tts, GET /v1/tts/voices, wss://api.x.ai/v1/tts 28 voices, language required; custom voices Enterprise
grok-embedding-small DOCUMENTED · ACCOUNT_RESTRICTED text → embedding not published POST /v1/embeddings, GET /v1/embedding-models (empty) 404 for this team; Collections index model

# 5. Capability table — current Gemini models

From capabilities, pricing and deprecation in generated/models.json (kinds stable, preview, alias, agent). Thinking = thinkingLevel values (default in bold; budget = 2.5-style thinkingBudget). Tools: FC function calling, Search (googleSearch), Maps, URL (urlContext), Code (codeExecution), CU (computerUse, preview), FS (fileSearch). Cache = implicit / explicit (cachedContents). Tiers = B batch (50 %), X flex (50 %), P priority (1.8×). Prices standard in / cached / out (>200k tier for Pro). Free = free-tier availability per the pricing page.

Model Status Lifecycle Release Ctx in / out Modalities Thinking SO FC Search / Maps / URL / Code / CU / FS Cache impl / expl Tiers Live Free In / cached / out Shutdown
gemini-3.8-flash DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED Stable (GA 2026-09-02); gemini-flash-latest → here (observed) 2026-09-02 1,048,576 / 65,536 text, image, audio, video, PDF → text low, medium, high (minimal → 400) ✔ ✔ ✔ ✔ ✔ ✔ ✔(P) ✔ 4,096 / ✔ B X P ✘ ✔ 0.75 / 0.075 / 3.75 (→ 1.50 / 0.15 / 7.50 on 2027-01-01) none announced
gemini-3.7-flash DOCUMENTED · LIVE_DISCOVERED Stable (GA 2026-08-13) 2026-08-13 1,048,576 / 65,536 same low, medium, high ✔ ✔ ✔ ✔ ✔ ✔ ✔(P) ✔ 4,096 / ✔ B X P ✘ ✔ 0.75 / 0.075 / 3.75 (intro) none
gemini-3.6-flash DOCUMENTED · LIVE_DISCOVERED Stable (GA 2026-07-21) 2026-07-21 1,048,576 / 65,536 same minimal, low, medium, high; thinkingBudget: 0 allowed ✔ ✔ ✔ ✔ ✔ ✔ ✔(P) ✔ 4,096 / ✔ B X P ✘ ✔ 0.75 / 0.075 / 3.75 (intro) none
gemini-3.5-flash DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED Stable (GA 2026-05-19) 2026-05-19 1,048,576 / 65,536 same minimal, low, medium, high; budget 0 allowed ✔ ✔ ✔ ✔ ✔ ✔ ✔(P) ✔ 4,096 / ✔ B X P ✘ ✔ (caching limit: 0 live) 1.50 / 0.15 / 9.00 none
gemini-3.5-flash-lite DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED Stable (GA 2026-07-21) 2026-07-21 1,048,576 / 65,536 same minimal, low, medium, high (budget 0 → 400) ✔ ✔ ✔ ✔ ✔ ✔ ✔(P) ✔ ? / ✔ B X P ✘ ✔ (no caching) 0.30 / 0.03 / 2.50 none
gemini-3.1-pro-preview (+ -customtools) DOCUMENTED · LIVE_DISCOVERED · PREVIEW · ACCOUNT_RESTRICTED Preview; gemini-pro-latest → here 2026-02-19 1,048,576 / 65,536 same low, medium, high ✔ ✔ ✔ ✔ ✔ ✔ ✘ ✔ (AI Studio only) 4,096 / ✔ B X P ✘ ✘ (paid only) 2.00 / 0.20 / 12.00; >200k 4.00 / 0.40 / 18.00 none
gemini-3.1-flash-lite DOCUMENTED · LIVE_DISCOVERED · DEPRECATED Deprecated 2026-05-07 1,048,576 / 65,536 same minimal, low, medium, high (default not stated) ✔ ✔ ✔ ✔ ✔ ✔ ✘ ✔ ? / ✔ B X P ✘ ✔ 0.25 (0.50 audio) / 0.025 / 1.50 2027-05-07 → gemini-3.5-flash-lite
gemini-3-flash-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW Preview (replacement gemini-3.6-flash, no date) 2025-12-17 1,048,576 / 65,536 same minimal, low, medium, high ✔ ✔ ✔ ✔ ✔ ✔ ✔ ✔ 4,096 / ✔ B X P ✘ ✔ 0.50 (1.00 audio) / 0.05 / 3.00 none
gemini-2.5-pro DOCUMENTED · LIVE_DISCOVERED (deprecations.json: ACCOUNT_RESTRICTED) 'no longer available to new users' (404 live) 2025-06-17 1,048,576 / 65,536 text, image, audio, video, PDF → text budget 128–32,768 (cannot disable) ✔ ✔ ✔ ✔ ✔ ✔ ✘ ✔ 2,048 / ✔ B X P ✘ listed 1.25 / 0.125 / 10.00; >200k 2.50 / 0.25 / 15.00 undated
gemini-2.5-flash DOCUMENTED · LIVE_DISCOVERED 'no longer available to new users' 2025-06-17 1,048,576 / 65,536 text, image, audio, video → text budget 0–24,576 ✔ ✔ ✔ ✔ ✔ ✔ ✘ ✔ 2,048 / ✔ B X P ✘ ✔ 0.30 (1.00 audio) / 0.03 / 2.50 undated
gemini-2.5-flash-lite DOCUMENTED · LIVE_DISCOVERED · ACCOUNT_RESTRICTED 404 live → gemini-3.5-flash-lite 2025-07-22 1,048,576 / 65,536 same off by default; budget 512–24,576 ✔ ✔ ✔ ✔ ✔ ✔ ✘ ✔ ? / ✔ B X P ✘ ✔ 0.10 (0.30 audio) / 0.01 / 0.40 undated
gemma-4-31b-it, gemma-4-26b-a4b-it DOCUMENTED · LIVE_DISCOVERED (26B LIVE_VERIFIED) Stable 2026-04-02 262,144 / 32,768 text → text live thinking: true (thoughtsTokenCount: 5) ? ? ? (grounding 'not available') ✘ / ✘ — ✘ free only free; paid tier not available none
gemini-3.8-live / -extended-thinking DOCUMENTED · LIVE_DISCOVERED Stable (GA 2026-09-15) 2026-09-15 131,072 / 65,536 text, image, audio, video → audio (+text) interleaved (omit config) / LOW-MEDIUM-HIGH ✘ ✔ (async) ✔ ✘ ✘ ✘ ✘ ✘ ? / ✘ — ✔ ✔ text 0.75 / audio 3.00 in; text 4.50 / audio 12.00 out none
gemini-3.1-flash-live-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW 'legacy Live preview' → 3.8-live 2026-03-11 131,072 / 65,536 same MINIMAL…HIGH ✘ ✔ (sync) ✔ ✘ ✘ ✘ ✘ ✘ — — ✔ ✔ same as 3.8-live none
gemini-2.5-flash-native-audio-preview-12-2025 DOCUMENTED · LIVE_DISCOVERED · PREVIEW Preview → 3.8-live 2025-12-12 131,072 / 8,192 text, audio, video → text, audio thinkingBudget ✘ ✔ ✔ ✘ ✘ ✘ ✘ ✘ — — ✔ ✔ 0.50 text / 3.00 audio-video in; 2.00 / 12.00 out none
gemini-embedding-2 DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED Stable (GA 2026-04-22) 2026-04-22 8,192 / 1 text, image, audio, video, PDF → embedding (128–3,072 dims) — — — — — B ✘ ✔ 0.20 text / 0.45 image / 6.50 audio / 12.00 video none
gemini-embedding-001 DOCUMENTED · LIVE_DISCOVERED · DEPRECATED Deprecated 2025-07-14 2,048 / 1 text → embedding — — — — — — ✘ — not on pricing page 2028-05-14 → gemini-embedding-2

# 5b. Gemini media and agent models

Model Status Lifecycle Modalities Price Notes
gemini-3.1-flash-image (Nano Banana 2) DOCUMENTED · LIVE_DISCOVERED Stable (GA 2026-05-28) text, image, video, PDF → text, image in 0.50 / text out 3.00 / image out 60.00 per 1M → $0.045 (0.5K) … $0.151 (4K) per image; batch 50 % imageConfig 512/1K/2K/4K, 14 aspect ratios, web + image search grounding; thinking MINIMAL/HIGH; no free tier; live limits 65,536 / 65,536 (docs 131,072 / 32,768)
gemini-3.1-flash-lite-image DOCUMENTED · LIVE_DISCOVERED Stable (2026-06-30) same 0.25 / 1.50 / 30.00 → $0.034 (1K) 1K only, no search; C2PA + SynthID
gemini-3-pro-image (Nano Banana Pro) DOCUMENTED · LIVE_DISCOVERED Stable (GA 2026-05-28); -preview / nano-banana-pro-preview RETIRED 2026-06-25 but still listed text, image → text, image 2.00 / 12.00 / 120.00 → $0.134 (1K/2K), $0.24 (4K) Google Search grounding; priority 216 image out
gemini-2.5-flash-image (Nano Banana) DOCUMENTED · LIVE_DISCOVERED · DEPRECATED shutdown 2026-10-02 text, image → text, image $0.039 / image (30.00 / 1M) replacement gemini-3.1-flash-image / -lite-image
veo-3.1-generate-preview / -fast- / -lite- DOCUMENTED · LIVE_DISCOVERED · PREVIEW (lite LIVE_VERIFIED GET) Preview; Veo 2.0/3.0 RETIRED 2026-06-30 text, image → video + audio $0.40 / $0.40 / $0.60 · $0.10 / $0.12 / $0.30 · $0.05 / $0.08 / — per second (720p / 1080p / 4K) :predictLongRunning + operations; 4/6/8 s; no free tier
gemini-omni-1.1-flash DOCUMENTED · LIVE_DISCOVERED Stable (GA 2026-08-27); gemini-omni-flash-preview DEPRECATED → 2026-09-30 text, image, video → video in 1.50 / text out 9.00 / video out 17.50 per 1M (≈ $0.10 / s at 720p) Interactions only
lyria-3.5 · lyria-3-clip-preview · lyria-3-pro-preview · lyria-realtime-exp DOCUMENTED · LIVE_DISCOVERED (3.5 LIVE_VERIFIED GET; RealTime LIVE_VERIFIED WS) Stable (GA 2026-09-03) · Preview · Preview → 3.5 · Experimental text, image → music + lyrics $0.08 / song · $0.04 / 30-s clip · $0.08 · unpriced no free tier; RealTime BidiGenerateMusic WebSocket
gemini-3.1-flash-tts-preview · gemini-2.5-flash-preview-tts · gemini-2.5-pro-preview-tts DOCUMENTED · LIVE_DISCOVERED · PREVIEW Preview (2.5 → 3.1) text → audio $1 / $20 · $0.50 / $10 · $1 / $20 per 1M (text in / audio out) 30 voices, ≤2 speakers; streaming TTS on 3.1 only; free tier 10 req/day observed
gemini-3.5-transcribe · gemini-3.5-transcribe-live · gemini-3.5-live-translate-preview DOCUMENTED · LIVE_DISCOVERED (unary LIVE_VERIFIED) Stable (Aug 2026) · Stable · Preview audio → text · audio → text · audio → audio + text ≈ $0.005 / min · ≈ $0.009 / min · ≈ $0.037 / min audioTranscriptionConfig (diarization, word timestamps, custom vocabulary)
deep-research-preview-04-2026 · deep-research-max-preview-04-2026 · deep-research-pro-preview-12-2025 DOCUMENTED · LIVE_DISCOVERED · PREVIEW (04-2026 LIVE_VERIFIED) Preview agents text, image, audio, video, PDF → text (+ image) list-rate tokens (incl. intermediate) + tool fees; est. $1–3 / $3–7 per task Interactions agent, background: true mandatory, ≤60 min; listed as models (generateContent advertised)
antigravity-preview-09-2026 · antigravity-preview-05-2026 DOCUMENTED · LIVE_DISCOVERED · PREVIEW (05-2026 DEPRECATED → 2026-10-05) Preview coding agent text → text list-rate tokens; sandbox compute unbilled in preview 09-2026 renamed built-in tools (PascalCase params)
gemini-2.5-computer-use-preview-10-2025 DOCUMENTED · LIVE_DISCOVERED · PREVIEW (deprecations.json: LEGACY) replaced by built-in computer use on 3.x text, image → text 1.00 / 5.00 browser only; 429 limit: 0 live
gemini-robotics-er-2-preview / -streaming-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW Preview (1.5 / 1.6 RETIRED) text, image, audio, video → text 1.00 / 0.10 / 5.00 (streaming: pricing section empty) embodied reasoning
aqa LIVE_DISCOVERED Stable (PaLM-era) text → text not priced :generateAnswer only (LIVE_VERIFIED), English

# 6. Naming and versioning conventions

Topic OpenAI Anthropic xAI Gemini
Id families gpt-<gen>[.<minor>][-<size>][-<variant>], o<N>[-mini|-pro], gpt-<gen>-codex, gpt-image-*, gpt-realtime-*, gpt-audio-*, gpt-live-*, text-embedding-*, omni-moderation-*, sora-*; codenames for GPT-5.6 tiers (sol/terra/luna) and GPT-6 (astra); cyber/Daybreak aliases claude-<family>-<gen>[-<minor>][-<YYYYMMDD>] (claude-opus-5, claude-sonnet-4-6); families Fable (frontier), Mythos (invite research line), Opus, Sonnet, Haiku grok-<major>.<minor> (grok-4.6, grok-4.3), dated variant ids grok-4.20-0309-reasoning / -non-reasoning / grok-4.20-multi-agent-0309, product lines grok-build-<v>, grok-imagine-image[-<v>|-quality], grok-imagine-video[-<v>], grok-voice-think-fast-<v>, grok-voice-transcribe-<v>, grok-embedding-small; TTS has no model id gemini-<major>.<minor>-<size>[-<capability>][-preview[-MM-YYYY]] (gemini-3.8-flash, gemini-3.1-pro-preview, gemini-3.1-flash-image, gemini-3.8-live, gemini-3.5-transcribe), gemini-embedding-<n>, gemma-4-<params>-it, veo-<v>-<tier>-generate-preview, lyria-<v>, imagen-* (retired); agents deep-research-*, antigravity-preview-MM-YYYY
Alias vs snapshot Undated alias (gpt-5.4) → dated snapshot (gpt-5.4-2026-03-05); the alias moves when a new snapshot ships; response model echoes the snapshot. GPT-5.6/6 tier ids have no dated snapshot; gpt-5.6 is an alias of gpt-5.6-sol that 404s on GET /v1/models but works on POST. *-latest ids (chat-latest, chatgpt-image-latest, codex-mini-latest) float. Since Claude 4.6 the canonical id is undated (claude-opus-4-6, claude-sonnet-5, claude-fable-5-1) and is itself the snapshot. Older lines keep dated snapshots with undated aliases (claude-haiku-4-5 → claude-haiku-4-5-20251001); aliases follow the newest snapshot. No -latest suffix. Cloud ids differ per platform (anthropic.claude-opus-5 on Bedrock, claude-opus-5@… on Vertex). Documented scheme: <model> = latest stable, <model>-latest, <model>-<date> pinned — in practice the dated 4.20 ids are canonical and the bare names (grok-4.20) are aliases (15 aliases on one model); grok-4.6 has no -latest. Retired ids redirect (grok-3, grok-4-0709, grok-4-fast*, grok-4-1-fast* → grok-4.3; grok-code-fast-1 → grok-build-0.1) and appear in the catalogue's aliases[]. grok-voice-latest routes to grok-voice-think-fast-2.0 since 2026-08-05. Stable ids are undated (gemini-3.8-flash, 'usually don't change'); -preview ids may carry a month suffix (gemini-2.5-flash-native-audio-preview-12-2025) and can be deprecated with two weeks' notice; -latest aliases (gemini-flash-latest, gemini-pro-latest, gemini-flash-lite-latest, gemini-2.5-flash-native-audio-latest) are hot-swapped (2-week e-mail notice) — gemini-flash-latest moved 3 Flash Preview → 3.5 Flash → 3.8 Flash (observed via modelVersion); -001 suffixes only on 2.0 / embedding-001 / Veo 3.0; -exp experimental; gemini-3-pro-preview was repointed to 3.1 Pro on 2026-03-09.
Version pinning advice pin the dated snapshot in production; GET /v1/models/{id} exposes shutdown_date pin the undated id for 4.6+ (already immutable); for 4.5 lines pin the dated snapshot pin the dated 4.20 ids or the bare grok-4.x id (no snapshots exist); check response.model — a retired id silently returns its replacement pin the stable id; avoid -latest in production; modelVersion in every response tells you what actually ran
Tool versioning tool type strings are undated (web_search, code_interpreter) with a few dated snapshots (web_search_2025_08_26) every server/Anthropic-defined tool type carries _YYYYMMDD (web_search_20260318, code_execution_20260521) undated OpenAI-style type strings (web_search, x_search, code_interpreter/code_execution, file_search/collections_search, mcp, shell, tool_search) undated camelCase keys (googleSearch, urlContext, codeExecution, fileSearch, computerUse); legacy googleSearchRetrieval
API surface versioning none (openai-version: 2020-10-01 response header); breaking changes via new endpoints and OpenAI-Beta surfaces anthropic-version: 2023-06-01 required; features gated by dated anthropic-beta values that graduate to GA none — no version or beta headers; docs.x.ai/developers/release-notes is the changelog; gRPC proto v6 path version /v1beta (86 methods) vs /v1 (47); Interactions optional Api-Revision: 2026-05-20; breaking schema change 2026-05-26/06-08 (outputs → steps)
Knowledge cutoff single date per model page (knowledge_cutoff), e.g. gpt-6-astra 2026-04-30 two dates: reliable knowledge and training data (knowledge_cutoff.reliable / .training_data) published only for grok-4.6 (2026-02-01); null for every other Grok record not stated on the API model pages; January 2025 from the Gemini 3 / 2.5 model cards; June 2025 for gemini-2.5-flash-image; null for 64 of 97 records
Discovery GET /v1/models lists 136 ids incl. snapshots, owned_by, created, shutdown_date; fine-tuned ids ft:… GET /v1/models lists 11 ids with capabilities (thinking types, effort levels, context_management strategies, structured_outputs, citations, pdf_input) GET /v1/models 12 ids; typed catalogues /v1/language-models, /v1/image-generation-models, /v1/video-generation-models, /v1/embedding-models expose live prices in ticks, aliases[], input_modalities, fingerprints; voice models absent GET /v1beta/models 58 ids (inputTokenLimit, outputTokenLimit, supportedGenerationMethods, thinking, sampling defaults, version); /v1/models 22; agents listed as models; shut-down previews still listed

# 7. Deprecation and retirement policies

Aspect OpenAI Anthropic xAI Gemini
Vocabulary Deprecated = retirement announced with a shutdown date; Legacy = no longer updated, will be deprecated later; Sunset/shut down = no longer accessible Active / Legacy (no more updates) / Deprecated (replacement + retirement date, ≥60 days) / Retired (requests fail) Available (on /developers/models + /pricing + GET /v1/models) / Deprecated (notice period) (migration guide; ~60-day notice observed) / Retired (removed from the catalogue but the slug keeps resolving and redirects to a replacement billed at the replacement's price) / undocumented Legacy (older ids vanished from docs) Stable (GA) / Preview (production allowed, tighter limits, ≥2 weeks notice) / Latest alias / Experimental / Legacy (existing customers) / Deprecated (earliest shutdown date announced) / Shut down; API enum ModelStatus.modelStage (EXPERIMENTAL, PREVIEW, STABLE, LEGACY, DEPRECATED, RETIRED)
Notice period dated per announcement (typically ≥ 6 months for flagship snapshots, shorter for previews); shutdown_date also exposed live on GET /v1/models ≥ 60 days from deprecation to retirement; every active model carries a "not sooner than" retirement commitment (Fable 5.1 ≥ 2027-09-01, Opus 5 ≥ 2027-07-24, Sonnet 5 ≥ 2027-06-30, Haiku 4.5 ≥ 2026-10-15) no published minimum; observed ~60 days (grok-imagine-image-quality announced 2026-09-02 → 2026-11-02); the May-15 batch was announced via a migration guide with automatic redirects previews ≥ 2 weeks; stable models get an earliest-shutdown date on the deprecations page (e.g. gemini-3.1-flash-lite 2026-05-07 → 2027-05-07 = 12 months); 'exact date communicated in advance'; Vertex has its own schedule
Replacement mapping each row names a replacement (o4-mini → gpt-5.6-terra, gpt-image-1 → gpt-image-2, whisper-1 → gpt-transcribe) each retirement names a replacement (claude-opus-4-1-20250805 → claude-opus-4-8, claude-3-haiku-20240307 → claude-haiku-4-5-20251001) migration guides name the replacement and the effort setting (grok-4-fast-reasoning → grok-4.3 reasoning_effort=low; grok-4-fast-non-reasoning → grok-4.3 none; grok-code-fast-1 → grok-build-0.1; grok-imagine-image-quality → grok-imagine-image-2.0 quality=low) deprecations page names a replacement (gemini-3.1-flash-lite → gemini-3.5-flash-lite, gemini-2.5-flash-image → gemini-3.1-flash-image, gemini-embedding-001 → gemini-embedding-2, text-embedding-004 → gemini-embedding-2); live 404 messages also name one (gemini-2.5-flash-lite → gemini-3.5-flash-lite)
Upcoming (after 2026-09-18) 2026-09-24 Videos API + Sora 2; 2026-09-28 gpt-3.5-turbo-instruct, babbage-002, davinci-002; 2026-10-01 gpt-5.4-cyber; 2026-10-23 o1, o1-pro, o3-mini, o4-mini, gpt-4, gpt-4-turbo, gpt-4.1-nano, gpt-image-1; 2026-11-30 Evals API, Agent Builder, reusable prompts; 2026-12-01 gpt-image-1-mini/1.5; 2026-12-11 gpt-5 2025 snapshots, o3, o3-pro; 2027-01-06 no new fine-tuning jobs; 2027-01-20 legacy audio/realtime; 2027-02-26 whisper-1, gpt-4o-transcribe no dated model retirements pending; claude-mythos-preview deprecated 2026-06-09 (date TBA). Feature-level: context-1m-2025-08-07 header retired 2026-04-30; fast mode removed on Opus 4.6/4.7; sampling params and manual thinking deprecated (400 on 4.7+); prefill deprecated (400 on 4.6+) 2026-09-21 12:00 PT x_search per-call billing → per post / per profile; 2026-11-02 grok-imagine-image-quality retires (redirect to 2.0 low); /v1/messages Anthropic compat 'fully deprecated' (no date); logprobs ignored on 4.20+ 2026-09-30 gemini-omni-flash-preview; 2026-10-02 gemini-2.5-flash-image; 2026-10-05 antigravity-preview-05-2026; September 2026 standard API keys rejected (auth keys only); 2027-05-07 gemini-3.1-flash-lite; 2028-05-14 gemini-embedding-001; replacement recommended without date: gemini-3-flash-preview, gemini-3.1-flash-live-preview, 2.5 native-audio / TTS previews, lyria-3-pro-preview; sampling params deprecated 2026-07-21
Recently retired Assistants API (2026-08-26), DALL·E 2/3 (2026-05-12), text-moderation-*, search-preview snapshots (2026-07-23), codex snapshots ≤ 5.2, deep-research models, computer-use-preview Opus 4.1 (2026-08-05), Sonnet 4 / Opus 4 (2026-06-15), Claude 3 Haiku (2026-04-20), Claude 3.5/3.7, Claude 3 Opus/Sonnet, 2.x, 1.x, Instant; /v1/complete effectively retired (400) 2026-05-15: grok-3, grok-4-0709, grok-4-fast-reasoning, grok-4-fast-non-reasoning, grok-4-1-fast-reasoning, grok-4-1-fast-non-reasoning, grok-code-fast-1, grok-imagine-image-pro (all redirect); grok-2-image 404; Live Search (chat search_parameters) 410; /v1/completions, /v1/complete 400 Gemini 2.0 Flash/-Lite (2026-06-01); 2.5 previews (2025-11 → 2026-03); gemini-3-pro-preview (2026-03-09, repointed); gemini-3.1-flash-lite-preview (2026-05-25); image previews gemini-3.1-flash-image-preview, gemini-3-pro-image-preview / nano-banana-pro-preview (2026-06-25); Imagen 3/4 (2026-08-17); Veo 2.0/3.0 (2026-06-30); half-cascade Live models (2025-12-09); text-embedding-004 (2026-01-14), embedding-001, embedding-gecko-001, gemini-embedding-exp* (2025-10-30); robotics ER 1.5/1.6; model tuning (May 2025); LearnLM
After retirement requests fail; fine-tuned models retire with their base 404 not_found_error on the Claude API; several retired ids remain available on Bedrock / Vertex slug still resolves: GET /v1/models/grok-3 returns the grok-4.3 object; requests are served and billed at the replacement's rates (response.model shows it); grok-2-image is a hard 404 404 NOT_FOUND on generation (Model is not found for api version v1beta), yet shut-down ids remain in GET /v1beta/models (e.g. gemini-3-pro-preview, gemini-3.1-flash-lite-preview, image previews) — listing ≠ availability
Machine-readable generated/deprecations.json (165 flat OpenAI records: model, shutdown_date, replacement, phase) generated/deprecations.json (1 Anthropic record with models[], active_models_retirement_commitments{}, api_features[]) generated/deprecations.json (1 xAI record: models[] with behavior_after and our_probe, api_features[], release_notes_digest[] by month); models.json kinds retired_redirect, legacy, alias generated/deprecations.json (1 Gemini record: 60+ models[] with released/shutdown/replacement/live_listed, alias_history[], api_features[], active_models_no_shutdown_announced[])

# 8. Live discovery vs docs

  • OpenAI: 136 ids in GET /v1/models; aliases gpt-5.6, gpt-5.5-cyber, gpt-5-search-api-2025-10-14 behave as id-only/alias records (record_kind: alias|id_only); gpt-5.6 echoes gpt-5.6-sol; gpt-5.4-nano echoes effort: none when omitted; gpt-6-astra rejects reasoning.effort: none.
  • Anthropic: 11 ids in GET /v1/models (Mythos ids absent → 404); claude-fable-5-1 returned Opus-class rate-limit headers (10M ITPM) although docs list 4M for the Fable class (account-specific); Haiku 4.5 tolerated edited/removed thinking blocks and prefill-with-thinking (docs say 400) — graceful degradation, do not rely on it.
  • xAI: 12 ids in GET /v1/models (7 language, 3 image, 2 video; voice/embedding models absent); redirects verified (GET /v1/models/grok-3, /grok-4-fast, /v1/language-models/grok-4-0709 → grok-4.3; grok-code-fast-1 → grok-build-0.1); grok-4.3 accepts reasoning_effort: none (undocumented) while grok-4.20-0309-reasoning and grok-build-0.1 reject reasoning_effort although docs/parameters list them as compatible; x_search emits custom_tool_call items (docs: x_search_call); GET /v1/responses/{id} returns 200 for store:false ids; rate-limit headers exist but are undocumented (7,200 / 1,800 requests per minute vs documented 150 / 37 RPS); eu-west-1.api.x.ai serves grok-4.3 undocumented; every text call bills 70–180 reasoning tokens for "Reply with OK.".
  • Gemini: 58 ids in GET /v1beta/models vs 22 in GET /v1/models (docs claim parity); gemini-flash-latest resolved to gemini-3.8-flash (changelog last said 3.5 Flash); Gemini 2.5 Pro/Flash/Flash-Lite return 200 on models.get but 404 'no longer available to new users' on generation (undocumented policy); Pro models, image/video/music models, cachedContents create and Batch create return 429 limit: 0 / 400 FAILED_PRECONDITION on the free tier (ACCOUNT_RESTRICTED); explicit-cache minimum is 1,024 tokens live (docs' 4,096 is the implicit-cache threshold); thoughtSignature validation can be bypassed with the documented dummy strings; token limits differ from the model pages for image, Lyria, Deep Research, translate and computer-use ids (see §9); gemini-3.8-live reports version: 3.1-flash-live-03-2026.

# 9. Data inconsistencies found during synthesis

Reported for the fragment owners (fix in generated/fragments/**, never in merged files). Items 1–16 carry over from the two-provider synthesis; 17+ were found while adding xAI and Gemini.

  1. models.json o3: pricing.standard = $2 / $0.5 / $8 but pricing["model_page:Text tokens"] = $1 / $0.25 / $4 (two sources disagree; other models agree).
  2. models.json OpenAI GPT-5.6 family, gpt-6-astra, gpt-5.6-cyber, gpt-daybreak-*, gpt-4o-mini-audio-preview*: capabilities.function_calling: true while capabilities.tool_function_calling: false.
  3. models.json Anthropic alias records (claude-opus-4-5, claude-sonnet-4-5, claude-haiku-4-5) store pricing as a string ("same as …") instead of the schema object; modalities, capabilities, thinking are null on aliases.
  4. models.json gpt-5.5-cyber and gpt-5-search-api(-2025-10-14) are id_only records without context/output/modalities although priced.
  5. pricing.json unit strings are heterogeneous (per 1M tokens vs model-page 1M tokens; image_output low 1024x1024 dimensions encode size in the dimension name) — now 12 distinct unit strings across four providers (per 1K requests, per 1K search queries, per 1K grounded prompts, per 1K calls, per call, per request, multiplier, per hour, per page, per 1M tokens per hour, per GiB per day, per 1M characters).
  6. endpoints.json: six OpenAI records carry verification.result: "success" without a live call; three LIVE_VERIFIED endpoints have result: "failure" — status and result should be reconciled.
  7. endpoints.json: POST /v1/responses?beta=true and POST /v1/messages (beta surface) encode a variant in the path field; auth is a string for some fragments and an object for others.
  8. streaming-events.json: Anthropic core stream recorded under two api labels; Managed Agents client→server events appear twice.
  9. webhook-events.json: the 28 OpenAI records have api: null while Anthropic records set api: managed-agents.
  10. headers.json: required is boolean for OpenAI rows and a free string for Anthropic/xAI/Gemini rows.
  11. errors.json: api_family is null for core Anthropic and OpenAI errors; several Anthropic rows have http_status null or composite; Gemini rows record observed_live: false for 400 FAILED_PRECONDITION / 403 / 501 although the Gemini docs pages report them observed.
  12. deprecations.json: OpenAI = 165 flat records, Anthropic / xAI / Gemini = 1 nested record each — two schemas in one file.
  13. rate-limits.json: OpenAI record has no status field and no documented key (its content is under concepts/usage_tiers); the Gemini enqueued-token table lists gemini-2.0-flash-image and shut-down 2.5 previews while omitting the GA image ids.
  14. tools.json OpenAI programmatic_tool_calling has an empty compatible_models list; OpenAI computer tool is UNVERIFIED although eleven model records claim tool_computer_use: true.
  15. Anthropic tool-search docs table omits claude-sonnet-5 while the model record lists both tool-search types.
  16. Live vs docs (Anthropic): mcp_servers error message advertises mcp-client-2026-09-15 which the API rejects; code_execution_requests usage counter absent live; Managed Agents webhook names differ from stream names.
  17. xAI models.json: max_output is null for all 17 model records (only the grok-4.6 page says "No text output limit"); knowledge_cutoff is null for every model except grok-4.6; parameters.json records max_output_tokens default: None while docs/xai/responses.md says 128,000; a documented sample echoed max_output_tokens: 2000 unexplained.
  18. xAI reasoning flags: parameters.json lists grok-4.20-0309-reasoning and grok-build-0.1 as compatible with reasoning_effort (Chat) / reasoning.effort (Responses) but both return 400 live; none appears in the enum only as "(LIVE_DISCOVERED, grok-4.3)"; grok-4.5 xhigh is listed without the "treated as high" caveat from docs/xai/reasoning.md.
  19. xAI prices: grok-imagine-image-2.0 per_image_default: 0.06 (catalogue, medium/1k) vs pricing page $0.04 (low/1k); the REST reference describes price units as "USD cents per 100M tokens" while the live catalogue and pricing.json use ticks of 1e-10 USD; grok-4.20-multi-agent-0309 has no priority rows in pricing.json while models.json says "2.0 (not documented per model)".
  20. xAI tools/streaming: x_search returns custom_tool_call items live vs documented x_search_call; streaming-events.json has no x_search events; image usage documents input_tokens/output_tokens that are absent live (only cost_in_usd_ticks).
  21. xAI endpoints: Collections are documented on management-api.x.ai with a Management key but the same paths answer on api.x.ai with an inference key (records carry both ACCOUNT_RESTRICTED and LIVE_VERIFIED); POST /v1/collections/{id}/documents multipart → 405; batch responses requests come back as chat_get_completion; deferred completions documented "retrievable exactly once" but a second GET returned 200; GET /v1/responses/{id} 200 for store:false; /v1/completions marked RETIRED in endpoints.json while docs/xai/legacy-completions.md says LEGACY "retired in practice"; the 4.20 non-reasoning model also rejects raw sampling although docs say "non-reasoning only".
  22. xAI rate limits / headers: documented Tier 0 RPS 150 / 37 vs observed x-ratelimit-limit-requests 7,200 / 1,800 per minute (= 120 / 30 RPS); the headers themselves are undocumented; STT default model is grok-voice-transcribe-1.0 in release notes and 2.0 on the model page (DOCUMENTATION_INCOMPLETE); realtime default voice xai_ara live vs eve in docs; response.done.usage and ping undocumented; response.audio.delta documented but not emitted; GET /v1/api-key create_time/modify_time returned empty strings; grok-imagine-video-1.5 live input_modalities include audio (model page: text, image); video pending polls return HTTP 202 (undocumented); eu-west-1.api.x.ai undocumented but live.
  23. Gemini pricing.json: 5 null-price rows (veo-3.1-lite 4k, gemini-embedding-001 input, tool:custom_tools_endpoint, agent:deep-research, service:document_tokens); 2 non-numeric prices ('1.8x', '0.50-8.10'); duplicate keys with conflicting values — tool:google_search standard = $14 "per 1K requests" and $35 "per 1K grounded prompts", tool:google_maps = $14 "per 1K search queries" and $25 "per 1K grounded prompts" (the 3.x vs 2.5 rows are not disambiguated in the key); Gemma rows use tier: free with per-dimension rows while other models use one free_tier dimension.
  24. Gemini models.json: knowledge_cutoff null for 64/97, context_window null 28, max_output null 33, release_date null 13; gemini-3.5-flash flex cached_input 0.08 vs batch 0.075 (not 50 % of 0.15); gemini-3.5-flash-lite batch/flex cached_input 0.02 not on the pricing page; gemini-2.5-pro/-flash carry ACCOUNT_RESTRICTED in deprecations.json but not in models.json (only -flash-lite does); thinking flag true live on gemini-3.5-transcribe and gemini-3.1-flash-tts-preview (docs: not supported) and absent on Live/robotics-streaming models (docs: supported).
  25. Gemini parameters.json: 11 duplicate (endpoint, parameter) pairs from different fragments with conflicting enums — generationConfig.thinkingConfig.thinkingLevel MINIMAL|LOW|MEDIUM|HIGH vs MINIMAL|HIGH (image models), responseModalities TEXT|IMAGE|AUDIO vs AUDIO, Interactions response_format ×3, generation_config.thinking_level ×2, transcription_config ×2; generationConfig._responseJsonSchema recorded as LIVE_DISCOVERED.
  26. Gemini endpoints.json: the core POST /v1beta/models/{model}:generateContent record sits under api_family: transcription (the generate-content family holds only /v1/…, streamGenerateContent and dynamic/{id}); dynamic endpoints duplicated across two families with different path spellings and statuses; Interactions /v1/* tagged GA but UNVERIFIED; webhooks documented on /v1 while everything else is /v1beta.
  27. Gemini streaming-events.json: the 7 lyria-realtime records have null status although docs/gemini/music-generation.md marks the protocol LIVE_VERIFIED.
  28. Gemini docs vs docs / live: Maps pricing $25/1k grounded prompts (tools table) vs $14/1k queries (Gemini 3 tables) → DOCUMENTATION_INCOMPLETE; File Search indexing $0.15/1M vs gemini-embedding-2 text $0.20/1M on the same page; priority "75–100 % more" (prose) vs 1.8× (tables); Batch/Flex "not available" on free tier for most models but "free of charge" for gemini-3.5-flash-lite; 3.5-flash caching "free of charge" on the page vs live limit: 0; explicit-cache minimum 4,096 (docs) vs 1,024 (live); audio token rate 32 tok/s (tokens guide) vs 25 tok/s (Live/pricing); image token counts 280/560/1,120/2,240 (docs) vs 256/529/1,089/2,209 (observed); PDF page 520 IMAGE tokens (generateContent) vs 560 DOCUMENT tokens (countTokens); token limits live vs docs for gemini-3.1-flash-image (65,536/65,536 vs 131,072/32,768), -lite-image (out 65,536 vs 4,096), gemini-3-pro-image (in 131,072 vs 65,536), gemini-2.5-flash-image (in 32,768 vs 65,536), lyria-3* (1,048,576 vs 131,072), deep-research-* (131,072 vs 1,048,576), gemini-3.5-live-translate-preview (16,384/32,768 vs 131,072/65,536), 2.5 computer use (131,072/65,536 vs 128,000/64,000); docs say every model is in /v1 and /v1beta (live 22 vs 58); gemini-3-pro-preview computer use "not supported" (model page) vs "launched" (changelog); interaction.status_update still emitted although the breaking-changes guide says superseded; transcription parts[].audioTranscription undocumented, speaker label spk:0 (REST) vs spk_1 (Interactions docs); Google Search ACCOUNT_RESTRICTED on this free-tier key although the pricing page lists 500 RPD free for 2.5 Flash/Flash-Lite; gemma-4-* absent from docs/models tables; deprecations.json replacement text differs from the docs table for gemini-2.5-flash-image and gemini-2.0-flash.
  29. Cross-provider: docs/errors/xai.md does not exist (xAI errors live in docs/xai/authentication-headers-errors.md) while docs/errors/{openai,anthropic,gemini}.md do; sdks.json has version: null on every xAI/Gemini record although the docs pages quote versions (xai-sdk 1.19, google-genai 2.24, @google/genai 2.23); models.json uses four different kind vocabularies (OpenAI model|snapshot|id_only|alias, Anthropic snapshot|alias, xAI model|legacy|retired_redirect|alias|service, Gemini stable|preview|alias|agent|experimental).

Related: index · features · pricing · caching and reasoning · FAQ.