Models — OpenAI ↔ Anthropic ↔ xAI ↔ Gemini equivalence and use-case table
Status: every value is copied from generated/models.json (statuses, context, output, cutoffs, capability flags, pricing blocks; live-verified on 2026-09-18/19 where marked), generated/pricing.json and generated/deprecations.json. Equivalence rows are use-case pairings by price/tier and capability, not benchmark claims. xAI probes ran on a Tier 0 team, Gemini probes on a free-tier key (paid-only Gemini models are ACCOUNT_RESTRICTED, not absent).
Sources: https://developers.openai.com/api/docs/models · https://developers.openai.com/api/docs/pricing · https://developers.openai.com/api/docs/deprecations · https://platform.claude.com/docs/en/models/overview · https://platform.claude.com/docs/en/about-claude/pricing · https://platform.claude.com/docs/en/about-claude/model-deprecations · https://docs.x.ai/developers/models · https://docs.x.ai/developers/pricing · https://docs.x.ai/developers/migration/may-15-retirement · https://ai.google.dev/gemini-api/docs/models · https://ai.google.dev/gemini-api/docs/pricing · https://ai.google.dev/gemini-api/docs/deprecations · docs/models/openai-models.md · docs/models/anthropic-models.md · docs/models/xai-models.md · docs/models/gemini-models.md
Last verified: 2026-09-18
1. Use-case pairing (current models, four providers)
Prices USD per 1M tokens, standard tier: in / cached-in / out. "Cached-in" is a cache read. Context = documented window; Out = documented max output. xAI prices double for requests whose prompt is ≥200k tokens (all tokens of that request); Gemini Pro prices rise for prompts >200k; Gemini 3.6–3.8 Flash prices are introductory (double on 2027-01-01).
| Use case | OpenAI | Anthropic | xAI | Gemini | Why they pair |
|---|---|---|---|---|---|
| Frontier reasoning, hardest tasks | gpt-6-astra — 1,050k ctx / 128k out · text+image → text · effort low…max (no none) · $10 / $1 / $50 · cutoff 2026-04-30 · rel. 2026-09-03 · LIVE_VERIFIED |
claude-fable-5-1 — 1M / 128k · text+image+PDF → text · adaptive thinking always on, effort low…max · $10 / $0.25 / $50 · cutoff Jun 2026 · rel. 2026-09-01 · LIVE_VERIFIED |
grok-4.6 — 500k / no published limit · text+image → text · reasoning always on, effort low…xhigh (default high) · $2 / $0.50 / $6 (≥200k: $4 / $1 / $12) · cutoff 2026-02-01 · rel. 2026-08 · LIVE_VERIFIED |
gemini-3.1-pro-preview — 1,048,576 / 65,536 · text+image+audio+video+PDF → text · thinkingLevel low/medium/high (default high, no minimal) · $2 / $0.20 / $12 (>200k: $4 / $0.40 / $18) · cutoff Jan 2025 (model card) · rel. 2026-02-19 · PREVIEW · ACCOUNT_RESTRICTED (paid tier only) |
Each provider's top reasoning line. Price spread is 5× ($2 vs $10 in). Astra and Fable are always-on reasoners; Grok 4.6 too; Gemini 3.1 Pro cannot go below low. Only Gemini's is still a preview. |
| Flagship general / agentic | gpt-5.6-sol (alias gpt-5.6) — 1,050k / 128k · effort none…max, reasoning.mode: pro · $4 / $0.40 / $20 (promo ≥ 2026-11-21) · cutoff 2026-02-16 · rel. 2026-07-09 |
claude-opus-5 — 1M / 128k · adaptive (on by default, can disable except xhigh/max) · $5 / $0.50 / $25 · cutoff May 2026 · rel. 2026-07-24 · fast mode $10 / $50 |
grok-4.5 — 500k · reasoning always on, effort low…high (xhigh treated as high) · $2 / $0.30 / $6 · rel. 2026-07 · priority 2× · no batch · aliases grok-4.5-latest, grok-build-latest |
gemini-3.8-flash — 1,048,576 / 65,536 · thinking low/medium/high (default medium; minimal → 400) · $0.75 / $0.075 / $3.75 (→ $1.50 / $0.15 / $7.50 in 2027) · rel. 2026-09-02 · GA · LIVE_VERIFIED · free tier |
Newest general-purpose line per provider. Gemini's Flash line is the current flagship for the Developer API (Pro is preview); Grok 4.5/4.6 sit between OpenAI's mini and flagship prices. |
| Balanced mid-tier | gpt-5.6-terra — 1,050k / 128k · $2 / $0.20 / $12; also gpt-5.5 $5 / $0.50 / $30, gpt-5.4 $2.5 / $0.25 / $15 |
claude-sonnet-5 — 1M / 128k · adaptive on by default · $2 / $0.20 / $10 · cutoff Jan 2026 · rel. 2026-06-30 |
grok-4.3 — 1M · reasoning optional (reasoning_effort: none LIVE_DISCOVERED → 0 reasoning tokens; default low) · $1.25 / $0.20 / $2.50 (≥200k ×2) · batch −20 % · rel. 2026-04 · redirect target of grok-3, grok-4-0709, grok-4-fast-* |
gemini-3.5-flash — 1,048,576 / 65,536 · thinking minimal…high (default medium; budget 0 disables) · $1.50 / $0.15 / $9 · rel. 2026-05-19 · GA · LIVE_VERIFIED |
Identical $2 input on Terra / Sonnet 5 / 3.1 Pro; grok-4.3 and gemini-3.5-flash undercut them. All support the full tool set of their provider (§3–§5). |
| Small / high-volume | gpt-5.6-luna — $0.20 / $0.02 / $1.20; gpt-5.4-mini — 400k / 128k · $0.75 / $0.075 / $4.50; gpt-5.4-nano — $0.20 / $0.02 / $1.25 |
claude-haiku-4-5-20251001 (alias claude-haiku-4-5) — 200k / 64k · manual extended thinking only, no effort · $1 / $0.10 / $5 · cutoff Feb 2025 · rel. 2025-10-15 |
grok-build-0.1 — 256k · $1 / $0.20 / $2 · PREVIEW ('early access'); grok-4.20-0309-non-reasoning — 1M · no reasoning · $1.25 / $0.20 / $2.50 · batch −20 % |
gemini-3.5-flash-lite — 1,048,576 / 65,536 · thinking default minimal · $0.30 / $0.03 / $2.50 · rel. 2026-07-21 · GA · LIVE_VERIFIED · free tier; gemini-3.6-flash / -3.7-flash $0.75 / $0.075 / $3.75 (introductory) |
Cheapest current text model per provider. OpenAI nano/Luna and Gemini Flash-Lite are ~5× cheaper than Haiku 4.5; xAI has no sub-$1 text model. |
| Deep-compute "pro" | gpt-5.5-pro / gpt-5.4-pro — 1,050k / 128k · effort medium…xhigh · no prompt caching, no fast · $30 / — / $180; gpt-5.6-sol with reasoning.mode: pro (standard rates) |
no pro SKU — output_config.effort: max on Opus 5 / Fable 5.1 |
grok-4.20-multi-agent-0309 — 1M · one Responses call fans out to 4 or 16 agents (reasoning.effort low/medium vs high/xhigh) · $1.25 / $0.20 / $2.50 per agent-token · BETA · Responses-only (Chat → 400) |
Deep Research agents deep-research-preview-04-2026 / -max- (Interactions, background: true, ≤60 min, est. $1–3 / $3–7 per task, list-rate tokens + tools) · PREVIEW; gemini-3.8-live-extended-thinking for voice |
OpenAI sells extra compute as a model/mode, xAI as a multi-agent model, Gemini as an agent, Anthropic as the top effort value. |
| Coding agents | gpt-5.3-codex — 400k / 128k · $1.75 / $0.175 / $14 · standard + fast only; GPT-5.6 family lists hosted shell / apply_patch / skills |
claude-opus-5, claude-sonnet-5 (bash, text_editor, code_execution, skills, computer/browser toolsets, programmatic tool calling); claude-opus-4-8 legacy |
grok-build-0.1 (grok-code-fast-1 redirect; 256k; $1 / $0.20 / $2; PREVIEW) + Grok Build CLI (grok, BETA: TUI, headless, ACP, MCP, hooks, skills, sandbox) |
antigravity-preview-09-2026 (Interactions coding agent with Linux sandbox, default model gemini-3.8-flash, max_total_tokens) · PREVIEW; gemini-3.1-pro-preview-customtools (bash-style custom tools, 3.1 Pro price) |
Only xAI and Gemini ship a coding-specific model id; OpenAI ships codex; Anthropic ships coding tools on its general models. |
| Previous-generation flagships still served | gpt-5.4 ($2.5 / $0.25 / $15), gpt-5.2 ($1.75 / $0.175 / $14), gpt-5.1 & gpt-5 ($1.25 / $0.125 / $10), o3 ($2 / $0.50 / $8), o3-pro ($20 / — / $80) |
claude-opus-4-8, -4-7, -4-6 ($5 / $0.50 / $25, LEGACY), claude-opus-4-5-20251101 ($5, 200k, LEGACY), claude-sonnet-4-6 ($3 / $0.30 / $15, LEGACY), claude-sonnet-4-5-20250929 ($3, 200k, LEGACY) |
grok-4.20-0309-reasoning — 1M · reasoning always on (reasoning_effort → 400) · $1.25 / $0.20 / $2.50 · batch −20 % · 15 aliases incl. grok-4.20, grok-4.20-beta*, grok-4.20-reasoning-gv2 |
gemini-3-flash-preview ($0.50 / $0.05 / $3, PREVIEW, replacement 3.6 Flash, no date), gemini-3.1-flash-lite ($0.25 / $0.025 / $1.50, DEPRECATED → 2027-05-07), gemini-2.5-pro / -flash / -flash-lite (priced and listed but 404 'no longer available to new users') |
OpenAI keeps old flagships "Active" until a dated deprecation; Anthropic marks "Active (legacy)"; xAI keeps 4.20 on the price list; Gemini keeps 2.5 listed while blocking new users. |
| Deprecated but callable | o4-mini, o3-mini, o1, o1-pro, gpt-4.1-nano (2026-10-23), gpt-5-*-2025-08-07, o3-2025-04-16, o3-pro-2025-06-10 (2026-12-11) |
claude-mythos-preview (invite, replacement claude-mythos-5, date TBA) |
grok-imagine-image-quality (→ 2026-11-02, redirects to grok-imagine-image-2.0 low); /v1/messages compat surface |
gemini-3.1-flash-lite (→ 2027-05-07), gemini-2.5-flash-image (→ 2026-10-02), gemini-omni-flash-preview (→ 2026-09-30), antigravity-preview-05-2026 (→ 2026-10-05), gemini-embedding-001 (→ 2028-05-14) |
see §7 |
| Non-reasoning models | gpt-4.1 (1,047k / 32k, $2 / $0.50 / $8, fine-tunable), gpt-4.1-mini, gpt-4o (128k / 16k), gpt-4o-mini, chat-latest (400k, $5 / $0.50 / $30) |
none current — every active Claude model reasons (Haiku/Sonnet 4.5 optionally) | grok-4.20-0309-non-reasoning (no reasoning_content, reasoning_tokens: 0); grok-4.3 with reasoning_effort: none |
none — every Gemini 3.x text model thinks (minimal on Flash-Lite / 3.6 / 3.5 / 3 Flash; thinkingBudget: 0 on 3.6 / 3.5 Flash only); Gemma 4 live returned thoughtsTokenCount: 5 |
OpenAI and xAI keep explicit non-reasoning ids; Anthropic and Gemini only expose reasoning-off switches on some models. |
| Gated / invite / paid-only | gpt-5.6-cyber, gpt-5.5-cyber ($12.5 / $1.25 / $75), gpt-daybreak-blue-latest, gpt-daybreak-red-latest, gpt-rosalind-research (ACCOUNT_RESTRICTED) |
claude-mythos-5-1, claude-mythos-5 ($10 / $0.25–$1 / $50, Project Glasswing, ACCOUNT_RESTRICTED, 404 for our key) |
grok-embedding-small (404, no price), Skills API (404), tool_search (403 alpha), custom-voice creation (Enterprise), Management API (needs a Management key) — all ACCOUNT_RESTRICTED |
Pro models and every image / video / music model are paid-tier only (limit: 0 on the free tier): gemini-3.1-pro-preview, gemini-pro-latest, gemini-3.1-flash-image, gemini-3-pro-image, Veo 3.1, Lyria 3.x; cachedContents and Batch also 429/400 on free |
Different gates: invite programs (OpenAI, Anthropic), ACLs/alpha (xAI), billing tier (Gemini). |
| Free of charge | none | none | none (prepaid credits) | free-tier rows on 3.8/3.7/3.6/3.5 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite, 3 Flash Preview, Live 3.8 / 3.1, TTS 3.1 / 2.5 Flash, gemini-3.5-transcribe, gemini-embedding-2; Gemma 4 (gemma-4-31b-it, gemma-4-26b-a4b-it, 262,144 / 32,768, text only) is free-only; content may be used to improve Google products |
Gemini is the only provider with a documented $0 tier. |
2. Capability table — current OpenAI text models
From capabilities in generated/models.json. ✔ true · ✘ false · ? "unknown". Effort = accepted reasoning.effort values (default in bold when recorded). Cache = prompt caching / 24 h retention. Tiers = service tiers listed as available (S standard, B batch, X flex, F fast).
| Model | Status | Ctx | Out | Cutoff | Modalities | Effort | SO | FC | Cache / 24h | Web | Code | MCP | Comp. | Shell | Patch | Skills | ToolSearch | FT | Tiers |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
gpt-6-astra |
DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2026-04-30 | text, image → text | low, medium, high, xhigh, max | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F |
gpt-5.6-sol (= gpt-5.6) |
DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2026-02-16 | text, image → text | none, low, medium, high, xhigh, max · mode pro | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F |
gpt-5.6-terra |
DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2026-02-16 | text, image → text | none…max · pro | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F |
gpt-5.6-luna |
DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2026-02-16 | text, image → text | none…max · pro | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F |
gpt-5.5 |
DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2025-12-01 | text, image → text | none, low, medium, high, xhigh | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F |
gpt-5.5-pro |
DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2025-12-01 | text, image → text | medium, high, xhigh | ✔ | ✔ | ✘ / ✔ | ✔ | ✔ | ✔ | ✘ | ✔ | ✘ | ✘ | ✘ | ✘ | S B X |
gpt-5.4 |
DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2025-08-31 | text, image → text | none, low, medium, high, xhigh | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F |
gpt-5.4-pro |
DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2025-08-31 | text, image → text | medium, high, xhigh | ✘ | ✔ | ✘ / ✘ | ✔ | ✘ | ✔ | ✔ | ✘ | ✔ | ✘ | ✔ | ✘ | S B X |
gpt-5.4-mini |
DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | none, low, medium, high, xhigh | ✔ | ✔ | ✔ / ? | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F |
gpt-5.4-nano |
DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | none, low, medium, high, xhigh | ✔ | ✔ | ✔ / ? | ✔ | ✔ | ✔ | ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | S B X |
gpt-5.3-codex |
DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | low, medium, high, xhigh | ✔ | ✔ | ✔ / ? | ✔ | ✘ | ✘ | ✘ | ✔ | ✘ | ✔ | ✘ | ✘ | S F |
gpt-5.2 |
DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | none, low, medium, high, xhigh | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | S B X F |
gpt-5.2-pro |
DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | (model default) | ✘ | ✔ | ✘ / ✘ | ✔ | ✘ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B |
gpt-5.1 |
DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2024-09-30 | text, image → text | none, low, medium, high | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✘ | ✘ | ✔ | ✘ | ✘ | ✘ | S B X F |
gpt-5 |
DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2024-09-30 | text, image → text | minimal, low, medium, high | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B X F |
gpt-5-mini |
DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2024-05-31 | text, image → text | (default) | ✔ | ✔ | ✘ / ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B X F |
gpt-5-nano |
DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2024-05-31 | text, image → text | (default) | ✔ | ✔ | ✔ / ? | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B X |
gpt-5-pro |
DOCUMENTED · LIVE_VERIFIED | 400,000 | 272,000 | 2024-09-30 | text, image → text | high | ✔ | ✔ | ✘ / ✘ | ✔ | ✘ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B |
o3 |
DOCUMENTED · LIVE_VERIFIED | 200,000 | 100,000 | 2024-06-01 | text, image → text | (default) | ✔ | ✔ | ✔ / ? | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B X F |
o3-pro |
DOCUMENTED · LIVE_VERIFIED | 200,000 | 100,000 | 2024-06-01 | text, image → text | (default) | ✔ | ✔ | ✘ / ✘ | ✔ | ✘ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B |
o4-mini |
DOCUMENTED · LIVE_VERIFIED · DEPRECATED (2026-10-23) | 200,000 | 100,000 | 2024-06-01 | text, image → text | (default) | ✔ | ✔ | ✔ / ? | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✔ | S B X F |
gpt-4.1 |
DOCUMENTED · LIVE_VERIFIED | 1,047,576 | 32,768 | 2024-06-01 | text, image → text | non-reasoning | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✔ | S B F |
gpt-4.1-mini |
DOCUMENTED · LIVE_VERIFIED | 1,047,576 | 32,768 | 2024-06-01 | text, image → text | non-reasoning | ✔ | ✔ | ✘ / ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✔ | S B F |
gpt-4o / gpt-4o-mini |
DOCUMENTED · LIVE_VERIFIED | 128,000 | 16,384 | 2023-10-01 | text, image → text | non-reasoning | ✔ | ✔ | ✘ / ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✔ | S B F |
chat-latest |
DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | non-reasoning | ✔ | ✔ | ✘ / ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S |
Data note: for the GPT-5.6 family, GPT-6 Astra and the Daybreak/cyber ids the records have function_calling: true and tool_function_calling: false — the second flag was read from a model-page tool list that omits "function calling" for those pages; treat function_calling as authoritative (function tools were used live on gpt-5.6-luna in the Agents API run).
3. Capability table — current Anthropic models
From capabilities and pricing in generated/models.json. Cache min = minimum cacheable prefix tokens. Prefill = assistant prefill accepted. Sampling = temperature/top_p/top_k behaviour. ZDR = zero-data-retention eligible.
| Model | Status | Lifecycle | Ctx | Out | Cutoff (reliable / training) | Thinking | Effort (default high) | SO / strict | Cache min · 1h | Fast | tool_choice any/tool | Prefill | Sampling | Web / Fetch / Code / MCP / PTC | Computer toolset GA / Browser | ZDR | Retirement not before |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
claude-fable-5-1 |
DOCUMENTED · LIVE_VERIFIED | Active (latest) | 1,000,000 | 128,000 | Jun 2026 / Jun 2026 | adaptive, always on | low, medium, high, xhigh, max · per-message (beta) | ✔ / ✔ | 512 · ✔ | ✘ | ✘ (400) | ✘ | 400 if non-default | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✘ | 2027-09-01 |
claude-mythos-5-1 |
DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW | Active (invite) | 1,000,000 | 128,000 | Jun 2026 | adaptive, always on | low…max · per-message | ✔ / ✔ | 512 · ✔ | ✘ | ✘ | ✘ | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✘ | 2027-09-01 |
claude-fable-5 |
DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 1,000,000 | 128,000 | Jan 2026 | adaptive, always on | low…max | ✔ / ✔ | 512 · ✔ | ✘ | ✔ | ✘ | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✘ | 2027-06-09 |
claude-mythos-5 |
DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW | Active (invite) | 1,000,000 | 128,000 | Jan 2026 | adaptive, always on | low…max | ✔ / ✔ | 512 · ✔ | ✘ | ✔ | ✘ | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✘ | 2027-06-09 |
claude-opus-5 |
DOCUMENTED · LIVE_VERIFIED | Active (latest) | 1,000,000 | 128,000 | May 2026 | adaptive, on by default (disable except xhigh/max) | low…max · per-message | ✔ / ✔ | 512 · ✔ | ✔ $10/$50 | ✔ | ✘ | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✔ | 2027-07-24 |
claude-opus-4-8 |
DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 1,000,000 | 128,000 | Jan 2026 | adaptive (off by default) | low…max | ✔ / ✔ | 1,024 · ✔ | ✔ $10/$50 | ✔ | ? | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✔ | 2027-05-28 |
claude-opus-4-7 |
DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 1,000,000 | 128,000 | Jan 2026 | adaptive | low…max | ✔ / ✔ | 2,048 · ✔ | ✘ (removed) | ✔ | ? | 400 | ✔ ✔ ✔ ✔ ✔ | ✘ (beta computer_20251124) / ✘ |
✔ | 2027-04-16 |
claude-opus-4-6 |
DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 1,000,000 | 128,000 | May 2025 / Aug 2025 | adaptive; manual enabled deprecated |
low, medium, high, max | ✔ / ✔ | 4,096 · ✔ | ✘ (runs standard) | ✔ | ✘ | accepted (not with thinking) | ✔ ✔ ✔ ✔ ✔ | ✘ / ✘ | ✔ | 2027-02-05 |
claude-opus-4-5-20251101 (alias claude-opus-4-5) |
DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 200,000 | 64,000 | May 2025 / Aug 2025 | extended (manual budget_tokens) |
low, medium, high | ✔ / ✔ | 4,096 · ✔ | ✘ | ✔ | ? | accepted | ✔ ✔ ✔ ✔ ✔ | ✘ / ✘ | ✔ | 2026-11-24 |
claude-sonnet-5 |
DOCUMENTED · LIVE_VERIFIED | Active (latest) | 1,000,000 | 128,000 | Jan 2026 | adaptive, on by default | low…max | ✔ / ✔ | 1,024 · ✔ | ✘ | ✔ | ✘ (400 live) | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✔ | 2027-06-30 |
claude-sonnet-4-6 |
DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 1,000,000 | 128,000 | Aug 2025 / Jan 2026 | adaptive; manual deprecated | low, medium, high, max | ✔ / ✔ | 1,024 · ✔ | ✘ | ✔ | ? | accepted | ✔ ✔ ✔ ✔ ✔ | ✘ / ✘ | ✔ | 2027-02-17 |
claude-sonnet-4-5-20250929 (alias claude-sonnet-4-5) |
DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 200,000 | 64,000 | Jan 2025 / Jul 2025 | extended (manual) | ✘ (400) | ✔ / ✔ | 1,024 · ✔ | ✘ | ✔ | ? | accepted | ✔ ✔ ✔ ✔ ✔ | ✘ / ✘ | ✔ | 2026-09-29 |
claude-haiku-4-5-20251001 (alias claude-haiku-4-5) |
DOCUMENTED · LIVE_VERIFIED | Active (latest) | 200,000 | 64,000 | Feb 2025 / Jul 2025 | extended (manual) | ✘ (400) | ✔ / ✔ | 4,096 · ✔ | ✘ | ✔ | ✔ (live) | accepted | ✔ ✔ ✔ ✔ ✘ (PTC 400) | ✘ / ✘ | ✔ | 2026-10-15 |
Modalities for every Claude model above: text + image + PDF in → text out. Tool support per exact tool type (e.g. web_search_20260318 vs 20250305) is in generated/compatibility/model-tool-matrix.json.
4. Capability table — current xAI Grok models
From capabilities, pricing and rate_limits in generated/models.json (kind: model). Reasoning = always-on / optional / none and accepted reasoning_effort values (default in bold). Cache = automatic prefix cache (no markers). Tools = all server tools (web_search, x_search, code_interpreter, file_search, mcp, image_generation, attachment search) — tools.json lists every Grok text model as compatible. Prices standard (< 200k prompt) in / cached / out; ≥200k = 2× on all tokens. Batch = 20 % discount where supported. Priority = service_tier: priority 2×. RPS/TPM = documented Tier 0 → Tier 4.
| Model | Status | Ctx | Out | Cutoff | Release | Modalities | Reasoning · effort | SO | FC | Cache | Tools | Compact / WS Responses | Batch | Priority | US 1.1× | In / cached / out | RPS · TPM (T0→T4) | Aliases |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
grok-4.6 |
DOCUMENTED · LIVE_VERIFIED | 500,000 | none published | 2026-02-01 | 2026-08 | text, image → text | always on · low, medium, high, xhigh | ✔ | ✔ | ✔ (0.25×) | ✔ all | ✔ / ✔ | ✘ | 2× | ✔ (only model) | 2.00 / 0.50 / 6.00 | 150→500 · 50M→100M | — |
grok-4.5 |
DOCUMENTED · LIVE_VERIFIED | 500,000 | null | null | 2026-07 | text, image → text | always on · low, medium, high (xhigh = high) |
✔ | ✔ | ✔ (0.15×) | ✔ all | ✔ / ✔ | ✘ | 2× | ✘ | 2.00 / 0.30 / 6.00 | 150→500 · 50M→100M | grok-4.5-latest, grok-build-latest |
grok-4.3 |
DOCUMENTED · LIVE_VERIFIED | 1,000,000 | null | null | 2026-04 | text, image → text | optional · low, medium, high, xhigh + none (LIVE_DISCOVERED) |
✔ | ✔ | ✔ (0.16×) | ✔ all | ✔ / ✔ | −20 % | 2× | ✘ (eu-west-1 host) | 1.25 / 0.20 / 2.50 | 37→208 · 10M→85M | grok-4.3-latest; redirects grok-3, grok-4-0709, grok-4-fast-*, grok-4-1-fast-* |
grok-4.20-0309-reasoning |
DOCUMENTED · LIVE_VERIFIED | 1,000,000 | null | null | 2026-03-09 | text, image → text | always on · reasoning_effort → 400 |
✔ | ✔ | ✔ | ✔ all | ✔ / ✔ | −20 % | 2× | ✘ | 1.25 / 0.20 / 2.50 | 37→208 · 10M→85M | grok-4.20, grok-4.20-reasoning(-latest), grok-4.20-0309, grok-4.20-beta*, grok-4.20-experimental-beta*, grok-4.20-reasoning-gv2 (15) |
grok-4.20-0309-non-reasoning |
DOCUMENTED · LIVE_VERIFIED | 1,000,000 | null | null | 2026-03-09 | text, image → text | none (0 reasoning tokens) | ✔ | ✔ | ✔ | ✔ all | ✔ / ✔ | −20 % | 2× | ✘ | 1.25 / 0.20 / 2.50 | 37→208 · 10M→85M | grok-4.20-non-reasoning(-latest), -beta-*, -gv2 (8) |
grok-4.20-multi-agent-0309 |
DOCUMENTED · BETA · LIVE_VERIFIED | 1,000,000 | null | null | 2026-03 | text, image → text | multi-agent · effort = agent count (low/medium 4, high/xhigh 16; default medium) | ✔ | ✔ | ✔ | ✔ all | Responses only (Chat → 400) | −20 % | not documented | ✘ | 1.25 / 0.20 / 2.50 (per agent token) | 9→56 · 2.5M→21M | grok-4.20-multi-agent(-latest), -beta-*, -experimental-beta-* (6) |
grok-build-0.1 |
DOCUMENTED · PREVIEW · LIVE_VERIFIED | 256,000 | null | null | 2026-05 ('early access') | text, image → text | always on · reasoning_effort → 400 |
✔ | ✔ | ✔ (0.2×) | ✔ all | ✔ / ✔ | ✘ | 2× | ✘ | 1.00 / 0.20 / 2.00 | 37→208 · 10M→85M | grok-code-fast-1, grok-code-fast, grok-code-fast-1-0825 |
Every Grok text model also accepts the Anthropic-compatible /v1/messages (DEPRECATED), POST /v1/responses/compact, wss://api.x.ai/v1/responses, deferred Chat Completions and /v1/tokenize-text. usage.cost_in_usd_ticks (1 USD = 10^10 ticks) is returned on every call. Fine-tuning: none offered. logprobs: ignored on 4.20+.
4b. xAI media, voice and embedding models
| Model | Status | Modalities | Price | Endpoints | Notes |
|---|---|---|---|---|---|
grok-imagine-image |
DOCUMENTED · LIVE_VERIFIED | text+image → image | $0.02 / image | POST /v1/images/generations, /edits (JSON only), Batch, gRPC |
prompt limit 16,000 tokens; alias grok-imagine-image-2026-03-02; JPEG only |
grok-imagine-image-2.0 |
DOCUMENTED · LIVE_VERIFIED | text+image → image | $0.04–$0.08 (quality low/medium × resolution 1k/1.5k/2k; catalogue default $0.06) | same | 5 reference images, 21:9 / 5:2 aspect ratios (Aug 2026); quality default moved to auto |
grok-imagine-image-quality |
DOCUMENTED · DEPRECATED · LIVE_VERIFIED | text+image → image | $0.05 | same (no Batch) | retires 2026-11-02 → grok-imagine-image-2.0 quality=low; aliases -20260403, -latest, grok-imagine-image-pro |
grok-imagine-video |
DOCUMENTED · LIVE_VERIFIED | text+image+video → video | $0.05 / s | POST /v1/videos/generations, /edits, /extensions, GET /v1/videos/{request_id} (202 pending) |
1–15 s, 480p–1080p, audio; batch URLs expire 1 h |
grok-imagine-video-1.5 |
DOCUMENTED · LIVE_VERIFIED | text+image+audio → video | $0.08 / s | same | native 1080p, reference-to-video, reference_audios, last_frame; aliases -preview, -2026-05-30 |
grok-voice-think-fast-2.0 (grok-voice-latest) |
DOCUMENTED | audio+text ↔ audio+text | $0.08 / min ($4.80 / h) + $0.004 per text item | wss://api.x.ai/v1/realtime, POST /v1/realtime/client_secrets, SIP /v2/phone-numbers, /v1/realtime/calls/{id}/refer and …/hangup |
OpenAI Realtime event vocabulary; tools in-session; 120-min sessions; -1.0 LEGACY; not in GET /v1/models |
grok-voice-transcribe-2.0 / -1.0 |
DOCUMENTED | audio → text | $0.10 / h REST, $0.20 / h streaming | POST /v1/stt, wss://api.x.ai/v1/stt |
25 languages, diarization, smart_turn; default model ambiguous in docs |
| text-to-speech (no model id) | DOCUMENTED | text → audio | $15 / 1M characters | POST /v1/tts, GET /v1/tts/voices, wss://api.x.ai/v1/tts |
28 voices, language required; custom voices Enterprise |
grok-embedding-small |
DOCUMENTED · ACCOUNT_RESTRICTED | text → embedding | not published | POST /v1/embeddings, GET /v1/embedding-models (empty) |
404 for this team; Collections index model |
5. Capability table — current Gemini models
From capabilities, pricing and deprecation in generated/models.json (kinds stable, preview, alias, agent). Thinking = thinkingLevel values (default in bold; budget = 2.5-style thinkingBudget). Tools: FC function calling, Search (googleSearch), Maps, URL (urlContext), Code (codeExecution), CU (computerUse, preview), FS (fileSearch). Cache = implicit / explicit (cachedContents). Tiers = B batch (50 %), X flex (50 %), P priority (1.8×). Prices standard in / cached / out (>200k tier for Pro). Free = free-tier availability per the pricing page.
| Model | Status | Lifecycle | Release | Ctx in / out | Modalities | Thinking | SO | FC | Search / Maps / URL / Code / CU / FS | Cache impl / expl | Tiers | Live | Free | In / cached / out | Shutdown |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
gemini-3.8-flash |
DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED | Stable (GA 2026-09-02); gemini-flash-latest → here (observed) |
2026-09-02 | 1,048,576 / 65,536 | text, image, audio, video, PDF → text | low, medium, high (minimal → 400) |
✔ | ✔ | ✔ ✔ ✔ ✔ ✔(P) ✔ | 4,096 / ✔ | B X P | ✘ | ✔ | 0.75 / 0.075 / 3.75 (→ 1.50 / 0.15 / 7.50 on 2027-01-01) | none announced |
gemini-3.7-flash |
DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-08-13) | 2026-08-13 | 1,048,576 / 65,536 | same | low, medium, high | ✔ | ✔ | ✔ ✔ ✔ ✔ ✔(P) ✔ | 4,096 / ✔ | B X P | ✘ | ✔ | 0.75 / 0.075 / 3.75 (intro) | none |
gemini-3.6-flash |
DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-07-21) | 2026-07-21 | 1,048,576 / 65,536 | same | minimal, low, medium, high; thinkingBudget: 0 allowed |
✔ | ✔ | ✔ ✔ ✔ ✔ ✔(P) ✔ | 4,096 / ✔ | B X P | ✘ | ✔ | 0.75 / 0.075 / 3.75 (intro) | none |
gemini-3.5-flash |
DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED | Stable (GA 2026-05-19) | 2026-05-19 | 1,048,576 / 65,536 | same | minimal, low, medium, high; budget 0 allowed | ✔ | ✔ | ✔ ✔ ✔ ✔ ✔(P) ✔ | 4,096 / ✔ | B X P | ✘ | ✔ (caching limit: 0 live) |
1.50 / 0.15 / 9.00 | none |
gemini-3.5-flash-lite |
DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED | Stable (GA 2026-07-21) | 2026-07-21 | 1,048,576 / 65,536 | same | minimal, low, medium, high (budget 0 → 400) | ✔ | ✔ | ✔ ✔ ✔ ✔ ✔(P) ✔ | ? / ✔ | B X P | ✘ | ✔ (no caching) | 0.30 / 0.03 / 2.50 | none |
gemini-3.1-pro-preview (+ -customtools) |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW · ACCOUNT_RESTRICTED | Preview; gemini-pro-latest → here |
2026-02-19 | 1,048,576 / 65,536 | same | low, medium, high | ✔ | ✔ | ✔ ✔ ✔ ✔ ✘ ✔ (AI Studio only) | 4,096 / ✔ | B X P | ✘ | ✘ (paid only) | 2.00 / 0.20 / 12.00; >200k 4.00 / 0.40 / 18.00 | none |
gemini-3.1-flash-lite |
DOCUMENTED · LIVE_DISCOVERED · DEPRECATED | Deprecated | 2026-05-07 | 1,048,576 / 65,536 | same | minimal, low, medium, high (default not stated) | ✔ | ✔ | ✔ ✔ ✔ ✔ ✘ ✔ | ? / ✔ | B X P | ✘ | ✔ | 0.25 (0.50 audio) / 0.025 / 1.50 | 2027-05-07 → gemini-3.5-flash-lite |
gemini-3-flash-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW | Preview (replacement gemini-3.6-flash, no date) |
2025-12-17 | 1,048,576 / 65,536 | same | minimal, low, medium, high | ✔ | ✔ | ✔ ✔ ✔ ✔ ✔ ✔ | 4,096 / ✔ | B X P | ✘ | ✔ | 0.50 (1.00 audio) / 0.05 / 3.00 | none |
gemini-2.5-pro |
DOCUMENTED · LIVE_DISCOVERED (deprecations.json: ACCOUNT_RESTRICTED) | 'no longer available to new users' (404 live) | 2025-06-17 | 1,048,576 / 65,536 | text, image, audio, video, PDF → text | budget 128–32,768 (cannot disable) | ✔ | ✔ | ✔ ✔ ✔ ✔ ✘ ✔ | 2,048 / ✔ | B X P | ✘ | listed | 1.25 / 0.125 / 10.00; >200k 2.50 / 0.25 / 15.00 | undated |
gemini-2.5-flash |
DOCUMENTED · LIVE_DISCOVERED | 'no longer available to new users' | 2025-06-17 | 1,048,576 / 65,536 | text, image, audio, video → text | budget 0–24,576 | ✔ | ✔ | ✔ ✔ ✔ ✔ ✘ ✔ | 2,048 / ✔ | B X P | ✘ | ✔ | 0.30 (1.00 audio) / 0.03 / 2.50 | undated |
gemini-2.5-flash-lite |
DOCUMENTED · LIVE_DISCOVERED · ACCOUNT_RESTRICTED | 404 live → gemini-3.5-flash-lite |
2025-07-22 | 1,048,576 / 65,536 | same | off by default; budget 512–24,576 | ✔ | ✔ | ✔ ✔ ✔ ✔ ✘ ✔ | ? / ✔ | B X P | ✘ | ✔ | 0.10 (0.30 audio) / 0.01 / 0.40 | undated |
gemma-4-31b-it, gemma-4-26b-a4b-it |
DOCUMENTED · LIVE_DISCOVERED (26B LIVE_VERIFIED) | Stable | 2026-04-02 | 262,144 / 32,768 | text → text | live thinking: true (thoughtsTokenCount: 5) |
? | ? | ? (grounding 'not available') | ✘ / ✘ | — | ✘ | free only | free; paid tier not available | none |
gemini-3.8-live / -extended-thinking |
DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-09-15) | 2026-09-15 | 131,072 / 65,536 | text, image, audio, video → audio (+text) | interleaved (omit config) / LOW-MEDIUM-HIGH | ✘ | ✔ (async) | ✔ ✘ ✘ ✘ ✘ ✘ | ? / ✘ | — | ✔ | ✔ | text 0.75 / audio 3.00 in; text 4.50 / audio 12.00 out | none |
gemini-3.1-flash-live-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW | 'legacy Live preview' → 3.8-live | 2026-03-11 | 131,072 / 65,536 | same | MINIMAL…HIGH | ✘ | ✔ (sync) | ✔ ✘ ✘ ✘ ✘ ✘ | — | — | ✔ | ✔ | same as 3.8-live | none |
gemini-2.5-flash-native-audio-preview-12-2025 |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW | Preview → 3.8-live | 2025-12-12 | 131,072 / 8,192 | text, audio, video → text, audio | thinkingBudget |
✘ | ✔ | ✔ ✘ ✘ ✘ ✘ ✘ | — | — | ✔ | ✔ | 0.50 text / 3.00 audio-video in; 2.00 / 12.00 out | none |
gemini-embedding-2 |
DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED | Stable (GA 2026-04-22) | 2026-04-22 | 8,192 / 1 | text, image, audio, video, PDF → embedding (128–3,072 dims) | — | — | — | — | — | B | ✘ | ✔ | 0.20 text / 0.45 image / 6.50 audio / 12.00 video | none |
gemini-embedding-001 |
DOCUMENTED · LIVE_DISCOVERED · DEPRECATED | Deprecated | 2025-07-14 | 2,048 / 1 | text → embedding | — | — | — | — | — | — | ✘ | — | not on pricing page | 2028-05-14 → gemini-embedding-2 |
5b. Gemini media and agent models
| Model | Status | Lifecycle | Modalities | Price | Notes |
|---|---|---|---|---|---|
gemini-3.1-flash-image (Nano Banana 2) |
DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-05-28) | text, image, video, PDF → text, image | in 0.50 / text out 3.00 / image out 60.00 per 1M → $0.045 (0.5K) … $0.151 (4K) per image; batch 50 % | imageConfig 512/1K/2K/4K, 14 aspect ratios, web + image search grounding; thinking MINIMAL/HIGH; no free tier; live limits 65,536 / 65,536 (docs 131,072 / 32,768) |
gemini-3.1-flash-lite-image |
DOCUMENTED · LIVE_DISCOVERED | Stable (2026-06-30) | same | 0.25 / 1.50 / 30.00 → $0.034 (1K) | 1K only, no search; C2PA + SynthID |
gemini-3-pro-image (Nano Banana Pro) |
DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-05-28); -preview / nano-banana-pro-preview RETIRED 2026-06-25 but still listed |
text, image → text, image | 2.00 / 12.00 / 120.00 → $0.134 (1K/2K), $0.24 (4K) | Google Search grounding; priority 216 image out |
gemini-2.5-flash-image (Nano Banana) |
DOCUMENTED · LIVE_DISCOVERED · DEPRECATED | shutdown 2026-10-02 | text, image → text, image | $0.039 / image (30.00 / 1M) | replacement gemini-3.1-flash-image / -lite-image |
veo-3.1-generate-preview / -fast- / -lite- |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW (lite LIVE_VERIFIED GET) | Preview; Veo 2.0/3.0 RETIRED 2026-06-30 | text, image → video + audio | $0.40 / $0.40 / $0.60 · $0.10 / $0.12 / $0.30 · $0.05 / $0.08 / — per second (720p / 1080p / 4K) | :predictLongRunning + operations; 4/6/8 s; no free tier |
gemini-omni-1.1-flash |
DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-08-27); gemini-omni-flash-preview DEPRECATED → 2026-09-30 |
text, image, video → video | in 1.50 / text out 9.00 / video out 17.50 per 1M (≈ $0.10 / s at 720p) | Interactions only |
lyria-3.5 · lyria-3-clip-preview · lyria-3-pro-preview · lyria-realtime-exp |
DOCUMENTED · LIVE_DISCOVERED (3.5 LIVE_VERIFIED GET; RealTime LIVE_VERIFIED WS) | Stable (GA 2026-09-03) · Preview · Preview → 3.5 · Experimental | text, image → music + lyrics | $0.08 / song · $0.04 / 30-s clip · $0.08 · unpriced | no free tier; RealTime BidiGenerateMusic WebSocket |
gemini-3.1-flash-tts-preview · gemini-2.5-flash-preview-tts · gemini-2.5-pro-preview-tts |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW | Preview (2.5 → 3.1) | text → audio | $1 / $20 · $0.50 / $10 · $1 / $20 per 1M (text in / audio out) | 30 voices, ≤2 speakers; streaming TTS on 3.1 only; free tier 10 req/day observed |
gemini-3.5-transcribe · gemini-3.5-transcribe-live · gemini-3.5-live-translate-preview |
DOCUMENTED · LIVE_DISCOVERED (unary LIVE_VERIFIED) | Stable (Aug 2026) · Stable · Preview | audio → text · audio → text · audio → audio + text | ≈ $0.005 / min · ≈ $0.009 / min · ≈ $0.037 / min | audioTranscriptionConfig (diarization, word timestamps, custom vocabulary) |
deep-research-preview-04-2026 · deep-research-max-preview-04-2026 · deep-research-pro-preview-12-2025 |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW (04-2026 LIVE_VERIFIED) | Preview agents | text, image, audio, video, PDF → text (+ image) | list-rate tokens (incl. intermediate) + tool fees; est. $1–3 / $3–7 per task | Interactions agent, background: true mandatory, ≤60 min; listed as models (generateContent advertised) |
antigravity-preview-09-2026 · antigravity-preview-05-2026 |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW (05-2026 DEPRECATED → 2026-10-05) | Preview coding agent | text → text | list-rate tokens; sandbox compute unbilled in preview | 09-2026 renamed built-in tools (PascalCase params) |
gemini-2.5-computer-use-preview-10-2025 |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW (deprecations.json: LEGACY) | replaced by built-in computer use on 3.x | text, image → text | 1.00 / 5.00 | browser only; 429 limit: 0 live |
gemini-robotics-er-2-preview / -streaming-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW | Preview (1.5 / 1.6 RETIRED) | text, image, audio, video → text | 1.00 / 0.10 / 5.00 (streaming: pricing section empty) | embodied reasoning |
aqa |
LIVE_DISCOVERED | Stable (PaLM-era) | text → text | not priced | :generateAnswer only (LIVE_VERIFIED), English |
6. Naming and versioning conventions
| Topic | OpenAI | Anthropic | xAI | Gemini |
|---|---|---|---|---|
| Id families | gpt-<gen>[.<minor>][-<size>][-<variant>], o<N>[-mini|-pro], gpt-<gen>-codex, gpt-image-*, gpt-realtime-*, gpt-audio-*, gpt-live-*, text-embedding-*, omni-moderation-*, sora-*; codenames for GPT-5.6 tiers (sol/terra/luna) and GPT-6 (astra); cyber/Daybreak aliases |
claude-<family>-<gen>[-<minor>][-<YYYYMMDD>] (claude-opus-5, claude-sonnet-4-6); families Fable (frontier), Mythos (invite research line), Opus, Sonnet, Haiku |
grok-<major>.<minor> (grok-4.6, grok-4.3), dated variant ids grok-4.20-0309-reasoning / -non-reasoning / grok-4.20-multi-agent-0309, product lines grok-build-<v>, grok-imagine-image[-<v>|-quality], grok-imagine-video[-<v>], grok-voice-think-fast-<v>, grok-voice-transcribe-<v>, grok-embedding-small; TTS has no model id |
gemini-<major>.<minor>-<size>[-<capability>][-preview[-MM-YYYY]] (gemini-3.8-flash, gemini-3.1-pro-preview, gemini-3.1-flash-image, gemini-3.8-live, gemini-3.5-transcribe), gemini-embedding-<n>, gemma-4-<params>-it, veo-<v>-<tier>-generate-preview, lyria-<v>, imagen-* (retired); agents deep-research-*, antigravity-preview-MM-YYYY |
| Alias vs snapshot | Undated alias (gpt-5.4) → dated snapshot (gpt-5.4-2026-03-05); the alias moves when a new snapshot ships; response model echoes the snapshot. GPT-5.6/6 tier ids have no dated snapshot; gpt-5.6 is an alias of gpt-5.6-sol that 404s on GET /v1/models but works on POST. *-latest ids (chat-latest, chatgpt-image-latest, codex-mini-latest) float. |
Since Claude 4.6 the canonical id is undated (claude-opus-4-6, claude-sonnet-5, claude-fable-5-1) and is itself the snapshot. Older lines keep dated snapshots with undated aliases (claude-haiku-4-5 → claude-haiku-4-5-20251001); aliases follow the newest snapshot. No -latest suffix. Cloud ids differ per platform (anthropic.claude-opus-5 on Bedrock, claude-opus-5@… on Vertex). |
Documented scheme: <model> = latest stable, <model>-latest, <model>-<date> pinned — in practice the dated 4.20 ids are canonical and the bare names (grok-4.20) are aliases (15 aliases on one model); grok-4.6 has no -latest. Retired ids redirect (grok-3, grok-4-0709, grok-4-fast*, grok-4-1-fast* → grok-4.3; grok-code-fast-1 → grok-build-0.1) and appear in the catalogue's aliases[]. grok-voice-latest routes to grok-voice-think-fast-2.0 since 2026-08-05. |
Stable ids are undated (gemini-3.8-flash, 'usually don't change'); -preview ids may carry a month suffix (gemini-2.5-flash-native-audio-preview-12-2025) and can be deprecated with two weeks' notice; -latest aliases (gemini-flash-latest, gemini-pro-latest, gemini-flash-lite-latest, gemini-2.5-flash-native-audio-latest) are hot-swapped (2-week e-mail notice) — gemini-flash-latest moved 3 Flash Preview → 3.5 Flash → 3.8 Flash (observed via modelVersion); -001 suffixes only on 2.0 / embedding-001 / Veo 3.0; -exp experimental; gemini-3-pro-preview was repointed to 3.1 Pro on 2026-03-09. |
| Version pinning advice | pin the dated snapshot in production; GET /v1/models/{id} exposes shutdown_date |
pin the undated id for 4.6+ (already immutable); for 4.5 lines pin the dated snapshot | pin the dated 4.20 ids or the bare grok-4.x id (no snapshots exist); check response.model — a retired id silently returns its replacement |
pin the stable id; avoid -latest in production; modelVersion in every response tells you what actually ran |
| Tool versioning | tool type strings are undated (web_search, code_interpreter) with a few dated snapshots (web_search_2025_08_26) |
every server/Anthropic-defined tool type carries _YYYYMMDD (web_search_20260318, code_execution_20260521) |
undated OpenAI-style type strings (web_search, x_search, code_interpreter/code_execution, file_search/collections_search, mcp, shell, tool_search) |
undated camelCase keys (googleSearch, urlContext, codeExecution, fileSearch, computerUse); legacy googleSearchRetrieval |
| API surface versioning | none (openai-version: 2020-10-01 response header); breaking changes via new endpoints and OpenAI-Beta surfaces |
anthropic-version: 2023-06-01 required; features gated by dated anthropic-beta values that graduate to GA |
none — no version or beta headers; docs.x.ai/developers/release-notes is the changelog; gRPC proto v6 |
path version /v1beta (86 methods) vs /v1 (47); Interactions optional Api-Revision: 2026-05-20; breaking schema change 2026-05-26/06-08 (outputs → steps) |
| Knowledge cutoff | single date per model page (knowledge_cutoff), e.g. gpt-6-astra 2026-04-30 |
two dates: reliable knowledge and training data (knowledge_cutoff.reliable / .training_data) |
published only for grok-4.6 (2026-02-01); null for every other Grok record |
not stated on the API model pages; January 2025 from the Gemini 3 / 2.5 model cards; June 2025 for gemini-2.5-flash-image; null for 64 of 97 records |
| Discovery | GET /v1/models lists 136 ids incl. snapshots, owned_by, created, shutdown_date; fine-tuned ids ft:… |
GET /v1/models lists 11 ids with capabilities (thinking types, effort levels, context_management strategies, structured_outputs, citations, pdf_input) |
GET /v1/models 12 ids; typed catalogues /v1/language-models, /v1/image-generation-models, /v1/video-generation-models, /v1/embedding-models expose live prices in ticks, aliases[], input_modalities, fingerprints; voice models absent |
GET /v1beta/models 58 ids (inputTokenLimit, outputTokenLimit, supportedGenerationMethods, thinking, sampling defaults, version); /v1/models 22; agents listed as models; shut-down previews still listed |
7. Deprecation and retirement policies
| Aspect | OpenAI | Anthropic | xAI | Gemini |
|---|---|---|---|---|
| Vocabulary | Deprecated = retirement announced with a shutdown date; Legacy = no longer updated, will be deprecated later; Sunset/shut down = no longer accessible | Active / Legacy (no more updates) / Deprecated (replacement + retirement date, ≥60 days) / Retired (requests fail) | Available (on /developers/models + /pricing + GET /v1/models) / Deprecated (notice period) (migration guide; ~60-day notice observed) / Retired (removed from the catalogue but the slug keeps resolving and redirects to a replacement billed at the replacement's price) / undocumented Legacy (older ids vanished from docs) |
Stable (GA) / Preview (production allowed, tighter limits, ≥2 weeks notice) / Latest alias / Experimental / Legacy (existing customers) / Deprecated (earliest shutdown date announced) / Shut down; API enum ModelStatus.modelStage (EXPERIMENTAL, PREVIEW, STABLE, LEGACY, DEPRECATED, RETIRED) |
| Notice period | dated per announcement (typically ≥ 6 months for flagship snapshots, shorter for previews); shutdown_date also exposed live on GET /v1/models |
≥ 60 days from deprecation to retirement; every active model carries a "not sooner than" retirement commitment (Fable 5.1 ≥ 2027-09-01, Opus 5 ≥ 2027-07-24, Sonnet 5 ≥ 2027-06-30, Haiku 4.5 ≥ 2026-10-15) | no published minimum; observed ~60 days (grok-imagine-image-quality announced 2026-09-02 → 2026-11-02); the May-15 batch was announced via a migration guide with automatic redirects |
previews ≥ 2 weeks; stable models get an earliest-shutdown date on the deprecations page (e.g. gemini-3.1-flash-lite 2026-05-07 → 2027-05-07 = 12 months); 'exact date communicated in advance'; Vertex has its own schedule |
| Replacement mapping | each row names a replacement (o4-mini → gpt-5.6-terra, gpt-image-1 → gpt-image-2, whisper-1 → gpt-transcribe) |
each retirement names a replacement (claude-opus-4-1-20250805 → claude-opus-4-8, claude-3-haiku-20240307 → claude-haiku-4-5-20251001) |
migration guides name the replacement and the effort setting (grok-4-fast-reasoning → grok-4.3 reasoning_effort=low; grok-4-fast-non-reasoning → grok-4.3 none; grok-code-fast-1 → grok-build-0.1; grok-imagine-image-quality → grok-imagine-image-2.0 quality=low) |
deprecations page names a replacement (gemini-3.1-flash-lite → gemini-3.5-flash-lite, gemini-2.5-flash-image → gemini-3.1-flash-image, gemini-embedding-001 → gemini-embedding-2, text-embedding-004 → gemini-embedding-2); live 404 messages also name one (gemini-2.5-flash-lite → gemini-3.5-flash-lite) |
| Upcoming (after 2026-09-18) | 2026-09-24 Videos API + Sora 2; 2026-09-28 gpt-3.5-turbo-instruct, babbage-002, davinci-002; 2026-10-01 gpt-5.4-cyber; 2026-10-23 o1, o1-pro, o3-mini, o4-mini, gpt-4, gpt-4-turbo, gpt-4.1-nano, gpt-image-1; 2026-11-30 Evals API, Agent Builder, reusable prompts; 2026-12-01 gpt-image-1-mini/1.5; 2026-12-11 gpt-5 2025 snapshots, o3, o3-pro; 2027-01-06 no new fine-tuning jobs; 2027-01-20 legacy audio/realtime; 2027-02-26 whisper-1, gpt-4o-transcribe | no dated model retirements pending; claude-mythos-preview deprecated 2026-06-09 (date TBA). Feature-level: context-1m-2025-08-07 header retired 2026-04-30; fast mode removed on Opus 4.6/4.7; sampling params and manual thinking deprecated (400 on 4.7+); prefill deprecated (400 on 4.6+) |
2026-09-21 12:00 PT x_search per-call billing → per post / per profile; 2026-11-02 grok-imagine-image-quality retires (redirect to 2.0 low); /v1/messages Anthropic compat 'fully deprecated' (no date); logprobs ignored on 4.20+ |
2026-09-30 gemini-omni-flash-preview; 2026-10-02 gemini-2.5-flash-image; 2026-10-05 antigravity-preview-05-2026; September 2026 standard API keys rejected (auth keys only); 2027-05-07 gemini-3.1-flash-lite; 2028-05-14 gemini-embedding-001; replacement recommended without date: gemini-3-flash-preview, gemini-3.1-flash-live-preview, 2.5 native-audio / TTS previews, lyria-3-pro-preview; sampling params deprecated 2026-07-21 |
| Recently retired | Assistants API (2026-08-26), DALL·E 2/3 (2026-05-12), text-moderation-*, search-preview snapshots (2026-07-23), codex snapshots ≤ 5.2, deep-research models, computer-use-preview |
Opus 4.1 (2026-08-05), Sonnet 4 / Opus 4 (2026-06-15), Claude 3 Haiku (2026-04-20), Claude 3.5/3.7, Claude 3 Opus/Sonnet, 2.x, 1.x, Instant; /v1/complete effectively retired (400) |
2026-05-15: grok-3, grok-4-0709, grok-4-fast-reasoning, grok-4-fast-non-reasoning, grok-4-1-fast-reasoning, grok-4-1-fast-non-reasoning, grok-code-fast-1, grok-imagine-image-pro (all redirect); grok-2-image 404; Live Search (chat search_parameters) 410; /v1/completions, /v1/complete 400 |
Gemini 2.0 Flash/-Lite (2026-06-01); 2.5 previews (2025-11 → 2026-03); gemini-3-pro-preview (2026-03-09, repointed); gemini-3.1-flash-lite-preview (2026-05-25); image previews gemini-3.1-flash-image-preview, gemini-3-pro-image-preview / nano-banana-pro-preview (2026-06-25); Imagen 3/4 (2026-08-17); Veo 2.0/3.0 (2026-06-30); half-cascade Live models (2025-12-09); text-embedding-004 (2026-01-14), embedding-001, embedding-gecko-001, gemini-embedding-exp* (2025-10-30); robotics ER 1.5/1.6; model tuning (May 2025); LearnLM |
| After retirement | requests fail; fine-tuned models retire with their base | 404 not_found_error on the Claude API; several retired ids remain available on Bedrock / Vertex |
slug still resolves: GET /v1/models/grok-3 returns the grok-4.3 object; requests are served and billed at the replacement's rates (response.model shows it); grok-2-image is a hard 404 |
404 NOT_FOUND on generation (Model is not found for api version v1beta), yet shut-down ids remain in GET /v1beta/models (e.g. gemini-3-pro-preview, gemini-3.1-flash-lite-preview, image previews) — listing ≠ availability |
| Machine-readable | generated/deprecations.json (165 flat OpenAI records: model, shutdown_date, replacement, phase) |
generated/deprecations.json (1 Anthropic record with models[], active_models_retirement_commitments{}, api_features[]) |
generated/deprecations.json (1 xAI record: models[] with behavior_after and our_probe, api_features[], release_notes_digest[] by month); models.json kinds retired_redirect, legacy, alias |
generated/deprecations.json (1 Gemini record: 60+ models[] with released/shutdown/replacement/live_listed, alias_history[], api_features[], active_models_no_shutdown_announced[]) |
8. Live discovery vs docs
- OpenAI: 136 ids in
GET /v1/models; aliasesgpt-5.6,gpt-5.5-cyber,gpt-5-search-api-2025-10-14behave as id-only/alias records (record_kind: alias|id_only);gpt-5.6echoesgpt-5.6-sol;gpt-5.4-nanoechoeseffort: nonewhen omitted;gpt-6-astrarejectsreasoning.effort: none. - Anthropic: 11 ids in
GET /v1/models(Mythos ids absent → 404);claude-fable-5-1returned Opus-class rate-limit headers (10M ITPM) although docs list 4M for the Fable class (account-specific); Haiku 4.5 tolerated edited/removed thinking blocks and prefill-with-thinking (docs say 400) — graceful degradation, do not rely on it. - xAI: 12 ids in
GET /v1/models(7 language, 3 image, 2 video; voice/embedding models absent); redirects verified (GET /v1/models/grok-3,/grok-4-fast,/v1/language-models/grok-4-0709→ grok-4.3;grok-code-fast-1→ grok-build-0.1);grok-4.3acceptsreasoning_effort: none(undocumented) whilegrok-4.20-0309-reasoningandgrok-build-0.1rejectreasoning_effortalthough docs/parameters list them as compatible;x_searchemitscustom_tool_callitems (docs:x_search_call);GET /v1/responses/{id}returns 200 forstore:falseids; rate-limit headers exist but are undocumented (7,200 / 1,800 requests per minute vs documented 150 / 37 RPS);eu-west-1.api.x.aiserves grok-4.3 undocumented; every text call bills 70–180 reasoning tokens for "Reply with OK.". - Gemini: 58 ids in
GET /v1beta/modelsvs 22 inGET /v1/models(docs claim parity);gemini-flash-latestresolved togemini-3.8-flash(changelog last said 3.5 Flash); Gemini 2.5 Pro/Flash/Flash-Lite return 200 onmodels.getbut 404 'no longer available to new users' on generation (undocumented policy); Pro models, image/video/music models,cachedContentscreate and Batch create return 429limit: 0/ 400 FAILED_PRECONDITION on the free tier (ACCOUNT_RESTRICTED); explicit-cache minimum is 1,024 tokens live (docs' 4,096 is the implicit-cache threshold);thoughtSignaturevalidation can be bypassed with the documented dummy strings; token limits differ from the model pages for image, Lyria, Deep Research, translate and computer-use ids (see §9);gemini-3.8-livereportsversion: 3.1-flash-live-03-2026.
9. Data inconsistencies found during synthesis
Reported for the fragment owners (fix in generated/fragments/**, never in merged files). Items 1–16 carry over from the two-provider synthesis; 17+ were found while adding xAI and Gemini.
models.jsono3:pricing.standard= $2 / $0.5 / $8 butpricing["model_page:Text tokens"]= $1 / $0.25 / $4 (two sources disagree; other models agree).models.jsonOpenAI GPT-5.6 family,gpt-6-astra,gpt-5.6-cyber,gpt-daybreak-*,gpt-4o-mini-audio-preview*:capabilities.function_calling: truewhilecapabilities.tool_function_calling: false.models.jsonAnthropic alias records (claude-opus-4-5,claude-sonnet-4-5,claude-haiku-4-5) storepricingas a string ("same as …") instead of the schema object;modalities,capabilities,thinkingare null on aliases.models.jsongpt-5.5-cyberandgpt-5-search-api(-2025-10-14)areid_onlyrecords without context/output/modalities although priced.pricing.jsonunit strings are heterogeneous (per 1M tokensvs model-page1M tokens;image_output low 1024x1024dimensions encode size in the dimension name) — now 12 distinct unit strings across four providers (per 1K requests,per 1K search queries,per 1K grounded prompts,per 1K calls,per call,per request,multiplier,per hour,per page,per 1M tokens per hour,per GiB per day,per 1M characters).endpoints.json: six OpenAI records carryverification.result: "success"without a live call; threeLIVE_VERIFIEDendpoints haveresult: "failure"— status and result should be reconciled.endpoints.json:POST /v1/responses?beta=trueandPOST /v1/messages (beta surface)encode a variant in thepathfield;authis a string for some fragments and an object for others.streaming-events.json: Anthropic core stream recorded under twoapilabels; Managed Agents client→server events appear twice.webhook-events.json: the 28 OpenAI records haveapi: nullwhile Anthropic records setapi: managed-agents.headers.json:requiredis boolean for OpenAI rows and a free string for Anthropic/xAI/Gemini rows.errors.json:api_familyis null for core Anthropic and OpenAI errors; several Anthropic rows havehttp_statusnull or composite; Gemini rows recordobserved_live: falsefor 400 FAILED_PRECONDITION / 403 / 501 although the Gemini docs pages report them observed.deprecations.json: OpenAI = 165 flat records, Anthropic / xAI / Gemini = 1 nested record each — two schemas in one file.rate-limits.json: OpenAI record has nostatusfield and nodocumentedkey (its content is underconcepts/usage_tiers); the Gemini enqueued-token table listsgemini-2.0-flash-imageand shut-down 2.5 previews while omitting the GA image ids.tools.jsonOpenAIprogrammatic_tool_callinghas an emptycompatible_modelslist; OpenAIcomputertool isUNVERIFIEDalthough eleven model records claimtool_computer_use: true.- Anthropic tool-search docs table omits
claude-sonnet-5while the model record lists both tool-search types. - Live vs docs (Anthropic):
mcp_serverserror message advertisesmcp-client-2026-09-15which the API rejects;code_execution_requestsusage counter absent live; Managed Agents webhook names differ from stream names. - xAI
models.json:max_outputis null for all 17 model records (only the grok-4.6 page says "No text output limit");knowledge_cutoffis null for every model except grok-4.6;parameters.jsonrecordsmax_output_tokensdefault: Nonewhile docs/xai/responses.md says 128,000; a documented sample echoedmax_output_tokens: 2000unexplained. - xAI reasoning flags:
parameters.jsonlistsgrok-4.20-0309-reasoningandgrok-build-0.1as compatible withreasoning_effort(Chat) /reasoning.effort(Responses) but both return 400 live;noneappears in the enum only as "(LIVE_DISCOVERED, grok-4.3)";grok-4.5xhighis listed without the "treated as high" caveat from docs/xai/reasoning.md. - xAI prices:
grok-imagine-image-2.0per_image_default: 0.06(catalogue, medium/1k) vs pricing page $0.04 (low/1k); the REST reference describes price units as "USD cents per 100M tokens" while the live catalogue andpricing.jsonuse ticks of 1e-10 USD;grok-4.20-multi-agent-0309has no priority rows inpricing.jsonwhilemodels.jsonsays "2.0 (not documented per model)". - xAI tools/streaming:
x_searchreturnscustom_tool_callitems live vs documentedx_search_call;streaming-events.jsonhas no x_search events; imageusagedocumentsinput_tokens/output_tokensthat are absent live (onlycost_in_usd_ticks). - xAI endpoints: Collections are documented on
management-api.x.aiwith a Management key but the same paths answer onapi.x.aiwith an inference key (records carry bothACCOUNT_RESTRICTEDandLIVE_VERIFIED);POST /v1/collections/{id}/documentsmultipart → 405; batchresponsesrequests come back aschat_get_completion; deferred completions documented "retrievable exactly once" but a second GET returned 200;GET /v1/responses/{id}200 forstore:false;/v1/completionsmarked RETIRED inendpoints.jsonwhile docs/xai/legacy-completions.md says LEGACY "retired in practice"; the 4.20 non-reasoning model also rejects raw sampling although docs say "non-reasoning only". - xAI rate limits / headers: documented Tier 0 RPS 150 / 37 vs observed
x-ratelimit-limit-requests7,200 / 1,800 per minute (= 120 / 30 RPS); the headers themselves are undocumented; STT default model isgrok-voice-transcribe-1.0in release notes and 2.0 on the model page (DOCUMENTATION_INCOMPLETE); realtime default voicexai_aralive vsevein docs;response.done.usageandpingundocumented;response.audio.deltadocumented but not emitted;GET /v1/api-keycreate_time/modify_timereturned empty strings;grok-imagine-video-1.5liveinput_modalitiesincludeaudio(model page: text, image); video pending polls return HTTP 202 (undocumented);eu-west-1.api.x.aiundocumented but live. - Gemini
pricing.json: 5 null-price rows (veo-3.1-lite4k,gemini-embedding-001input,tool:custom_tools_endpoint,agent:deep-research,service:document_tokens); 2 non-numeric prices ('1.8x','0.50-8.10'); duplicate keys with conflicting values —tool:google_searchstandard = $14 "per 1K requests" and $35 "per 1K grounded prompts",tool:google_maps= $14 "per 1K search queries" and $25 "per 1K grounded prompts" (the 3.x vs 2.5 rows are not disambiguated in the key); Gemma rows usetier: freewith per-dimension rows while other models use onefree_tierdimension. - Gemini
models.json:knowledge_cutoffnull for 64/97,context_windownull 28,max_outputnull 33,release_datenull 13;gemini-3.5-flashflexcached_input0.08 vs batch 0.075 (not 50 % of 0.15);gemini-3.5-flash-litebatch/flexcached_input0.02 not on the pricing page;gemini-2.5-pro/-flashcarryACCOUNT_RESTRICTEDindeprecations.jsonbut not inmodels.json(only-flash-litedoes);thinkingflag true live ongemini-3.5-transcribeandgemini-3.1-flash-tts-preview(docs: not supported) and absent on Live/robotics-streaming models (docs: supported). - Gemini
parameters.json: 11 duplicate (endpoint, parameter) pairs from different fragments with conflicting enums —generationConfig.thinkingConfig.thinkingLevelMINIMAL|LOW|MEDIUM|HIGHvsMINIMAL|HIGH(image models),responseModalitiesTEXT|IMAGE|AUDIOvsAUDIO, Interactionsresponse_format×3,generation_config.thinking_level×2,transcription_config×2;generationConfig._responseJsonSchemarecorded asLIVE_DISCOVERED. - Gemini
endpoints.json: the corePOST /v1beta/models/{model}:generateContentrecord sits underapi_family: transcription(thegenerate-contentfamily holds only/v1/…,streamGenerateContentanddynamic/{id});dynamicendpoints duplicated across two families with different path spellings and statuses; Interactions/v1/*taggedGAbutUNVERIFIED; webhooks documented on/v1while everything else is/v1beta. - Gemini
streaming-events.json: the 7lyria-realtimerecords have nullstatusalthough docs/gemini/music-generation.md marks the protocolLIVE_VERIFIED. - Gemini docs vs docs / live: Maps pricing $25/1k grounded prompts (tools table) vs $14/1k queries (Gemini 3 tables) →
DOCUMENTATION_INCOMPLETE; File Search indexing $0.15/1M vs gemini-embedding-2 text $0.20/1M on the same page; priority "75–100 % more" (prose) vs 1.8× (tables); Batch/Flex "not available" on free tier for most models but "free of charge" forgemini-3.5-flash-lite; 3.5-flash caching "free of charge" on the page vs livelimit: 0; explicit-cache minimum 4,096 (docs) vs 1,024 (live); audio token rate 32 tok/s (tokens guide) vs 25 tok/s (Live/pricing); image token counts 280/560/1,120/2,240 (docs) vs 256/529/1,089/2,209 (observed); PDF page 520 IMAGE tokens (generateContent) vs 560 DOCUMENT tokens (countTokens); token limits live vs docs forgemini-3.1-flash-image(65,536/65,536 vs 131,072/32,768),-lite-image(out 65,536 vs 4,096),gemini-3-pro-image(in 131,072 vs 65,536),gemini-2.5-flash-image(in 32,768 vs 65,536),lyria-3*(1,048,576 vs 131,072),deep-research-*(131,072 vs 1,048,576),gemini-3.5-live-translate-preview(16,384/32,768 vs 131,072/65,536), 2.5 computer use (131,072/65,536 vs 128,000/64,000); docs say every model is in/v1and/v1beta(live 22 vs 58);gemini-3-pro-previewcomputer use "not supported" (model page) vs "launched" (changelog);interaction.status_updatestill emitted although the breaking-changes guide says superseded; transcriptionparts[].audioTranscriptionundocumented, speaker labelspk:0(REST) vsspk_1(Interactions docs); Google SearchACCOUNT_RESTRICTEDon this free-tier key although the pricing page lists 500 RPD free for 2.5 Flash/Flash-Lite;gemma-4-*absent from docs/models tables;deprecations.jsonreplacement text differs from the docs table forgemini-2.5-flash-imageandgemini-2.0-flash. - Cross-provider:
docs/errors/xai.mddoes not exist (xAI errors live in docs/xai/authentication-headers-errors.md) whiledocs/errors/{openai,anthropic,gemini}.mddo;sdks.jsonhasversion: nullon every xAI/Gemini record although the docs pages quote versions (xai-sdk 1.19, google-genai 2.24, @google/genai 2.23);models.jsonuses four differentkindvocabularies (OpenAImodel|snapshot|id_only|alias, Anthropicsnapshot|alias, xAImodel|legacy|retired_redirect|alias|service, Geministable|preview|alias|agent|experimental).
Related: index · features · pricing · caching and reasoning · FAQ.