# Models — OpenAI ↔ Anthropic ↔ xAI ↔ Gemini equivalence and use-case table **Status:** every value is copied from `generated/models.json` (statuses, context, output, cutoffs, capability flags, pricing blocks; live-verified on 2026-09-18/19 where marked), `generated/pricing.json` and `generated/deprecations.json`. Equivalence rows are **use-case pairings by price/tier and capability**, not benchmark claims. xAI probes ran on a Tier 0 team, Gemini probes on a free-tier key (paid-only Gemini models are `ACCOUNT_RESTRICTED`, not absent). **Sources:** https://developers.openai.com/api/docs/models · https://developers.openai.com/api/docs/pricing · https://developers.openai.com/api/docs/deprecations · https://platform.claude.com/docs/en/models/overview · https://platform.claude.com/docs/en/about-claude/pricing · https://platform.claude.com/docs/en/about-claude/model-deprecations · https://docs.x.ai/developers/models · https://docs.x.ai/developers/pricing · https://docs.x.ai/developers/migration/may-15-retirement · https://ai.google.dev/gemini-api/docs/models · https://ai.google.dev/gemini-api/docs/pricing · https://ai.google.dev/gemini-api/docs/deprecations · docs/models/openai-models.md · docs/models/anthropic-models.md · docs/models/xai-models.md · docs/models/gemini-models.md **Last verified:** 2026-09-18 ## 1. Use-case pairing (current models, four providers) Prices USD per 1M tokens, standard tier: **in / cached-in / out**. "Cached-in" is a cache **read**. Context = documented window; Out = documented max output. xAI prices double for requests whose prompt is ≥200k tokens (all tokens of that request); Gemini Pro prices rise for prompts >200k; Gemini 3.6–3.8 Flash prices are introductory (double on 2027-01-01). | Use case | OpenAI | Anthropic | xAI | Gemini | Why they pair | |---|---|---|---|---|---| | Frontier reasoning, hardest tasks | `gpt-6-astra` — 1,050k ctx / 128k out · text+image → text · effort low…max (no `none`) · $10 / $1 / $50 · cutoff 2026-04-30 · rel. 2026-09-03 · `LIVE_VERIFIED` | `claude-fable-5-1` — 1M / 128k · text+image+PDF → text · adaptive thinking always on, effort low…max · $10 / $0.25 / $50 · cutoff Jun 2026 · rel. 2026-09-01 · `LIVE_VERIFIED` | `grok-4.6` — 500k / no published limit · text+image → text · reasoning always on, effort low…xhigh (default high) · $2 / $0.50 / $6 (≥200k: $4 / $1 / $12) · cutoff 2026-02-01 · rel. 2026-08 · `LIVE_VERIFIED` | `gemini-3.1-pro-preview` — 1,048,576 / 65,536 · text+image+audio+video+PDF → text · `thinkingLevel` low/medium/high (default high, no minimal) · $2 / $0.20 / $12 (>200k: $4 / $0.40 / $18) · cutoff Jan 2025 (model card) · rel. 2026-02-19 · `PREVIEW` · `ACCOUNT_RESTRICTED` (paid tier only) | Each provider's top reasoning line. Price spread is 5× ($2 vs $10 in). Astra and Fable are always-on reasoners; Grok 4.6 too; Gemini 3.1 Pro cannot go below `low`. Only Gemini's is still a preview. | | Flagship general / agentic | `gpt-5.6-sol` (alias `gpt-5.6`) — 1,050k / 128k · effort none…max, `reasoning.mode: pro` · $4 / $0.40 / $20 (promo ≥ 2026-11-21) · cutoff 2026-02-16 · rel. 2026-07-09 | `claude-opus-5` — 1M / 128k · adaptive (on by default, can disable except xhigh/max) · $5 / $0.50 / $25 · cutoff May 2026 · rel. 2026-07-24 · fast mode $10 / $50 | `grok-4.5` — 500k · reasoning always on, effort low…high (`xhigh` treated as high) · $2 / $0.30 / $6 · rel. 2026-07 · priority 2× · no batch · aliases `grok-4.5-latest`, `grok-build-latest` | `gemini-3.8-flash` — 1,048,576 / 65,536 · thinking low/medium/high (default medium; `minimal` → 400) · $0.75 / $0.075 / $3.75 (→ $1.50 / $0.15 / $7.50 in 2027) · rel. 2026-09-02 · GA · `LIVE_VERIFIED` · free tier | Newest general-purpose line per provider. Gemini's Flash line is the current flagship for the Developer API (Pro is preview); Grok 4.5/4.6 sit between OpenAI's mini and flagship prices. | | Balanced mid-tier | `gpt-5.6-terra` — 1,050k / 128k · $2 / $0.20 / $12; also `gpt-5.5` $5 / $0.50 / $30, `gpt-5.4` $2.5 / $0.25 / $15 | `claude-sonnet-5` — 1M / 128k · adaptive on by default · $2 / $0.20 / $10 · cutoff Jan 2026 · rel. 2026-06-30 | `grok-4.3` — 1M · reasoning **optional** (`reasoning_effort: none` LIVE_DISCOVERED → 0 reasoning tokens; default low) · $1.25 / $0.20 / $2.50 (≥200k ×2) · batch −20 % · rel. 2026-04 · redirect target of `grok-3`, `grok-4-0709`, `grok-4-fast-*` | `gemini-3.5-flash` — 1,048,576 / 65,536 · thinking minimal…high (default medium; budget 0 disables) · $1.50 / $0.15 / $9 · rel. 2026-05-19 · GA · `LIVE_VERIFIED` | Identical $2 input on Terra / Sonnet 5 / 3.1 Pro; grok-4.3 and gemini-3.5-flash undercut them. All support the full tool set of their provider (§3–§5). | | Small / high-volume | `gpt-5.6-luna` — $0.20 / $0.02 / $1.20; `gpt-5.4-mini` — 400k / 128k · $0.75 / $0.075 / $4.50; `gpt-5.4-nano` — $0.20 / $0.02 / $1.25 | `claude-haiku-4-5-20251001` (alias `claude-haiku-4-5`) — 200k / 64k · manual extended thinking only, no `effort` · $1 / $0.10 / $5 · cutoff Feb 2025 · rel. 2025-10-15 | `grok-build-0.1` — 256k · $1 / $0.20 / $2 · `PREVIEW` ('early access'); `grok-4.20-0309-non-reasoning` — 1M · no reasoning · $1.25 / $0.20 / $2.50 · batch −20 % | `gemini-3.5-flash-lite` — 1,048,576 / 65,536 · thinking default minimal · $0.30 / $0.03 / $2.50 · rel. 2026-07-21 · GA · `LIVE_VERIFIED` · free tier; `gemini-3.6-flash` / `-3.7-flash` $0.75 / $0.075 / $3.75 (introductory) | Cheapest current text model per provider. OpenAI nano/Luna and Gemini Flash-Lite are ~5× cheaper than Haiku 4.5; xAI has no sub-$1 text model. | | Deep-compute "pro" | `gpt-5.5-pro` / `gpt-5.4-pro` — 1,050k / 128k · effort medium…xhigh · **no prompt caching**, no `fast` · $30 / — / $180; `gpt-5.6-sol` with `reasoning.mode: pro` (standard rates) | no pro SKU — `output_config.effort: max` on Opus 5 / Fable 5.1 | `grok-4.20-multi-agent-0309` — 1M · one Responses call fans out to **4 or 16 agents** (`reasoning.effort` low/medium vs high/xhigh) · $1.25 / $0.20 / $2.50 per agent-token · `BETA` · Responses-only (Chat → 400) | Deep Research agents `deep-research-preview-04-2026` / `-max-` (Interactions, `background: true`, ≤60 min, est. $1–3 / $3–7 per task, list-rate tokens + tools) · `PREVIEW`; `gemini-3.8-live-extended-thinking` for voice | OpenAI sells extra compute as a model/mode, xAI as a multi-agent model, Gemini as an agent, Anthropic as the top effort value. | | Coding agents | `gpt-5.3-codex` — 400k / 128k · $1.75 / $0.175 / $14 · standard + fast only; GPT-5.6 family lists hosted shell / apply_patch / skills | `claude-opus-5`, `claude-sonnet-5` (bash, text_editor, code_execution, skills, computer/browser toolsets, programmatic tool calling); `claude-opus-4-8` legacy | `grok-build-0.1` (`grok-code-fast-1` redirect; 256k; $1 / $0.20 / $2; `PREVIEW`) + **Grok Build CLI** (`grok`, BETA: TUI, headless, ACP, MCP, hooks, skills, sandbox) | `antigravity-preview-09-2026` (Interactions coding agent with Linux sandbox, default model gemini-3.8-flash, `max_total_tokens`) · `PREVIEW`; `gemini-3.1-pro-preview-customtools` (bash-style custom tools, 3.1 Pro price) | Only xAI and Gemini ship a coding-specific model id; OpenAI ships codex; Anthropic ships coding tools on its general models. | | Previous-generation flagships still served | `gpt-5.4` ($2.5 / $0.25 / $15), `gpt-5.2` ($1.75 / $0.175 / $14), `gpt-5.1` & `gpt-5` ($1.25 / $0.125 / $10), `o3` ($2 / $0.50 / $8), `o3-pro` ($20 / — / $80) | `claude-opus-4-8`, `-4-7`, `-4-6` ($5 / $0.50 / $25, LEGACY), `claude-opus-4-5-20251101` ($5, 200k, LEGACY), `claude-sonnet-4-6` ($3 / $0.30 / $15, LEGACY), `claude-sonnet-4-5-20250929` ($3, 200k, LEGACY) | `grok-4.20-0309-reasoning` — 1M · reasoning always on (`reasoning_effort` → 400) · $1.25 / $0.20 / $2.50 · batch −20 % · 15 aliases incl. `grok-4.20`, `grok-4.20-beta*`, `grok-4.20-reasoning-gv2` | `gemini-3-flash-preview` ($0.50 / $0.05 / $3, `PREVIEW`, replacement 3.6 Flash, no date), `gemini-3.1-flash-lite` ($0.25 / $0.025 / $1.50, **DEPRECATED** → 2027-05-07), `gemini-2.5-pro` / `-flash` / `-flash-lite` (priced and listed but **404 'no longer available to new users'**) | OpenAI keeps old flagships "Active" until a dated deprecation; Anthropic marks "Active (legacy)"; xAI keeps 4.20 on the price list; Gemini keeps 2.5 listed while blocking new users. | | Deprecated but callable | `o4-mini`, `o3-mini`, `o1`, `o1-pro`, `gpt-4.1-nano` (2026-10-23), `gpt-5-*-2025-08-07`, `o3-2025-04-16`, `o3-pro-2025-06-10` (2026-12-11) | `claude-mythos-preview` (invite, replacement `claude-mythos-5`, date TBA) | `grok-imagine-image-quality` (→ 2026-11-02, redirects to `grok-imagine-image-2.0` low); `/v1/messages` compat surface | `gemini-3.1-flash-lite` (→ 2027-05-07), `gemini-2.5-flash-image` (→ 2026-10-02), `gemini-omni-flash-preview` (→ 2026-09-30), `antigravity-preview-05-2026` (→ 2026-10-05), `gemini-embedding-001` (→ 2028-05-14) | see §7 | | Non-reasoning models | `gpt-4.1` (1,047k / 32k, $2 / $0.50 / $8, fine-tunable), `gpt-4.1-mini`, `gpt-4o` (128k / 16k), `gpt-4o-mini`, `chat-latest` (400k, $5 / $0.50 / $30) | none current — every active Claude model reasons (Haiku/Sonnet 4.5 optionally) | `grok-4.20-0309-non-reasoning` (no `reasoning_content`, `reasoning_tokens: 0`); `grok-4.3` with `reasoning_effort: none` | none — every Gemini 3.x text model thinks (`minimal` on Flash-Lite / 3.6 / 3.5 / 3 Flash; `thinkingBudget: 0` on 3.6 / 3.5 Flash only); Gemma 4 live returned `thoughtsTokenCount: 5` | OpenAI and xAI keep explicit non-reasoning ids; Anthropic and Gemini only expose reasoning-off switches on some models. | | Gated / invite / paid-only | `gpt-5.6-cyber`, `gpt-5.5-cyber` ($12.5 / $1.25 / $75), `gpt-daybreak-blue-latest`, `gpt-daybreak-red-latest`, `gpt-rosalind-research` (`ACCOUNT_RESTRICTED`) | `claude-mythos-5-1`, `claude-mythos-5` ($10 / $0.25–$1 / $50, Project Glasswing, `ACCOUNT_RESTRICTED`, 404 for our key) | `grok-embedding-small` (404, no price), Skills API (404), `tool_search` (403 alpha), custom-voice creation (Enterprise), Management API (needs a Management key) — all `ACCOUNT_RESTRICTED` | Pro models and every image / video / music model are **paid-tier only** (`limit: 0` on the free tier): `gemini-3.1-pro-preview`, `gemini-pro-latest`, `gemini-3.1-flash-image`, `gemini-3-pro-image`, Veo 3.1, Lyria 3.x; `cachedContents` and Batch also 429/400 on free | Different gates: invite programs (OpenAI, Anthropic), ACLs/alpha (xAI), billing tier (Gemini). | | Free of charge | none | none | none (prepaid credits) | free-tier rows on 3.8/3.7/3.6/3.5 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite, 3 Flash Preview, Live 3.8 / 3.1, TTS 3.1 / 2.5 Flash, `gemini-3.5-transcribe`, `gemini-embedding-2`; **Gemma 4** (`gemma-4-31b-it`, `gemma-4-26b-a4b-it`, 262,144 / 32,768, text only) is free-only; content may be used to improve Google products | Gemini is the only provider with a documented $0 tier. | ## 2. Capability table — current OpenAI text models From `capabilities` in `generated/models.json`. ✔ true · ✘ false · ? `"unknown"`. Effort = accepted `reasoning.effort` values (default in bold when recorded). Cache = prompt caching / 24 h retention. Tiers = service tiers listed as available (S standard, B batch, X flex, F fast). | Model | Status | Ctx | Out | Cutoff | Modalities | Effort | SO | FC | Cache / 24h | Web | Code | MCP | Comp. | Shell | Patch | Skills | ToolSearch | FT | Tiers | |---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| | `gpt-6-astra` | DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2026-04-30 | text, image → text | low, medium, high, xhigh, max | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F | | `gpt-5.6-sol` (= `gpt-5.6`) | DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2026-02-16 | text, image → text | none, low, **medium**, high, xhigh, max · mode pro | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F | | `gpt-5.6-terra` | DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2026-02-16 | text, image → text | none…max · pro | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F | | `gpt-5.6-luna` | DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2026-02-16 | text, image → text | none…max · pro | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F | | `gpt-5.5` | DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2025-12-01 | text, image → text | none, low, **medium**, high, xhigh | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F | | `gpt-5.5-pro` | DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2025-12-01 | text, image → text | medium, **high**, xhigh | ✔ | ✔ | ✘ / ✔ | ✔ | ✔ | ✔ | ✘ | ✔ | ✘ | ✘ | ✘ | ✘ | S B X | | `gpt-5.4` | DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2025-08-31 | text, image → text | **none**, low, medium, high, xhigh | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F | | `gpt-5.4-pro` | DOCUMENTED · LIVE_VERIFIED | 1,050,000 | 128,000 | 2025-08-31 | text, image → text | **medium**, high, xhigh | ✘ | ✔ | ✘ / ✘ | ✔ | ✘ | ✔ | ✔ | ✘ | ✔ | ✘ | ✔ | ✘ | S B X | | `gpt-5.4-mini` | DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | **none**, low, medium, high, xhigh | ✔ | ✔ | ✔ / ? | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | S B X F | | `gpt-5.4-nano` | DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | **none**, low, medium, high, xhigh | ✔ | ✔ | ✔ / ? | ✔ | ✔ | ✔ | ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | S B X | | `gpt-5.3-codex` | DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | low, medium, high, xhigh | ✔ | ✔ | ✔ / ? | ✔ | ✘ | ✘ | ✘ | ✔ | ✘ | ✔ | ✘ | ✘ | S F | | `gpt-5.2` | DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | **none**, low, medium, high, xhigh | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | S B X F | | `gpt-5.2-pro` | DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | (model default) | ✘ | ✔ | ✘ / ✘ | ✔ | ✘ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B | | `gpt-5.1` | DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2024-09-30 | text, image → text | **none**, low, medium, high | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✘ | ✘ | ✔ | ✘ | ✘ | ✘ | S B X F | | `gpt-5` | DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2024-09-30 | text, image → text | minimal, low, medium, high | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B X F | | `gpt-5-mini` | DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2024-05-31 | text, image → text | (default) | ✔ | ✔ | ✘ / ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B X F | | `gpt-5-nano` | DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2024-05-31 | text, image → text | (default) | ✔ | ✔ | ✔ / ? | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B X | | `gpt-5-pro` | DOCUMENTED · LIVE_VERIFIED | 400,000 | 272,000 | 2024-09-30 | text, image → text | **high** | ✔ | ✔ | ✘ / ✘ | ✔ | ✘ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B | | `o3` | DOCUMENTED · LIVE_VERIFIED | 200,000 | 100,000 | 2024-06-01 | text, image → text | (default) | ✔ | ✔ | ✔ / ? | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B X F | | `o3-pro` | DOCUMENTED · LIVE_VERIFIED | 200,000 | 100,000 | 2024-06-01 | text, image → text | (default) | ✔ | ✔ | ✘ / ✘ | ✔ | ✘ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S B | | `o4-mini` | DOCUMENTED · LIVE_VERIFIED · **DEPRECATED** (2026-10-23) | 200,000 | 100,000 | 2024-06-01 | text, image → text | (default) | ✔ | ✔ | ✔ / ? | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✔ | S B X F | | `gpt-4.1` | DOCUMENTED · LIVE_VERIFIED | 1,047,576 | 32,768 | 2024-06-01 | text, image → text | non-reasoning | ✔ | ✔ | ✔ / ✔ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✔ | S B F | | `gpt-4.1-mini` | DOCUMENTED · LIVE_VERIFIED | 1,047,576 | 32,768 | 2024-06-01 | text, image → text | non-reasoning | ✔ | ✔ | ✘ / ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✔ | S B F | | `gpt-4o` / `gpt-4o-mini` | DOCUMENTED · LIVE_VERIFIED | 128,000 | 16,384 | 2023-10-01 | text, image → text | non-reasoning | ✔ | ✔ | ✘ / ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✔ | S B F | | `chat-latest` | DOCUMENTED · LIVE_VERIFIED | 400,000 | 128,000 | 2025-08-31 | text, image → text | non-reasoning | ✔ | ✔ | ✘ / ✘ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | S | Data note: for the GPT-5.6 family, GPT-6 Astra and the Daybreak/cyber ids the records have `function_calling: true` **and** `tool_function_calling: false` — the second flag was read from a model-page tool list that omits "function calling" for those pages; treat `function_calling` as authoritative (function tools were used live on `gpt-5.6-luna` in the Agents API run). ## 3. Capability table — current Anthropic models From `capabilities` and `pricing` in `generated/models.json`. Cache min = minimum cacheable prefix tokens. Prefill = assistant prefill accepted. Sampling = `temperature`/`top_p`/`top_k` behaviour. ZDR = zero-data-retention eligible. | Model | Status | Lifecycle | Ctx | Out | Cutoff (reliable / training) | Thinking | Effort (default high) | SO / strict | Cache min · 1h | Fast | tool_choice any/tool | Prefill | Sampling | Web / Fetch / Code / MCP / PTC | Computer toolset GA / Browser | ZDR | Retirement not before | |---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| | `claude-fable-5-1` | DOCUMENTED · LIVE_VERIFIED | Active (latest) | 1,000,000 | 128,000 | Jun 2026 / Jun 2026 | adaptive, always on | low, medium, high, xhigh, max · per-message (beta) | ✔ / ✔ | 512 · ✔ | ✘ | ✘ (400) | ✘ | 400 if non-default | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✘ | 2027-09-01 | | `claude-mythos-5-1` | DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW | Active (invite) | 1,000,000 | 128,000 | Jun 2026 | adaptive, always on | low…max · per-message | ✔ / ✔ | 512 · ✔ | ✘ | ✘ | ✘ | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✘ | 2027-09-01 | | `claude-fable-5` | DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 1,000,000 | 128,000 | Jan 2026 | adaptive, always on | low…max | ✔ / ✔ | 512 · ✔ | ✘ | ✔ | ✘ | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✘ | 2027-06-09 | | `claude-mythos-5` | DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW | Active (invite) | 1,000,000 | 128,000 | Jan 2026 | adaptive, always on | low…max | ✔ / ✔ | 512 · ✔ | ✘ | ✔ | ✘ | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✘ | 2027-06-09 | | `claude-opus-5` | DOCUMENTED · LIVE_VERIFIED | Active (latest) | 1,000,000 | 128,000 | May 2026 | adaptive, on by default (disable except xhigh/max) | low…max · per-message | ✔ / ✔ | 512 · ✔ | ✔ $10/$50 | ✔ | ✘ | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✔ | 2027-07-24 | | `claude-opus-4-8` | DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 1,000,000 | 128,000 | Jan 2026 | adaptive (off by default) | low…max | ✔ / ✔ | 1,024 · ✔ | ✔ $10/$50 | ✔ | ? | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✔ | 2027-05-28 | | `claude-opus-4-7` | DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 1,000,000 | 128,000 | Jan 2026 | adaptive | low…max | ✔ / ✔ | 2,048 · ✔ | ✘ (removed) | ✔ | ? | 400 | ✔ ✔ ✔ ✔ ✔ | ✘ (beta `computer_20251124`) / ✘ | ✔ | 2027-04-16 | | `claude-opus-4-6` | DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 1,000,000 | 128,000 | May 2025 / Aug 2025 | adaptive; manual `enabled` deprecated | low, medium, high, max | ✔ / ✔ | 4,096 · ✔ | ✘ (runs standard) | ✔ | ✘ | accepted (not with thinking) | ✔ ✔ ✔ ✔ ✔ | ✘ / ✘ | ✔ | 2027-02-05 | | `claude-opus-4-5-20251101` (alias `claude-opus-4-5`) | DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 200,000 | 64,000 | May 2025 / Aug 2025 | extended (manual `budget_tokens`) | low, medium, high | ✔ / ✔ | 4,096 · ✔ | ✘ | ✔ | ? | accepted | ✔ ✔ ✔ ✔ ✔ | ✘ / ✘ | ✔ | 2026-11-24 | | `claude-sonnet-5` | DOCUMENTED · LIVE_VERIFIED | Active (latest) | 1,000,000 | 128,000 | Jan 2026 | adaptive, on by default | low…max | ✔ / ✔ | 1,024 · ✔ | ✘ | ✔ | ✘ (400 live) | 400 | ✔ ✔ ✔ ✔ ✔ | ✔ / ✔ | ✔ | 2027-06-30 | | `claude-sonnet-4-6` | DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 1,000,000 | 128,000 | Aug 2025 / Jan 2026 | adaptive; manual deprecated | low, medium, high, max | ✔ / ✔ | 1,024 · ✔ | ✘ | ✔ | ? | accepted | ✔ ✔ ✔ ✔ ✔ | ✘ / ✘ | ✔ | 2027-02-17 | | `claude-sonnet-4-5-20250929` (alias `claude-sonnet-4-5`) | DOCUMENTED · LIVE_VERIFIED · LEGACY | Active (legacy) | 200,000 | 64,000 | Jan 2025 / Jul 2025 | extended (manual) | ✘ (400) | ✔ / ✔ | 1,024 · ✔ | ✘ | ✔ | ? | accepted | ✔ ✔ ✔ ✔ ✔ | ✘ / ✘ | ✔ | 2026-09-29 | | `claude-haiku-4-5-20251001` (alias `claude-haiku-4-5`) | DOCUMENTED · LIVE_VERIFIED | Active (latest) | 200,000 | 64,000 | Feb 2025 / Jul 2025 | extended (manual) | ✘ (400) | ✔ / ✔ | 4,096 · ✔ | ✘ | ✔ | ✔ (live) | accepted | ✔ ✔ ✔ ✔ ✘ (PTC 400) | ✘ / ✘ | ✔ | 2026-10-15 | Modalities for every Claude model above: text + image + PDF in → text out. Tool support per exact tool `type` (e.g. `web_search_20260318` vs `20250305`) is in `generated/compatibility/model-tool-matrix.json`. ## 4. Capability table — current xAI Grok models From `capabilities`, `pricing` and `rate_limits` in `generated/models.json` (`kind: model`). Reasoning = always-on / optional / none and accepted `reasoning_effort` values (default in bold). Cache = automatic prefix cache (no markers). Tools = all server tools (`web_search`, `x_search`, `code_interpreter`, `file_search`, `mcp`, `image_generation`, attachment search) — `tools.json` lists every Grok text model as compatible. Prices standard (< 200k prompt) in / cached / out; ≥200k = 2× on all tokens. Batch = 20 % discount where supported. Priority = `service_tier: priority` 2×. RPS/TPM = documented Tier 0 → Tier 4. | Model | Status | Ctx | Out | Cutoff | Release | Modalities | Reasoning · effort | SO | FC | Cache | Tools | Compact / WS Responses | Batch | Priority | US 1.1× | In / cached / out | RPS · TPM (T0→T4) | Aliases | |---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| | `grok-4.6` | DOCUMENTED · LIVE_VERIFIED | 500,000 | none published | 2026-02-01 | 2026-08 | text, image → text | always on · low, medium, **high**, xhigh | ✔ | ✔ | ✔ (0.25×) | ✔ all | ✔ / ✔ | ✘ | 2× | ✔ (only model) | 2.00 / 0.50 / 6.00 | 150→500 · 50M→100M | — | | `grok-4.5` | DOCUMENTED · LIVE_VERIFIED | 500,000 | null | null | 2026-07 | text, image → text | always on · low, medium, high (`xhigh` = high) | ✔ | ✔ | ✔ (0.15×) | ✔ all | ✔ / ✔ | ✘ | 2× | ✘ | 2.00 / 0.30 / 6.00 | 150→500 · 50M→100M | `grok-4.5-latest`, `grok-build-latest` | | `grok-4.3` | DOCUMENTED · LIVE_VERIFIED | 1,000,000 | null | null | 2026-04 | text, image → text | optional · **low**, medium, high, xhigh + `none` (LIVE_DISCOVERED) | ✔ | ✔ | ✔ (0.16×) | ✔ all | ✔ / ✔ | −20 % | 2× | ✘ (eu-west-1 host) | 1.25 / 0.20 / 2.50 | 37→208 · 10M→85M | `grok-4.3-latest`; redirects `grok-3`, `grok-4-0709`, `grok-4-fast-*`, `grok-4-1-fast-*` | | `grok-4.20-0309-reasoning` | DOCUMENTED · LIVE_VERIFIED | 1,000,000 | null | null | 2026-03-09 | text, image → text | always on · `reasoning_effort` → **400** | ✔ | ✔ | ✔ | ✔ all | ✔ / ✔ | −20 % | 2× | ✘ | 1.25 / 0.20 / 2.50 | 37→208 · 10M→85M | `grok-4.20`, `grok-4.20-reasoning(-latest)`, `grok-4.20-0309`, `grok-4.20-beta*`, `grok-4.20-experimental-beta*`, `grok-4.20-reasoning-gv2` (15) | | `grok-4.20-0309-non-reasoning` | DOCUMENTED · LIVE_VERIFIED | 1,000,000 | null | null | 2026-03-09 | text, image → text | none (0 reasoning tokens) | ✔ | ✔ | ✔ | ✔ all | ✔ / ✔ | −20 % | 2× | ✘ | 1.25 / 0.20 / 2.50 | 37→208 · 10M→85M | `grok-4.20-non-reasoning(-latest)`, `-beta-*`, `-gv2` (8) | | `grok-4.20-multi-agent-0309` | DOCUMENTED · **BETA** · LIVE_VERIFIED | 1,000,000 | null | null | 2026-03 | text, image → text | multi-agent · effort = agent count (low/medium 4, high/xhigh 16; default medium) | ✔ | ✔ | ✔ | ✔ all | Responses only (Chat → 400) | −20 % | not documented | ✘ | 1.25 / 0.20 / 2.50 (per agent token) | 9→56 · 2.5M→21M | `grok-4.20-multi-agent(-latest)`, `-beta-*`, `-experimental-beta-*` (6) | | `grok-build-0.1` | DOCUMENTED · **PREVIEW** · LIVE_VERIFIED | 256,000 | null | null | 2026-05 ('early access') | text, image → text | always on · `reasoning_effort` → 400 | ✔ | ✔ | ✔ (0.2×) | ✔ all | ✔ / ✔ | ✘ | 2× | ✘ | 1.00 / 0.20 / 2.00 | 37→208 · 10M→85M | `grok-code-fast-1`, `grok-code-fast`, `grok-code-fast-1-0825` | Every Grok text model also accepts the Anthropic-compatible `/v1/messages` (DEPRECATED), `POST /v1/responses/compact`, `wss://api.x.ai/v1/responses`, deferred Chat Completions and `/v1/tokenize-text`. `usage.cost_in_usd_ticks` (1 USD = 10^10 ticks) is returned on every call. Fine-tuning: none offered. `logprobs`: ignored on 4.20+. ### 4b. xAI media, voice and embedding models | Model | Status | Modalities | Price | Endpoints | Notes | |---|---|---|---|---|---| | `grok-imagine-image` | DOCUMENTED · LIVE_VERIFIED | text+image → image | $0.02 / image | `POST /v1/images/generations`, `/edits` (JSON only), Batch, gRPC | prompt limit 16,000 tokens; alias `grok-imagine-image-2026-03-02`; JPEG only | | `grok-imagine-image-2.0` | DOCUMENTED · LIVE_VERIFIED | text+image → image | $0.04–$0.08 (quality low/medium × resolution 1k/1.5k/2k; catalogue default $0.06) | same | 5 reference images, 21:9 / 5:2 aspect ratios (Aug 2026); `quality` default moved to `auto` | | `grok-imagine-image-quality` | DOCUMENTED · **DEPRECATED** · LIVE_VERIFIED | text+image → image | $0.05 | same (no Batch) | retires **2026-11-02** → `grok-imagine-image-2.0` quality=low; aliases `-20260403`, `-latest`, `grok-imagine-image-pro` | | `grok-imagine-video` | DOCUMENTED · LIVE_VERIFIED | text+image+video → video | $0.05 / s | `POST /v1/videos/generations`, `/edits`, `/extensions`, `GET /v1/videos/{request_id}` (202 pending) | 1–15 s, 480p–1080p, audio; batch URLs expire 1 h | | `grok-imagine-video-1.5` | DOCUMENTED · LIVE_VERIFIED | text+image+audio → video | $0.08 / s | same | native 1080p, reference-to-video, `reference_audios`, `last_frame`; aliases `-preview`, `-2026-05-30` | | `grok-voice-think-fast-2.0` (`grok-voice-latest`) | DOCUMENTED | audio+text ↔ audio+text | $0.08 / min ($4.80 / h) + $0.004 per text item | `wss://api.x.ai/v1/realtime`, `POST /v1/realtime/client_secrets`, SIP `/v2/phone-numbers`, `/v1/realtime/calls/{id}/refer` and `…/hangup` | OpenAI Realtime event vocabulary; tools in-session; 120-min sessions; `-1.0` LEGACY; not in `GET /v1/models` | | `grok-voice-transcribe-2.0` / `-1.0` | DOCUMENTED | audio → text | $0.10 / h REST, $0.20 / h streaming | `POST /v1/stt`, `wss://api.x.ai/v1/stt` | 25 languages, diarization, `smart_turn`; default model ambiguous in docs | | text-to-speech (no model id) | DOCUMENTED | text → audio | $15 / 1M characters | `POST /v1/tts`, `GET /v1/tts/voices`, `wss://api.x.ai/v1/tts` | 28 voices, `language` required; custom voices Enterprise | | `grok-embedding-small` | DOCUMENTED · **ACCOUNT_RESTRICTED** | text → embedding | not published | `POST /v1/embeddings`, `GET /v1/embedding-models` (empty) | 404 for this team; Collections index model | ## 5. Capability table — current Gemini models From `capabilities`, `pricing` and `deprecation` in `generated/models.json` (kinds `stable`, `preview`, `alias`, `agent`). Thinking = `thinkingLevel` values (default in bold; `budget` = 2.5-style `thinkingBudget`). Tools: FC function calling, Search (`googleSearch`), Maps, URL (`urlContext`), Code (`codeExecution`), CU (`computerUse`, preview), FS (`fileSearch`). Cache = implicit / explicit (`cachedContents`). Tiers = B batch (50 %), X flex (50 %), P priority (1.8×). Prices standard in / cached / out (>200k tier for Pro). Free = free-tier availability per the pricing page. | Model | Status | Lifecycle | Release | Ctx in / out | Modalities | Thinking | SO | FC | Search / Maps / URL / Code / CU / FS | Cache impl / expl | Tiers | Live | Free | In / cached / out | Shutdown | |---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| | `gemini-3.8-flash` | DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED | Stable (GA 2026-09-02); `gemini-flash-latest` → here (observed) | 2026-09-02 | 1,048,576 / 65,536 | text, image, audio, video, PDF → text | low, **medium**, high (`minimal` → 400) | ✔ | ✔ | ✔ ✔ ✔ ✔ ✔(P) ✔ | 4,096 / ✔ | B X P | ✘ | ✔ | 0.75 / 0.075 / 3.75 (→ 1.50 / 0.15 / 7.50 on 2027-01-01) | none announced | | `gemini-3.7-flash` | DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-08-13) | 2026-08-13 | 1,048,576 / 65,536 | same | low, **medium**, high | ✔ | ✔ | ✔ ✔ ✔ ✔ ✔(P) ✔ | 4,096 / ✔ | B X P | ✘ | ✔ | 0.75 / 0.075 / 3.75 (intro) | none | | `gemini-3.6-flash` | DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-07-21) | 2026-07-21 | 1,048,576 / 65,536 | same | minimal, low, **medium**, high; `thinkingBudget: 0` allowed | ✔ | ✔ | ✔ ✔ ✔ ✔ ✔(P) ✔ | 4,096 / ✔ | B X P | ✘ | ✔ | 0.75 / 0.075 / 3.75 (intro) | none | | `gemini-3.5-flash` | DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED | Stable (GA 2026-05-19) | 2026-05-19 | 1,048,576 / 65,536 | same | minimal, low, **medium**, high; budget 0 allowed | ✔ | ✔ | ✔ ✔ ✔ ✔ ✔(P) ✔ | 4,096 / ✔ | B X P | ✘ | ✔ (caching `limit: 0` live) | 1.50 / 0.15 / 9.00 | none | | `gemini-3.5-flash-lite` | DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED | Stable (GA 2026-07-21) | 2026-07-21 | 1,048,576 / 65,536 | same | **minimal**, low, medium, high (budget 0 → 400) | ✔ | ✔ | ✔ ✔ ✔ ✔ ✔(P) ✔ | ? / ✔ | B X P | ✘ | ✔ (no caching) | 0.30 / 0.03 / 2.50 | none | | `gemini-3.1-pro-preview` (+ `-customtools`) | DOCUMENTED · LIVE_DISCOVERED · PREVIEW · ACCOUNT_RESTRICTED | Preview; `gemini-pro-latest` → here | 2026-02-19 | 1,048,576 / 65,536 | same | low, medium, **high** | ✔ | ✔ | ✔ ✔ ✔ ✔ ✘ ✔ (AI Studio only) | 4,096 / ✔ | B X P | ✘ | ✘ (paid only) | 2.00 / 0.20 / 12.00; >200k 4.00 / 0.40 / 18.00 | none | | `gemini-3.1-flash-lite` | DOCUMENTED · LIVE_DISCOVERED · **DEPRECATED** | Deprecated | 2026-05-07 | 1,048,576 / 65,536 | same | minimal, low, medium, high (default not stated) | ✔ | ✔ | ✔ ✔ ✔ ✔ ✘ ✔ | ? / ✔ | B X P | ✘ | ✔ | 0.25 (0.50 audio) / 0.025 / 1.50 | **2027-05-07** → `gemini-3.5-flash-lite` | | `gemini-3-flash-preview` | DOCUMENTED · LIVE_DISCOVERED · PREVIEW | Preview (replacement `gemini-3.6-flash`, no date) | 2025-12-17 | 1,048,576 / 65,536 | same | minimal, low, medium, **high** | ✔ | ✔ | ✔ ✔ ✔ ✔ ✔ ✔ | 4,096 / ✔ | B X P | ✘ | ✔ | 0.50 (1.00 audio) / 0.05 / 3.00 | none | | `gemini-2.5-pro` | DOCUMENTED · LIVE_DISCOVERED (deprecations.json: ACCOUNT_RESTRICTED) | 'no longer available to new users' (404 live) | 2025-06-17 | 1,048,576 / 65,536 | text, image, audio, video, PDF → text | budget 128–32,768 (cannot disable) | ✔ | ✔ | ✔ ✔ ✔ ✔ ✘ ✔ | 2,048 / ✔ | B X P | ✘ | listed | 1.25 / 0.125 / 10.00; >200k 2.50 / 0.25 / 15.00 | undated | | `gemini-2.5-flash` | DOCUMENTED · LIVE_DISCOVERED | 'no longer available to new users' | 2025-06-17 | 1,048,576 / 65,536 | text, image, audio, video → text | budget 0–24,576 | ✔ | ✔ | ✔ ✔ ✔ ✔ ✘ ✔ | 2,048 / ✔ | B X P | ✘ | ✔ | 0.30 (1.00 audio) / 0.03 / 2.50 | undated | | `gemini-2.5-flash-lite` | DOCUMENTED · LIVE_DISCOVERED · ACCOUNT_RESTRICTED | 404 live → `gemini-3.5-flash-lite` | 2025-07-22 | 1,048,576 / 65,536 | same | off by default; budget 512–24,576 | ✔ | ✔ | ✔ ✔ ✔ ✔ ✘ ✔ | ? / ✔ | B X P | ✘ | ✔ | 0.10 (0.30 audio) / 0.01 / 0.40 | undated | | `gemma-4-31b-it`, `gemma-4-26b-a4b-it` | DOCUMENTED · LIVE_DISCOVERED (26B LIVE_VERIFIED) | Stable | 2026-04-02 | 262,144 / 32,768 | text → text | live `thinking: true` (`thoughtsTokenCount: 5`) | ? | ? | ? (grounding 'not available') | ✘ / ✘ | — | ✘ | **free only** | free; paid tier not available | none | | `gemini-3.8-live` / `-extended-thinking` | DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-09-15) | 2026-09-15 | 131,072 / 65,536 | text, image, audio, video → audio (+text) | interleaved (omit config) / LOW-MEDIUM-HIGH | ✘ | ✔ (async) | ✔ ✘ ✘ ✘ ✘ ✘ | ? / ✘ | — | ✔ | ✔ | text 0.75 / audio 3.00 in; text 4.50 / audio 12.00 out | none | | `gemini-3.1-flash-live-preview` | DOCUMENTED · LIVE_DISCOVERED · PREVIEW | 'legacy Live preview' → 3.8-live | 2026-03-11 | 131,072 / 65,536 | same | MINIMAL…HIGH | ✘ | ✔ (sync) | ✔ ✘ ✘ ✘ ✘ ✘ | — | — | ✔ | ✔ | same as 3.8-live | none | | `gemini-2.5-flash-native-audio-preview-12-2025` | DOCUMENTED · LIVE_DISCOVERED · PREVIEW | Preview → 3.8-live | 2025-12-12 | 131,072 / 8,192 | text, audio, video → text, audio | `thinkingBudget` | ✘ | ✔ | ✔ ✘ ✘ ✘ ✘ ✘ | — | — | ✔ | ✔ | 0.50 text / 3.00 audio-video in; 2.00 / 12.00 out | none | | `gemini-embedding-2` | DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED | Stable (GA 2026-04-22) | 2026-04-22 | 8,192 / 1 | text, image, audio, video, PDF → embedding (128–3,072 dims) | — | — | — | — | — | B | ✘ | ✔ | 0.20 text / 0.45 image / 6.50 audio / 12.00 video | none | | `gemini-embedding-001` | DOCUMENTED · LIVE_DISCOVERED · DEPRECATED | Deprecated | 2025-07-14 | 2,048 / 1 | text → embedding | — | — | — | — | — | — | ✘ | — | not on pricing page | **2028-05-14** → `gemini-embedding-2` | ### 5b. Gemini media and agent models | Model | Status | Lifecycle | Modalities | Price | Notes | |---|---|---|---|---|---| | `gemini-3.1-flash-image` (Nano Banana 2) | DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-05-28) | text, image, video, PDF → text, image | in 0.50 / text out 3.00 / image out 60.00 per 1M → $0.045 (0.5K) … $0.151 (4K) per image; batch 50 % | `imageConfig` 512/1K/2K/4K, 14 aspect ratios, web + image search grounding; thinking MINIMAL/HIGH; no free tier; live limits 65,536 / 65,536 (docs 131,072 / 32,768) | | `gemini-3.1-flash-lite-image` | DOCUMENTED · LIVE_DISCOVERED | Stable (2026-06-30) | same | 0.25 / 1.50 / 30.00 → $0.034 (1K) | 1K only, no search; C2PA + SynthID | | `gemini-3-pro-image` (Nano Banana Pro) | DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-05-28); `-preview` / `nano-banana-pro-preview` RETIRED 2026-06-25 but still listed | text, image → text, image | 2.00 / 12.00 / 120.00 → $0.134 (1K/2K), $0.24 (4K) | Google Search grounding; priority 216 image out | | `gemini-2.5-flash-image` (Nano Banana) | DOCUMENTED · LIVE_DISCOVERED · **DEPRECATED** | shutdown **2026-10-02** | text, image → text, image | $0.039 / image (30.00 / 1M) | replacement `gemini-3.1-flash-image` / `-lite-image` | | `veo-3.1-generate-preview` / `-fast-` / `-lite-` | DOCUMENTED · LIVE_DISCOVERED · PREVIEW (lite LIVE_VERIFIED GET) | Preview; Veo 2.0/3.0 RETIRED 2026-06-30 | text, image → video + audio | $0.40 / $0.40 / $0.60 · $0.10 / $0.12 / $0.30 · $0.05 / $0.08 / — per second (720p / 1080p / 4K) | `:predictLongRunning` + operations; 4/6/8 s; no free tier | | `gemini-omni-1.1-flash` | DOCUMENTED · LIVE_DISCOVERED | Stable (GA 2026-08-27); `gemini-omni-flash-preview` DEPRECATED → 2026-09-30 | text, image, video → video | in 1.50 / text out 9.00 / video out 17.50 per 1M (≈ $0.10 / s at 720p) | Interactions only | | `lyria-3.5` · `lyria-3-clip-preview` · `lyria-3-pro-preview` · `lyria-realtime-exp` | DOCUMENTED · LIVE_DISCOVERED (3.5 LIVE_VERIFIED GET; RealTime LIVE_VERIFIED WS) | Stable (GA 2026-09-03) · Preview · Preview → 3.5 · Experimental | text, image → music + lyrics | $0.08 / song · $0.04 / 30-s clip · $0.08 · unpriced | no free tier; RealTime `BidiGenerateMusic` WebSocket | | `gemini-3.1-flash-tts-preview` · `gemini-2.5-flash-preview-tts` · `gemini-2.5-pro-preview-tts` | DOCUMENTED · LIVE_DISCOVERED · PREVIEW | Preview (2.5 → 3.1) | text → audio | $1 / $20 · $0.50 / $10 · $1 / $20 per 1M (text in / audio out) | 30 voices, ≤2 speakers; streaming TTS on 3.1 only; free tier 10 req/day observed | | `gemini-3.5-transcribe` · `gemini-3.5-transcribe-live` · `gemini-3.5-live-translate-preview` | DOCUMENTED · LIVE_DISCOVERED (unary LIVE_VERIFIED) | Stable (Aug 2026) · Stable · Preview | audio → text · audio → text · audio → audio + text | ≈ $0.005 / min · ≈ $0.009 / min · ≈ $0.037 / min | `audioTranscriptionConfig` (diarization, word timestamps, custom vocabulary) | | `deep-research-preview-04-2026` · `deep-research-max-preview-04-2026` · `deep-research-pro-preview-12-2025` | DOCUMENTED · LIVE_DISCOVERED · PREVIEW (04-2026 LIVE_VERIFIED) | Preview agents | text, image, audio, video, PDF → text (+ image) | list-rate tokens (incl. intermediate) + tool fees; est. $1–3 / $3–7 per task | Interactions `agent`, `background: true` mandatory, ≤60 min; listed as models (`generateContent` advertised) | | `antigravity-preview-09-2026` · `antigravity-preview-05-2026` | DOCUMENTED · LIVE_DISCOVERED · PREVIEW (05-2026 DEPRECATED → 2026-10-05) | Preview coding agent | text → text | list-rate tokens; sandbox compute unbilled in preview | 09-2026 renamed built-in tools (PascalCase params) | | `gemini-2.5-computer-use-preview-10-2025` | DOCUMENTED · LIVE_DISCOVERED · PREVIEW (deprecations.json: LEGACY) | replaced by built-in computer use on 3.x | text, image → text | 1.00 / 5.00 | browser only; 429 `limit: 0` live | | `gemini-robotics-er-2-preview` / `-streaming-preview` | DOCUMENTED · LIVE_DISCOVERED · PREVIEW | Preview (1.5 / 1.6 RETIRED) | text, image, audio, video → text | 1.00 / 0.10 / 5.00 (streaming: pricing section empty) | embodied reasoning | | `aqa` | LIVE_DISCOVERED | Stable (PaLM-era) | text → text | not priced | `:generateAnswer` only (LIVE_VERIFIED), English | ## 6. Naming and versioning conventions | Topic | OpenAI | Anthropic | xAI | Gemini | |---|---|---|---|---| | Id families | `gpt-[.][-][-]`, `o[-mini\|-pro]`, `gpt--codex`, `gpt-image-*`, `gpt-realtime-*`, `gpt-audio-*`, `gpt-live-*`, `text-embedding-*`, `omni-moderation-*`, `sora-*`; codenames for GPT-5.6 tiers (`sol`/`terra`/`luna`) and GPT-6 (`astra`); cyber/Daybreak aliases | `claude--[-][-]` (`claude-opus-5`, `claude-sonnet-4-6`); families Fable (frontier), Mythos (invite research line), Opus, Sonnet, Haiku | `grok-.` (`grok-4.6`, `grok-4.3`), dated variant ids `grok-4.20-0309-reasoning` / `-non-reasoning` / `grok-4.20-multi-agent-0309`, product lines `grok-build-`, `grok-imagine-image[-\|-quality]`, `grok-imagine-video[-]`, `grok-voice-think-fast-`, `grok-voice-transcribe-`, `grok-embedding-small`; TTS has no model id | `gemini-.-[-][-preview[-MM-YYYY]]` (`gemini-3.8-flash`, `gemini-3.1-pro-preview`, `gemini-3.1-flash-image`, `gemini-3.8-live`, `gemini-3.5-transcribe`), `gemini-embedding-`, `gemma-4--it`, `veo---generate-preview`, `lyria-`, `imagen-*` (retired); agents `deep-research-*`, `antigravity-preview-MM-YYYY` | | Alias vs snapshot | **Undated alias** (`gpt-5.4`) → **dated snapshot** (`gpt-5.4-2026-03-05`); the alias moves when a new snapshot ships; response `model` echoes the snapshot. GPT-5.6/6 tier ids have **no dated snapshot**; `gpt-5.6` is an alias of `gpt-5.6-sol` that **404s on `GET /v1/models`** but works on POST. `*-latest` ids (`chat-latest`, `chatgpt-image-latest`, `codex-mini-latest`) float. | Since Claude 4.6 the canonical id is **undated** (`claude-opus-4-6`, `claude-sonnet-5`, `claude-fable-5-1`) and is itself the snapshot. Older lines keep **dated snapshots** with undated **aliases** (`claude-haiku-4-5` → `claude-haiku-4-5-20251001`); aliases follow the newest snapshot. No `-latest` suffix. Cloud ids differ per platform (`anthropic.claude-opus-5` on Bedrock, `claude-opus-5@…` on Vertex). | Documented scheme: `` = latest stable, `-latest`, `-` pinned — in practice the **dated 4.20 ids are canonical** and the bare names (`grok-4.20`) are aliases (15 aliases on one model); `grok-4.6` has no `-latest`. **Retired ids redirect** (`grok-3`, `grok-4-0709`, `grok-4-fast*`, `grok-4-1-fast*` → `grok-4.3`; `grok-code-fast-1` → `grok-build-0.1`) and appear in the catalogue's `aliases[]`. `grok-voice-latest` routes to `grok-voice-think-fast-2.0` since 2026-08-05. | **Stable** ids are undated (`gemini-3.8-flash`, 'usually don't change'); **`-preview`** ids may carry a month suffix (`gemini-2.5-flash-native-audio-preview-12-2025`) and can be deprecated with two weeks' notice; **`-latest` aliases** (`gemini-flash-latest`, `gemini-pro-latest`, `gemini-flash-lite-latest`, `gemini-2.5-flash-native-audio-latest`) are hot-swapped (2-week e-mail notice) — `gemini-flash-latest` moved 3 Flash Preview → 3.5 Flash → 3.8 Flash (observed via `modelVersion`); `-001` suffixes only on 2.0 / embedding-001 / Veo 3.0; `-exp` experimental; `gemini-3-pro-preview` was **repointed** to 3.1 Pro on 2026-03-09. | | Version pinning advice | pin the dated snapshot in production; `GET /v1/models/{id}` exposes `shutdown_date` | pin the undated id for 4.6+ (already immutable); for 4.5 lines pin the dated snapshot | pin the dated 4.20 ids or the bare `grok-4.x` id (no snapshots exist); check `response.model` — a retired id silently returns its replacement | pin the stable id; avoid `-latest` in production; `modelVersion` in every response tells you what actually ran | | Tool versioning | tool `type` strings are undated (`web_search`, `code_interpreter`) with a few dated snapshots (`web_search_2025_08_26`) | every server/Anthropic-defined tool `type` carries `_YYYYMMDD` (`web_search_20260318`, `code_execution_20260521`) | undated OpenAI-style `type` strings (`web_search`, `x_search`, `code_interpreter`/`code_execution`, `file_search`/`collections_search`, `mcp`, `shell`, `tool_search`) | undated camelCase keys (`googleSearch`, `urlContext`, `codeExecution`, `fileSearch`, `computerUse`); legacy `googleSearchRetrieval` | | API surface versioning | none (`openai-version: 2020-10-01` response header); breaking changes via new endpoints and `OpenAI-Beta` surfaces | `anthropic-version: 2023-06-01` required; features gated by dated `anthropic-beta` values that graduate to GA | none — no version or beta headers; `docs.x.ai/developers/release-notes` is the changelog; gRPC proto v6 | path version `/v1beta` (86 methods) vs `/v1` (47); Interactions optional `Api-Revision: 2026-05-20`; breaking schema change 2026-05-26/06-08 (`outputs` → `steps`) | | Knowledge cutoff | single date per model page (`knowledge_cutoff`), e.g. gpt-6-astra 2026-04-30 | two dates: **reliable knowledge** and **training data** (`knowledge_cutoff.reliable` / `.training_data`) | published only for `grok-4.6` (2026-02-01); null for every other Grok record | not stated on the API model pages; **January 2025** from the Gemini 3 / 2.5 model cards; June 2025 for `gemini-2.5-flash-image`; null for 64 of 97 records | | Discovery | `GET /v1/models` lists 136 ids incl. snapshots, `owned_by`, `created`, `shutdown_date`; fine-tuned ids `ft:…` | `GET /v1/models` lists 11 ids with `capabilities` (thinking types, effort levels, context_management strategies, structured_outputs, citations, pdf_input) | `GET /v1/models` 12 ids; typed catalogues `/v1/language-models`, `/v1/image-generation-models`, `/v1/video-generation-models`, `/v1/embedding-models` expose **live prices in ticks**, `aliases[]`, `input_modalities`, fingerprints; voice models absent | `GET /v1beta/models` 58 ids (`inputTokenLimit`, `outputTokenLimit`, `supportedGenerationMethods`, `thinking`, sampling defaults, `version`); `/v1/models` 22; agents listed as models; shut-down previews still listed | ## 7. Deprecation and retirement policies | Aspect | OpenAI | Anthropic | xAI | Gemini | |---|---|---|---|---| | Vocabulary | **Deprecated** = retirement announced with a shutdown date; **Legacy** = no longer updated, will be deprecated later; **Sunset/shut down** = no longer accessible | **Active** / **Legacy** (no more updates) / **Deprecated** (replacement + retirement date, ≥60 days) / **Retired** (requests fail) | **Available** (on `/developers/models` + `/pricing` + `GET /v1/models`) / **Deprecated (notice period)** (migration guide; ~60-day notice observed) / **Retired** (removed from the catalogue but the slug **keeps resolving and redirects** to a replacement billed at the replacement's price) / undocumented **Legacy** (older ids vanished from docs) | **Stable (GA)** / **Preview** (production allowed, tighter limits, ≥2 weeks notice) / **Latest alias** / **Experimental** / **Legacy** (existing customers) / **Deprecated** (earliest shutdown date announced) / **Shut down**; API enum `ModelStatus.modelStage` (EXPERIMENTAL, PREVIEW, STABLE, LEGACY, DEPRECATED, RETIRED) | | Notice period | dated per announcement (typically ≥ 6 months for flagship snapshots, shorter for previews); `shutdown_date` also exposed live on `GET /v1/models` | ≥ 60 days from deprecation to retirement; every active model carries a **"not sooner than"** retirement commitment (Fable 5.1 ≥ 2027-09-01, Opus 5 ≥ 2027-07-24, Sonnet 5 ≥ 2027-06-30, Haiku 4.5 ≥ 2026-10-15) | **no published minimum**; observed ~60 days (`grok-imagine-image-quality` announced 2026-09-02 → 2026-11-02); the May-15 batch was announced via a migration guide with automatic redirects | previews ≥ 2 weeks; stable models get an earliest-shutdown date on the deprecations page (e.g. `gemini-3.1-flash-lite` 2026-05-07 → 2027-05-07 = 12 months); 'exact date communicated in advance'; Vertex has its own schedule | | Replacement mapping | each row names a replacement (`o4-mini` → `gpt-5.6-terra`, `gpt-image-1` → `gpt-image-2`, `whisper-1` → `gpt-transcribe`) | each retirement names a replacement (`claude-opus-4-1-20250805` → `claude-opus-4-8`, `claude-3-haiku-20240307` → `claude-haiku-4-5-20251001`) | migration guides name the replacement **and the effort setting** (`grok-4-fast-reasoning` → `grok-4.3` `reasoning_effort=low`; `grok-4-fast-non-reasoning` → `grok-4.3` `none`; `grok-code-fast-1` → `grok-build-0.1`; `grok-imagine-image-quality` → `grok-imagine-image-2.0` `quality=low`) | deprecations page names a replacement (`gemini-3.1-flash-lite` → `gemini-3.5-flash-lite`, `gemini-2.5-flash-image` → `gemini-3.1-flash-image`, `gemini-embedding-001` → `gemini-embedding-2`, `text-embedding-004` → `gemini-embedding-2`); live 404 messages also name one (`gemini-2.5-flash-lite` → `gemini-3.5-flash-lite`) | | Upcoming (after 2026-09-18) | 2026-09-24 Videos API + Sora 2; 2026-09-28 gpt-3.5-turbo-instruct, babbage-002, davinci-002; 2026-10-01 gpt-5.4-cyber; 2026-10-23 o1, o1-pro, o3-mini, o4-mini, gpt-4, gpt-4-turbo, gpt-4.1-nano, gpt-image-1; 2026-11-30 Evals API, Agent Builder, reusable prompts; 2026-12-01 gpt-image-1-mini/1.5; 2026-12-11 gpt-5 2025 snapshots, o3, o3-pro; 2027-01-06 no new fine-tuning jobs; 2027-01-20 legacy audio/realtime; 2027-02-26 whisper-1, gpt-4o-transcribe | no dated model retirements pending; `claude-mythos-preview` deprecated 2026-06-09 (date TBA). Feature-level: `context-1m-2025-08-07` header retired 2026-04-30; fast mode removed on Opus 4.6/4.7; sampling params and manual thinking deprecated (400 on 4.7+); prefill deprecated (400 on 4.6+) | **2026-09-21 12:00 PT** `x_search` per-call billing → per post / per profile; **2026-11-02** `grok-imagine-image-quality` retires (redirect to 2.0 low); `/v1/messages` Anthropic compat 'fully deprecated' (no date); `logprobs` ignored on 4.20+ | **2026-09-30** `gemini-omni-flash-preview`; **2026-10-02** `gemini-2.5-flash-image`; **2026-10-05** `antigravity-preview-05-2026`; **September 2026** standard API keys rejected (auth keys only); **2027-05-07** `gemini-3.1-flash-lite`; **2028-05-14** `gemini-embedding-001`; replacement recommended without date: `gemini-3-flash-preview`, `gemini-3.1-flash-live-preview`, 2.5 native-audio / TTS previews, `lyria-3-pro-preview`; sampling params deprecated 2026-07-21 | | Recently retired | Assistants API (2026-08-26), DALL·E 2/3 (2026-05-12), `text-moderation-*`, search-preview snapshots (2026-07-23), codex snapshots ≤ 5.2, deep-research models, computer-use-preview | Opus 4.1 (2026-08-05), Sonnet 4 / Opus 4 (2026-06-15), Claude 3 Haiku (2026-04-20), Claude 3.5/3.7, Claude 3 Opus/Sonnet, 2.x, 1.x, Instant; `/v1/complete` effectively retired (400) | **2026-05-15**: `grok-3`, `grok-4-0709`, `grok-4-fast-reasoning`, `grok-4-fast-non-reasoning`, `grok-4-1-fast-reasoning`, `grok-4-1-fast-non-reasoning`, `grok-code-fast-1`, `grok-imagine-image-pro` (all redirect); `grok-2-image` 404; Live Search (chat `search_parameters`) 410; `/v1/completions`, `/v1/complete` 400 | Gemini 2.0 Flash/-Lite (2026-06-01); 2.5 previews (2025-11 → 2026-03); `gemini-3-pro-preview` (2026-03-09, repointed); `gemini-3.1-flash-lite-preview` (2026-05-25); image previews `gemini-3.1-flash-image-preview`, `gemini-3-pro-image-preview` / `nano-banana-pro-preview` (2026-06-25); Imagen 3/4 (2026-08-17); Veo 2.0/3.0 (2026-06-30); half-cascade Live models (2025-12-09); `text-embedding-004` (2026-01-14), `embedding-001`, `embedding-gecko-001`, `gemini-embedding-exp*` (2025-10-30); robotics ER 1.5/1.6; model tuning (May 2025); LearnLM | | After retirement | requests fail; fine-tuned models retire with their base | 404 `not_found_error` on the Claude API; several retired ids remain **available on Bedrock / Vertex** | slug still resolves: `GET /v1/models/grok-3` returns the **grok-4.3** object; requests are served and billed at the replacement's rates (`response.model` shows it); `grok-2-image` is a hard 404 | 404 `NOT_FOUND` on generation (`Model is not found for api version v1beta`), yet shut-down ids **remain in `GET /v1beta/models`** (e.g. `gemini-3-pro-preview`, `gemini-3.1-flash-lite-preview`, image previews) — listing ≠ availability | | Machine-readable | `generated/deprecations.json` (165 flat OpenAI records: model, shutdown_date, replacement, phase) | `generated/deprecations.json` (1 Anthropic record with `models[]`, `active_models_retirement_commitments{}`, `api_features[]`) | `generated/deprecations.json` (1 xAI record: `models[]` with `behavior_after` and `our_probe`, `api_features[]`, `release_notes_digest[]` by month); `models.json` kinds `retired_redirect`, `legacy`, `alias` | `generated/deprecations.json` (1 Gemini record: 60+ `models[]` with `released`/`shutdown`/`replacement`/`live_listed`, `alias_history[]`, `api_features[]`, `active_models_no_shutdown_announced[]`) | ## 8. Live discovery vs docs - **OpenAI**: 136 ids in `GET /v1/models`; aliases `gpt-5.6`, `gpt-5.5-cyber`, `gpt-5-search-api-2025-10-14` behave as id-only/alias records (`record_kind: alias|id_only`); `gpt-5.6` echoes `gpt-5.6-sol`; `gpt-5.4-nano` echoes `effort: none` when omitted; `gpt-6-astra` rejects `reasoning.effort: none`. - **Anthropic**: 11 ids in `GET /v1/models` (Mythos ids absent → 404); `claude-fable-5-1` returned Opus-class rate-limit headers (10M ITPM) although docs list 4M for the Fable class (account-specific); Haiku 4.5 tolerated edited/removed thinking blocks and prefill-with-thinking (docs say 400) — graceful degradation, do not rely on it. - **xAI**: 12 ids in `GET /v1/models` (7 language, 3 image, 2 video; voice/embedding models absent); redirects verified (`GET /v1/models/grok-3`, `/grok-4-fast`, `/v1/language-models/grok-4-0709` → grok-4.3; `grok-code-fast-1` → grok-build-0.1); `grok-4.3` accepts `reasoning_effort: none` (undocumented) while `grok-4.20-0309-reasoning` and `grok-build-0.1` reject `reasoning_effort` although docs/parameters list them as compatible; `x_search` emits `custom_tool_call` items (docs: `x_search_call`); `GET /v1/responses/{id}` returns 200 for `store:false` ids; rate-limit headers exist but are undocumented (7,200 / 1,800 requests per minute vs documented 150 / 37 RPS); `eu-west-1.api.x.ai` serves grok-4.3 undocumented; every text call bills 70–180 reasoning tokens for "Reply with OK.". - **Gemini**: 58 ids in `GET /v1beta/models` vs 22 in `GET /v1/models` (docs claim parity); `gemini-flash-latest` resolved to `gemini-3.8-flash` (changelog last said 3.5 Flash); Gemini 2.5 Pro/Flash/Flash-Lite return 200 on `models.get` but **404 'no longer available to new users'** on generation (undocumented policy); Pro models, image/video/music models, `cachedContents` create and Batch create return 429 `limit: 0` / 400 FAILED_PRECONDITION on the free tier (`ACCOUNT_RESTRICTED`); explicit-cache minimum is 1,024 tokens live (docs' 4,096 is the implicit-cache threshold); `thoughtSignature` validation can be bypassed with the documented dummy strings; token limits differ from the model pages for image, Lyria, Deep Research, translate and computer-use ids (see §9); `gemini-3.8-live` reports `version: 3.1-flash-live-03-2026`. ## 9. Data inconsistencies found during synthesis Reported for the fragment owners (fix in `generated/fragments/**`, never in merged files). Items 1–16 carry over from the two-provider synthesis; 17+ were found while adding xAI and Gemini. 1. `models.json` `o3`: `pricing.standard` = $2 / $0.5 / $8 but `pricing["model_page:Text tokens"]` = $1 / $0.25 / $4 (two sources disagree; other models agree). 2. `models.json` OpenAI GPT-5.6 family, `gpt-6-astra`, `gpt-5.6-cyber`, `gpt-daybreak-*`, `gpt-4o-mini-audio-preview*`: `capabilities.function_calling: true` while `capabilities.tool_function_calling: false`. 3. `models.json` Anthropic alias records (`claude-opus-4-5`, `claude-sonnet-4-5`, `claude-haiku-4-5`) store `pricing` as a string ("same as …") instead of the schema object; `modalities`, `capabilities`, `thinking` are null on aliases. 4. `models.json` `gpt-5.5-cyber` and `gpt-5-search-api(-2025-10-14)` are `id_only` records without context/output/modalities although priced. 5. `pricing.json` unit strings are heterogeneous (`per 1M tokens` vs model-page `1M tokens`; `image_output low 1024x1024` dimensions encode size in the dimension name) — now 12 distinct unit strings across four providers (`per 1K requests`, `per 1K search queries`, `per 1K grounded prompts`, `per 1K calls`, `per call`, `per request`, `multiplier`, `per hour`, `per page`, `per 1M tokens per hour`, `per GiB per day`, `per 1M characters`). 6. `endpoints.json`: six OpenAI records carry `verification.result: "success"` without a live call; three `LIVE_VERIFIED` endpoints have `result: "failure"` — status and result should be reconciled. 7. `endpoints.json`: `POST /v1/responses?beta=true` and `POST /v1/messages (beta surface)` encode a variant in the `path` field; `auth` is a string for some fragments and an object for others. 8. `streaming-events.json`: Anthropic core stream recorded under two `api` labels; Managed Agents client→server events appear twice. 9. `webhook-events.json`: the 28 OpenAI records have `api: null` while Anthropic records set `api: managed-agents`. 10. `headers.json`: `required` is boolean for OpenAI rows and a free string for Anthropic/xAI/Gemini rows. 11. `errors.json`: `api_family` is null for core Anthropic and OpenAI errors; several Anthropic rows have `http_status` null or composite; Gemini rows record `observed_live: false` for 400 FAILED_PRECONDITION / 403 / 501 although the Gemini docs pages report them observed. 12. `deprecations.json`: OpenAI = 165 flat records, Anthropic / xAI / Gemini = 1 nested record each — two schemas in one file. 13. `rate-limits.json`: OpenAI record has no `status` field and no `documented` key (its content is under `concepts`/`usage_tiers`); the Gemini enqueued-token table lists `gemini-2.0-flash-image` and shut-down 2.5 previews while omitting the GA image ids. 14. `tools.json` OpenAI `programmatic_tool_calling` has an empty `compatible_models` list; OpenAI `computer` tool is `UNVERIFIED` although eleven model records claim `tool_computer_use: true`. 15. Anthropic tool-search docs table omits `claude-sonnet-5` while the model record lists both tool-search types. 16. Live vs docs (Anthropic): `mcp_servers` error message advertises `mcp-client-2026-09-15` which the API rejects; `code_execution_requests` usage counter absent live; Managed Agents webhook names differ from stream names. 17. **xAI `models.json`**: `max_output` is null for all 17 model records (only the grok-4.6 page says "No text output limit"); `knowledge_cutoff` is null for every model except grok-4.6; `parameters.json` records `max_output_tokens` `default: None` while docs/xai/responses.md says 128,000; a documented sample echoed `max_output_tokens: 2000` unexplained. 18. **xAI reasoning flags**: `parameters.json` lists `grok-4.20-0309-reasoning` and `grok-build-0.1` as compatible with `reasoning_effort` (Chat) / `reasoning.effort` (Responses) but both return 400 live; `none` appears in the enum only as "(LIVE_DISCOVERED, grok-4.3)"; `grok-4.5` `xhigh` is listed without the "treated as high" caveat from docs/xai/reasoning.md. 19. **xAI prices**: `grok-imagine-image-2.0` `per_image_default: 0.06` (catalogue, medium/1k) vs pricing page $0.04 (low/1k); the REST reference describes price units as "USD cents per 100M tokens" while the live catalogue and `pricing.json` use ticks of 1e-10 USD; `grok-4.20-multi-agent-0309` has no priority rows in `pricing.json` while `models.json` says "2.0 (not documented per model)". 20. **xAI tools/streaming**: `x_search` returns `custom_tool_call` items live vs documented `x_search_call`; `streaming-events.json` has no x_search events; image `usage` documents `input_tokens`/`output_tokens` that are absent live (only `cost_in_usd_ticks`). 21. **xAI endpoints**: Collections are documented on `management-api.x.ai` with a Management key but the same paths answer on `api.x.ai` with an inference key (records carry both `ACCOUNT_RESTRICTED` and `LIVE_VERIFIED`); `POST /v1/collections/{id}/documents` multipart → 405; batch `responses` requests come back as `chat_get_completion`; deferred completions documented "retrievable exactly once" but a second GET returned 200; `GET /v1/responses/{id}` 200 for `store:false`; `/v1/completions` marked RETIRED in `endpoints.json` while docs/xai/legacy-completions.md says LEGACY "retired in practice"; the 4.20 non-reasoning model also rejects raw sampling although docs say "non-reasoning only". 22. **xAI rate limits / headers**: documented Tier 0 RPS 150 / 37 vs observed `x-ratelimit-limit-requests` 7,200 / 1,800 per minute (= 120 / 30 RPS); the headers themselves are undocumented; STT default model is `grok-voice-transcribe-1.0` in release notes and 2.0 on the model page (`DOCUMENTATION_INCOMPLETE`); realtime default voice `xai_ara` live vs `eve` in docs; `response.done.usage` and `ping` undocumented; `response.audio.delta` documented but not emitted; `GET /v1/api-key` `create_time`/`modify_time` returned empty strings; `grok-imagine-video-1.5` live `input_modalities` include `audio` (model page: text, image); video pending polls return HTTP 202 (undocumented); `eu-west-1.api.x.ai` undocumented but live. 23. **Gemini `pricing.json`**: 5 null-price rows (`veo-3.1-lite` 4k, `gemini-embedding-001` input, `tool:custom_tools_endpoint`, `agent:deep-research`, `service:document_tokens`); 2 non-numeric prices (`'1.8x'`, `'0.50-8.10'`); duplicate keys with conflicting values — `tool:google_search` standard = $14 "per 1K requests" **and** $35 "per 1K grounded prompts", `tool:google_maps` = $14 "per 1K search queries" **and** $25 "per 1K grounded prompts" (the 3.x vs 2.5 rows are not disambiguated in the key); Gemma rows use `tier: free` with per-dimension rows while other models use one `free_tier` dimension. 24. **Gemini `models.json`**: `knowledge_cutoff` null for 64/97, `context_window` null 28, `max_output` null 33, `release_date` null 13; `gemini-3.5-flash` flex `cached_input` 0.08 vs batch 0.075 (not 50 % of 0.15); `gemini-3.5-flash-lite` batch/flex `cached_input` 0.02 not on the pricing page; `gemini-2.5-pro`/`-flash` carry `ACCOUNT_RESTRICTED` in `deprecations.json` but not in `models.json` (only `-flash-lite` does); `thinking` flag true live on `gemini-3.5-transcribe` and `gemini-3.1-flash-tts-preview` (docs: not supported) and absent on Live/robotics-streaming models (docs: supported). 25. **Gemini `parameters.json`**: 11 duplicate (endpoint, parameter) pairs from different fragments with conflicting enums — `generationConfig.thinkingConfig.thinkingLevel` `MINIMAL|LOW|MEDIUM|HIGH` vs `MINIMAL|HIGH` (image models), `responseModalities` `TEXT|IMAGE|AUDIO` vs `AUDIO`, Interactions `response_format` ×3, `generation_config.thinking_level` ×2, `transcription_config` ×2; `generationConfig._responseJsonSchema` recorded as `LIVE_DISCOVERED`. 26. **Gemini `endpoints.json`**: the core `POST /v1beta/models/{model}:generateContent` record sits under `api_family: transcription` (the `generate-content` family holds only `/v1/…`, `streamGenerateContent` and `dynamic/{id}`); `dynamic` endpoints duplicated across two families with different path spellings and statuses; Interactions `/v1/*` tagged `GA` but `UNVERIFIED`; webhooks documented on `/v1` while everything else is `/v1beta`. 27. **Gemini `streaming-events.json`**: the 7 `lyria-realtime` records have null `status` although docs/gemini/music-generation.md marks the protocol `LIVE_VERIFIED`. 28. **Gemini docs vs docs / live**: Maps pricing $25/1k grounded prompts (tools table) vs $14/1k queries (Gemini 3 tables) → `DOCUMENTATION_INCOMPLETE`; File Search indexing $0.15/1M vs gemini-embedding-2 text $0.20/1M on the same page; priority "75–100 % more" (prose) vs 1.8× (tables); Batch/Flex "not available" on free tier for most models but "free of charge" for `gemini-3.5-flash-lite`; 3.5-flash caching "free of charge" on the page vs live `limit: 0`; explicit-cache minimum 4,096 (docs) vs 1,024 (live); audio token rate 32 tok/s (tokens guide) vs 25 tok/s (Live/pricing); image token counts 280/560/1,120/2,240 (docs) vs 256/529/1,089/2,209 (observed); PDF page 520 IMAGE tokens (generateContent) vs 560 DOCUMENT tokens (countTokens); token limits live vs docs for `gemini-3.1-flash-image` (65,536/65,536 vs 131,072/32,768), `-lite-image` (out 65,536 vs 4,096), `gemini-3-pro-image` (in 131,072 vs 65,536), `gemini-2.5-flash-image` (in 32,768 vs 65,536), `lyria-3*` (1,048,576 vs 131,072), `deep-research-*` (131,072 vs 1,048,576), `gemini-3.5-live-translate-preview` (16,384/32,768 vs 131,072/65,536), 2.5 computer use (131,072/65,536 vs 128,000/64,000); docs say every model is in `/v1` and `/v1beta` (live 22 vs 58); `gemini-3-pro-preview` computer use "not supported" (model page) vs "launched" (changelog); `interaction.status_update` still emitted although the breaking-changes guide says superseded; transcription `parts[].audioTranscription` undocumented, speaker label `spk:0` (REST) vs `spk_1` (Interactions docs); Google Search `ACCOUNT_RESTRICTED` on this free-tier key although the pricing page lists 500 RPD free for 2.5 Flash/Flash-Lite; `gemma-4-*` absent from docs/models tables; `deprecations.json` replacement text differs from the docs table for `gemini-2.5-flash-image` and `gemini-2.0-flash`. 29. **Cross-provider**: `docs/errors/xai.md` does not exist (xAI errors live in docs/xai/authentication-headers-errors.md) while `docs/errors/{openai,anthropic,gemini}.md` do; `sdks.json` has `version: null` on every xAI/Gemini record although the docs pages quote versions (xai-sdk 1.19, google-genai 2.24, @google/genai 2.23); `models.json` uses four different `kind` vocabularies (OpenAI `model|snapshot|id_only|alias`, Anthropic `snapshot|alias`, xAI `model|legacy|retired_redirect|alias|service`, Gemini `stable|preview|alias|agent|experimental`). Related: [index](index.md) · [features](features.md) · [pricing](pricing.md) · [caching and reasoning](caching-and-reasoning.md) · [FAQ](../faq.md).