# xAI (Grok) models catalogue **Status:** `LIVE_VERIFIED` for the 12 ids returned by `GET https://api.x.ai/v1/models` on 2026-09-19 (all 7 language models also exercised with one minimal completion each; image/video models listed only, no media generated); `DOCUMENTED` for voice models (not in the catalogue endpoint), `ACCOUNT_RESTRICTED` for `grok-embedding-small`, `RETIRED` (+ redirect verified live) for the May-15-2026 batch, `LEGACY/UNVERIFIED` for grok-2/grok-3-mini/grok-beta era ids. **Sources:** https://docs.x.ai/developers/models · /developers/models/ (per-model pages, fetched live 2026-09-19 — absent from llms-full.txt) · /developers/pricing · /developers/rate-limits · /developers/release-notes · /developers/grok-4-6 · /developers/model-capabilities/text/{reasoning,multi-agent} · /developers/migration/{may-15-retirement,imagine-image-quality-nov-2} · /developers/rest-api-reference/inference/models · `sources/xai/{models-api-raw,language-models-raw,image-generation-models-raw}.json` · `tmp-live/xai-models/catalogue-probe*.json`. **Last verified:** 2026-09-18 (UTC 2026-09-19) · Machine-readable twin: `generated/fragments/models/xai-models.json` (built by `scripts/gen_xai_fragments.py`); prices in `generated/fragments/pricing/xai-pricing.json`. ## 1. Live catalogue (GET /v1/models, 12 ids) | id | Display | Family | Status | created (live) | Release (docs) | Cutoff | Context | Long-context threshold | Modalities | Reasoning | `reasoning_effort` | Batch API | Regions (model page) | |---|---|---|---|---|---|---|---|---|---|---|---|---|---| | `grok-4.6` | Grok 4.6 | Grok 4.x | DOCUMENTED, LIVE_VERIFIED | 2026-08-06 | Aug 2026 | 2026-02-01 | 500,000 | 200k | text+image → text | always on | low / medium / **high** / xhigh | no | us-east-1, us-west-2, us-central-1 (+ `us.api.x.ai`) | | `grok-4.5` | Grok 4.5 | Grok 4.x | DOCUMENTED, LIVE_VERIFIED | 2026-06-29 | Jul 2026 | — | 500,000 | 200k | text+image → text | always on | low / medium / **high** (+ xhigh on model page; reasoning page: treated as high) | no | us-east-1, us-west-2 | | `grok-4.3` | Grok 4.3 | Grok 4.x | DOCUMENTED, LIVE_VERIFIED | 2026-04-17 | Apr/May 2026 | — | 1,000,000 | 200k | text+image → text | optional | none / **low** / medium / high / xhigh (`none` accepted live) | yes (−20 %) | us-east-1, eu-west-1, us-west-2 (+ `eu-west-1.api.x.ai`, live-discovered) | | `grok-4.20-0309-reasoning` | Grok 4.20 | Grok 4.20 | DOCUMENTED, LIVE_VERIFIED | 2026-03-09 | Mar 2026 | — | 1,000,000 | 200k | text+image → text | always on | **not supported** (400 live) | yes (−20 %) | us-east-1, us-west-2 | | `grok-4.20-0309-non-reasoning` | Grok 4.20 (non-reasoning) | Grok 4.20 | DOCUMENTED, LIVE_VERIFIED | 2026-03-09 | Mar 2026 | — | 1,000,000 | 200k | text+image → text | none | not supported (400 live) | yes (−20 %) | us-east-1, us-west-2 | | `grok-4.20-multi-agent-0309` | Grok 4.20 Multi-Agent | Grok 4.20 | DOCUMENTED, BETA, LIVE_VERIFIED | 2026-03-09 | Mar 2026 | — | 1,000,000 | 200k | text+image → text | always on (multi-agent) | `reasoning.effort` = agent count: low/medium → 4, high/xhigh → 16 (observed default `medium`) | yes (−20 %) | us-east-1, us-west-2 | | `grok-build-0.1` | Grok Build 0.1 | Build | DOCUMENTED, PREVIEW, LIVE_VERIFIED | 2026-04-16 | May 2026 (early access) | — | 256,000 | 200k | text+image → text | always on | not supported (400 live) | no | us-east-1, us-west-2 | | `grok-imagine-image` | Imagine Image 1.0 | Imagine | DOCUMENTED, LIVE_VERIFIED | 2026-01-28 | Jan 2026 | — | prompt 16,000 | — | text+image → image | — | — | yes (us-east-1, us-west-2) | + us-saltlake-2 | | `grok-imagine-image-2.0` | Imagine Image 2.0 | Imagine | DOCUMENTED, LIVE_VERIFIED | 2026-08-08 | Aug 2026 | — | prompt 64,000 | — | text+image → image | — | — | yes | us-east-1, us-west-2 | | `grok-imagine-image-quality` | Imagine Image Quality | Imagine | DOCUMENTED, **DEPRECATED** (retires 2026-11-02), LIVE_VERIFIED | 2026-04-03 | Apr 2026 | — | prompt 16,000 | — | text+image → image | — | — | no | us-east-1, us-west-2 | | `grok-imagine-video` | Imagine Video 1.0 | Imagine | DOCUMENTED, LIVE_VERIFIED | 2026-01-28 | Jan 2026 | — | — | — | text+image+video → video | — | — | yes (us-east-1, us-west-2) | + us-saltlake-2 | | `grok-imagine-video-1.5` | Imagine Video 1.5 | Imagine | DOCUMENTED, LIVE_VERIFIED | 2026-05-27 | May–Jul 2026 | — | — | — | text+image+**audio** → video (live; page says text, image) | — | — | yes | us-east-1, us-west-2 | `max_output`: xAI publishes no output-token limit ("No text output limit" on the grok-4.6 page); none is returned by the catalogue. ### Aliases (live `aliases[]`) | Canonical id | Aliases | |---|---| | `grok-4.6` | *(none — no `grok-4.6-latest` exists; `grok-latest` → 404 on us.api.x.ai per docs)* | | `grok-4.5` | `grok-4.5-latest`, `grok-build-latest` | | `grok-4.3` | `grok-4.3-latest` (+ retired redirects: `grok-3`, `grok-4-0709`, `grok-4-fast(-reasoning/-non-reasoning)`, `grok-4-1-fast-*` — not listed in `aliases[]` but `GET /v1/models/grok-3` returns the grok-4.3 object) | | `grok-4.20-0309-reasoning` | `grok-4.20`, `grok-4.20-reasoning`, `grok-4.20-reasoning-latest`, `grok-4.20-0309`, `grok-4.20-beta`, `grok-4.20-beta-0309`, `grok-4.20-beta-latest`, `grok-4.20-beta-latest-reasoning`, `grok-4.20-beta-reasoning`, `grok-4.20-beta-0309-reasoning`, `grok-4.20-experimental-beta-0304`, `grok-4.20-experimental-beta-0304-reasoning`, `grok-4.20-experimental-beta-latest`, `grok-4.20-experimental-beta-reasoning-latest`, `grok-4.20-reasoning-gv2` | | `grok-4.20-0309-non-reasoning` | `grok-4.20-non-reasoning`, `grok-4.20-non-reasoning-latest`, `grok-4.20-beta-non-reasoning`, `grok-4.20-beta-latest-non-reasoning`, `grok-4.20-beta-0309-non-reasoning`, `grok-4.20-experimental-beta-0304-non-reasoning`, `grok-4.20-experimental-beta-non-reasoning-latest`, `grok-4.20-non-reasoning-gv2` | | `grok-4.20-multi-agent-0309` | `grok-4.20-multi-agent`, `grok-4.20-multi-agent-latest`, `grok-4.20-multi-agent-beta-latest`, `grok-4.20-multi-agent-beta-0309`, `grok-4.20-multi-agent-experimental-beta-0304`, `grok-4.20-multi-agent-experimental-beta-latest` | | `grok-build-0.1` | `grok-code-fast-1`, `grok-code-fast`, `grok-code-fast-1-0825` | | `grok-imagine-image` | `grok-imagine-image-2026-03-02` | | `grok-imagine-image-quality` | `grok-imagine-image-quality-20260403`, `grok-imagine-image-quality-latest`, `grok-imagine-image-pro` | | `grok-imagine-video-1.5` | `grok-imagine-video-1.5-preview`, `grok-imagine-video-1.5-2026-05-30` | Alias convention (docs): `` = latest stable, `-latest` = latest version, `-` = pinned. In practice the dated 4.20 ids are the canonical ids and the bare names are aliases. ## 2. Documented ids outside the catalogue | id | Kind | Status | Notes | |---|---|---|---| | `grok-voice-think-fast-2.0` (`grok-voice-latest` since 2026-08-05) | speech-to-speech (wss `/v1/realtime`) | DOCUMENTED | `GET /v1/models/grok-voice-think-fast-2.0` → 404 (voice models are not in the catalogue). $0.08/min audio + $0.004/text item; concurrent sessions T0 10; 120 min max session. Regions: us-east-1 (S2S page) / us-east-1, eu-west-1, us-saltlake-2 (model page). | | `grok-voice-think-fast-1.0` | speech-to-speech | DOCUMENTED, LEGACY | April 2026; superseded, no retirement date. | | `grok-voice-transcribe-2.0` / `-1.0` | speech-to-text (`POST /v1/stt`, wss `/v1/stt`) | DOCUMENTED | **Default discrepancy:** release notes (Sept 2026) say default = 1.0; the Speech-to-Text model page says default = 2.0. $0.10/h REST, $0.20/h streaming. | | Text-to-Speech (no model id) | `POST /v1/tts`, wss `/v1/tts`, `GET /v1/tts/voices` | DOCUMENTED | $15 / 1M characters; voice selection instead of model id. | | `grok-embedding-small` | embeddings (Collections `index_configuration.model_name`) | DOCUMENTED, ACCOUNT_RESTRICTED | `GET /v1/embedding-models` → `{"models": []}`; `GET /v1/embedding-models/grok-embedding-small` → 404 with our key. `POST /v1/embeddings` exists in the OpenAPI spec. | | `grok-3`, `grok-4-0709`, `grok-4-fast-*`, `grok-4-1-fast-*` | retired 2026-05-15 | RETIRED (redirect → `grok-4.3`, live-verified for grok-3, grok-4-fast, grok-4-0709) | Billed at grok-4.3 prices; reasoning slugs get `low`, non-reasoning slugs `none`. | | `grok-code-fast-1` | retired 2026-05-15 | RETIRED → alias of `grok-build-0.1` | `GET /v1/models/grok-code-fast-1` returns the grok-build-0.1 object. | | `grok-imagine-image-pro` | retired 2026-05-15 | RETIRED → alias of `grok-imagine-image-quality` → 2.0 `low` from 2026-11-02 | | | `grok-2-image(-1212)`, `grok-2-vision-1212`, `grok-2-1212`, `grok-3-mini(-fast)`, `grok-beta`, `grok-vision-beta`, `grok-4(-latest)` | historical | LEGACY, UNVERIFIED | Release-notes era ids; `GET /v1/models/grok-2-image` → 404 `not-found`. | ## 3. Live vs docs discrepancies (2026-09-19) - **Price units:** the REST reference says token prices are "USD cents per 100 million tokens" and `image_price` is "USD cents"; live values are consistent with **ticks of 1e-10 USD** for both (`12500` → $1.25/1M tokens, `200000000` → $0.02/image). `price_per_image` is documented as "1/100,000,000ths of a USD cent" = the same unit. `cost_in_usd_ticks` (usage) uses the same 1e10 ticks/USD. - `grok-4.5`: model page lists `xhigh` as supported; reasoning page says only low/medium/high and `xhigh` is treated as `high`. - `grok-4.20-0309-reasoning`, `grok-4.20-0309-non-reasoning`, `grok-build-0.1`: model pages show "Reasoning: Yes/No" but `reasoning_effort` → `400 invalid-argument "Model … does not support parameter reasoningEffort"`. Only grok-4.6, grok-4.5, grok-4.3 (and multi-agent via `reasoning.effort`) accept it. - `grok-4.20-multi-agent-0309`: `POST /v1/chat/completions` → `400 "Multi Agent requests are not allowed on chat completions"` (bare string body); works on `POST /v1/responses`. The response's `reasoning.effort` defaulted to `"medium"`; output text carried a trailing `\confidence{80}` / `\confidence{90}` marker; usage contained undocumented `context_details{input_tokens,output_tokens}`. - `grok-imagine-video-1.5`: live `input_modalities` = text, image, **audio**; model page says text, image. - `grok-4.5` created 2026-06-29 (live) vs July 2026 release note; `grok-build-0.1` created 2026-04-16 vs May 2026 note. - Regional catalogues: `GET https://us.api.x.ai/v1/models` → only `grok-4.6` with prices ×1.1 (22000 / 5500 / 66000 ticks, long-context 44000 / 11000 / 132000). **`GET https://eu-west-1.api.x.ai/v1/models` → 200, only `grok-4.3` at global prices — this host is not documented** (model pages mention the `eu-west-1` cluster only). - `search_price` is `0` for every language model (Live Search superseded by tools). - `GET /v1/api-key` returned empty strings for `create_time` / `modify_time` (docs: Unix timestamp; example: RFC 3339). - Documentation stubs: the gRPC reference pages and per-model pages are rendered client-side — `llms-full.txt` contains only their titles. Per-model pages were fetched live as `.md`; gRPC methods were taken from the `xai-sdk` 1.19 proto stubs (see `docs/xai/grpc-api.md`). ## 4. Model × Capability matrix (text models) | Capability | 4.6 | 4.5 | 4.3 | 4.20 R | 4.20 NR | 4.20 MA | build-0.1 | |---|---|---|---|---|---|---|---| | image input | yes | yes | yes | yes | yes | yes | yes | | reasoning | always | always | optional (`none`) | always | no | multi-agent | always | | `reasoning_effort` | low/med/high/xhigh | low/med/high (xhigh→high) | none/low/med/high/xhigh | **400** | **400** | agent count | **400** | | `reasoning_content` in chat response (live) | yes | yes | (none → absent) | n/a | no | n/a | yes | | function calling / parallel tools | yes | yes | yes | yes | yes | yes | yes | | structured outputs | yes | yes | yes | yes | yes | yes | yes | | server-side tools (web/x search, code exec…) | yes (documented) | unknown | unknown | unknown | unknown | unknown (deep research) | unknown | | `logprobs` | ignored | ignored | ignored | ignored | ignored | ignored | unknown | | automatic prompt caching (`cached_tokens`) | yes (512 cached on a 640-token prompt) | yes | yes | yes | yes | yes (2560/2655) | yes | | Batch API | no | no | yes −20 % | yes −20 % | yes −20 % | yes −20 % | no | | Priority (`service_tier: priority`, 2×) | yes | yes | yes | yes | yes | not documented | yes | | deferred completions | yes | yes | yes | yes | yes | no (Responses only) | yes | | context compaction / WebSocket Responses | yes | yes | yes | yes | yes | yes | yes | | `us.api.x.ai` regional | **only model** | no | no | no | no | no | no | | `eu-west-1` cluster | no | no | yes | no | no | no | no | | fine-tuning | none offered by xAI | | | | | | | ## 5. Model × Endpoint matrix | Endpoint | text models | 4.20 multi-agent | image models | video models | voice | |---|---|---|---|---|---| | `POST /v1/chat/completions` | yes | **400** | — | — | — | | `POST /v1/responses` (+ `wss://api.x.ai/v1/responses`, `/v1/responses/compact`) | yes | yes (only) | (image_generation tool) | — | — | | `POST /v1/messages` (Anthropic-compatible) / `POST /v1/complete` | yes / legacy non-reasoning | ? | — | — | — | | `POST /v1/completions` (legacy) | non-reasoning only | — | — | — | — | | `GET /v1/chat/deferred-completion/{id}` | yes | — | — | — | — | | `POST /v1/images/generations`, `/v1/images/edits` | — | — | yes | — | — | | `POST /v1/videos/{generations,edits,extensions}`, `GET /v1/videos/{id}` | — | — | — | yes | — | | Batch API | 4.3 / 4.20 R / 4.20 NR | yes | 1.0, 2.0 (not quality) | yes | — | | `POST /v1/tokenize-text` | yes | yes | — | — | — | | `wss://api.x.ai/v1/realtime`, `POST /v1/stt`, `POST /v1/tts` | — | — | — | — | yes | ## 6. Model × Tool matrix (type strings) Responses API (`tools[].type`): `web_search`, `x_search`, `code_execution` (alias `code_interpreter`), `image_generation`, `attachment_search`, `collections_search` (alias `file_search`), `view_image` / `view_x_video` (implicit, on search results), `mcp` (remote MCP), `function` (client). gRPC / xai-sdk: `web_search()`, `x_search()`, `code_execution()`, `collections_search()`, `mcp()` — `code_interpreter` and `file_search` aliases unsupported. Documentation lists tool support explicitly only for grok-4.6 (function calling, web search, X search, code execution); all text models accept function tools. Details: `docs/tools/xai/index.md`. ## 7. Pricing summary (USD per 1M tokens; full detail in `docs/xai/pricing.md`) | Model | Input | Cached | Output | ≥ 200k prompt: input / cached / output | Batch | Priority | |---|---|---|---|---|---|---| | grok-4.6 | 2.00 | 0.50 | 6.00 | 4.00 / 1.00 / 12.00 | — | ×2 | | grok-4.5 | 2.00 | 0.30 | 6.00 | 4.00 / 0.60 / 12.00 | — | ×2 | | grok-4.3 | 1.25 | 0.20 | 2.50 | 2.50 / 0.40 / 5.00 | −20 % | ×2 | | grok-4.20 (R, NR, multi-agent) | 1.25 | 0.20 | 2.50 | 2.50 / 0.40 / 5.00 | −20 % | ×2 | | grok-build-0.1 | 1.00 | 0.20 | 2.00 | 2.00 / 0.40 / 4.00 | — | ×2 | Tiered rule: when the prompt reaches 200k tokens **all** tokens of the request are billed at the long-context rate. Image input tokens cost the same as text. Reasoning tokens are output tokens. Images: 1.0 $0.02, quality $0.05, 2.0 $0.04 default (low 1k/1.5k/2k = 0.04/0.05/0.06; medium = 0.06/0.07/0.08). Video: 1.0 $0.05/s, 1.5 $0.08/s. ## 8. Rate limits (Tier 0, documented → observed) | Model | Doc RPS T0 | Doc TPM T0 | Observed `x-ratelimit-limit-requests` | Observed `x-ratelimit-limit-tokens` | |---|---|---|---|---| | grok-4.6, grok-4.5 | 150 | 50M | 7200 | 50000000 | | grok-4.3, 4.20 R/NR, build-0.1 | 37 | 10M | 1800 | 10000000 | | grok-4.20-multi-agent | 9 | 2.5M | (no headers on Responses) | — | | image models | 6 | — | — | — | | video models | 10 | — | — | — | Full tier tables: `docs/xai/rate-limits.md`. ## 9. Observed usage shapes (minimal "Reply with OK.", max_tokens 8) | Model | prompt / cached | completion | reasoning | cost (ticks → USD) | |---|---|---|---|---| | grok-4.6 | 640 / 512 | 1 | 91 | 10,640,000 → $0.00106 | | grok-4.5 | 498 / 384 | 1 | 55 | 6,792,000 → $0.00068 | | grok-4.20-0309-non-reasoning | 188 / 128 | 2 | 0 | 1,056,000 → $0.00011 | | grok-4.3 (`reasoning_effort: none`) | 188 / 128 | 2 | 0 | 1,056,000 → $0.00011 | | grok-build-0.1 | 190 / 128 | 1 | 162 | 4,136,000 → $0.00041 | | grok-4.20-multi-agent-0309 (Responses) | 2655 / 2560 | 1645 (1618 reasoning) | — | 47,432,500 → $0.0047 | Chat `usage` = `{prompt_tokens, completion_tokens, total_tokens, prompt_tokens_details{text_tokens, audio_tokens, image_tokens, cached_tokens}, completion_tokens_details{reasoning_tokens, audio_tokens, accepted_prediction_tokens, rejected_prediction_tokens}, num_sources_used, cost_in_usd_ticks}` plus top-level `system_fingerprint`, `service_tier: "default"`, `message.reasoning_content`, `message.refusal`. Responses `usage` adds `num_server_side_tools_used`, `input_tokens_details.cached_tokens`, `output_tokens_details.reasoning_tokens`, `context_details`. Note the ~190–640 prompt tokens for a 4-token prompt (system template overhead; FAQ acknowledges API counts exceed the tokenizer).