# xAI (Grok) models catalogue
Status: LIVE_VERIFIED for the 12 ids returned by GET https://api.x.ai/v1/models on 2026-09-19 (all 7 language models also exercised with one minimal completion each; image/video models listed only, no media generated); DOCUMENTED for voice models (not in the catalogue endpoint), ACCOUNT_RESTRICTED for grok-embedding-small, RETIRED (+ redirect verified live) for the May-15-2026 batch, LEGACY/UNVERIFIED for grok-2/grok-3-mini/grok-beta era ids.
Sources: https://docs.x.ai/developers/models · /developers/models/ (per-model pages, fetched live 2026-09-19 — absent from llms-full.txt) · /developers/pricing · /developers/rate-limits · /developers/release-notes · /developers/grok-4-6 · /developers/model-capabilities/text/{reasoning,multi-agent} · /developers/migration/{may-15-retirement,imagine-image-quality-nov-2} · /developers/rest-api-reference/inference/models · sources/xai/{models-api-raw,language-models-raw,image-generation-models-raw}.json · tmp-live/xai-models/catalogue-probe*.json.
Last verified: 2026-09-18 (UTC 2026-09-19) · Machine-readable twin: generated/fragments/models/xai-models.json (built by scripts/gen_xai_fragments.py); prices in generated/fragments/pricing/xai-pricing.json.
# 1. Live catalogue (GET /v1/models, 12 ids)
| id |
Display |
Family |
Status |
created (live) |
Release (docs) |
Cutoff |
Context |
Long-context threshold |
Modalities |
Reasoning |
reasoning_effort |
Batch API |
Regions (model page) |
grok-4.6 |
Grok 4.6 |
Grok 4.x |
DOCUMENTED, LIVE_VERIFIED |
2026-08-06 |
Aug 2026 |
2026-02-01 |
500,000 |
200k |
text+image → text |
always on |
low / medium / high / xhigh |
no |
us-east-1, us-west-2, us-central-1 (+ us.api.x.ai) |
grok-4.5 |
Grok 4.5 |
Grok 4.x |
DOCUMENTED, LIVE_VERIFIED |
2026-06-29 |
Jul 2026 |
— |
500,000 |
200k |
text+image → text |
always on |
low / medium / high (+ xhigh on model page; reasoning page: treated as high) |
no |
us-east-1, us-west-2 |
grok-4.3 |
Grok 4.3 |
Grok 4.x |
DOCUMENTED, LIVE_VERIFIED |
2026-04-17 |
Apr/May 2026 |
— |
1,000,000 |
200k |
text+image → text |
optional |
none / low / medium / high / xhigh (none accepted live) |
yes (−20 %) |
us-east-1, eu-west-1, us-west-2 (+ eu-west-1.api.x.ai, live-discovered) |
grok-4.20-0309-reasoning |
Grok 4.20 |
Grok 4.20 |
DOCUMENTED, LIVE_VERIFIED |
2026-03-09 |
Mar 2026 |
— |
1,000,000 |
200k |
text+image → text |
always on |
not supported (400 live) |
yes (−20 %) |
us-east-1, us-west-2 |
grok-4.20-0309-non-reasoning |
Grok 4.20 (non-reasoning) |
Grok 4.20 |
DOCUMENTED, LIVE_VERIFIED |
2026-03-09 |
Mar 2026 |
— |
1,000,000 |
200k |
text+image → text |
none |
not supported (400 live) |
yes (−20 %) |
us-east-1, us-west-2 |
grok-4.20-multi-agent-0309 |
Grok 4.20 Multi-Agent |
Grok 4.20 |
DOCUMENTED, BETA, LIVE_VERIFIED |
2026-03-09 |
Mar 2026 |
— |
1,000,000 |
200k |
text+image → text |
always on (multi-agent) |
reasoning.effort = agent count: low/medium → 4, high/xhigh → 16 (observed default medium) |
yes (−20 %) |
us-east-1, us-west-2 |
grok-build-0.1 |
Grok Build 0.1 |
Build |
DOCUMENTED, PREVIEW, LIVE_VERIFIED |
2026-04-16 |
May 2026 (early access) |
— |
256,000 |
200k |
text+image → text |
always on |
not supported (400 live) |
no |
us-east-1, us-west-2 |
grok-imagine-image |
Imagine Image 1.0 |
Imagine |
DOCUMENTED, LIVE_VERIFIED |
2026-01-28 |
Jan 2026 |
— |
prompt 16,000 |
— |
text+image → image |
— |
— |
yes (us-east-1, us-west-2) |
+ us-saltlake-2 |
grok-imagine-image-2.0 |
Imagine Image 2.0 |
Imagine |
DOCUMENTED, LIVE_VERIFIED |
2026-08-08 |
Aug 2026 |
— |
prompt 64,000 |
— |
text+image → image |
— |
— |
yes |
us-east-1, us-west-2 |
grok-imagine-image-quality |
Imagine Image Quality |
Imagine |
DOCUMENTED, DEPRECATED (retires 2026-11-02), LIVE_VERIFIED |
2026-04-03 |
Apr 2026 |
— |
prompt 16,000 |
— |
text+image → image |
— |
— |
no |
us-east-1, us-west-2 |
grok-imagine-video |
Imagine Video 1.0 |
Imagine |
DOCUMENTED, LIVE_VERIFIED |
2026-01-28 |
Jan 2026 |
— |
— |
— |
text+image+video → video |
— |
— |
yes (us-east-1, us-west-2) |
+ us-saltlake-2 |
grok-imagine-video-1.5 |
Imagine Video 1.5 |
Imagine |
DOCUMENTED, LIVE_VERIFIED |
2026-05-27 |
May–Jul 2026 |
— |
— |
— |
text+image+audio → video (live; page says text, image) |
— |
— |
yes |
us-east-1, us-west-2 |
max_output: xAI publishes no output-token limit ("No text output limit" on the grok-4.6 page); none is returned by the catalogue.
# Aliases (live aliases[])
| Canonical id |
Aliases |
grok-4.6 |
(none — no grok-4.6-latest exists; grok-latest → 404 on us.api.x.ai per docs) |
grok-4.5 |
grok-4.5-latest, grok-build-latest |
grok-4.3 |
grok-4.3-latest (+ retired redirects: grok-3, grok-4-0709, grok-4-fast(-reasoning/-non-reasoning), grok-4-1-fast-* — not listed in aliases[] but GET /v1/models/grok-3 returns the grok-4.3 object) |
grok-4.20-0309-reasoning |
grok-4.20, grok-4.20-reasoning, grok-4.20-reasoning-latest, grok-4.20-0309, grok-4.20-beta, grok-4.20-beta-0309, grok-4.20-beta-latest, grok-4.20-beta-latest-reasoning, grok-4.20-beta-reasoning, grok-4.20-beta-0309-reasoning, grok-4.20-experimental-beta-0304, grok-4.20-experimental-beta-0304-reasoning, grok-4.20-experimental-beta-latest, grok-4.20-experimental-beta-reasoning-latest, grok-4.20-reasoning-gv2 |
grok-4.20-0309-non-reasoning |
grok-4.20-non-reasoning, grok-4.20-non-reasoning-latest, grok-4.20-beta-non-reasoning, grok-4.20-beta-latest-non-reasoning, grok-4.20-beta-0309-non-reasoning, grok-4.20-experimental-beta-0304-non-reasoning, grok-4.20-experimental-beta-non-reasoning-latest, grok-4.20-non-reasoning-gv2 |
grok-4.20-multi-agent-0309 |
grok-4.20-multi-agent, grok-4.20-multi-agent-latest, grok-4.20-multi-agent-beta-latest, grok-4.20-multi-agent-beta-0309, grok-4.20-multi-agent-experimental-beta-0304, grok-4.20-multi-agent-experimental-beta-latest |
grok-build-0.1 |
grok-code-fast-1, grok-code-fast, grok-code-fast-1-0825 |
grok-imagine-image |
grok-imagine-image-2026-03-02 |
grok-imagine-image-quality |
grok-imagine-image-quality-20260403, grok-imagine-image-quality-latest, grok-imagine-image-pro |
grok-imagine-video-1.5 |
grok-imagine-video-1.5-preview, grok-imagine-video-1.5-2026-05-30 |
Alias convention (docs): <model> = latest stable, <model>-latest = latest version, <model>-<date> = pinned. In practice the dated 4.20 ids are the canonical ids and the bare names are aliases.
# 2. Documented ids outside the catalogue
| id |
Kind |
Status |
Notes |
grok-voice-think-fast-2.0 (grok-voice-latest since 2026-08-05) |
speech-to-speech (wss /v1/realtime) |
DOCUMENTED |
GET /v1/models/grok-voice-think-fast-2.0 → 404 (voice models are not in the catalogue). $0.08/min audio + $0.004/text item; concurrent sessions T0 10; 120 min max session. Regions: us-east-1 (S2S page) / us-east-1, eu-west-1, us-saltlake-2 (model page). |
grok-voice-think-fast-1.0 |
speech-to-speech |
DOCUMENTED, LEGACY |
April 2026; superseded, no retirement date. |
grok-voice-transcribe-2.0 / -1.0 |
speech-to-text (POST /v1/stt, wss /v1/stt) |
DOCUMENTED |
Default discrepancy: release notes (Sept 2026) say default = 1.0; the Speech-to-Text model page says default = 2.0. $0.10/h REST, $0.20/h streaming. |
| Text-to-Speech (no model id) |
POST /v1/tts, wss /v1/tts, GET /v1/tts/voices |
DOCUMENTED |
$15 / 1M characters; voice selection instead of model id. |
grok-embedding-small |
embeddings (Collections index_configuration.model_name) |
DOCUMENTED, ACCOUNT_RESTRICTED |
GET /v1/embedding-models → {"models": []}; GET /v1/embedding-models/grok-embedding-small → 404 with our key. POST /v1/embeddings exists in the OpenAPI spec. |
grok-3, grok-4-0709, grok-4-fast-*, grok-4-1-fast-* |
retired 2026-05-15 |
RETIRED (redirect → grok-4.3, live-verified for grok-3, grok-4-fast, grok-4-0709) |
Billed at grok-4.3 prices; reasoning slugs get low, non-reasoning slugs none. |
grok-code-fast-1 |
retired 2026-05-15 |
RETIRED → alias of grok-build-0.1 |
GET /v1/models/grok-code-fast-1 returns the grok-build-0.1 object. |
grok-imagine-image-pro |
retired 2026-05-15 |
RETIRED → alias of grok-imagine-image-quality → 2.0 low from 2026-11-02 |
|
grok-2-image(-1212), grok-2-vision-1212, grok-2-1212, grok-3-mini(-fast), grok-beta, grok-vision-beta, grok-4(-latest) |
historical |
LEGACY, UNVERIFIED |
Release-notes era ids; GET /v1/models/grok-2-image → 404 not-found. |
# 3. Live vs docs discrepancies (2026-09-19)
- Price units: the REST reference says token prices are "USD cents per 100 million tokens" and
image_price is "USD cents"; live values are consistent with ticks of 1e-10 USD for both (12500 → $1.25/1M tokens, 200000000 → $0.02/image). price_per_image is documented as "1/100,000,000ths of a USD cent" = the same unit. cost_in_usd_ticks (usage) uses the same 1e10 ticks/USD.
grok-4.5: model page lists xhigh as supported; reasoning page says only low/medium/high and xhigh is treated as high.
grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning, grok-build-0.1: model pages show "Reasoning: Yes/No" but reasoning_effort → 400 invalid-argument "Model … does not support parameter reasoningEffort". Only grok-4.6, grok-4.5, grok-4.3 (and multi-agent via reasoning.effort) accept it.
grok-4.20-multi-agent-0309: POST /v1/chat/completions → 400 "Multi Agent requests are not allowed on chat completions" (bare string body); works on POST /v1/responses. The response's reasoning.effort defaulted to "medium"; output text carried a trailing \confidence{80} / \confidence{90} marker; usage contained undocumented context_details{input_tokens,output_tokens}.
grok-imagine-video-1.5: live input_modalities = text, image, audio; model page says text, image.
grok-4.5 created 2026-06-29 (live) vs July 2026 release note; grok-build-0.1 created 2026-04-16 vs May 2026 note.
- Regional catalogues:
GET https://us.api.x.ai/v1/models → only grok-4.6 with prices ×1.1 (22000 / 5500 / 66000 ticks, long-context 44000 / 11000 / 132000). GET https://eu-west-1.api.x.ai/v1/models → 200, only grok-4.3 at global prices — this host is not documented (model pages mention the eu-west-1 cluster only).
search_price is 0 for every language model (Live Search superseded by tools).
GET /v1/api-key returned empty strings for create_time / modify_time (docs: Unix timestamp; example: RFC 3339).
- Documentation stubs: the gRPC reference pages and per-model pages are rendered client-side —
llms-full.txt contains only their titles. Per-model pages were fetched live as .md; gRPC methods were taken from the xai-sdk 1.19 proto stubs (see docs/xai/grpc-api.md).
# 4. Model × Capability matrix (text models)
| Capability |
4.6 |
4.5 |
4.3 |
4.20 R |
4.20 NR |
4.20 MA |
build-0.1 |
| image input |
yes |
yes |
yes |
yes |
yes |
yes |
yes |
| reasoning |
always |
always |
optional (none) |
always |
no |
multi-agent |
always |
reasoning_effort |
low/med/high/xhigh |
low/med/high (xhigh→high) |
none/low/med/high/xhigh |
400 |
400 |
agent count |
400 |
reasoning_content in chat response (live) |
yes |
yes |
(none → absent) |
n/a |
no |
n/a |
yes |
| function calling / parallel tools |
yes |
yes |
yes |
yes |
yes |
yes |
yes |
| structured outputs |
yes |
yes |
yes |
yes |
yes |
yes |
yes |
| server-side tools (web/x search, code exec…) |
yes (documented) |
unknown |
unknown |
unknown |
unknown |
unknown (deep research) |
unknown |
logprobs |
ignored |
ignored |
ignored |
ignored |
ignored |
ignored |
unknown |
automatic prompt caching (cached_tokens) |
yes (512 cached on a 640-token prompt) |
yes |
yes |
yes |
yes |
yes (2560/2655) |
yes |
| Batch API |
no |
no |
yes −20 % |
yes −20 % |
yes −20 % |
yes −20 % |
no |
Priority (service_tier: priority, 2×) |
yes |
yes |
yes |
yes |
yes |
not documented |
yes |
| deferred completions |
yes |
yes |
yes |
yes |
yes |
no (Responses only) |
yes |
| context compaction / WebSocket Responses |
yes |
yes |
yes |
yes |
yes |
yes |
yes |
us.api.x.ai regional |
only model |
no |
no |
no |
no |
no |
no |
eu-west-1 cluster |
no |
no |
yes |
no |
no |
no |
no |
| fine-tuning |
none offered by xAI |
|
|
|
|
|
|
# 5. Model × Endpoint matrix
| Endpoint |
text models |
4.20 multi-agent |
image models |
video models |
voice |
POST /v1/chat/completions |
yes |
400 |
— |
— |
— |
POST /v1/responses (+ wss://api.x.ai/v1/responses, /v1/responses/compact) |
yes |
yes (only) |
(image_generation tool) |
— |
— |
POST /v1/messages (Anthropic-compatible) / POST /v1/complete |
yes / legacy non-reasoning |
? |
— |
— |
— |
POST /v1/completions (legacy) |
non-reasoning only |
— |
— |
— |
— |
GET /v1/chat/deferred-completion/{id} |
yes |
— |
— |
— |
— |
POST /v1/images/generations, /v1/images/edits |
— |
— |
yes |
— |
— |
POST /v1/videos/{generations,edits,extensions}, GET /v1/videos/{id} |
— |
— |
— |
yes |
— |
| Batch API |
4.3 / 4.20 R / 4.20 NR |
yes |
1.0, 2.0 (not quality) |
yes |
— |
POST /v1/tokenize-text |
yes |
yes |
— |
— |
— |
wss://api.x.ai/v1/realtime, POST /v1/stt, POST /v1/tts |
— |
— |
— |
— |
yes |
Responses API (tools[].type): web_search, x_search, code_execution (alias code_interpreter), image_generation, attachment_search, collections_search (alias file_search), view_image / view_x_video (implicit, on search results), mcp (remote MCP), function (client). gRPC / xai-sdk: web_search(), x_search(), code_execution(), collections_search(), mcp() — code_interpreter and file_search aliases unsupported. Documentation lists tool support explicitly only for grok-4.6 (function calling, web search, X search, code execution); all text models accept function tools. Details: docs/tools/xai/index.md.
# 7. Pricing summary (USD per 1M tokens; full detail in docs/xai/pricing.md)
| Model |
Input |
Cached |
Output |
≥ 200k prompt: input / cached / output |
Batch |
Priority |
| grok-4.6 |
2.00 |
0.50 |
6.00 |
4.00 / 1.00 / 12.00 |
— |
×2 |
| grok-4.5 |
2.00 |
0.30 |
6.00 |
4.00 / 0.60 / 12.00 |
— |
×2 |
| grok-4.3 |
1.25 |
0.20 |
2.50 |
2.50 / 0.40 / 5.00 |
−20 % |
×2 |
| grok-4.20 (R, NR, multi-agent) |
1.25 |
0.20 |
2.50 |
2.50 / 0.40 / 5.00 |
−20 % |
×2 |
| grok-build-0.1 |
1.00 |
0.20 |
2.00 |
2.00 / 0.40 / 4.00 |
— |
×2 |
Tiered rule: when the prompt reaches 200k tokens all tokens of the request are billed at the long-context rate. Image input tokens cost the same as text. Reasoning tokens are output tokens. Images: 1.0 $0.02, quality $0.05, 2.0 $0.04 default (low 1k/1.5k/2k = 0.04/0.05/0.06; medium = 0.06/0.07/0.08). Video: 1.0 $0.05/s, 1.5 $0.08/s.
# 8. Rate limits (Tier 0, documented → observed)
| Model |
Doc RPS T0 |
Doc TPM T0 |
Observed x-ratelimit-limit-requests |
Observed x-ratelimit-limit-tokens |
| grok-4.6, grok-4.5 |
150 |
50M |
7200 |
50000000 |
| grok-4.3, 4.20 R/NR, build-0.1 |
37 |
10M |
1800 |
10000000 |
| grok-4.20-multi-agent |
9 |
2.5M |
(no headers on Responses) |
— |
| image models |
6 |
— |
— |
— |
| video models |
10 |
— |
— |
— |
Full tier tables: docs/xai/rate-limits.md.
# 9. Observed usage shapes (minimal "Reply with OK.", max_tokens 8)
| Model |
prompt / cached |
completion |
reasoning |
cost (ticks → USD) |
| grok-4.6 |
640 / 512 |
1 |
91 |
10,640,000 → $0.00106 |
| grok-4.5 |
498 / 384 |
1 |
55 |
6,792,000 → $0.00068 |
| grok-4.20-0309-non-reasoning |
188 / 128 |
2 |
0 |
1,056,000 → $0.00011 |
grok-4.3 (reasoning_effort: none) |
188 / 128 |
2 |
0 |
1,056,000 → $0.00011 |
| grok-build-0.1 |
190 / 128 |
1 |
162 |
4,136,000 → $0.00041 |
| grok-4.20-multi-agent-0309 (Responses) |
2655 / 2560 |
1645 (1618 reasoning) |
— |
47,432,500 → $0.0047 |
Chat usage = {prompt_tokens, completion_tokens, total_tokens, prompt_tokens_details{text_tokens, audio_tokens, image_tokens, cached_tokens}, completion_tokens_details{reasoning_tokens, audio_tokens, accepted_prediction_tokens, rejected_prediction_tokens}, num_sources_used, cost_in_usd_ticks} plus top-level system_fingerprint, service_tier: "default", message.reasoning_content, message.refusal. Responses usage adds num_server_side_tools_used, input_tokens_details.cached_tokens, output_tokens_details.reasoning_tokens, context_details. Note the ~190–640 prompt tokens for a 4-token prompt (system template overhead; FAQ acknowledges API counts exceed the tokenizer).