SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
16.9 KB

# xAI (Grok) models catalogue

Status: LIVE_VERIFIED for the 12 ids returned by GET https://api.x.ai/v1/models on 2026-09-19 (all 7 language models also exercised with one minimal completion each; image/video models listed only, no media generated); DOCUMENTED for voice models (not in the catalogue endpoint), ACCOUNT_RESTRICTED for grok-embedding-small, RETIRED (+ redirect verified live) for the May-15-2026 batch, LEGACY/UNVERIFIED for grok-2/grok-3-mini/grok-beta era ids.
Sources: https://docs.x.ai/developers/models · /developers/models/ (per-model pages, fetched live 2026-09-19 — absent from llms-full.txt) · /developers/pricing · /developers/rate-limits · /developers/release-notes · /developers/grok-4-6 · /developers/model-capabilities/text/{reasoning,multi-agent} · /developers/migration/{may-15-retirement,imagine-image-quality-nov-2} · /developers/rest-api-reference/inference/models · sources/xai/{models-api-raw,language-models-raw,image-generation-models-raw}.json · tmp-live/xai-models/catalogue-probe*.json.
Last verified: 2026-09-18 (UTC 2026-09-19) · Machine-readable twin: generated/fragments/models/xai-models.json (built by scripts/gen_xai_fragments.py); prices in generated/fragments/pricing/xai-pricing.json.

# 1. Live catalogue (GET /v1/models, 12 ids)

id Display Family Status created (live) Release (docs) Cutoff Context Long-context threshold Modalities Reasoning reasoning_effort Batch API Regions (model page)
grok-4.6 Grok 4.6 Grok 4.x DOCUMENTED, LIVE_VERIFIED 2026-08-06 Aug 2026 2026-02-01 500,000 200k text+image → text always on low / medium / high / xhigh no us-east-1, us-west-2, us-central-1 (+ us.api.x.ai)
grok-4.5 Grok 4.5 Grok 4.x DOCUMENTED, LIVE_VERIFIED 2026-06-29 Jul 2026 — 500,000 200k text+image → text always on low / medium / high (+ xhigh on model page; reasoning page: treated as high) no us-east-1, us-west-2
grok-4.3 Grok 4.3 Grok 4.x DOCUMENTED, LIVE_VERIFIED 2026-04-17 Apr/May 2026 — 1,000,000 200k text+image → text optional none / low / medium / high / xhigh (none accepted live) yes (−20 %) us-east-1, eu-west-1, us-west-2 (+ eu-west-1.api.x.ai, live-discovered)
grok-4.20-0309-reasoning Grok 4.20 Grok 4.20 DOCUMENTED, LIVE_VERIFIED 2026-03-09 Mar 2026 — 1,000,000 200k text+image → text always on not supported (400 live) yes (−20 %) us-east-1, us-west-2
grok-4.20-0309-non-reasoning Grok 4.20 (non-reasoning) Grok 4.20 DOCUMENTED, LIVE_VERIFIED 2026-03-09 Mar 2026 — 1,000,000 200k text+image → text none not supported (400 live) yes (−20 %) us-east-1, us-west-2
grok-4.20-multi-agent-0309 Grok 4.20 Multi-Agent Grok 4.20 DOCUMENTED, BETA, LIVE_VERIFIED 2026-03-09 Mar 2026 — 1,000,000 200k text+image → text always on (multi-agent) reasoning.effort = agent count: low/medium → 4, high/xhigh → 16 (observed default medium) yes (−20 %) us-east-1, us-west-2
grok-build-0.1 Grok Build 0.1 Build DOCUMENTED, PREVIEW, LIVE_VERIFIED 2026-04-16 May 2026 (early access) — 256,000 200k text+image → text always on not supported (400 live) no us-east-1, us-west-2
grok-imagine-image Imagine Image 1.0 Imagine DOCUMENTED, LIVE_VERIFIED 2026-01-28 Jan 2026 — prompt 16,000 — text+image → image — — yes (us-east-1, us-west-2) + us-saltlake-2
grok-imagine-image-2.0 Imagine Image 2.0 Imagine DOCUMENTED, LIVE_VERIFIED 2026-08-08 Aug 2026 — prompt 64,000 — text+image → image — — yes us-east-1, us-west-2
grok-imagine-image-quality Imagine Image Quality Imagine DOCUMENTED, DEPRECATED (retires 2026-11-02), LIVE_VERIFIED 2026-04-03 Apr 2026 — prompt 16,000 — text+image → image — — no us-east-1, us-west-2
grok-imagine-video Imagine Video 1.0 Imagine DOCUMENTED, LIVE_VERIFIED 2026-01-28 Jan 2026 — — — text+image+video → video — — yes (us-east-1, us-west-2) + us-saltlake-2
grok-imagine-video-1.5 Imagine Video 1.5 Imagine DOCUMENTED, LIVE_VERIFIED 2026-05-27 May–Jul 2026 — — — text+image+audio → video (live; page says text, image) — — yes us-east-1, us-west-2

max_output: xAI publishes no output-token limit ("No text output limit" on the grok-4.6 page); none is returned by the catalogue.

# Aliases (live aliases[])

Canonical id Aliases
grok-4.6 (none — no grok-4.6-latest exists; grok-latest → 404 on us.api.x.ai per docs)
grok-4.5 grok-4.5-latest, grok-build-latest
grok-4.3 grok-4.3-latest (+ retired redirects: grok-3, grok-4-0709, grok-4-fast(-reasoning/-non-reasoning), grok-4-1-fast-* — not listed in aliases[] but GET /v1/models/grok-3 returns the grok-4.3 object)
grok-4.20-0309-reasoning grok-4.20, grok-4.20-reasoning, grok-4.20-reasoning-latest, grok-4.20-0309, grok-4.20-beta, grok-4.20-beta-0309, grok-4.20-beta-latest, grok-4.20-beta-latest-reasoning, grok-4.20-beta-reasoning, grok-4.20-beta-0309-reasoning, grok-4.20-experimental-beta-0304, grok-4.20-experimental-beta-0304-reasoning, grok-4.20-experimental-beta-latest, grok-4.20-experimental-beta-reasoning-latest, grok-4.20-reasoning-gv2
grok-4.20-0309-non-reasoning grok-4.20-non-reasoning, grok-4.20-non-reasoning-latest, grok-4.20-beta-non-reasoning, grok-4.20-beta-latest-non-reasoning, grok-4.20-beta-0309-non-reasoning, grok-4.20-experimental-beta-0304-non-reasoning, grok-4.20-experimental-beta-non-reasoning-latest, grok-4.20-non-reasoning-gv2
grok-4.20-multi-agent-0309 grok-4.20-multi-agent, grok-4.20-multi-agent-latest, grok-4.20-multi-agent-beta-latest, grok-4.20-multi-agent-beta-0309, grok-4.20-multi-agent-experimental-beta-0304, grok-4.20-multi-agent-experimental-beta-latest
grok-build-0.1 grok-code-fast-1, grok-code-fast, grok-code-fast-1-0825
grok-imagine-image grok-imagine-image-2026-03-02
grok-imagine-image-quality grok-imagine-image-quality-20260403, grok-imagine-image-quality-latest, grok-imagine-image-pro
grok-imagine-video-1.5 grok-imagine-video-1.5-preview, grok-imagine-video-1.5-2026-05-30

Alias convention (docs): <model> = latest stable, <model>-latest = latest version, <model>-<date> = pinned. In practice the dated 4.20 ids are the canonical ids and the bare names are aliases.

# 2. Documented ids outside the catalogue

id Kind Status Notes
grok-voice-think-fast-2.0 (grok-voice-latest since 2026-08-05) speech-to-speech (wss /v1/realtime) DOCUMENTED GET /v1/models/grok-voice-think-fast-2.0 → 404 (voice models are not in the catalogue). $0.08/min audio + $0.004/text item; concurrent sessions T0 10; 120 min max session. Regions: us-east-1 (S2S page) / us-east-1, eu-west-1, us-saltlake-2 (model page).
grok-voice-think-fast-1.0 speech-to-speech DOCUMENTED, LEGACY April 2026; superseded, no retirement date.
grok-voice-transcribe-2.0 / -1.0 speech-to-text (POST /v1/stt, wss /v1/stt) DOCUMENTED Default discrepancy: release notes (Sept 2026) say default = 1.0; the Speech-to-Text model page says default = 2.0. $0.10/h REST, $0.20/h streaming.
Text-to-Speech (no model id) POST /v1/tts, wss /v1/tts, GET /v1/tts/voices DOCUMENTED $15 / 1M characters; voice selection instead of model id.
grok-embedding-small embeddings (Collections index_configuration.model_name) DOCUMENTED, ACCOUNT_RESTRICTED GET /v1/embedding-models → {"models": []}; GET /v1/embedding-models/grok-embedding-small → 404 with our key. POST /v1/embeddings exists in the OpenAPI spec.
grok-3, grok-4-0709, grok-4-fast-*, grok-4-1-fast-* retired 2026-05-15 RETIRED (redirect → grok-4.3, live-verified for grok-3, grok-4-fast, grok-4-0709) Billed at grok-4.3 prices; reasoning slugs get low, non-reasoning slugs none.
grok-code-fast-1 retired 2026-05-15 RETIRED → alias of grok-build-0.1 GET /v1/models/grok-code-fast-1 returns the grok-build-0.1 object.
grok-imagine-image-pro retired 2026-05-15 RETIRED → alias of grok-imagine-image-quality → 2.0 low from 2026-11-02
grok-2-image(-1212), grok-2-vision-1212, grok-2-1212, grok-3-mini(-fast), grok-beta, grok-vision-beta, grok-4(-latest) historical LEGACY, UNVERIFIED Release-notes era ids; GET /v1/models/grok-2-image → 404 not-found.

# 3. Live vs docs discrepancies (2026-09-19)

  • Price units: the REST reference says token prices are "USD cents per 100 million tokens" and image_price is "USD cents"; live values are consistent with ticks of 1e-10 USD for both (12500 → $1.25/1M tokens, 200000000 → $0.02/image). price_per_image is documented as "1/100,000,000ths of a USD cent" = the same unit. cost_in_usd_ticks (usage) uses the same 1e10 ticks/USD.
  • grok-4.5: model page lists xhigh as supported; reasoning page says only low/medium/high and xhigh is treated as high.
  • grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning, grok-build-0.1: model pages show "Reasoning: Yes/No" but reasoning_effort → 400 invalid-argument "Model … does not support parameter reasoningEffort". Only grok-4.6, grok-4.5, grok-4.3 (and multi-agent via reasoning.effort) accept it.
  • grok-4.20-multi-agent-0309: POST /v1/chat/completions → 400 "Multi Agent requests are not allowed on chat completions" (bare string body); works on POST /v1/responses. The response's reasoning.effort defaulted to "medium"; output text carried a trailing \confidence{80} / \confidence{90} marker; usage contained undocumented context_details{input_tokens,output_tokens}.
  • grok-imagine-video-1.5: live input_modalities = text, image, audio; model page says text, image.
  • grok-4.5 created 2026-06-29 (live) vs July 2026 release note; grok-build-0.1 created 2026-04-16 vs May 2026 note.
  • Regional catalogues: GET https://us.api.x.ai/v1/models → only grok-4.6 with prices ×1.1 (22000 / 5500 / 66000 ticks, long-context 44000 / 11000 / 132000). GET https://eu-west-1.api.x.ai/v1/models → 200, only grok-4.3 at global prices — this host is not documented (model pages mention the eu-west-1 cluster only).
  • search_price is 0 for every language model (Live Search superseded by tools).
  • GET /v1/api-key returned empty strings for create_time / modify_time (docs: Unix timestamp; example: RFC 3339).
  • Documentation stubs: the gRPC reference pages and per-model pages are rendered client-side — llms-full.txt contains only their titles. Per-model pages were fetched live as .md; gRPC methods were taken from the xai-sdk 1.19 proto stubs (see docs/xai/grpc-api.md).

# 4. Model × Capability matrix (text models)

Capability 4.6 4.5 4.3 4.20 R 4.20 NR 4.20 MA build-0.1
image input yes yes yes yes yes yes yes
reasoning always always optional (none) always no multi-agent always
reasoning_effort low/med/high/xhigh low/med/high (xhigh→high) none/low/med/high/xhigh 400 400 agent count 400
reasoning_content in chat response (live) yes yes (none → absent) n/a no n/a yes
function calling / parallel tools yes yes yes yes yes yes yes
structured outputs yes yes yes yes yes yes yes
server-side tools (web/x search, code exec…) yes (documented) unknown unknown unknown unknown unknown (deep research) unknown
logprobs ignored ignored ignored ignored ignored ignored unknown
automatic prompt caching (cached_tokens) yes (512 cached on a 640-token prompt) yes yes yes yes yes (2560/2655) yes
Batch API no no yes −20 % yes −20 % yes −20 % yes −20 % no
Priority (service_tier: priority, 2×) yes yes yes yes yes not documented yes
deferred completions yes yes yes yes yes no (Responses only) yes
context compaction / WebSocket Responses yes yes yes yes yes yes yes
us.api.x.ai regional only model no no no no no no
eu-west-1 cluster no no yes no no no no
fine-tuning none offered by xAI

# 5. Model × Endpoint matrix

Endpoint text models 4.20 multi-agent image models video models voice
POST /v1/chat/completions yes 400 — — —
POST /v1/responses (+ wss://api.x.ai/v1/responses, /v1/responses/compact) yes yes (only) (image_generation tool) — —
POST /v1/messages (Anthropic-compatible) / POST /v1/complete yes / legacy non-reasoning ? — — —
POST /v1/completions (legacy) non-reasoning only — — — —
GET /v1/chat/deferred-completion/{id} yes — — — —
POST /v1/images/generations, /v1/images/edits — — yes — —
POST /v1/videos/{generations,edits,extensions}, GET /v1/videos/{id} — — — yes —
Batch API 4.3 / 4.20 R / 4.20 NR yes 1.0, 2.0 (not quality) yes —
POST /v1/tokenize-text yes yes — — —
wss://api.x.ai/v1/realtime, POST /v1/stt, POST /v1/tts — — — — yes

# 6. Model × Tool matrix (type strings)

Responses API (tools[].type): web_search, x_search, code_execution (alias code_interpreter), image_generation, attachment_search, collections_search (alias file_search), view_image / view_x_video (implicit, on search results), mcp (remote MCP), function (client). gRPC / xai-sdk: web_search(), x_search(), code_execution(), collections_search(), mcp() — code_interpreter and file_search aliases unsupported. Documentation lists tool support explicitly only for grok-4.6 (function calling, web search, X search, code execution); all text models accept function tools. Details: docs/tools/xai/index.md.

# 7. Pricing summary (USD per 1M tokens; full detail in docs/xai/pricing.md)

Model Input Cached Output ≥ 200k prompt: input / cached / output Batch Priority
grok-4.6 2.00 0.50 6.00 4.00 / 1.00 / 12.00 — ×2
grok-4.5 2.00 0.30 6.00 4.00 / 0.60 / 12.00 — ×2
grok-4.3 1.25 0.20 2.50 2.50 / 0.40 / 5.00 −20 % ×2
grok-4.20 (R, NR, multi-agent) 1.25 0.20 2.50 2.50 / 0.40 / 5.00 −20 % ×2
grok-build-0.1 1.00 0.20 2.00 2.00 / 0.40 / 4.00 — ×2

Tiered rule: when the prompt reaches 200k tokens all tokens of the request are billed at the long-context rate. Image input tokens cost the same as text. Reasoning tokens are output tokens. Images: 1.0 $0.02, quality $0.05, 2.0 $0.04 default (low 1k/1.5k/2k = 0.04/0.05/0.06; medium = 0.06/0.07/0.08). Video: 1.0 $0.05/s, 1.5 $0.08/s.

# 8. Rate limits (Tier 0, documented → observed)

Model Doc RPS T0 Doc TPM T0 Observed x-ratelimit-limit-requests Observed x-ratelimit-limit-tokens
grok-4.6, grok-4.5 150 50M 7200 50000000
grok-4.3, 4.20 R/NR, build-0.1 37 10M 1800 10000000
grok-4.20-multi-agent 9 2.5M (no headers on Responses) —
image models 6 — — —
video models 10 — — —

Full tier tables: docs/xai/rate-limits.md.

# 9. Observed usage shapes (minimal "Reply with OK.", max_tokens 8)

Model prompt / cached completion reasoning cost (ticks → USD)
grok-4.6 640 / 512 1 91 10,640,000 → $0.00106
grok-4.5 498 / 384 1 55 6,792,000 → $0.00068
grok-4.20-0309-non-reasoning 188 / 128 2 0 1,056,000 → $0.00011
grok-4.3 (reasoning_effort: none) 188 / 128 2 0 1,056,000 → $0.00011
grok-build-0.1 190 / 128 1 162 4,136,000 → $0.00041
grok-4.20-multi-agent-0309 (Responses) 2655 / 2560 1645 (1618 reasoning) — 47,432,500 → $0.0047

Chat usage = {prompt_tokens, completion_tokens, total_tokens, prompt_tokens_details{text_tokens, audio_tokens, image_tokens, cached_tokens}, completion_tokens_details{reasoning_tokens, audio_tokens, accepted_prediction_tokens, rejected_prediction_tokens}, num_sources_used, cost_in_usd_ticks} plus top-level system_fingerprint, service_tier: "default", message.reasoning_content, message.refusal. Responses usage adds num_server_side_tools_used, input_tokens_details.cached_tokens, output_tokens_details.reasoning_tokens, context_details. Note the ~190–640 prompt tokens for a 4-token prompt (system template overhead; FAQ acknowledges API counts exceed the tokenizer).