SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
14 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
16.9 KB · 146 lines markdown
Rendered Raw Blame History
1# xAI (Grok) models catalogue23**Status:** `LIVE_VERIFIED` for the 12 ids returned by `GET https://api.x.ai/v1/models` on 2026-09-19 (all 7 language models also exercised with one minimal completion each; image/video models listed only, no media generated); `DOCUMENTED` for voice models (not in the catalogue endpoint), `ACCOUNT_RESTRICTED` for `grok-embedding-small`, `RETIRED` (+ redirect verified live) for the May-15-2026 batch, `LEGACY/UNVERIFIED` for grok-2/grok-3-mini/grok-beta era ids.  4**Sources:** https://docs.x.ai/developers/models · /developers/models/<id> (per-model pages, fetched live 2026-09-19 — absent from llms-full.txt) · /developers/pricing · /developers/rate-limits · /developers/release-notes · /developers/grok-4-6 · /developers/model-capabilities/text/{reasoning,multi-agent} · /developers/migration/{may-15-retirement,imagine-image-quality-nov-2} · /developers/rest-api-reference/inference/models · `sources/xai/{models-api-raw,language-models-raw,image-generation-models-raw}.json` · `tmp-live/xai-models/catalogue-probe*.json`.  5**Last verified:** 2026-09-18 (UTC 2026-09-19) · Machine-readable twin: `generated/fragments/models/xai-models.json` (built by `scripts/gen_xai_fragments.py`); prices in `generated/fragments/pricing/xai-pricing.json`.67## 1. Live catalogue (GET /v1/models, 12 ids)89| id | Display | Family | Status | created (live) | Release (docs) | Cutoff | Context | Long-context threshold | Modalities | Reasoning | `reasoning_effort` | Batch API | Regions (model page) |10|---|---|---|---|---|---|---|---|---|---|---|---|---|---|11| `grok-4.6` | Grok 4.6 | Grok 4.x | DOCUMENTED, LIVE_VERIFIED | 2026-08-06 | Aug 2026 | 2026-02-01 | 500,000 | 200k | text+image → text | always on | low / medium / **high** / xhigh | no | us-east-1, us-west-2, us-central-1 (+ `us.api.x.ai`) |12| `grok-4.5` | Grok 4.5 | Grok 4.x | DOCUMENTED, LIVE_VERIFIED | 2026-06-29 | Jul 2026 | — | 500,000 | 200k | text+image → text | always on | low / medium / **high** (+ xhigh on model page; reasoning page: treated as high) | no | us-east-1, us-west-2 |13| `grok-4.3` | Grok 4.3 | Grok 4.x | DOCUMENTED, LIVE_VERIFIED | 2026-04-17 | Apr/May 2026 | — | 1,000,000 | 200k | text+image → text | optional | none / **low** / medium / high / xhigh (`none` accepted live) | yes (−20 %) | us-east-1, eu-west-1, us-west-2 (+ `eu-west-1.api.x.ai`, live-discovered) |14| `grok-4.20-0309-reasoning` | Grok 4.20 | Grok 4.20 | DOCUMENTED, LIVE_VERIFIED | 2026-03-09 | Mar 2026 | — | 1,000,000 | 200k | text+image → text | always on | **not supported** (400 live) | yes (−20 %) | us-east-1, us-west-2 |15| `grok-4.20-0309-non-reasoning` | Grok 4.20 (non-reasoning) | Grok 4.20 | DOCUMENTED, LIVE_VERIFIED | 2026-03-09 | Mar 2026 | — | 1,000,000 | 200k | text+image → text | none | not supported (400 live) | yes (−20 %) | us-east-1, us-west-2 |16| `grok-4.20-multi-agent-0309` | Grok 4.20 Multi-Agent | Grok 4.20 | DOCUMENTED, BETA, LIVE_VERIFIED | 2026-03-09 | Mar 2026 | — | 1,000,000 | 200k | text+image → text | always on (multi-agent) | `reasoning.effort` = agent count: low/medium → 4, high/xhigh → 16 (observed default `medium`) | yes (−20 %) | us-east-1, us-west-2 |17| `grok-build-0.1` | Grok Build 0.1 | Build | DOCUMENTED, PREVIEW, LIVE_VERIFIED | 2026-04-16 | May 2026 (early access) | — | 256,000 | 200k | text+image → text | always on | not supported (400 live) | no | us-east-1, us-west-2 |18| `grok-imagine-image` | Imagine Image 1.0 | Imagine | DOCUMENTED, LIVE_VERIFIED | 2026-01-28 | Jan 2026 | — | prompt 16,000 | — | text+image → image | — | — | yes (us-east-1, us-west-2) | + us-saltlake-2 |19| `grok-imagine-image-2.0` | Imagine Image 2.0 | Imagine | DOCUMENTED, LIVE_VERIFIED | 2026-08-08 | Aug 2026 | — | prompt 64,000 | — | text+image → image | — | — | yes | us-east-1, us-west-2 |20| `grok-imagine-image-quality` | Imagine Image Quality | Imagine | DOCUMENTED, **DEPRECATED** (retires 2026-11-02), LIVE_VERIFIED | 2026-04-03 | Apr 2026 | — | prompt 16,000 | — | text+image → image | — | — | no | us-east-1, us-west-2 |21| `grok-imagine-video` | Imagine Video 1.0 | Imagine | DOCUMENTED, LIVE_VERIFIED | 2026-01-28 | Jan 2026 | — | — | — | text+image+video → video | — | — | yes (us-east-1, us-west-2) | + us-saltlake-2 |22| `grok-imagine-video-1.5` | Imagine Video 1.5 | Imagine | DOCUMENTED, LIVE_VERIFIED | 2026-05-27 | May–Jul 2026 | — | — | — | text+image+**audio** → video (live; page says text, image) | — | — | yes | us-east-1, us-west-2 |2324`max_output`: xAI publishes no output-token limit ("No text output limit" on the grok-4.6 page); none is returned by the catalogue.2526### Aliases (live `aliases[]`)2728| Canonical id | Aliases |29|---|---|30| `grok-4.6` | *(none — no `grok-4.6-latest` exists; `grok-latest` → 404 on us.api.x.ai per docs)* |31| `grok-4.5` | `grok-4.5-latest`, `grok-build-latest` |32| `grok-4.3` | `grok-4.3-latest` (+ retired redirects: `grok-3`, `grok-4-0709`, `grok-4-fast(-reasoning/-non-reasoning)`, `grok-4-1-fast-*` — not listed in `aliases[]` but `GET /v1/models/grok-3` returns the grok-4.3 object) |33| `grok-4.20-0309-reasoning` | `grok-4.20`, `grok-4.20-reasoning`, `grok-4.20-reasoning-latest`, `grok-4.20-0309`, `grok-4.20-beta`, `grok-4.20-beta-0309`, `grok-4.20-beta-latest`, `grok-4.20-beta-latest-reasoning`, `grok-4.20-beta-reasoning`, `grok-4.20-beta-0309-reasoning`, `grok-4.20-experimental-beta-0304`, `grok-4.20-experimental-beta-0304-reasoning`, `grok-4.20-experimental-beta-latest`, `grok-4.20-experimental-beta-reasoning-latest`, `grok-4.20-reasoning-gv2` |34| `grok-4.20-0309-non-reasoning` | `grok-4.20-non-reasoning`, `grok-4.20-non-reasoning-latest`, `grok-4.20-beta-non-reasoning`, `grok-4.20-beta-latest-non-reasoning`, `grok-4.20-beta-0309-non-reasoning`, `grok-4.20-experimental-beta-0304-non-reasoning`, `grok-4.20-experimental-beta-non-reasoning-latest`, `grok-4.20-non-reasoning-gv2` |35| `grok-4.20-multi-agent-0309` | `grok-4.20-multi-agent`, `grok-4.20-multi-agent-latest`, `grok-4.20-multi-agent-beta-latest`, `grok-4.20-multi-agent-beta-0309`, `grok-4.20-multi-agent-experimental-beta-0304`, `grok-4.20-multi-agent-experimental-beta-latest` |36| `grok-build-0.1` | `grok-code-fast-1`, `grok-code-fast`, `grok-code-fast-1-0825` |37| `grok-imagine-image` | `grok-imagine-image-2026-03-02` |38| `grok-imagine-image-quality` | `grok-imagine-image-quality-20260403`, `grok-imagine-image-quality-latest`, `grok-imagine-image-pro` |39| `grok-imagine-video-1.5` | `grok-imagine-video-1.5-preview`, `grok-imagine-video-1.5-2026-05-30` |4041Alias convention (docs): `<model>` = latest stable, `<model>-latest` = latest version, `<model>-<date>` = pinned. In practice the dated 4.20 ids are the canonical ids and the bare names are aliases.4243## 2. Documented ids outside the catalogue4445| id | Kind | Status | Notes |46|---|---|---|---|47| `grok-voice-think-fast-2.0` (`grok-voice-latest` since 2026-08-05) | speech-to-speech (wss `/v1/realtime`) | DOCUMENTED | `GET /v1/models/grok-voice-think-fast-2.0` → 404 (voice models are not in the catalogue). $0.08/min audio + $0.004/text item; concurrent sessions T0 10; 120 min max session. Regions: us-east-1 (S2S page) / us-east-1, eu-west-1, us-saltlake-2 (model page). |48| `grok-voice-think-fast-1.0` | speech-to-speech | DOCUMENTED, LEGACY | April 2026; superseded, no retirement date. |49| `grok-voice-transcribe-2.0` / `-1.0` | speech-to-text (`POST /v1/stt`, wss `/v1/stt`) | DOCUMENTED | **Default discrepancy:** release notes (Sept 2026) say default = 1.0; the Speech-to-Text model page says default = 2.0. $0.10/h REST, $0.20/h streaming. |50| Text-to-Speech (no model id) | `POST /v1/tts`, wss `/v1/tts`, `GET /v1/tts/voices` | DOCUMENTED | $15 / 1M characters; voice selection instead of model id. |51| `grok-embedding-small` | embeddings (Collections `index_configuration.model_name`) | DOCUMENTED, ACCOUNT_RESTRICTED | `GET /v1/embedding-models` → `{"models": []}`; `GET /v1/embedding-models/grok-embedding-small` → 404 with our key. `POST /v1/embeddings` exists in the OpenAPI spec. |52| `grok-3`, `grok-4-0709`, `grok-4-fast-*`, `grok-4-1-fast-*` | retired 2026-05-15 | RETIRED (redirect → `grok-4.3`, live-verified for grok-3, grok-4-fast, grok-4-0709) | Billed at grok-4.3 prices; reasoning slugs get `low`, non-reasoning slugs `none`. |53| `grok-code-fast-1` | retired 2026-05-15 | RETIRED → alias of `grok-build-0.1` | `GET /v1/models/grok-code-fast-1` returns the grok-build-0.1 object. |54| `grok-imagine-image-pro` | retired 2026-05-15 | RETIRED → alias of `grok-imagine-image-quality` → 2.0 `low` from 2026-11-02 | |55| `grok-2-image(-1212)`, `grok-2-vision-1212`, `grok-2-1212`, `grok-3-mini(-fast)`, `grok-beta`, `grok-vision-beta`, `grok-4(-latest)` | historical | LEGACY, UNVERIFIED | Release-notes era ids; `GET /v1/models/grok-2-image` → 404 `not-found`. |5657## 3. Live vs docs discrepancies (2026-09-19)5859- **Price units:** the REST reference says token prices are "USD cents per 100 million tokens" and `image_price` is "USD cents"; live values are consistent with **ticks of 1e-10 USD** for both (`12500` → $1.25/1M tokens, `200000000` → $0.02/image). `price_per_image` is documented as "1/100,000,000ths of a USD cent" = the same unit. `cost_in_usd_ticks` (usage) uses the same 1e10 ticks/USD.60- `grok-4.5`: model page lists `xhigh` as supported; reasoning page says only low/medium/high and `xhigh` is treated as `high`.61- `grok-4.20-0309-reasoning`, `grok-4.20-0309-non-reasoning`, `grok-build-0.1`: model pages show "Reasoning: Yes/No" but `reasoning_effort` → `400 invalid-argument "Model … does not support parameter reasoningEffort"`. Only grok-4.6, grok-4.5, grok-4.3 (and multi-agent via `reasoning.effort`) accept it.62- `grok-4.20-multi-agent-0309`: `POST /v1/chat/completions` → `400 "Multi Agent requests are not allowed on chat completions"` (bare string body); works on `POST /v1/responses`. The response's `reasoning.effort` defaulted to `"medium"`; output text carried a trailing `\confidence{80}` / `\confidence{90}` marker; usage contained undocumented `context_details{input_tokens,output_tokens}`.63- `grok-imagine-video-1.5`: live `input_modalities` = text, image, **audio**; model page says text, image.64- `grok-4.5` created 2026-06-29 (live) vs July 2026 release note; `grok-build-0.1` created 2026-04-16 vs May 2026 note.65- Regional catalogues: `GET https://us.api.x.ai/v1/models` → only `grok-4.6` with prices ×1.1 (22000 / 5500 / 66000 ticks, long-context 44000 / 11000 / 132000). **`GET https://eu-west-1.api.x.ai/v1/models` → 200, only `grok-4.3` at global prices — this host is not documented** (model pages mention the `eu-west-1` cluster only).66- `search_price` is `0` for every language model (Live Search superseded by tools).67- `GET /v1/api-key` returned empty strings for `create_time` / `modify_time` (docs: Unix timestamp; example: RFC 3339).68- Documentation stubs: the gRPC reference pages and per-model pages are rendered client-side — `llms-full.txt` contains only their titles. Per-model pages were fetched live as `.md`; gRPC methods were taken from the `xai-sdk` 1.19 proto stubs (see `docs/xai/grpc-api.md`).6970## 4. Model × Capability matrix (text models)7172| Capability | 4.6 | 4.5 | 4.3 | 4.20 R | 4.20 NR | 4.20 MA | build-0.1 |73|---|---|---|---|---|---|---|---|74| image input | yes | yes | yes | yes | yes | yes | yes |75| reasoning | always | always | optional (`none`) | always | no | multi-agent | always |76| `reasoning_effort` | low/med/high/xhigh | low/med/high (xhigh→high) | none/low/med/high/xhigh | **400** | **400** | agent count | **400** |77| `reasoning_content` in chat response (live) | yes | yes | (none → absent) | n/a | no | n/a | yes |78| function calling / parallel tools | yes | yes | yes | yes | yes | yes | yes |79| structured outputs | yes | yes | yes | yes | yes | yes | yes |80| server-side tools (web/x search, code exec…) | yes (documented) | unknown | unknown | unknown | unknown | unknown (deep research) | unknown |81| `logprobs` | ignored | ignored | ignored | ignored | ignored | ignored | unknown |82| automatic prompt caching (`cached_tokens`) | yes (512 cached on a 640-token prompt) | yes | yes | yes | yes | yes (2560/2655) | yes |83| Batch API | no | no | yes −20 % | yes −20 % | yes −20 % | yes −20 % | no |84| Priority (`service_tier: priority`, 2×) | yes | yes | yes | yes | yes | not documented | yes |85| deferred completions | yes | yes | yes | yes | yes | no (Responses only) | yes |86| context compaction / WebSocket Responses | yes | yes | yes | yes | yes | yes | yes |87| `us.api.x.ai` regional | **only model** | no | no | no | no | no | no |88| `eu-west-1` cluster | no | no | yes | no | no | no | no |89| fine-tuning | none offered by xAI | | | | | | |9091## 5. Model × Endpoint matrix9293| Endpoint | text models | 4.20 multi-agent | image models | video models | voice |94|---|---|---|---|---|---|95| `POST /v1/chat/completions` | yes | **400** | — | — | — |96| `POST /v1/responses` (+ `wss://api.x.ai/v1/responses`, `/v1/responses/compact`) | yes | yes (only) | (image_generation tool) | — | — |97| `POST /v1/messages` (Anthropic-compatible) / `POST /v1/complete` | yes / legacy non-reasoning | ? | — | — | — |98| `POST /v1/completions` (legacy) | non-reasoning only | — | — | — | — |99| `GET /v1/chat/deferred-completion/{id}` | yes | — | — | — | — |100| `POST /v1/images/generations`, `/v1/images/edits` | — | — | yes | — | — |101| `POST /v1/videos/{generations,edits,extensions}`, `GET /v1/videos/{id}` | — | — | — | yes | — |102| Batch API | 4.3 / 4.20 R / 4.20 NR | yes | 1.0, 2.0 (not quality) | yes | — |103| `POST /v1/tokenize-text` | yes | yes | — | — | — |104| `wss://api.x.ai/v1/realtime`, `POST /v1/stt`, `POST /v1/tts` | — | — | — | — | yes |105106## 6. Model × Tool matrix (type strings)107108Responses API (`tools[].type`): `web_search`, `x_search`, `code_execution` (alias `code_interpreter`), `image_generation`, `attachment_search`, `collections_search` (alias `file_search`), `view_image` / `view_x_video` (implicit, on search results), `mcp` (remote MCP), `function` (client). gRPC / xai-sdk: `web_search()`, `x_search()`, `code_execution()`, `collections_search()`, `mcp()` — `code_interpreter` and `file_search` aliases unsupported. Documentation lists tool support explicitly only for grok-4.6 (function calling, web search, X search, code execution); all text models accept function tools. Details: `docs/tools/xai/index.md`.109110## 7. Pricing summary (USD per 1M tokens; full detail in `docs/xai/pricing.md`)111112| Model | Input | Cached | Output | ≥ 200k prompt: input / cached / output | Batch | Priority |113|---|---|---|---|---|---|---|114| grok-4.6 | 2.00 | 0.50 | 6.00 | 4.00 / 1.00 / 12.00 | — | ×2 |115| grok-4.5 | 2.00 | 0.30 | 6.00 | 4.00 / 0.60 / 12.00 | — | ×2 |116| grok-4.3 | 1.25 | 0.20 | 2.50 | 2.50 / 0.40 / 5.00 | −20 % | ×2 |117| grok-4.20 (R, NR, multi-agent) | 1.25 | 0.20 | 2.50 | 2.50 / 0.40 / 5.00 | −20 % | ×2 |118| grok-build-0.1 | 1.00 | 0.20 | 2.00 | 2.00 / 0.40 / 4.00 | — | ×2 |119120Tiered rule: when the prompt reaches 200k tokens **all** tokens of the request are billed at the long-context rate. Image input tokens cost the same as text. Reasoning tokens are output tokens. Images: 1.0 $0.02, quality $0.05, 2.0 $0.04 default (low 1k/1.5k/2k = 0.04/0.05/0.06; medium = 0.06/0.07/0.08). Video: 1.0 $0.05/s, 1.5 $0.08/s.121122## 8. Rate limits (Tier 0, documented → observed)123124| Model | Doc RPS T0 | Doc TPM T0 | Observed `x-ratelimit-limit-requests` | Observed `x-ratelimit-limit-tokens` |125|---|---|---|---|---|126| grok-4.6, grok-4.5 | 150 | 50M | 7200 | 50000000 |127| grok-4.3, 4.20 R/NR, build-0.1 | 37 | 10M | 1800 | 10000000 |128| grok-4.20-multi-agent | 9 | 2.5M | (no headers on Responses) | — |129| image models | 6 | — | — | — |130| video models | 10 | — | — | — |131132Full tier tables: `docs/xai/rate-limits.md`.133134## 9. Observed usage shapes (minimal "Reply with OK.", max_tokens 8)135136| Model | prompt / cached | completion | reasoning | cost (ticks → USD) |137|---|---|---|---|---|138| grok-4.6 | 640 / 512 | 1 | 91 | 10,640,000 → $0.00106 |139| grok-4.5 | 498 / 384 | 1 | 55 | 6,792,000 → $0.00068 |140| grok-4.20-0309-non-reasoning | 188 / 128 | 2 | 0 | 1,056,000 → $0.00011 |141| grok-4.3 (`reasoning_effort: none`) | 188 / 128 | 2 | 0 | 1,056,000 → $0.00011 |142| grok-build-0.1 | 190 / 128 | 1 | 162 | 4,136,000 → $0.00041 |143| grok-4.20-multi-agent-0309 (Responses) | 2655 / 2560 | 1645 (1618 reasoning) | — | 47,432,500 → $0.0047 |144145Chat `usage` = `{prompt_tokens, completion_tokens, total_tokens, prompt_tokens_details{text_tokens, audio_tokens, image_tokens, cached_tokens}, completion_tokens_details{reasoning_tokens, audio_tokens, accepted_prediction_tokens, rejected_prediction_tokens}, num_sources_used, cost_in_usd_ticks}` plus top-level `system_fingerprint`, `service_tier: "default"`, `message.reasoning_content`, `message.refusal`. Responses `usage` adds `num_server_side_tools_used`, `input_tokens_details.cached_tokens`, `output_tokens_details.reasoning_tokens`, `context_details`. Note the ~190–640 prompt tokens for a 4-token prompt (system template overhead; FAQ acknowledges API counts exceed the tokenizer).146