xAI reasoning — always-on reasoning models, reasoning_effort, reasoning_content, encrypted reasoning
Status: DOCUMENTED + LIVE_VERIFIED (2026-09-19). Machine-readable: parameters/xai-chat-completions.json (reasoning_effort), parameters/xai-responses.json (reasoning.*, include), objects/xai-inference-objects.json (Reasoning item).
Sources: https://docs.x.ai/developers/model-capabilities/text/reasoning · https://docs.x.ai/developers/model-capabilities/text/generate-text#returning-encrypted-thinking-content · https://docs.x.ai/developers/model-capabilities/text/multi-agent · https://docs.x.ai/developers/models (live /v1/language-models)
Last verified: 2026-09-19
Model classes (live GET /v1/language-models, 7 text models)
| Model | Reasoning | reasoning_effort values |
Live notes |
|---|---|---|---|
grok-4.6 (aliases none) |
always on | low medium high (default) xhigh |
knowledge cutoff 2026-02-01; summarized reasoning exposed |
grok-4.5 (grok-4.5-latest, grok-build-latest) |
always on | low medium high; xhigh treated as high |
|
grok-4.3 (grok-4.3-latest) |
always on | live: low/medium/high/xhigh all 200; none → 200 with reasoning_tokens: 0 (LIVE_DISCOVERED, undocumented) |
70–180 reasoning tokens on "Reply with OK."; response echo reasoning.effort:"low" when unset |
grok-4.20-0309-reasoning (grok-4.20, …) |
always on | docs: see model page | |
grok-4.20-0309-non-reasoning (grok-4.20-non-reasoning, …) |
none | ❌ reasoning_effort → 400 "does not support parameter reasoningEffort" |
reasoning_tokens: 0, no reasoning_content key; Responses echo reasoning:{effort:null, summary:null} |
grok-4.20-multi-agent-0309 |
multi-agent | effort controls agent count (low/medium → 4, high/xhigh → 16) |
not called |
grok-build-0.1 (grok-code-fast-1, …) |
reasoning (coding) | not called |
Reasoning cannot be disabled per docs (except via the live none value on grok-4.3). presence_penalty, frequency_penalty, stop return 400 on reasoning models.
Where the reasoning shows up
| Surface | Field | Content |
|---|---|---|
| Chat Completions | choices[].message.reasoning_content (string) |
short trace/summary, e.g. The user requested a reply of "OK."; streamed as delta.reasoning_content before content |
| Chat usage | usage.completion_tokens_details.reasoning_tokens |
billed at output price; total_tokens includes them |
| Responses | output[] item {type:"reasoning", id:"rs_…", summary:[{type:"summary_text", text}], status, encrypted_content?} |
summary always "detailed" per docs; item sometimes omitted; usage.output_tokens_details.reasoning_tokens |
| Responses stream | response.reasoning_summary_part.added/…text.delta/…text.done/…part.done |
response.reasoning_text.delta documented, not observed |
/v1/messages |
content[] {type:"thinking", thinking, signature:""} (+ stream thinking_delta) |
appeared with tool_use and in streams; usage.output_tokens includes reasoning |
Encrypted reasoning (stateless multi-turn / ZDR)
include: ["reasoning.encrypted_content"] (or xai-sdk use_encrypted_content=True) adds encrypted_content to the reasoning item (also visible in the stream's output_item.done). Replay ...response.output in the next input — live 200. Required for cache hits and correct agentic state when not using previous_response_id. Also the mechanism behind context compaction. Vercel AI SDK does this automatically unless store:false.
Cost
Reasoning tokens bill at the output rate (grok-4.3: $2.50/M). A minimal call ≈ 200 prompt + 2 output + ~130 reasoning ≈ $0.0006 (usage.cost_in_usd_ticks 3 809 000 = $0.00038 with cache). Batch API discounts (20 % on grok-4.3 / 4.20) apply to reasoning tokens too.