SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
3.8 KB

# xAI reasoning — always-on reasoning models, reasoning_effort, reasoning_content, encrypted reasoning

Status: DOCUMENTED + LIVE_VERIFIED (2026-09-19). Machine-readable: parameters/xai-chat-completions.json (reasoning_effort), parameters/xai-responses.json (reasoning.*, include), objects/xai-inference-objects.json (Reasoning item).

Sources: https://docs.x.ai/developers/model-capabilities/text/reasoning · https://docs.x.ai/developers/model-capabilities/text/generate-text#returning-encrypted-thinking-content · https://docs.x.ai/developers/model-capabilities/text/multi-agent · https://docs.x.ai/developers/models (live /v1/language-models) Last verified: 2026-09-19

# Model classes (live GET /v1/language-models, 7 text models)

Model Reasoning reasoning_effort values Live notes
grok-4.6 (aliases none) always on low medium high (default) xhigh knowledge cutoff 2026-02-01; summarized reasoning exposed
grok-4.5 (grok-4.5-latest, grok-build-latest) always on low medium high; xhigh treated as high
grok-4.3 (grok-4.3-latest) always on live: low/medium/high/xhigh all 200; none → 200 with reasoning_tokens: 0 (LIVE_DISCOVERED, undocumented) 70–180 reasoning tokens on "Reply with OK."; response echo reasoning.effort:"low" when unset
grok-4.20-0309-reasoning (grok-4.20, …) always on docs: see model page
grok-4.20-0309-non-reasoning (grok-4.20-non-reasoning, …) none ❌ reasoning_effort → 400 "does not support parameter reasoningEffort" reasoning_tokens: 0, no reasoning_content key; Responses echo reasoning:{effort:null, summary:null}
grok-4.20-multi-agent-0309 multi-agent effort controls agent count (low/medium → 4, high/xhigh → 16) not called
grok-build-0.1 (grok-code-fast-1, …) reasoning (coding) not called

Reasoning cannot be disabled per docs (except via the live none value on grok-4.3). presence_penalty, frequency_penalty, stop return 400 on reasoning models.

# Where the reasoning shows up

Surface Field Content
Chat Completions choices[].message.reasoning_content (string) short trace/summary, e.g. The user requested a reply of "OK."; streamed as delta.reasoning_content before content
Chat usage usage.completion_tokens_details.reasoning_tokens billed at output price; total_tokens includes them
Responses output[] item {type:"reasoning", id:"rs_…", summary:[{type:"summary_text", text}], status, encrypted_content?} summary always "detailed" per docs; item sometimes omitted; usage.output_tokens_details.reasoning_tokens
Responses stream response.reasoning_summary_part.added/…text.delta/…text.done/…part.done response.reasoning_text.delta documented, not observed
/v1/messages content[] {type:"thinking", thinking, signature:""} (+ stream thinking_delta) appeared with tool_use and in streams; usage.output_tokens includes reasoning

# Encrypted reasoning (stateless multi-turn / ZDR)

include: ["reasoning.encrypted_content"] (or xai-sdk use_encrypted_content=True) adds encrypted_content to the reasoning item (also visible in the stream's output_item.done). Replay ...response.output in the next input — live 200. Required for cache hits and correct agentic state when not using previous_response_id. Also the mechanism behind context compaction. Vercel AI SDK does this automatically unless store:false.

# Cost

Reasoning tokens bill at the output rate (grok-4.3: $2.50/M). A minimal call ≈ 200 prompt + 2 output + ~130 reasoning ≈ $0.0006 (usage.cost_in_usd_ticks 3 809 000 = $0.00038 with cache). Batch API discounts (20 % on grok-4.3 / 4.20) apply to reasoning tokens too.