SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
5.6 KB

# xAI streaming — Chat Completions chunks, Responses events, Messages-compat events

Status: DOCUMENTED + LIVE_VERIFIED (2026-09-19, grok-4.3; 5 streams captured: chat minimal, chat with reasoning, chat tool call, Responses minimal/reasoning/function/code_interpreter, Messages minimal). Machine-readable: generated/fragments/streaming-events/xai-inference.json.

Sources: https://docs.x.ai/developers/model-capabilities/text/streaming · https://docs.x.ai/developers/model-capabilities/text/reasoning (summary deltas) · https://docs.x.ai/developers/tools/streaming-and-sync · https://docs.x.ai/developers/tools/function-calling (tool call in one chunk) · OpenAPI ChatResponseChunk, Delta Last verified: 2026-09-19

Set "stream": true. Streaming is supported by every text-output model, not by image/video generation. Reasoning models can take long: raise client timeouts (docs use 3600 s).

# 1. Chat Completions (data:-only SSE, data: [DONE] terminator)

Observed sequence for "What is 17*23?" on grok-4.3:

text
data: {"id":"d8a0c652-…","object":"chat.completion.chunk","created":0,"model":"grok-4.3","choices":[{"index":0,"delta":{"reasoning_content":"Calculating 17 times","role":"assistant"}}],"system_fingerprint":"fp_…","service_tier":"default"}
data: {…"choices":[{"index":0,"delta":{"reasoning_content":" 23."}}]…}
data: {…"choices":[{"index":0,"delta":{"content":"391"}}]…}
data: {…"choices":[{"index":0,"delta":{},"finish_reason":"stop"}]…}
data: {…"choices":[],"usage":{"prompt_tokens":196,"completion_tokens":2,"total_tokens":319,"prompt_tokens_details":{…"cached_tokens":192},"completion_tokens_details":{"reasoning_tokens":121,…},"num_sources_used":0,"cost_in_usd_ticks":3509000}}   ← only with stream_options.include_usage
data: [DONE]
Field Notes
delta.reasoning_content reasoning deltas stream before content (only when the model produces a summary; the "Reply with OK." stream had none)
delta.content text deltas
delta.tool_calls[] whole call in one chunk: [{index:0, id:"call-…", type:"function", function:{name, arguments:"{…}"}}]; then finish_reason:"tool_calls"
delta.images[] image-generation tool output (documented)
created 0 in every chunk (live)
usage null/absent except the final usage chunk (choices: [])
citations last chunk only (Live Search, retired)

Errors during streaming are returned as a normal JSON error before any chunk (e.g. 400).

# 2. Responses API (event: + data:, sequence_number, no [DONE])

Minimal text (10 events): response.created → response.in_progress → response.output_item.added(message) → response.content_part.added → response.output_text.delta ×N → response.output_text.done → response.content_part.done → response.output_item.done → response.completed.

With reasoning summary + function call (17 events):

text
response.created, response.in_progress,
response.output_item.added {item:{type:"reasoning", id:"rs_…", summary:[], status:"in_progress"}, output_index:0},
response.reasoning_summary_part.added {part:{type:"summary_text", text:""}, summary_index:0},
response.reasoning_summary_text.delta {delta:"I need to calculate "} …,
response.reasoning_summary_text.done {text:"…"}, response.reasoning_summary_part.done,
response.output_item.done {item:{type:"reasoning", …, "encrypted_content":"…"}}     ← encrypted_content when include requested
response.output_item.added {item:{type:"function_call", arguments:"", call_id:"call-…", name:"get_weather", status:"in_progress"}, output_index:1},
response.function_call_arguments.delta {delta:"{\"city\":\"Paris\"}"},               ← one delta with the full JSON
response.function_call_arguments.done {arguments, name},
response.output_item.done, response.completed

Code interpreter adds: response.code_interpreter_call.in_progress → response.code_interpreter_call_code.delta → response.code_interpreter_call_code.done → response.code_interpreter_call.interpreting → response.code_interpreter_call.completed between output_item.added/done (outputs only in the final item with include:["code_interpreter_call.outputs"]).

Documented but not observed: response.reasoning_text.delta (raw reasoning; the guide lists it next to reasoning_summary_text.delta). Not exercised in stream: web_search / x_search / mcp / file_search progress events. response.created snapshots show reasoning:{effort:null, summary:null}, created_at:0.

# 3. /v1/messages (Anthropic-compatible, DEPRECATED)

event: message_start → content_block_start {index:0, content_block:{type:"thinking", thinking:"", signature:""}} → content_block_delta {delta:{type:"thinking_delta", thinking}} ×N → content_block_stop {index:0} → content_block_start {index:0, content_block:{type:"text", text:""}} (index is 0 again — Anthropic would use 1) → content_block_delta {delta:{type:"text_delta", text}} ×N (no index on deltas) → content_block_stop → message_delta {delta:{stop_reason:"end_turn", stop_sequence:null}, usage:{output_tokens:147}} → message_stop. No ping, no [DONE].

# 4. Client notes

  • Python: openai SDK client.chat.completions.create(stream=True) / client.responses.create(stream=True) parse both framings; examples/xai/streaming/*.py.
  • Raw: scripts/live.py xai_request(..., stream=True) yields lines; split on event:/data:.
  • xai-sdk (gRPC): for response, chunk in chat.stream(): chunk.content / chunk.reasoning_content / chunk.tool_calls; include=["verbose_streaming"] surfaces server-side tool calls as they happen.