xAI streaming — Chat Completions chunks, Responses events, Messages-compat events
Status: DOCUMENTED + LIVE_VERIFIED (2026-09-19, grok-4.3; 5 streams captured: chat minimal, chat with reasoning, chat tool call, Responses minimal/reasoning/function/code_interpreter, Messages minimal). Machine-readable: generated/fragments/streaming-events/xai-inference.json.
Sources: https://docs.x.ai/developers/model-capabilities/text/streaming · https://docs.x.ai/developers/model-capabilities/text/reasoning (summary deltas) · https://docs.x.ai/developers/tools/streaming-and-sync · https://docs.x.ai/developers/tools/function-calling (tool call in one chunk) · OpenAPI ChatResponseChunk, Delta
Last verified: 2026-09-19
Set "stream": true. Streaming is supported by every text-output model, not by image/video generation. Reasoning models can take long: raise client timeouts (docs use 3600 s).
1. Chat Completions (data:-only SSE, data: [DONE] terminator)
Observed sequence for "What is 17*23?" on grok-4.3:
data: {"id":"d8a0c652-…","object":"chat.completion.chunk","created":0,"model":"grok-4.3","choices":[{"index":0,"delta":{"reasoning_content":"Calculating 17 times","role":"assistant"}}],"system_fingerprint":"fp_…","service_tier":"default"}
data: {…"choices":[{"index":0,"delta":{"reasoning_content":" 23."}}]…}
data: {…"choices":[{"index":0,"delta":{"content":"391"}}]…}
data: {…"choices":[{"index":0,"delta":{},"finish_reason":"stop"}]…}
data: {…"choices":[],"usage":{"prompt_tokens":196,"completion_tokens":2,"total_tokens":319,"prompt_tokens_details":{…"cached_tokens":192},"completion_tokens_details":{"reasoning_tokens":121,…},"num_sources_used":0,"cost_in_usd_ticks":3509000}} ← only with stream_options.include_usage
data: [DONE]| Field | Notes |
|---|---|
delta.reasoning_content |
reasoning deltas stream before content (only when the model produces a summary; the "Reply with OK." stream had none) |
delta.content |
text deltas |
delta.tool_calls[] |
whole call in one chunk: [{index:0, id:"call-…", type:"function", function:{name, arguments:"{…}"}}]; then finish_reason:"tool_calls" |
delta.images[] |
image-generation tool output (documented) |
created |
0 in every chunk (live) |
usage |
null/absent except the final usage chunk (choices: []) |
citations |
last chunk only (Live Search, retired) |
Errors during streaming are returned as a normal JSON error before any chunk (e.g. 400).
2. Responses API (event: + data:, sequence_number, no [DONE])
Minimal text (10 events): response.created → response.in_progress → response.output_item.added(message) → response.content_part.added → response.output_text.delta ×N → response.output_text.done → response.content_part.done → response.output_item.done → response.completed.
With reasoning summary + function call (17 events):
response.created, response.in_progress,
response.output_item.added {item:{type:"reasoning", id:"rs_…", summary:[], status:"in_progress"}, output_index:0},
response.reasoning_summary_part.added {part:{type:"summary_text", text:""}, summary_index:0},
response.reasoning_summary_text.delta {delta:"I need to calculate "} …,
response.reasoning_summary_text.done {text:"…"}, response.reasoning_summary_part.done,
response.output_item.done {item:{type:"reasoning", …, "encrypted_content":"…"}} ← encrypted_content when include requested
response.output_item.added {item:{type:"function_call", arguments:"", call_id:"call-…", name:"get_weather", status:"in_progress"}, output_index:1},
response.function_call_arguments.delta {delta:"{\"city\":\"Paris\"}"}, ← one delta with the full JSON
response.function_call_arguments.done {arguments, name},
response.output_item.done, response.completedCode interpreter adds: response.code_interpreter_call.in_progress → response.code_interpreter_call_code.delta → response.code_interpreter_call_code.done → response.code_interpreter_call.interpreting → response.code_interpreter_call.completed between output_item.added/done (outputs only in the final item with include:["code_interpreter_call.outputs"]).
Documented but not observed: response.reasoning_text.delta (raw reasoning; the guide lists it next to reasoning_summary_text.delta). Not exercised in stream: web_search / x_search / mcp / file_search progress events. response.created snapshots show reasoning:{effort:null, summary:null}, created_at:0.
3. /v1/messages (Anthropic-compatible, DEPRECATED)
event: message_start → content_block_start {index:0, content_block:{type:"thinking", thinking:"", signature:""}} → content_block_delta {delta:{type:"thinking_delta", thinking}} ×N → content_block_stop {index:0} → content_block_start {index:0, content_block:{type:"text", text:""}} (index is 0 again — Anthropic would use 1) → content_block_delta {delta:{type:"text_delta", text}} ×N (no index on deltas) → content_block_stop → message_delta {delta:{stop_reason:"end_turn", stop_sequence:null}, usage:{output_tokens:147}} → message_stop. No ping, no [DONE].
4. Client notes
- Python:
openaiSDKclient.chat.completions.create(stream=True)/client.responses.create(stream=True)parse both framings;examples/xai/streaming/*.py. - Raw:
scripts/live.py xai_request(..., stream=True)yields lines; split onevent:/data:. - xai-sdk (gRPC):
for response, chunk in chat.stream(): chunk.content / chunk.reasoning_content / chunk.tool_calls;include=["verbose_streaming"]surfaces server-side tool calls as they happen.