Streaming — four SSE dialects and five WebSocket families
Status: event names and statuses from generated/streaming-events.json (506 records: OpenAI 282, Anthropic 68, xAI 94, Gemini 62); observed sequences quoted from docs/openai/streaming-events.md, docs/anthropic/streaming.md, docs/xai/streaming.md, docs/gemini/streaming.md, docs/gemini/interactions-api.md, docs/gemini/live-events.md (LIVE_VERIFIED 2026-09-18/19 on gpt-5.4-nano, claude-haiku-4-5-20251001, grok-4.3, gemini-3.5-flash-lite / gemini-3.8-flash).
Sources: https://developers.openai.com/api/docs/guides/streaming-responses · https://developers.openai.com/api/reference/resources/responses/streaming-events · https://platform.claude.com/docs/en/build-with-claude/streaming · https://docs.x.ai/developers/model-capabilities/text/streaming · https://docs.x.ai/developers/model-capabilities/audio/speech-to-speech · https://ai.google.dev/api/generate-content#method:-models.streamgeneratecontent · https://ai.google.dev/gemini-api/docs/live · https://ai.google.dev/gemini-api/docs/interactions
Last verified: 2026-09-18
1. Wire format
| OpenAI Responses | OpenAI Chat Completions | Anthropic Messages | xAI Responses | xAI Chat Completions | xAI /v1/messages (deprecated) |
Gemini streamGenerateContent |
Gemini Interactions | |
|---|---|---|---|---|---|---|---|---|
| Framing | event: <type> + data:; data.type == event |
data: <chunk> only |
event: <type> + data:; data.type == event |
event: <type> + data: (same names as OpenAI) |
data: only (chat.completion.chunk) |
event: + data: (Anthropic names) |
?alt=sse: data: <GenerateContentResponse> lines, no event: names; default: one pretty-printed JSON array streamed incrementally ([{…} , {…} ]) |
event: <type> + data: with event_type, event_id |
| Ordering guard | sequence_number |
none | none (fixed skeleton) | sequence_number |
none | none | none (chunk order; usageMetadata on every chunk) |
event_id (resume with last_event_id) |
| Terminal | response.completed|failed|incomplete | error — no [DONE] |
data: [DONE] after an optional usage chunk |
message_stop — no [DONE] |
response.completed (incomplete/failed/error UNVERIFIED) — no [DONE] |
data: [DONE] (usage chunk with choices: [] when stream_options.include_usage) |
message_stop |
last chunk carries finishReason (and on Gemini 3 an empty-text part with thoughtSignature — keep it); stream closes |
interaction.completed then done (data: [DONE]) |
| Padding | obfuscation on delta events |
obfuscation |
none | none observed | none | none | none | none |
| Full object in stream | lifecycle events embed the Response |
none | message_start skeleton; cumulative usage in message_delta |
response.created snapshot (reasoning: {effort: null}, created_at: 0) and full response in response.completed |
none (created: 0 in every chunk) |
message_start skeleton |
every chunk is a partial GenerateContentResponse with modelVersion, responseId, usageMetadata (last authoritative) |
step.start/step.stop carry step objects; interaction.completed the final interaction |
| Keep-alive | none | none | ping |
none | none | none (no ping) |
none | none |
| Errors after 200 | error event (terminal) |
error chunk | error event (overloaded_error…), no request_id |
error (UNVERIFIED) |
— | — | undocumented (treat a close without finishReason as an error); prompt block → single chunk with promptFeedback.blockReason, no candidates |
event: error {error: {code, message}} |
| Content type | text/event-stream |
text/event-stream |
text/event-stream |
text/event-stream |
text/event-stream |
text/event-stream |
text/event-stream (alt=sse) or application/json |
text/event-stream |
2. Skeletons (observed live)
OpenAI Responses — "Reply with OK."
response.created(0) → response.in_progress(1) → response.output_item.added(2, message) → response.content_part.added(3)
→ response.output_text.delta(4 "OK") → response.output_text.done(5) → response.content_part.done(6) → response.output_item.done(7) → response.completed(8)Anthropic Messages
message_start → content_block_start(0, text) → ping → content_block_delta(text_delta "OK") → content_block_delta(".") → content_block_stop(0)
→ message_delta(stop_reason end_turn, usage) → message_stopxAI Responses (grok-4.3) — reasoning item first
response.created → response.in_progress → response.output_item.added(reasoning) → response.reasoning_summary_part.added → response.reasoning_summary_text.delta… → response.reasoning_summary_text.done
→ response.reasoning_summary_part.done → response.output_item.done → response.output_item.added(message) → response.content_part.added → response.output_text.delta("OK") → response.output_text.done
→ response.content_part.done → response.output_item.done → response.completedxAI Chat Completions
data: {chunk delta.reasoning_content…} → data: {chunk delta.content "OK"} → data: {chunk finish_reason: stop} → [data: {choices: [], usage}] → data: [DONE]xAI /v1/messages — message_start → content_block_start(thinking) → content_block_delta(thinking_delta)… → content_block_stop → content_block_start(text, index 0 again) → content_block_delta(text_delta) → content_block_stop → message_delta → message_stop (no ping; index reuse is a quirk).
Gemini streamGenerateContent?alt=sse (gemini-3.8-flash)
data: {candidates:[{content:{parts:[{text:"OK", thought?}]}}], usageMetadata, modelVersion, responseId}
data: {candidates:[{content:{parts:[{text:"", thoughtSignature:"…"}]}, finishReason:"STOP"}], usageMetadata:{…thoughtsTokenCount}}With includeThoughts: true, {text, thought: true} parts stream first (not guaranteed on trivial prompts). JSON-mode deltas are partial JSON strings that concatenate.
Gemini Interactions (stream: true)
interaction.created → interaction.status_update → step.start(thought) → step.delta{thought_signature} → step.stop → step.start(model_output) → step.delta{text "OK"} → step.stop → interaction.completed → done [DONE]3. Event-name mapping — the four text streams
| Concern | OpenAI Responses | Anthropic Messages | xAI Responses / Chat | Gemini streamGenerateContent / Interactions |
|---|---|---|---|---|
| lifecycle | response.created, response.queued, response.in_progress, response.completed|failed|incomplete |
message_start, message_delta (stop_reason, usage), message_stop |
same as OpenAI (response.queued absent — no background mode); Chat: first chunk / finish_reason / [DONE] |
no lifecycle events — first chunk / chunk with finishReason; Interactions interaction.created, interaction.status_update (legacy but still emitted), interaction.in_progress, interaction.requires_action, interaction.completed, done |
| new output unit | response.output_item.added / .done |
content_block_start / content_block_stop (index) |
same as OpenAI; Chat: implicit | new parts[] entries in a chunk; Interactions step.start / step.stop |
| text | response.content_part.added/done, response.output_text.delta/done, response.output_text.annotation.added, response.refusal.delta/done |
content_block_delta {type: text_delta}; citations_delta |
same as OpenAI (response.output_text.delta, .done); Chat delta.content |
parts[].text fragments; Interactions step.delta {type: text}, text_annotation_delta |
| function / tool args | response.function_call_arguments.delta/done, response.custom_tool_call_input.delta/done |
content_block_delta {type: input_json_delta, partial_json} (buffered; eager_input_streaming) |
response.function_call_arguments.delta (one delta with the full JSON) + .done; Chat delta.tool_calls[] whole call in one chunk → finish_reason: tool_calls |
parts[].functionCall arrives complete in one chunk (+ thoughtSignature); Interactions step.delta {type: arguments_delta} (partial JSON) |
| reasoning | response.reasoning_summary_part.added/done, response.reasoning_summary_text.delta/done, response.reasoning_text.delta/done |
content_block_start(thinking), thinking_delta, signature_delta, redacted_thinking |
same summary events as OpenAI (reasoning_text.delta UNVERIFIED); Chat delta.reasoning_content (full text) |
parts[] {text, thought: true} chunks, final thoughtSignature; Interactions step.start(thought), step.delta {type: thought_summary | thought_signature} |
| web search | response.web_search_call.in_progress/searching/completed |
server_tool_use + input_json_delta, web_search_tool_result block |
response.output_item.added/done(web_search_call) (.searching events documented, not observed) |
groundingMetadata on candidates (chunks carry only new groundingChunks — accumulate); Interactions google_search_call / google_search_result steps |
| code execution | response.code_interpreter_call.in_progress/interpreting/completed, response.code_interpreter_call_code.delta/done |
server_tool_use + bash_code_execution_tool_result |
identical event names to OpenAI (all LIVE_VERIFIED) | executableCode / codeExecutionResult / inlineData parts as produced; Interactions code_execution_call / code_execution_result |
| file search / RAG | response.file_search_call.in_progress/searching/completed |
— | response.file_search_call.* (UNVERIFIED) |
groundingMetadata.groundingChunks[].retrievedContext; Interactions file_search_call/result |
| image generation | response.image_generation_call.in_progress/generating/partial_image/completed |
— | image_generation_call item events (not exercised) |
inlineData {mimeType: image/png} parts (image models); Interactions step.delta {type: image} |
| MCP | response.mcp_list_tools.*, response.mcp_call.*, response.mcp_call_arguments.delta/done |
mcp_tool_use / mcp_tool_result blocks |
response.mcp_call.* (UNVERIFIED) |
Interactions mcp call/result steps (UNVERIFIED) |
| shell / patch | response.shell_call_command.*, response.shell_call_output_content.*, response.apply_patch_call_operation_diff.* (LIVE_DISCOVERED) |
tool_use + input_json_delta |
shell_call items (not exercised) |
— |
| compaction | response.compaction.compacting |
content_block_*(compaction) (beta) |
— (compact endpoint is unary) | — |
| audio | response.audio.delta/done, response.audio.transcript.delta/done |
— | — (voice WebSocket) | inlineData audio parts (TTS models; streaming TTS on 3.1); Interactions step.delta {type: audio} |
| errors | error |
error |
error (UNVERIFIED) |
pre-stream JSON error; promptFeedback.blockReason; Interactions error |
| keep-alive | — | ping |
— | — |
Counts: OpenAI Responses SSE 63 + WebSocket 69/3, Chat 2; Anthropic Messages 13 core + 10 advanced; xAI Responses 24, Chat 3, Messages 6; Gemini generate-content 8 (framing + chunk kinds), Interactions 15. Full lists: FAQ Q4.
4. Usage and stop information
| OpenAI | Anthropic | xAI | Gemini | |
|---|---|---|---|---|
| where usage arrives | response.completed.response.usage (+ prompt_cache_diagnostics 5.6+) |
message_start.message.usage then cumulative message_delta.usage |
response.completed.response.usage (incl. cost_in_usd_ticks, server_side_tool_usage_details); Chat: final chunk with choices: [] when stream_options.include_usage: true |
usageMetadata on every chunk (last one authoritative: thoughtsTokenCount, toolUsePromptTokenCount, cachedContentTokenCount); Interactions interaction.completed.usage |
| stop reason | response.status + incomplete_details.reason |
message_delta.delta.stop_reason / stop_sequence / stop_details |
response.status + incomplete_details.reason (max_output_tokens | max_prompt_tokens | max_time_limit); Chat finish_reason |
candidates[].finishReason on the last chunk (21 values); Interactions status |
| tool call boundary | response.output_item.done for the function_call item |
content_block_stop of the tool_use block; stop_reason: tool_use |
response.output_item.done(function_call); Chat finish_reason: tool_calls |
chunk containing the functionCall part (usually the last, finishReason: STOP); Interactions interaction.requires_action |
5. The same streamed call on all four providers
# OpenAI
curl -N https://api.openai.com/v1/responses -H "Authorization: Bearer $OPENAI_API_KEY" -H "Content-Type: application/json" \
-d '{"model":"gpt-5.4-nano","input":"Reply with OK.","max_output_tokens":32,"stream":true}'
# Anthropic
curl -N https://api.anthropic.com/v1/messages -H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01" -H "content-type: application/json" \
-d '{"model":"claude-haiku-4-5-20251001","max_tokens":32,"messages":[{"role":"user","content":"Reply with OK."}],"stream":true}'
# xAI (OpenAI event names; reasoning summary events come first)
curl -N https://api.x.ai/v1/responses -H "Authorization: Bearer $XAI_API_KEY" -H "Content-Type: application/json" \
-d '{"model":"grok-4.3","input":"Reply with OK.","reasoning":{"effort":"low"},"max_output_tokens":64,"stream":true}'
# Gemini (SSE framing requires ?alt=sse; without it you get a JSON array)
curl -N "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash-lite:streamGenerateContent?alt=sse" \
-H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" \
-d '{"contents":[{"role":"user","parts":[{"text":"Reply with OK."}]}],"generationConfig":{"maxOutputTokens":64,"thinkingConfig":{"thinkingLevel":"minimal"}}}'Client algorithm that works for all four: dispatch on the event name where one exists (OpenAI/xAI event:, Anthropic event:, Interactions event:), otherwise on chunk shape (Gemini candidates[0].content.parts[]); accumulate text per (OpenAI/xAI output_index,content_index) / (Anthropic block index) / (Gemini part position); accumulate tool-argument fragments and parse at …arguments.done / content_block_stop / the complete functionCall part; take the final object from the terminal event (response.completed.response, interaction.completed) or rebuild it (message_start + deltas; last Gemini chunk + concatenated parts). SDK helpers: client.responses.stream().get_final_response() (OpenAI, and xAI via the OpenAI SDK), client.messages.stream().get_final_message() (Anthropic), client.chat.create(...).stream() (xai-sdk, gRPC), client.models.generate_content_stream() / client.interactions.create(stream=True) (google-genai).
6. Beyond SSE — WebSocket message families
| Capability | OpenAI | Anthropic | xAI | Gemini |
|---|---|---|---|---|
| WebSocket request mode (text) | wss://api.openai.com/v1/responses — response.create (+ stream_id, generate:false), response.steer, response.inject (beta); 32 lanes / 60 min; background unsupported |
— | wss://api.x.ai/v1/responses (WebSocket Responses mode, May 2026; DOCUMENTED) |
— (Live API is the only WebSocket for text+audio) |
| Resume a stream | GET /v1/responses/{id}?stream=true&starting_after=N (background) |
— | — (deferred Chat is unary polling) | Interactions GET /v1beta/interactions/{id}?stream=true&last_event_id=<id> |
| Agent session streams | Agents API SSE agent.session.* (31) |
Managed Agents SSE (37 events incl. event_start/event_delta) |
the Responses stream itself | Interactions SSE (15 events) |
| Voice streams | Realtime WebSocket/WebRTC/SIP (45 server + 11 client events), Live (20 + 11), translation (6 + 3) | — | Realtime wss://api.x.ai/v1/realtime (39 server + 9 client events, OpenAI vocabulary: session.update, input_audio_buffer.append/commit, conversation.item.create, response.create → response.output_audio.delta, response.output_audio_transcript.delta, response.done, error, LIVE_DISCOVERED ping); TTS wss://api.x.ai/v1/tts (text.delta/done → audio.delta/done/error); STT wss://api.x.ai/v1/stt (binary frames, finalize, audio.done → transcript.created/partial/done); SIP webhook realtime.call.incoming |
Live API BidiGenerateContent (21 server + 11 client message types: setup → setupComplete; clientContent, realtimeInput.{audio,video,text,activityStart,activityEnd,audioStreamEnd}, toolResponse → serverContent.{modelTurn,generationComplete,turnComplete,interrupted,inputTranscription,outputTranscription,interactionStatus,groundingMetadata,urlContextMetadata}, toolCall, toolCallCancellation, usageMetadata, goAway, sessionResumptionUpdate); Lyria RealTime BidiGenerateMusic (setup, clientContent.weightedPrompts, musicGenerationConfig, playbackControl → setupComplete, serverContent.audioChunks[], filteredPrompt) |
| Media streams | POST /v1/audio/speech SSE (speech.audio.delta/done), transcriptions (transcript.text.delta/segment/done), images (image_generation.partial_image/completed) |
— | none for images/videos (unary + 202 polling for videos) | streaming TTS on gemini-3.1-flash-tts-preview (audio inlineData chunks); Veo/Batch are long-running Operations, not streams |
| Batch outputs | JSONL file | JSONL results file | GET /v1/batches/{id}/results (paginated JSON) |
inlinedResponses[] or responsesFile (no streaming) |
Realtime / Live message-family mapping
| Purpose | OpenAI Realtime | xAI realtime | Gemini Live |
|---|---|---|---|
| open / configure | session.update {session:{type, model, instructions, audio, tools, turn_detection}} → session.created, session.updated |
session.update {instructions, voice, reasoning.effort, turn_detection, audio.input/output.format, tools[]} → session.created, session.updated, conversation.created |
setup {model, generationConfig, systemInstruction, tools, realtimeInputConfig, sessionResumption, contextWindowCompression, inputAudioTranscription, outputAudioTranscription} → setupComplete |
| send audio | input_audio_buffer.append {audio} / .commit / .clear; VAD input_audio_buffer.speech_started/stopped/committed |
same names (input_audio_buffer.append/commit/clear; speech_started/stopped/committed, timeout_triggered, dtmf_event_received) |
realtimeInput {audio:{data, mimeType: "audio/pcm;rate=16000"}} (+ activityStart/activityEnd when automatic VAD is disabled, audioStreamEnd) |
| send text / items | conversation.item.create {item} → conversation.item.created; .truncate, .delete |
conversation.item.create → conversation.item.added; .truncate, .delete (each text item $0.004) |
clientContent {turns[], turnComplete} or realtimeInput {text} |
| request a turn | response.create / response.cancel → response.created … response.done |
response.create / response.cancel → response.created … response.done (top-level usage {billable_audio_seconds}) |
implicit on turnComplete / VAD; no explicit create; serverContent.interrupted on barge-in |
| audio out | response.output_audio.delta/done, response.output_audio_transcript.delta/done |
response.output_audio.delta/done, response.output_audio_transcript.delta/done (response.audio.delta documented, not emitted) |
serverContent.modelTurn.parts[].inlineData (24 kHz PCM), serverContent.outputTranscription, generationComplete, turnComplete |
| tools | response.function_call_arguments.delta/done, mcp_list_tools.*, response.mcp_call.* |
same names (function + mcp_list_tools.*, response.mcp_call_arguments.*, response.mcp_call.*) |
toolCall {functionCalls[]} → toolResponse {functionResponses[] {id, response, scheduling, willContinue}}; toolCallCancellation |
| session lifetime | 60 min; session.expired-style errors |
120 min; error {type: timeout|max_duration} |
≈10-min connections (goAway {timeLeft}), 15-min audio sessions (unlimited with compression); sessionResumptionUpdate {newHandle} (2 h) |
| auth for browsers | POST /v1/realtime/client_secrets → ek_… |
POST /v1/realtime/client_secrets → xai-realtime… (subprotocol xai-client-secret.<token>) |
POST /v1beta/auth_tokens → auth_tokens/… on …BidiGenerateContentConstrained?access_token= |
7. Streaming constraints and gotchas
| Topic | OpenAI | Anthropic | xAI | Gemini |
|---|---|---|---|---|
| long generations | background mode recommended for long runs | SDKs require streaming above ~10 min of expected output; thinking budgets > 32k → Batches | deferred Chat Completions (deferred: true) for long runs; Responses background → 400 |
Interactions background: true (Deep Research); X-Server-Timeout hint; Flex requests may take 1–15 min |
| tool-argument JSON | fragments per delta; parse on .done |
fragments may be "" first; eager_input_streaming unvalidated |
the whole JSON arrives in one delta / one chunk | the functionCall part arrives complete; keep the thoughtSignature |
| structured outputs | text deltas of the JSON | text_delta fragments of one text block |
text deltas of the JSON | partial JSON strings across chunks (concatenate) |
| thinking display | summaries only | display: omitted → one empty thinking_delta + signature_delta |
Chat streams the full reasoning_content; Responses streams summaries |
includeThoughts summaries not guaranteed; an empty-text final part carries the signature |
| rate-limit headers | on the 200 | on the 200 | on the 200 (undocumented x-ratelimit-*) |
none |
| SSE resume | yes (background) | no | no | Interactions only (last_event_id) |
| framing traps | obfuscation padding |
ping events |
created: 0 in chunks; /v1/messages reuses block index 0 |
default (non-SSE) response is a JSON array — needs an incremental array parser; generateContent?alt=sse also emits one event |
Related: features · tool-execution · realtime-and-media · agents-platforms · docs/openai/streaming-events.md · docs/anthropic/streaming.md · docs/xai/streaming.md · docs/xai/voice.md · docs/gemini/streaming.md · docs/gemini/live-events.md.