# Streaming — four SSE dialects and five WebSocket families **Status:** event names and statuses from `generated/streaming-events.json` (506 records: OpenAI 282, Anthropic 68, xAI 94, Gemini 62); observed sequences quoted from docs/openai/streaming-events.md, docs/anthropic/streaming.md, docs/xai/streaming.md, docs/gemini/streaming.md, docs/gemini/interactions-api.md, docs/gemini/live-events.md (`LIVE_VERIFIED` 2026-09-18/19 on `gpt-5.4-nano`, `claude-haiku-4-5-20251001`, `grok-4.3`, `gemini-3.5-flash-lite` / `gemini-3.8-flash`). **Sources:** https://developers.openai.com/api/docs/guides/streaming-responses · https://developers.openai.com/api/reference/resources/responses/streaming-events · https://platform.claude.com/docs/en/build-with-claude/streaming · https://docs.x.ai/developers/model-capabilities/text/streaming · https://docs.x.ai/developers/model-capabilities/audio/speech-to-speech · https://ai.google.dev/api/generate-content#method:-models.streamgeneratecontent · https://ai.google.dev/gemini-api/docs/live · https://ai.google.dev/gemini-api/docs/interactions **Last verified:** 2026-09-18 ## 1. Wire format | | OpenAI Responses | OpenAI Chat Completions | Anthropic Messages | xAI Responses | xAI Chat Completions | xAI `/v1/messages` (deprecated) | Gemini `streamGenerateContent` | Gemini Interactions | |---|---|---|---|---|---|---|---|---| | Framing | `event: ` + `data:`; `data.type == event` | `data: ` only | `event: ` + `data:`; `data.type == event` | `event: ` + `data:` (same names as OpenAI) | `data:` only (`chat.completion.chunk`) | `event:` + `data:` (Anthropic names) | **`?alt=sse`**: `data: ` lines, **no `event:` names**; **default**: one pretty-printed **JSON array** streamed incrementally (`[{…}` `,` `{…}` `]`) | `event: ` + `data:` with `event_type`, `event_id` | | Ordering guard | `sequence_number` | none | none (fixed skeleton) | `sequence_number` | none | none | none (chunk order; `usageMetadata` on every chunk) | `event_id` (resume with `last_event_id`) | | Terminal | `response.completed\|failed\|incomplete` \| `error` — no `[DONE]` | `data: [DONE]` after an optional usage chunk | `message_stop` — no `[DONE]` | `response.completed` (`incomplete`/`failed`/`error` UNVERIFIED) — no `[DONE]` | `data: [DONE]` (usage chunk with `choices: []` when `stream_options.include_usage`) | `message_stop` | last chunk carries `finishReason` (and on Gemini 3 an **empty-text part with `thoughtSignature`** — keep it); stream closes | `interaction.completed` then `done` (`data: [DONE]`) | | Padding | `obfuscation` on delta events | `obfuscation` | none | none observed | none | none | none | none | | Full object in stream | lifecycle events embed the `Response` | none | `message_start` skeleton; cumulative `usage` in `message_delta` | `response.created` snapshot (`reasoning: {effort: null}`, `created_at: 0`) and full response in `response.completed` | none (`created: 0` in every chunk) | `message_start` skeleton | every chunk is a partial `GenerateContentResponse` with `modelVersion`, `responseId`, `usageMetadata` (last authoritative) | `step.start`/`step.stop` carry step objects; `interaction.completed` the final interaction | | Keep-alive | none | none | `ping` | none | none | none (no `ping`) | none | none | | Errors after 200 | `error` event (terminal) | error chunk | `error` event (`overloaded_error`…), no `request_id` | `error` (UNVERIFIED) | — | — | undocumented (treat a close without `finishReason` as an error); prompt block → single chunk with `promptFeedback.blockReason`, no candidates | `event: error` `{error: {code, message}}` | | Content type | `text/event-stream` | `text/event-stream` | `text/event-stream` | `text/event-stream` | `text/event-stream` | `text/event-stream` | `text/event-stream` (alt=sse) or `application/json` | `text/event-stream` | ## 2. Skeletons (observed live) **OpenAI Responses — "Reply with OK."** ``` response.created(0) → response.in_progress(1) → response.output_item.added(2, message) → response.content_part.added(3) → response.output_text.delta(4 "OK") → response.output_text.done(5) → response.content_part.done(6) → response.output_item.done(7) → response.completed(8) ``` **Anthropic Messages** ``` message_start → content_block_start(0, text) → ping → content_block_delta(text_delta "OK") → content_block_delta(".") → content_block_stop(0) → message_delta(stop_reason end_turn, usage) → message_stop ``` **xAI Responses (grok-4.3) — reasoning item first** ``` response.created → response.in_progress → response.output_item.added(reasoning) → response.reasoning_summary_part.added → response.reasoning_summary_text.delta… → response.reasoning_summary_text.done → response.reasoning_summary_part.done → response.output_item.done → response.output_item.added(message) → response.content_part.added → response.output_text.delta("OK") → response.output_text.done → response.content_part.done → response.output_item.done → response.completed ``` **xAI Chat Completions** ``` data: {chunk delta.reasoning_content…} → data: {chunk delta.content "OK"} → data: {chunk finish_reason: stop} → [data: {choices: [], usage}] → data: [DONE] ``` **xAI `/v1/messages`** — `message_start → content_block_start(thinking) → content_block_delta(thinking_delta)… → content_block_stop → content_block_start(text, index 0 again) → content_block_delta(text_delta) → content_block_stop → message_delta → message_stop` (no `ping`; index reuse is a quirk). **Gemini `streamGenerateContent?alt=sse` (gemini-3.8-flash)** ``` data: {candidates:[{content:{parts:[{text:"OK", thought?}]}}], usageMetadata, modelVersion, responseId} data: {candidates:[{content:{parts:[{text:"", thoughtSignature:"…"}]}, finishReason:"STOP"}], usageMetadata:{…thoughtsTokenCount}} ``` With `includeThoughts: true`, `{text, thought: true}` parts stream first (not guaranteed on trivial prompts). JSON-mode deltas are partial JSON strings that concatenate. **Gemini Interactions (`stream: true`)** ``` interaction.created → interaction.status_update → step.start(thought) → step.delta{thought_signature} → step.stop → step.start(model_output) → step.delta{text "OK"} → step.stop → interaction.completed → done [DONE] ``` ## 3. Event-name mapping — the four text streams | Concern | OpenAI Responses | Anthropic Messages | xAI Responses / Chat | Gemini `streamGenerateContent` / Interactions | |---|---|---|---|---| | lifecycle | `response.created`, `response.queued`, `response.in_progress`, `response.completed\|failed\|incomplete` | `message_start`, `message_delta` (stop_reason, usage), `message_stop` | same as OpenAI (`response.queued` absent — no background mode); Chat: first chunk / `finish_reason` / `[DONE]` | no lifecycle events — first chunk / chunk with `finishReason`; Interactions `interaction.created`, `interaction.status_update` (legacy but still emitted), `interaction.in_progress`, `interaction.requires_action`, `interaction.completed`, `done` | | new output unit | `response.output_item.added` / `.done` | `content_block_start` / `content_block_stop` (index) | same as OpenAI; Chat: implicit | new `parts[]` entries in a chunk; Interactions `step.start` / `step.stop` | | text | `response.content_part.added/done`, `response.output_text.delta/done`, `response.output_text.annotation.added`, `response.refusal.delta/done` | `content_block_delta {type: text_delta}`; `citations_delta` | same as OpenAI (`response.output_text.delta`, `.done`); Chat `delta.content` | `parts[].text` fragments; Interactions `step.delta {type: text}`, `text_annotation_delta` | | function / tool args | `response.function_call_arguments.delta/done`, `response.custom_tool_call_input.delta/done` | `content_block_delta {type: input_json_delta, partial_json}` (buffered; `eager_input_streaming`) | `response.function_call_arguments.delta` (**one delta with the full JSON**) + `.done`; Chat `delta.tool_calls[]` whole call in one chunk → `finish_reason: tool_calls` | `parts[].functionCall` arrives **complete** in one chunk (+ `thoughtSignature`); Interactions `step.delta {type: arguments_delta}` (partial JSON) | | reasoning | `response.reasoning_summary_part.added/done`, `response.reasoning_summary_text.delta/done`, `response.reasoning_text.delta/done` | `content_block_start(thinking)`, `thinking_delta`, `signature_delta`, `redacted_thinking` | same summary events as OpenAI (`reasoning_text.delta` UNVERIFIED); Chat `delta.reasoning_content` (full text) | `parts[] {text, thought: true}` chunks, final `thoughtSignature`; Interactions `step.start(thought)`, `step.delta {type: thought_summary \| thought_signature}` | | web search | `response.web_search_call.in_progress/searching/completed` | `server_tool_use` + `input_json_delta`, `web_search_tool_result` block | `response.output_item.added/done(web_search_call)` (`.searching` events documented, not observed) | `groundingMetadata` on candidates (chunks carry only **new** `groundingChunks` — accumulate); Interactions `google_search_call` / `google_search_result` steps | | code execution | `response.code_interpreter_call.in_progress/interpreting/completed`, `response.code_interpreter_call_code.delta/done` | `server_tool_use` + `bash_code_execution_tool_result` | **identical** event names to OpenAI (all LIVE_VERIFIED) | `executableCode` / `codeExecutionResult` / `inlineData` parts as produced; Interactions `code_execution_call` / `code_execution_result` | | file search / RAG | `response.file_search_call.in_progress/searching/completed` | — | `response.file_search_call.*` (UNVERIFIED) | `groundingMetadata.groundingChunks[].retrievedContext`; Interactions `file_search_call/result` | | image generation | `response.image_generation_call.in_progress/generating/partial_image/completed` | — | `image_generation_call` item events (not exercised) | `inlineData {mimeType: image/png}` parts (image models); Interactions `step.delta {type: image}` | | MCP | `response.mcp_list_tools.*`, `response.mcp_call.*`, `response.mcp_call_arguments.delta/done` | `mcp_tool_use` / `mcp_tool_result` blocks | `response.mcp_call.*` (UNVERIFIED) | Interactions mcp call/result steps (UNVERIFIED) | | shell / patch | `response.shell_call_command.*`, `response.shell_call_output_content.*`, `response.apply_patch_call_operation_diff.*` (LIVE_DISCOVERED) | `tool_use` + `input_json_delta` | `shell_call` items (not exercised) | — | | compaction | `response.compaction.compacting` | `content_block_*(compaction)` (beta) | — (compact endpoint is unary) | — | | audio | `response.audio.delta/done`, `response.audio.transcript.delta/done` | — | — (voice WebSocket) | `inlineData` audio parts (TTS models; streaming TTS on 3.1); Interactions `step.delta {type: audio}` | | errors | `error` | `error` | `error` (UNVERIFIED) | pre-stream JSON error; `promptFeedback.blockReason`; Interactions `error` | | keep-alive | — | `ping` | — | — | Counts: OpenAI Responses SSE 63 + WebSocket 69/3, Chat 2; Anthropic Messages 13 core + 10 advanced; xAI Responses 24, Chat 3, Messages 6; Gemini generate-content 8 (framing + chunk kinds), Interactions 15. Full lists: [FAQ Q4](../faq.md#q4). ## 4. Usage and stop information | | OpenAI | Anthropic | xAI | Gemini | |---|---|---|---|---| | where usage arrives | `response.completed.response.usage` (+ `prompt_cache_diagnostics` 5.6+) | `message_start.message.usage` then cumulative `message_delta.usage` | `response.completed.response.usage` (incl. `cost_in_usd_ticks`, `server_side_tool_usage_details`); Chat: final chunk with `choices: []` when `stream_options.include_usage: true` | `usageMetadata` on **every** chunk (last one authoritative: `thoughtsTokenCount`, `toolUsePromptTokenCount`, `cachedContentTokenCount`); Interactions `interaction.completed.usage` | | stop reason | `response.status` + `incomplete_details.reason` | `message_delta.delta.stop_reason` / `stop_sequence` / `stop_details` | `response.status` + `incomplete_details.reason` (max_output_tokens \| max_prompt_tokens \| max_time_limit); Chat `finish_reason` | `candidates[].finishReason` on the last chunk (21 values); Interactions `status` | | tool call boundary | `response.output_item.done` for the `function_call` item | `content_block_stop` of the `tool_use` block; `stop_reason: tool_use` | `response.output_item.done(function_call)`; Chat `finish_reason: tool_calls` | chunk containing the `functionCall` part (usually the last, `finishReason: STOP`); Interactions `interaction.requires_action` | ## 5. The same streamed call on all four providers ```bash # OpenAI curl -N https://api.openai.com/v1/responses -H "Authorization: Bearer $OPENAI_API_KEY" -H "Content-Type: application/json" \ -d '{"model":"gpt-5.4-nano","input":"Reply with OK.","max_output_tokens":32,"stream":true}' # Anthropic curl -N https://api.anthropic.com/v1/messages -H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01" -H "content-type: application/json" \ -d '{"model":"claude-haiku-4-5-20251001","max_tokens":32,"messages":[{"role":"user","content":"Reply with OK."}],"stream":true}' # xAI (OpenAI event names; reasoning summary events come first) curl -N https://api.x.ai/v1/responses -H "Authorization: Bearer $XAI_API_KEY" -H "Content-Type: application/json" \ -d '{"model":"grok-4.3","input":"Reply with OK.","reasoning":{"effort":"low"},"max_output_tokens":64,"stream":true}' # Gemini (SSE framing requires ?alt=sse; without it you get a JSON array) curl -N "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash-lite:streamGenerateContent?alt=sse" \ -H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" \ -d '{"contents":[{"role":"user","parts":[{"text":"Reply with OK."}]}],"generationConfig":{"maxOutputTokens":64,"thinkingConfig":{"thinkingLevel":"minimal"}}}' ``` Client algorithm that works for all four: dispatch on the event name where one exists (OpenAI/xAI `event:`, Anthropic `event:`, Interactions `event:`), otherwise on chunk shape (Gemini `candidates[0].content.parts[]`); accumulate text per (OpenAI/xAI `output_index`,`content_index`) / (Anthropic block `index`) / (Gemini part position); accumulate tool-argument fragments and parse at `…arguments.done` / `content_block_stop` / the complete `functionCall` part; take the final object from the terminal event (`response.completed.response`, `interaction.completed`) or rebuild it (`message_start` + deltas; last Gemini chunk + concatenated parts). SDK helpers: `client.responses.stream().get_final_response()` (OpenAI, and xAI via the OpenAI SDK), `client.messages.stream().get_final_message()` (Anthropic), `client.chat.create(...).stream()` (xai-sdk, gRPC), `client.models.generate_content_stream()` / `client.interactions.create(stream=True)` (google-genai). ## 6. Beyond SSE — WebSocket message families | Capability | OpenAI | Anthropic | xAI | Gemini | |---|---|---|---|---| | WebSocket request mode (text) | `wss://api.openai.com/v1/responses` — `response.create` (+ `stream_id`, `generate:false`), `response.steer`, `response.inject` (beta); 32 lanes / 60 min; `background` unsupported | — | `wss://api.x.ai/v1/responses` (WebSocket Responses mode, May 2026; DOCUMENTED) | — (Live API is the only WebSocket for text+audio) | | Resume a stream | `GET /v1/responses/{id}?stream=true&starting_after=N` (background) | — | — (deferred Chat is unary polling) | Interactions `GET /v1beta/interactions/{id}?stream=true&last_event_id=` | | Agent session streams | Agents API SSE `agent.session.*` (31) | Managed Agents SSE (37 events incl. `event_start`/`event_delta`) | the Responses stream itself | Interactions SSE (15 events) | | Voice streams | Realtime WebSocket/WebRTC/SIP (45 server + 11 client events), Live (20 + 11), translation (6 + 3) | — | Realtime `wss://api.x.ai/v1/realtime` (**39 server + 9 client events, OpenAI vocabulary**: `session.update`, `input_audio_buffer.append/commit`, `conversation.item.create`, `response.create` → `response.output_audio.delta`, `response.output_audio_transcript.delta`, `response.done`, `error`, LIVE_DISCOVERED `ping`); TTS `wss://api.x.ai/v1/tts` (`text.delta/done` → `audio.delta/done/error`); STT `wss://api.x.ai/v1/stt` (binary frames, `finalize`, `audio.done` → `transcript.created/partial/done`); SIP webhook `realtime.call.incoming` | Live API `BidiGenerateContent` (**21 server + 11 client message types**: `setup` → `setupComplete`; `clientContent`, `realtimeInput.{audio,video,text,activityStart,activityEnd,audioStreamEnd}`, `toolResponse` → `serverContent.{modelTurn,generationComplete,turnComplete,interrupted,inputTranscription,outputTranscription,interactionStatus,groundingMetadata,urlContextMetadata}`, `toolCall`, `toolCallCancellation`, `usageMetadata`, `goAway`, `sessionResumptionUpdate`); Lyria RealTime `BidiGenerateMusic` (`setup`, `clientContent.weightedPrompts`, `musicGenerationConfig`, `playbackControl` → `setupComplete`, `serverContent.audioChunks[]`, `filteredPrompt`) | | Media streams | `POST /v1/audio/speech` SSE (`speech.audio.delta/done`), transcriptions (`transcript.text.delta/segment/done`), images (`image_generation.partial_image/completed`) | — | none for images/videos (unary + 202 polling for videos) | streaming TTS on `gemini-3.1-flash-tts-preview` (audio `inlineData` chunks); Veo/Batch are long-running `Operation`s, not streams | | Batch outputs | JSONL file | JSONL results file | `GET /v1/batches/{id}/results` (paginated JSON) | `inlinedResponses[]` or `responsesFile` (no streaming) | ### Realtime / Live message-family mapping | Purpose | OpenAI Realtime | xAI realtime | Gemini Live | |---|---|---|---| | open / configure | `session.update {session:{type, model, instructions, audio, tools, turn_detection}}` → `session.created`, `session.updated` | `session.update {instructions, voice, reasoning.effort, turn_detection, audio.input/output.format, tools[]}` → `session.created`, `session.updated`, `conversation.created` | `setup {model, generationConfig, systemInstruction, tools, realtimeInputConfig, sessionResumption, contextWindowCompression, inputAudioTranscription, outputAudioTranscription}` → `setupComplete` | | send audio | `input_audio_buffer.append {audio}` / `.commit` / `.clear`; VAD `input_audio_buffer.speech_started/stopped/committed` | same names (`input_audio_buffer.append/commit/clear`; `speech_started/stopped/committed`, `timeout_triggered`, `dtmf_event_received`) | `realtimeInput {audio:{data, mimeType: "audio/pcm;rate=16000"}}` (+ `activityStart`/`activityEnd` when automatic VAD is disabled, `audioStreamEnd`) | | send text / items | `conversation.item.create {item}` → `conversation.item.created`; `.truncate`, `.delete` | `conversation.item.create` → `conversation.item.added`; `.truncate`, `.delete` (each text item $0.004) | `clientContent {turns[], turnComplete}` or `realtimeInput {text}` | | request a turn | `response.create` / `response.cancel` → `response.created` … `response.done` | `response.create` / `response.cancel` → `response.created` … `response.done` (top-level `usage {billable_audio_seconds}`) | implicit on `turnComplete` / VAD; no explicit create; `serverContent.interrupted` on barge-in | | audio out | `response.output_audio.delta/done`, `response.output_audio_transcript.delta/done` | `response.output_audio.delta/done`, `response.output_audio_transcript.delta/done` (`response.audio.delta` documented, not emitted) | `serverContent.modelTurn.parts[].inlineData` (24 kHz PCM), `serverContent.outputTranscription`, `generationComplete`, `turnComplete` | | tools | `response.function_call_arguments.delta/done`, `mcp_list_tools.*`, `response.mcp_call.*` | same names (function + `mcp_list_tools.*`, `response.mcp_call_arguments.*`, `response.mcp_call.*`) | `toolCall {functionCalls[]}` → `toolResponse {functionResponses[] {id, response, scheduling, willContinue}}`; `toolCallCancellation` | | session lifetime | 60 min; `session.expired`-style errors | 120 min; `error {type: timeout\|max_duration}` | ≈10-min connections (`goAway {timeLeft}`), 15-min audio sessions (unlimited with compression); `sessionResumptionUpdate {newHandle}` (2 h) | | auth for browsers | `POST /v1/realtime/client_secrets` → `ek_…` | `POST /v1/realtime/client_secrets` → `xai-realtime…` (subprotocol `xai-client-secret.`) | `POST /v1beta/auth_tokens` → `auth_tokens/…` on `…BidiGenerateContentConstrained?access_token=` | ## 7. Streaming constraints and gotchas | Topic | OpenAI | Anthropic | xAI | Gemini | |---|---|---|---|---| | long generations | background mode recommended for long runs | SDKs require streaming above ~10 min of expected output; thinking budgets > 32k → Batches | deferred Chat Completions (`deferred: true`) for long runs; Responses `background` → 400 | Interactions `background: true` (Deep Research); `X-Server-Timeout` hint; Flex requests may take 1–15 min | | tool-argument JSON | fragments per delta; parse on `.done` | fragments may be `""` first; `eager_input_streaming` unvalidated | the whole JSON arrives in one delta / one chunk | the `functionCall` part arrives complete; keep the `thoughtSignature` | | structured outputs | text deltas of the JSON | `text_delta` fragments of one text block | text deltas of the JSON | partial JSON strings across chunks (concatenate) | | thinking display | summaries only | `display: omitted` → one empty `thinking_delta` + `signature_delta` | Chat streams the full `reasoning_content`; Responses streams summaries | `includeThoughts` summaries not guaranteed; an empty-text final part carries the signature | | rate-limit headers | on the 200 | on the 200 | on the 200 (undocumented `x-ratelimit-*`) | none | | SSE resume | yes (background) | no | no | Interactions only (`last_event_id`) | | framing traps | `obfuscation` padding | `ping` events | `created: 0` in chunks; `/v1/messages` reuses block index 0 | default (non-SSE) response is a JSON array — needs an incremental array parser; `generateContent?alt=sse` also emits one event | Related: [features](features.md) · [tool-execution](tool-execution.md) · [realtime-and-media](realtime-and-media.md) · [agents-platforms](agents-platforms.md) · docs/openai/streaming-events.md · docs/anthropic/streaming.md · docs/xai/streaming.md · docs/xai/voice.md · docs/gemini/streaming.md · docs/gemini/live-events.md.