# OpenAI streaming events — Responses SSE, Responses WebSocket, Chat Completions chunks **Status:** `DOCUMENTED` (59 SSE event types from the OpenAPI `ResponseStreamEvent` union, 2 GA + 1 beta WebSocket client events, 62 GA + 2 beta server events) · `LIVE_VERIFIED` for the 11 lifecycle/text events observed on 2026-09-18 and for Chat Completions chunks. Machine-readable: `generated/fragments/streaming-events/openai-responses.json`, `openai-responses-websocket.json`, `openai-chat-completions.json`. **Sources** - https://developers.openai.com/api/reference/resources/responses/streaming-events · https://developers.openai.com/api/reference/resources/responses/websocket-events · https://developers.openai.com/api/reference/resources/beta/subresources/responses/streaming-events - https://developers.openai.com/api/reference/resources/chat/subresources/completions/streaming-events - https://developers.openai.com/api/docs/guides/streaming-responses · https://developers.openai.com/api/docs/guides/websocket-mode · https://developers.openai.com/api/docs/guides/background#streaming-a-background-response - OpenAPI `ResponseStreamEvent`, `ResponsesClientEvent`, `ResponsesServerEvent`, `CreateChatCompletionStreamResponse` **Last verified:** 2026-09-18 ## 1. Responses SSE (`stream: true`) Transport: `Content-Type: text/event-stream`. Each event is ``` event: response.output_text.delta data: {"type":"response.output_text.delta","item_id":"msg_…","output_index":0,"content_index":0,"delta":"OK","logprobs":[],"obfuscation":"tcI6fDMDcuuDp0","sequence_number":4} ``` Invariants (live): `event:` name **equals** `data.type`; `sequence_number` starts at 0 and increments by 1 per event; delta events carry an `obfuscation` pad (disable with `stream_options.include_obfuscation:false`); lifecycle events embed the full `Response` object (`response.created` has `status:"in_progress"`, `response.completed` carries final `usage`). No `data: [DONE]` sentinel — the stream ends after the terminal event (`response.completed` | `response.failed` | `response.incomplete` | `error`). ### Observed sequences (gpt-5.4-nano, "Reply with OK.", max_output_tokens 32) | Case | Ordered event types | |---|---| | plain text (b) | `response.created`(0) → `response.in_progress`(1) → `response.output_item.added`(2, message in_progress, `phase:"final_answer"`) → `response.content_part.added`(3, empty output_text) → `response.output_text.delta`(4 "OK") → `response.output_text.delta`(5 ".") → `response.output_text.done`(6) → `response.content_part.done`(7) → `response.output_item.done`(8) → `response.completed`(9) | | background + stream (j5) | `response.created`(0) → **`response.queued`**(1) → `response.in_progress`(2) → … same … → `response.completed`(9) | | resume `GET …?stream=true&starting_after=1` (j6) | replays `response.in_progress`(2) … `response.completed`(9) | | chat-completions style (l2) | see §3 | ### Complete SSE event catalogue (59) | Category | Event types | Key payload fields | |---|---|---| | Response lifecycle | `response.created`, `response.queued`, `response.in_progress`, `response.completed`, `response.failed`, `response.incomplete` | `response` (full Response), `sequence_number` | | Item lifecycle | `response.output_item.added`, `response.output_item.done` | `item` (OutputItem), `output_index` | | Content parts | `response.content_part.added`, `response.content_part.done` | `item_id`, `output_index`, `content_index`, `part` (output_text \| refusal \| reasoning_text) | | Text | `response.output_text.delta` (`delta`, `logprobs[]`), `response.output_text.done` (`text`, `logprobs[]`), `response.output_text.annotation.added` (`annotation`, `annotation_index`) | + `item_id`, `output_index`, `content_index` | | Refusal | `response.refusal.delta` (`delta`), `response.refusal.done` (`refusal`) | | | Reasoning | `response.reasoning_summary_part.added/.done` (`part{type:summary_text,text}`, `summary_index`), `response.reasoning_summary_text.delta/.done` (`delta`/`text`, `summary_index`), `response.reasoning_text.delta/.done` (`delta`/`text`, `content_index`) | | | Function tools | `response.function_call_arguments.delta` (`delta`), `response.function_call_arguments.done` (`arguments`, `name`) | `item_id`, `output_index` | | Custom tools | `response.custom_tool_call_input.delta` (`delta`), `response.custom_tool_call_input.done` (`input`) | | | File search | `response.file_search_call.in_progress/.searching/.completed` | `item_id`, `output_index` | | Web search | `response.web_search_call.in_progress/.searching/.completed` | | | Code interpreter | `response.code_interpreter_call.in_progress/.interpreting/.completed`, `response.code_interpreter_call_code.delta/.done` (`delta`/`code`) | | | Image generation | `response.image_generation_call.in_progress/.generating/.completed`, `response.image_generation_call.partial_image` (`partial_image_b64`, `partial_image_index`, `size`, `quality`, `background`, `output_format`) | | | MCP | `response.mcp_call.in_progress/.completed/.failed`, `response.mcp_call_arguments.delta/.done` (`delta`/`arguments`), `response.mcp_list_tools.in_progress/.completed/.failed` | | | Shell | `response.shell_call_command.added/.delta/.done`, `response.shell_call_output_content.delta/.done` | | | Compaction | `response.compaction.compacting` | emitted when server-side compaction runs mid-response | | Audio (chat-style audio output) | `response.audio.delta/.done`, `response.audio.transcript.delta/.done` | `delta` base64 / transcript text | | Error | `error` (`code`, `message`, `param`, `sequence_number`) | terminal | Not in the SSE union but present as output items / other tools' events: apply_patch has no dedicated streaming events in this spec snapshot (its calls appear via `output_item.*`). The beta SSE union (`BetaResponseStreamEvent`) is identical to the GA list. Handling guidance: branch on `type`; accumulate `output_text.delta` per (`output_index`,`content_index`); treat `output_item.done` as the canonical item (its `encrypted_content` is complete, `output_item.added` may be partial); read `usage` and `prompt_cache_diagnostics` from `response.completed.response`; persist the last `sequence_number` for background resume. ## 2. Responses WebSocket events (`wss://api.openai.com/v1/responses`) **Client → server** | Event | Fields | Notes | |---|---|---| | `response.create` | all `POST /v1/responses` fields + `stream_id` (1–256, `[A-Za-z0-9_.-]`); `generate:false` warm-up | `stream` implicit; `background` unsupported | | `response.steer` | `previous_response_id`, `input` (messages / `function_call_output`) | mid-turn steering; the target stops at a safe boundary (`incomplete_details.reason:"steered"`) and a successor response is created | | `response.inject` (beta) | `response_id`, `input[]` | inject client tool outputs into an active response | **Server → client**: every SSE event above (with an optional `stream_id` echo for named lanes) plus: | Event | Fields | Meaning | |---|---|---| | `response.steer.accepted` | `steer{…}`, `stream_id` | input validated and queued (commit point = successor `response.created`) | | `response.steer.pending` | `steer`, `reason` (e.g. `waiting_for_required_input`), `required_input[]` | queued input still owned by server after target completed — do not resend | | `response.steer.failed` | `steer`, `error` | returns the uncommitted input | | `response.inject.created` / `response.inject.failed` (beta) | | | | `error` | `status`, `stream_id?`, `error{type, code, message, param}` | codes: `previous_response_not_found`, `invalid_stream_id`, `websocket_stream_limit_reached` (32 lanes), `websocket_connection_limit_reached` (60 min) | Concurrency model: same `stream_id` → FIFO; different lanes → concurrent (≤16 in flight). `previous_response_id` = lineage (fork a lane from another lane's response; wait for the fork's `response.in_progress` before advancing the parent lane with `store:false`). Connection cache is lost on reconnect. ## 3. Chat Completions chunks (`stream: true`) Data-only SSE (no `event:` line): `data: {chat.completion.chunk}` … `data: [DONE]`. Observed (gpt-4.1-nano, `stream_options.include_usage:true`): ``` data: {"id":"chatcmpl-…","object":"chat.completion.chunk","created":…,"model":"gpt-4.1-nano-2025-04-14","service_tier":"default","system_fingerprint":"fp_…","choices":[{"index":0,"delta":{"role":"assistant","content":"","refusal":null},"logprobs":null,"finish_reason":null}],"usage":null,"obfuscation":"GO7oM1GH"} data: {…,"choices":[{"index":0,"delta":{"content":"OK"},"logprobs":null,"finish_reason":null}],"usage":null,"obfuscation":"4EFX0fab"} data: {…,"choices":[{"index":0,"delta":{},"logprobs":null,"finish_reason":"stop"}],"usage":null,"obfuscation":"9Pwv"} data: {…,"choices":[],"usage":{"prompt_tokens":11,"completion_tokens":1,"total_tokens":12,"prompt_tokens_details":{…},"completion_tokens_details":{…}}} data: [DONE] ``` - `delta` fields: `role` (first chunk), `content`, `refusal`, `tool_calls[]{index, id, type, function:{name, arguments}}` (arguments streamed as string fragments to merge by `index`), `function_call` (deprecated), `audio{id, data, transcript, expires_at}`. - `finish_reason` non-null only on the last content chunk; usage chunk has `choices: []`; `obfuscation` present unless `include_obfuscation:false`; moderation results (if requested) arrive on a final moderation chunk. - Legacy `/v1/completions` streams the same `text_completion` object per chunk (`choices[].text`), usage chunk with `choices:[]`, then `[DONE]`. ## 4. Responses vs Chat streaming at a glance | | Responses | Chat Completions | |---|---|---| | Framing | `event:` + `data:` typed events | `data:` only | | Terminal | `response.completed/failed/incomplete` or `error` | `data: [DONE]` | | Cursor / resume | `sequence_number` + `GET ?stream=true&starting_after` (background) | none | | Usage | in `response.completed.response.usage` | only with `stream_options.include_usage` | | Tool args | `response.function_call_arguments.delta/.done` per item | `delta.tool_calls[i].function.arguments` fragments | | Obfuscation | `obfuscation` on delta events | `obfuscation` on chunks |