OpenAI streaming events — Responses SSE, Responses WebSocket, Chat Completions chunks
Status: DOCUMENTED (59 SSE event types from the OpenAPI ResponseStreamEvent union, 2 GA + 1 beta WebSocket client events, 62 GA + 2 beta server events) · LIVE_VERIFIED for the 11 lifecycle/text events observed on 2026-09-18 and for Chat Completions chunks. Machine-readable: generated/fragments/streaming-events/openai-responses.json, openai-responses-websocket.json, openai-chat-completions.json.
Sources
- https://developers.openai.com/api/reference/resources/responses/streaming-events · https://developers.openai.com/api/reference/resources/responses/websocket-events · https://developers.openai.com/api/reference/resources/beta/subresources/responses/streaming-events
- https://developers.openai.com/api/reference/resources/chat/subresources/completions/streaming-events
- https://developers.openai.com/api/docs/guides/streaming-responses · https://developers.openai.com/api/docs/guides/websocket-mode · https://developers.openai.com/api/docs/guides/background#streaming-a-background-response
- OpenAPI
ResponseStreamEvent,ResponsesClientEvent,ResponsesServerEvent,CreateChatCompletionStreamResponse
Last verified: 2026-09-18
1. Responses SSE (stream: true)
Transport: Content-Type: text/event-stream. Each event is
event: response.output_text.delta
data: {"type":"response.output_text.delta","item_id":"msg_…","output_index":0,"content_index":0,"delta":"OK","logprobs":[],"obfuscation":"tcI6fDMDcuuDp0","sequence_number":4}Invariants (live): event: name equals data.type; sequence_number starts at 0 and increments by 1 per event; delta events carry an obfuscation pad (disable with stream_options.include_obfuscation:false); lifecycle events embed the full Response object (response.created has status:"in_progress", response.completed carries final usage). No data: [DONE] sentinel — the stream ends after the terminal event (response.completed | response.failed | response.incomplete | error).
Observed sequences (gpt-5.4-nano, "Reply with OK.", max_output_tokens 32)
| Case | Ordered event types |
|---|---|
| plain text (b) | response.created(0) → response.in_progress(1) → response.output_item.added(2, message in_progress, phase:"final_answer") → response.content_part.added(3, empty output_text) → response.output_text.delta(4 "OK") → response.output_text.delta(5 ".") → response.output_text.done(6) → response.content_part.done(7) → response.output_item.done(8) → response.completed(9) |
| background + stream (j5) | response.created(0) → response.queued(1) → response.in_progress(2) → … same … → response.completed(9) |
resume GET …?stream=true&starting_after=1 (j6) |
replays response.in_progress(2) … response.completed(9) |
| chat-completions style (l2) | see §3 |
Complete SSE event catalogue (59)
| Category | Event types | Key payload fields |
|---|---|---|
| Response lifecycle | response.created, response.queued, response.in_progress, response.completed, response.failed, response.incomplete |
response (full Response), sequence_number |
| Item lifecycle | response.output_item.added, response.output_item.done |
item (OutputItem), output_index |
| Content parts | response.content_part.added, response.content_part.done |
item_id, output_index, content_index, part (output_text | refusal | reasoning_text) |
| Text | response.output_text.delta (delta, logprobs[]), response.output_text.done (text, logprobs[]), response.output_text.annotation.added (annotation, annotation_index) |
+ item_id, output_index, content_index |
| Refusal | response.refusal.delta (delta), response.refusal.done (refusal) |
|
| Reasoning | response.reasoning_summary_part.added/.done (part{type:summary_text,text}, summary_index), response.reasoning_summary_text.delta/.done (delta/text, summary_index), response.reasoning_text.delta/.done (delta/text, content_index) |
|
| Function tools | response.function_call_arguments.delta (delta), response.function_call_arguments.done (arguments, name) |
item_id, output_index |
| Custom tools | response.custom_tool_call_input.delta (delta), response.custom_tool_call_input.done (input) |
|
| File search | response.file_search_call.in_progress/.searching/.completed |
item_id, output_index |
| Web search | response.web_search_call.in_progress/.searching/.completed |
|
| Code interpreter | response.code_interpreter_call.in_progress/.interpreting/.completed, response.code_interpreter_call_code.delta/.done (delta/code) |
|
| Image generation | response.image_generation_call.in_progress/.generating/.completed, response.image_generation_call.partial_image (partial_image_b64, partial_image_index, size, quality, background, output_format) |
|
| MCP | response.mcp_call.in_progress/.completed/.failed, response.mcp_call_arguments.delta/.done (delta/arguments), response.mcp_list_tools.in_progress/.completed/.failed |
|
| Shell | response.shell_call_command.added/.delta/.done, response.shell_call_output_content.delta/.done |
|
| Compaction | response.compaction.compacting |
emitted when server-side compaction runs mid-response |
| Audio (chat-style audio output) | response.audio.delta/.done, response.audio.transcript.delta/.done |
delta base64 / transcript text |
| Error | error (code, message, param, sequence_number) |
terminal |
Not in the SSE union but present as output items / other tools' events: apply_patch has no dedicated streaming events in this spec snapshot (its calls appear via output_item.*). The beta SSE union (BetaResponseStreamEvent) is identical to the GA list.
Handling guidance: branch on type; accumulate output_text.delta per (output_index,content_index); treat output_item.done as the canonical item (its encrypted_content is complete, output_item.added may be partial); read usage and prompt_cache_diagnostics from response.completed.response; persist the last sequence_number for background resume.
2. Responses WebSocket events (wss://api.openai.com/v1/responses)
Client → server
| Event | Fields | Notes |
|---|---|---|
response.create |
all POST /v1/responses fields + stream_id (1–256, [A-Za-z0-9_.-]); generate:false warm-up |
stream implicit; background unsupported |
response.steer |
previous_response_id, input (messages / function_call_output) |
mid-turn steering; the target stops at a safe boundary (incomplete_details.reason:"steered") and a successor response is created |
response.inject (beta) |
response_id, input[] |
inject client tool outputs into an active response |
Server → client: every SSE event above (with an optional stream_id echo for named lanes) plus:
| Event | Fields | Meaning |
|---|---|---|
response.steer.accepted |
steer{…}, stream_id |
input validated and queued (commit point = successor response.created) |
response.steer.pending |
steer, reason (e.g. waiting_for_required_input), required_input[] |
queued input still owned by server after target completed — do not resend |
response.steer.failed |
steer, error |
returns the uncommitted input |
response.inject.created / response.inject.failed (beta) |
||
error |
status, stream_id?, error{type, code, message, param} |
codes: previous_response_not_found, invalid_stream_id, websocket_stream_limit_reached (32 lanes), websocket_connection_limit_reached (60 min) |
Concurrency model: same stream_id → FIFO; different lanes → concurrent (≤16 in flight). previous_response_id = lineage (fork a lane from another lane's response; wait for the fork's response.in_progress before advancing the parent lane with store:false). Connection cache is lost on reconnect.
3. Chat Completions chunks (stream: true)
Data-only SSE (no event: line): data: {chat.completion.chunk} … data: [DONE].
Observed (gpt-4.1-nano, stream_options.include_usage:true):
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","created":…,"model":"gpt-4.1-nano-2025-04-14","service_tier":"default","system_fingerprint":"fp_…","choices":[{"index":0,"delta":{"role":"assistant","content":"","refusal":null},"logprobs":null,"finish_reason":null}],"usage":null,"obfuscation":"GO7oM1GH"}
data: {…,"choices":[{"index":0,"delta":{"content":"OK"},"logprobs":null,"finish_reason":null}],"usage":null,"obfuscation":"4EFX0fab"}
data: {…,"choices":[{"index":0,"delta":{},"logprobs":null,"finish_reason":"stop"}],"usage":null,"obfuscation":"9Pwv"}
data: {…,"choices":[],"usage":{"prompt_tokens":11,"completion_tokens":1,"total_tokens":12,"prompt_tokens_details":{…},"completion_tokens_details":{…}}}
data: [DONE]deltafields:role(first chunk),content,refusal,tool_calls[]{index, id, type, function:{name, arguments}}(arguments streamed as string fragments to merge byindex),function_call(deprecated),audio{id, data, transcript, expires_at}.finish_reasonnon-null only on the last content chunk; usage chunk haschoices: [];obfuscationpresent unlessinclude_obfuscation:false; moderation results (if requested) arrive on a final moderation chunk.- Legacy
/v1/completionsstreams the sametext_completionobject per chunk (choices[].text), usage chunk withchoices:[], then[DONE].
4. Responses vs Chat streaming at a glance
| Responses | Chat Completions | |
|---|---|---|
| Framing | event: + data: typed events |
data: only |
| Terminal | response.completed/failed/incomplete or error |
data: [DONE] |
| Cursor / resume | sequence_number + GET ?stream=true&starting_after (background) |
none |
| Usage | in response.completed.response.usage |
only with stream_options.include_usage |
| Tool args | response.function_call_arguments.delta/.done per item |
delta.tool_calls[i].function.arguments fragments |
| Obfuscation | obfuscation on delta events |
obfuscation on chunks |