SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
10.0 KB

# OpenAI streaming events — Responses SSE, Responses WebSocket, Chat Completions chunks

Status: DOCUMENTED (59 SSE event types from the OpenAPI ResponseStreamEvent union, 2 GA + 1 beta WebSocket client events, 62 GA + 2 beta server events) · LIVE_VERIFIED for the 11 lifecycle/text events observed on 2026-09-18 and for Chat Completions chunks. Machine-readable: generated/fragments/streaming-events/openai-responses.json, openai-responses-websocket.json, openai-chat-completions.json.

Sources

Last verified: 2026-09-18

# 1. Responses SSE (stream: true)

Transport: Content-Type: text/event-stream. Each event is

text
event: response.output_text.delta
data: {"type":"response.output_text.delta","item_id":"msg_…","output_index":0,"content_index":0,"delta":"OK","logprobs":[],"obfuscation":"tcI6fDMDcuuDp0","sequence_number":4}

Invariants (live): event: name equals data.type; sequence_number starts at 0 and increments by 1 per event; delta events carry an obfuscation pad (disable with stream_options.include_obfuscation:false); lifecycle events embed the full Response object (response.created has status:"in_progress", response.completed carries final usage). No data: [DONE] sentinel — the stream ends after the terminal event (response.completed | response.failed | response.incomplete | error).

# Observed sequences (gpt-5.4-nano, "Reply with OK.", max_output_tokens 32)

Case Ordered event types
plain text (b) response.created(0) → response.in_progress(1) → response.output_item.added(2, message in_progress, phase:"final_answer") → response.content_part.added(3, empty output_text) → response.output_text.delta(4 "OK") → response.output_text.delta(5 ".") → response.output_text.done(6) → response.content_part.done(7) → response.output_item.done(8) → response.completed(9)
background + stream (j5) response.created(0) → response.queued(1) → response.in_progress(2) → … same … → response.completed(9)
resume GET …?stream=true&starting_after=1 (j6) replays response.in_progress(2) … response.completed(9)
chat-completions style (l2) see §3

# Complete SSE event catalogue (59)

Category Event types Key payload fields
Response lifecycle response.created, response.queued, response.in_progress, response.completed, response.failed, response.incomplete response (full Response), sequence_number
Item lifecycle response.output_item.added, response.output_item.done item (OutputItem), output_index
Content parts response.content_part.added, response.content_part.done item_id, output_index, content_index, part (output_text | refusal | reasoning_text)
Text response.output_text.delta (delta, logprobs[]), response.output_text.done (text, logprobs[]), response.output_text.annotation.added (annotation, annotation_index) + item_id, output_index, content_index
Refusal response.refusal.delta (delta), response.refusal.done (refusal)
Reasoning response.reasoning_summary_part.added/.done (part{type:summary_text,text}, summary_index), response.reasoning_summary_text.delta/.done (delta/text, summary_index), response.reasoning_text.delta/.done (delta/text, content_index)
Function tools response.function_call_arguments.delta (delta), response.function_call_arguments.done (arguments, name) item_id, output_index
Custom tools response.custom_tool_call_input.delta (delta), response.custom_tool_call_input.done (input)
File search response.file_search_call.in_progress/.searching/.completed item_id, output_index
Web search response.web_search_call.in_progress/.searching/.completed
Code interpreter response.code_interpreter_call.in_progress/.interpreting/.completed, response.code_interpreter_call_code.delta/.done (delta/code)
Image generation response.image_generation_call.in_progress/.generating/.completed, response.image_generation_call.partial_image (partial_image_b64, partial_image_index, size, quality, background, output_format)
MCP response.mcp_call.in_progress/.completed/.failed, response.mcp_call_arguments.delta/.done (delta/arguments), response.mcp_list_tools.in_progress/.completed/.failed
Shell response.shell_call_command.added/.delta/.done, response.shell_call_output_content.delta/.done
Compaction response.compaction.compacting emitted when server-side compaction runs mid-response
Audio (chat-style audio output) response.audio.delta/.done, response.audio.transcript.delta/.done delta base64 / transcript text
Error error (code, message, param, sequence_number) terminal

Not in the SSE union but present as output items / other tools' events: apply_patch has no dedicated streaming events in this spec snapshot (its calls appear via output_item.*). The beta SSE union (BetaResponseStreamEvent) is identical to the GA list.

Handling guidance: branch on type; accumulate output_text.delta per (output_index,content_index); treat output_item.done as the canonical item (its encrypted_content is complete, output_item.added may be partial); read usage and prompt_cache_diagnostics from response.completed.response; persist the last sequence_number for background resume.

# 2. Responses WebSocket events (wss://api.openai.com/v1/responses)

Client → server

Event Fields Notes
response.create all POST /v1/responses fields + stream_id (1–256, [A-Za-z0-9_.-]); generate:false warm-up stream implicit; background unsupported
response.steer previous_response_id, input (messages / function_call_output) mid-turn steering; the target stops at a safe boundary (incomplete_details.reason:"steered") and a successor response is created
response.inject (beta) response_id, input[] inject client tool outputs into an active response

Server → client: every SSE event above (with an optional stream_id echo for named lanes) plus:

Event Fields Meaning
response.steer.accepted steer{…}, stream_id input validated and queued (commit point = successor response.created)
response.steer.pending steer, reason (e.g. waiting_for_required_input), required_input[] queued input still owned by server after target completed — do not resend
response.steer.failed steer, error returns the uncommitted input
response.inject.created / response.inject.failed (beta)
error status, stream_id?, error{type, code, message, param} codes: previous_response_not_found, invalid_stream_id, websocket_stream_limit_reached (32 lanes), websocket_connection_limit_reached (60 min)

Concurrency model: same stream_id → FIFO; different lanes → concurrent (≤16 in flight). previous_response_id = lineage (fork a lane from another lane's response; wait for the fork's response.in_progress before advancing the parent lane with store:false). Connection cache is lost on reconnect.

# 3. Chat Completions chunks (stream: true)

Data-only SSE (no event: line): data: {chat.completion.chunk} … data: [DONE].

Observed (gpt-4.1-nano, stream_options.include_usage:true):

text
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","created":…,"model":"gpt-4.1-nano-2025-04-14","service_tier":"default","system_fingerprint":"fp_…","choices":[{"index":0,"delta":{"role":"assistant","content":"","refusal":null},"logprobs":null,"finish_reason":null}],"usage":null,"obfuscation":"GO7oM1GH"}
data: {…,"choices":[{"index":0,"delta":{"content":"OK"},"logprobs":null,"finish_reason":null}],"usage":null,"obfuscation":"4EFX0fab"}
data: {…,"choices":[{"index":0,"delta":{},"logprobs":null,"finish_reason":"stop"}],"usage":null,"obfuscation":"9Pwv"}
data: {…,"choices":[],"usage":{"prompt_tokens":11,"completion_tokens":1,"total_tokens":12,"prompt_tokens_details":{…},"completion_tokens_details":{…}}}
data: [DONE]
  • delta fields: role (first chunk), content, refusal, tool_calls[]{index, id, type, function:{name, arguments}} (arguments streamed as string fragments to merge by index), function_call (deprecated), audio{id, data, transcript, expires_at}.
  • finish_reason non-null only on the last content chunk; usage chunk has choices: []; obfuscation present unless include_obfuscation:false; moderation results (if requested) arrive on a final moderation chunk.
  • Legacy /v1/completions streams the same text_completion object per chunk (choices[].text), usage chunk with choices:[], then [DONE].

# 4. Responses vs Chat streaming at a glance

Responses Chat Completions
Framing event: + data: typed events data: only
Terminal response.completed/failed/incomplete or error data: [DONE]
Cursor / resume sequence_number + GET ?stream=true&starting_after (background) none
Usage in response.completed.response.usage only with stream_options.include_usage
Tool args response.function_call_arguments.delta/.done per item delta.tool_calls[i].function.arguments fragments
Obfuscation obfuscation on delta events obfuscation on chunks