SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
22.0 KB

# Streaming — four SSE dialects and five WebSocket families

Status: event names and statuses from generated/streaming-events.json (506 records: OpenAI 282, Anthropic 68, xAI 94, Gemini 62); observed sequences quoted from docs/openai/streaming-events.md, docs/anthropic/streaming.md, docs/xai/streaming.md, docs/gemini/streaming.md, docs/gemini/interactions-api.md, docs/gemini/live-events.md (LIVE_VERIFIED 2026-09-18/19 on gpt-5.4-nano, claude-haiku-4-5-20251001, grok-4.3, gemini-3.5-flash-lite / gemini-3.8-flash). Sources: https://developers.openai.com/api/docs/guides/streaming-responses · https://developers.openai.com/api/reference/resources/responses/streaming-events · https://platform.claude.com/docs/en/build-with-claude/streaming · https://docs.x.ai/developers/model-capabilities/text/streaming · https://docs.x.ai/developers/model-capabilities/audio/speech-to-speech · https://ai.google.dev/api/generate-content#method:-models.streamgeneratecontent · https://ai.google.dev/gemini-api/docs/live · https://ai.google.dev/gemini-api/docs/interactions Last verified: 2026-09-18

# 1. Wire format

OpenAI Responses OpenAI Chat Completions Anthropic Messages xAI Responses xAI Chat Completions xAI /v1/messages (deprecated) Gemini streamGenerateContent Gemini Interactions
Framing event: <type> + data:; data.type == event data: <chunk> only event: <type> + data:; data.type == event event: <type> + data: (same names as OpenAI) data: only (chat.completion.chunk) event: + data: (Anthropic names) ?alt=sse: data: <GenerateContentResponse> lines, no event: names; default: one pretty-printed JSON array streamed incrementally ([{…} , {…} ]) event: <type> + data: with event_type, event_id
Ordering guard sequence_number none none (fixed skeleton) sequence_number none none none (chunk order; usageMetadata on every chunk) event_id (resume with last_event_id)
Terminal response.completed|failed|incomplete | error — no [DONE] data: [DONE] after an optional usage chunk message_stop — no [DONE] response.completed (incomplete/failed/error UNVERIFIED) — no [DONE] data: [DONE] (usage chunk with choices: [] when stream_options.include_usage) message_stop last chunk carries finishReason (and on Gemini 3 an empty-text part with thoughtSignature — keep it); stream closes interaction.completed then done (data: [DONE])
Padding obfuscation on delta events obfuscation none none observed none none none none
Full object in stream lifecycle events embed the Response none message_start skeleton; cumulative usage in message_delta response.created snapshot (reasoning: {effort: null}, created_at: 0) and full response in response.completed none (created: 0 in every chunk) message_start skeleton every chunk is a partial GenerateContentResponse with modelVersion, responseId, usageMetadata (last authoritative) step.start/step.stop carry step objects; interaction.completed the final interaction
Keep-alive none none ping none none none (no ping) none none
Errors after 200 error event (terminal) error chunk error event (overloaded_error…), no request_id error (UNVERIFIED) — — undocumented (treat a close without finishReason as an error); prompt block → single chunk with promptFeedback.blockReason, no candidates event: error {error: {code, message}}
Content type text/event-stream text/event-stream text/event-stream text/event-stream text/event-stream text/event-stream text/event-stream (alt=sse) or application/json text/event-stream

# 2. Skeletons (observed live)

OpenAI Responses — "Reply with OK."

text
response.created(0) → response.in_progress(1) → response.output_item.added(2, message) → response.content_part.added(3)
→ response.output_text.delta(4 "OK") → response.output_text.done(5) → response.content_part.done(6) → response.output_item.done(7) → response.completed(8)

Anthropic Messages

text
message_start → content_block_start(0, text) → ping → content_block_delta(text_delta "OK") → content_block_delta(".") → content_block_stop(0)
→ message_delta(stop_reason end_turn, usage) → message_stop

xAI Responses (grok-4.3) — reasoning item first

text
response.created → response.in_progress → response.output_item.added(reasoning) → response.reasoning_summary_part.added → response.reasoning_summary_text.delta… → response.reasoning_summary_text.done
→ response.reasoning_summary_part.done → response.output_item.done → response.output_item.added(message) → response.content_part.added → response.output_text.delta("OK") → response.output_text.done
→ response.content_part.done → response.output_item.done → response.completed

xAI Chat Completions

text
data: {chunk delta.reasoning_content…} → data: {chunk delta.content "OK"} → data: {chunk finish_reason: stop} → [data: {choices: [], usage}] → data: [DONE]

xAI /v1/messages — message_start → content_block_start(thinking) → content_block_delta(thinking_delta)… → content_block_stop → content_block_start(text, index 0 again) → content_block_delta(text_delta) → content_block_stop → message_delta → message_stop (no ping; index reuse is a quirk).

Gemini streamGenerateContent?alt=sse (gemini-3.8-flash)

text
data: {candidates:[{content:{parts:[{text:"OK", thought?}]}}], usageMetadata, modelVersion, responseId}
data: {candidates:[{content:{parts:[{text:"", thoughtSignature:"…"}]}, finishReason:"STOP"}], usageMetadata:{…thoughtsTokenCount}}

With includeThoughts: true, {text, thought: true} parts stream first (not guaranteed on trivial prompts). JSON-mode deltas are partial JSON strings that concatenate.

Gemini Interactions (stream: true)

text
interaction.created → interaction.status_update → step.start(thought) → step.delta{thought_signature} → step.stop → step.start(model_output) → step.delta{text "OK"} → step.stop → interaction.completed → done [DONE]

# 3. Event-name mapping — the four text streams

Concern OpenAI Responses Anthropic Messages xAI Responses / Chat Gemini streamGenerateContent / Interactions
lifecycle response.created, response.queued, response.in_progress, response.completed|failed|incomplete message_start, message_delta (stop_reason, usage), message_stop same as OpenAI (response.queued absent — no background mode); Chat: first chunk / finish_reason / [DONE] no lifecycle events — first chunk / chunk with finishReason; Interactions interaction.created, interaction.status_update (legacy but still emitted), interaction.in_progress, interaction.requires_action, interaction.completed, done
new output unit response.output_item.added / .done content_block_start / content_block_stop (index) same as OpenAI; Chat: implicit new parts[] entries in a chunk; Interactions step.start / step.stop
text response.content_part.added/done, response.output_text.delta/done, response.output_text.annotation.added, response.refusal.delta/done content_block_delta {type: text_delta}; citations_delta same as OpenAI (response.output_text.delta, .done); Chat delta.content parts[].text fragments; Interactions step.delta {type: text}, text_annotation_delta
function / tool args response.function_call_arguments.delta/done, response.custom_tool_call_input.delta/done content_block_delta {type: input_json_delta, partial_json} (buffered; eager_input_streaming) response.function_call_arguments.delta (one delta with the full JSON) + .done; Chat delta.tool_calls[] whole call in one chunk → finish_reason: tool_calls parts[].functionCall arrives complete in one chunk (+ thoughtSignature); Interactions step.delta {type: arguments_delta} (partial JSON)
reasoning response.reasoning_summary_part.added/done, response.reasoning_summary_text.delta/done, response.reasoning_text.delta/done content_block_start(thinking), thinking_delta, signature_delta, redacted_thinking same summary events as OpenAI (reasoning_text.delta UNVERIFIED); Chat delta.reasoning_content (full text) parts[] {text, thought: true} chunks, final thoughtSignature; Interactions step.start(thought), step.delta {type: thought_summary | thought_signature}
web search response.web_search_call.in_progress/searching/completed server_tool_use + input_json_delta, web_search_tool_result block response.output_item.added/done(web_search_call) (.searching events documented, not observed) groundingMetadata on candidates (chunks carry only new groundingChunks — accumulate); Interactions google_search_call / google_search_result steps
code execution response.code_interpreter_call.in_progress/interpreting/completed, response.code_interpreter_call_code.delta/done server_tool_use + bash_code_execution_tool_result identical event names to OpenAI (all LIVE_VERIFIED) executableCode / codeExecutionResult / inlineData parts as produced; Interactions code_execution_call / code_execution_result
file search / RAG response.file_search_call.in_progress/searching/completed — response.file_search_call.* (UNVERIFIED) groundingMetadata.groundingChunks[].retrievedContext; Interactions file_search_call/result
image generation response.image_generation_call.in_progress/generating/partial_image/completed — image_generation_call item events (not exercised) inlineData {mimeType: image/png} parts (image models); Interactions step.delta {type: image}
MCP response.mcp_list_tools.*, response.mcp_call.*, response.mcp_call_arguments.delta/done mcp_tool_use / mcp_tool_result blocks response.mcp_call.* (UNVERIFIED) Interactions mcp call/result steps (UNVERIFIED)
shell / patch response.shell_call_command.*, response.shell_call_output_content.*, response.apply_patch_call_operation_diff.* (LIVE_DISCOVERED) tool_use + input_json_delta shell_call items (not exercised) —
compaction response.compaction.compacting content_block_*(compaction) (beta) — (compact endpoint is unary) —
audio response.audio.delta/done, response.audio.transcript.delta/done — — (voice WebSocket) inlineData audio parts (TTS models; streaming TTS on 3.1); Interactions step.delta {type: audio}
errors error error error (UNVERIFIED) pre-stream JSON error; promptFeedback.blockReason; Interactions error
keep-alive — ping — —

Counts: OpenAI Responses SSE 63 + WebSocket 69/3, Chat 2; Anthropic Messages 13 core + 10 advanced; xAI Responses 24, Chat 3, Messages 6; Gemini generate-content 8 (framing + chunk kinds), Interactions 15. Full lists: FAQ Q4.

# 4. Usage and stop information

OpenAI Anthropic xAI Gemini
where usage arrives response.completed.response.usage (+ prompt_cache_diagnostics 5.6+) message_start.message.usage then cumulative message_delta.usage response.completed.response.usage (incl. cost_in_usd_ticks, server_side_tool_usage_details); Chat: final chunk with choices: [] when stream_options.include_usage: true usageMetadata on every chunk (last one authoritative: thoughtsTokenCount, toolUsePromptTokenCount, cachedContentTokenCount); Interactions interaction.completed.usage
stop reason response.status + incomplete_details.reason message_delta.delta.stop_reason / stop_sequence / stop_details response.status + incomplete_details.reason (max_output_tokens | max_prompt_tokens | max_time_limit); Chat finish_reason candidates[].finishReason on the last chunk (21 values); Interactions status
tool call boundary response.output_item.done for the function_call item content_block_stop of the tool_use block; stop_reason: tool_use response.output_item.done(function_call); Chat finish_reason: tool_calls chunk containing the functionCall part (usually the last, finishReason: STOP); Interactions interaction.requires_action

# 5. The same streamed call on all four providers

bash
# OpenAI
curl -N https://api.openai.com/v1/responses -H "Authorization: Bearer $OPENAI_API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"gpt-5.4-nano","input":"Reply with OK.","max_output_tokens":32,"stream":true}'
# Anthropic
curl -N https://api.anthropic.com/v1/messages -H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01" -H "content-type: application/json" \
  -d '{"model":"claude-haiku-4-5-20251001","max_tokens":32,"messages":[{"role":"user","content":"Reply with OK."}],"stream":true}'
# xAI (OpenAI event names; reasoning summary events come first)
curl -N https://api.x.ai/v1/responses -H "Authorization: Bearer $XAI_API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"grok-4.3","input":"Reply with OK.","reasoning":{"effort":"low"},"max_output_tokens":64,"stream":true}'
# Gemini (SSE framing requires ?alt=sse; without it you get a JSON array)
curl -N "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash-lite:streamGenerateContent?alt=sse" \
  -H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" \
  -d '{"contents":[{"role":"user","parts":[{"text":"Reply with OK."}]}],"generationConfig":{"maxOutputTokens":64,"thinkingConfig":{"thinkingLevel":"minimal"}}}'

Client algorithm that works for all four: dispatch on the event name where one exists (OpenAI/xAI event:, Anthropic event:, Interactions event:), otherwise on chunk shape (Gemini candidates[0].content.parts[]); accumulate text per (OpenAI/xAI output_index,content_index) / (Anthropic block index) / (Gemini part position); accumulate tool-argument fragments and parse at …arguments.done / content_block_stop / the complete functionCall part; take the final object from the terminal event (response.completed.response, interaction.completed) or rebuild it (message_start + deltas; last Gemini chunk + concatenated parts). SDK helpers: client.responses.stream().get_final_response() (OpenAI, and xAI via the OpenAI SDK), client.messages.stream().get_final_message() (Anthropic), client.chat.create(...).stream() (xai-sdk, gRPC), client.models.generate_content_stream() / client.interactions.create(stream=True) (google-genai).

# 6. Beyond SSE — WebSocket message families

Capability OpenAI Anthropic xAI Gemini
WebSocket request mode (text) wss://api.openai.com/v1/responses — response.create (+ stream_id, generate:false), response.steer, response.inject (beta); 32 lanes / 60 min; background unsupported — wss://api.x.ai/v1/responses (WebSocket Responses mode, May 2026; DOCUMENTED) — (Live API is the only WebSocket for text+audio)
Resume a stream GET /v1/responses/{id}?stream=true&starting_after=N (background) — — (deferred Chat is unary polling) Interactions GET /v1beta/interactions/{id}?stream=true&last_event_id=<id>
Agent session streams Agents API SSE agent.session.* (31) Managed Agents SSE (37 events incl. event_start/event_delta) the Responses stream itself Interactions SSE (15 events)
Voice streams Realtime WebSocket/WebRTC/SIP (45 server + 11 client events), Live (20 + 11), translation (6 + 3) — Realtime wss://api.x.ai/v1/realtime (39 server + 9 client events, OpenAI vocabulary: session.update, input_audio_buffer.append/commit, conversation.item.create, response.create → response.output_audio.delta, response.output_audio_transcript.delta, response.done, error, LIVE_DISCOVERED ping); TTS wss://api.x.ai/v1/tts (text.delta/done → audio.delta/done/error); STT wss://api.x.ai/v1/stt (binary frames, finalize, audio.done → transcript.created/partial/done); SIP webhook realtime.call.incoming Live API BidiGenerateContent (21 server + 11 client message types: setup → setupComplete; clientContent, realtimeInput.{audio,video,text,activityStart,activityEnd,audioStreamEnd}, toolResponse → serverContent.{modelTurn,generationComplete,turnComplete,interrupted,inputTranscription,outputTranscription,interactionStatus,groundingMetadata,urlContextMetadata}, toolCall, toolCallCancellation, usageMetadata, goAway, sessionResumptionUpdate); Lyria RealTime BidiGenerateMusic (setup, clientContent.weightedPrompts, musicGenerationConfig, playbackControl → setupComplete, serverContent.audioChunks[], filteredPrompt)
Media streams POST /v1/audio/speech SSE (speech.audio.delta/done), transcriptions (transcript.text.delta/segment/done), images (image_generation.partial_image/completed) — none for images/videos (unary + 202 polling for videos) streaming TTS on gemini-3.1-flash-tts-preview (audio inlineData chunks); Veo/Batch are long-running Operations, not streams
Batch outputs JSONL file JSONL results file GET /v1/batches/{id}/results (paginated JSON) inlinedResponses[] or responsesFile (no streaming)

# Realtime / Live message-family mapping

Purpose OpenAI Realtime xAI realtime Gemini Live
open / configure session.update {session:{type, model, instructions, audio, tools, turn_detection}} → session.created, session.updated session.update {instructions, voice, reasoning.effort, turn_detection, audio.input/output.format, tools[]} → session.created, session.updated, conversation.created setup {model, generationConfig, systemInstruction, tools, realtimeInputConfig, sessionResumption, contextWindowCompression, inputAudioTranscription, outputAudioTranscription} → setupComplete
send audio input_audio_buffer.append {audio} / .commit / .clear; VAD input_audio_buffer.speech_started/stopped/committed same names (input_audio_buffer.append/commit/clear; speech_started/stopped/committed, timeout_triggered, dtmf_event_received) realtimeInput {audio:{data, mimeType: "audio/pcm;rate=16000"}} (+ activityStart/activityEnd when automatic VAD is disabled, audioStreamEnd)
send text / items conversation.item.create {item} → conversation.item.created; .truncate, .delete conversation.item.create → conversation.item.added; .truncate, .delete (each text item $0.004) clientContent {turns[], turnComplete} or realtimeInput {text}
request a turn response.create / response.cancel → response.created … response.done response.create / response.cancel → response.created … response.done (top-level usage {billable_audio_seconds}) implicit on turnComplete / VAD; no explicit create; serverContent.interrupted on barge-in
audio out response.output_audio.delta/done, response.output_audio_transcript.delta/done response.output_audio.delta/done, response.output_audio_transcript.delta/done (response.audio.delta documented, not emitted) serverContent.modelTurn.parts[].inlineData (24 kHz PCM), serverContent.outputTranscription, generationComplete, turnComplete
tools response.function_call_arguments.delta/done, mcp_list_tools.*, response.mcp_call.* same names (function + mcp_list_tools.*, response.mcp_call_arguments.*, response.mcp_call.*) toolCall {functionCalls[]} → toolResponse {functionResponses[] {id, response, scheduling, willContinue}}; toolCallCancellation
session lifetime 60 min; session.expired-style errors 120 min; error {type: timeout|max_duration} ≈10-min connections (goAway {timeLeft}), 15-min audio sessions (unlimited with compression); sessionResumptionUpdate {newHandle} (2 h)
auth for browsers POST /v1/realtime/client_secrets → ek_… POST /v1/realtime/client_secrets → xai-realtime… (subprotocol xai-client-secret.<token>) POST /v1beta/auth_tokens → auth_tokens/… on …BidiGenerateContentConstrained?access_token=

# 7. Streaming constraints and gotchas

Topic OpenAI Anthropic xAI Gemini
long generations background mode recommended for long runs SDKs require streaming above ~10 min of expected output; thinking budgets > 32k → Batches deferred Chat Completions (deferred: true) for long runs; Responses background → 400 Interactions background: true (Deep Research); X-Server-Timeout hint; Flex requests may take 1–15 min
tool-argument JSON fragments per delta; parse on .done fragments may be "" first; eager_input_streaming unvalidated the whole JSON arrives in one delta / one chunk the functionCall part arrives complete; keep the thoughtSignature
structured outputs text deltas of the JSON text_delta fragments of one text block text deltas of the JSON partial JSON strings across chunks (concatenate)
thinking display summaries only display: omitted → one empty thinking_delta + signature_delta Chat streams the full reasoning_content; Responses streams summaries includeThoughts summaries not guaranteed; an empty-text final part carries the signature
rate-limit headers on the 200 on the 200 on the 200 (undocumented x-ratelimit-*) none
SSE resume yes (background) no no Interactions only (last_event_id)
framing traps obfuscation padding ping events created: 0 in chunks; /v1/messages reuses block index 0 default (non-SSE) response is a JSON array — needs an incremental array parser; generateContent?alt=sse also emits one event

Related: features · tool-execution · realtime-and-media · agents-platforms · docs/openai/streaming-events.md · docs/anthropic/streaming.md · docs/xai/streaming.md · docs/xai/voice.md · docs/gemini/streaming.md · docs/gemini/live-events.md.