SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
7.0 KB

# Anthropic streaming (stream: true) — SSE event reference + observed sequences

Status: DOCUMENTED + LIVE_VERIFIED 2026-09-18 (3 streams on claude-haiku-4-5-20251001: text, forced tool_use, extended thinking; raws tmp-live/anthropic-core/{b_stream,m2_stream_tool_use,m3_stream_thinking}.json). citations_delta, error event and beta blocks DOCUMENTED only. Sources: https://platform.claude.com/docs/en/build-with-claude/streaming · https://platform.claude.com/docs/en/api/messages (Raw*Event types) · https://platform.claude.com/docs/en/agents-and-tools/tool-use/fine-grained-tool-streaming · https://platform.claude.com/docs/en/api/versioning Machine-readable: generated/fragments/streaming-events/anthropic-messages.json Last verified: 2026-09-18

# Wire format

content-type: text/event-stream. Since version 2023-06-01 every frame is a named event: event: <type>\n data: <json>\n\n, where data.type == event name; deltas are incremental; there is no data: [DONE]. Rate-limit and request-id headers arrive with the 200 as usual.

# Event flow (documented ordering rules)

  1. message_start — exactly once, first. message has content: [], stop_reason: null, initial usage.
  2. Zero or more content blocks, each content_block_start → content_block_delta* → content_block_stop, with increasing index. Server-tool result blocks and beta fallback blocks arrive as a start/stop pair with no deltas.
  3. One or more message_delta — delta{stop_reason, stop_sequence, stop_details, container} and cumulative usage.
  4. message_stop — exactly once, last. ping may appear anywhere; error may appear after the 200; unknown event types must be ignored (versioning policy).

# Event catalogue

Event data Live example
message_start {type, message: Message} usage: {input_tokens: 11, cache_creation_input_tokens: 0, cache_read_input_tokens: 0, cache_creation: {…}, output_tokens: 1, service_tier: "standard", inference_geo: "not_available"}
content_block_start {type, index, content_block} {"type":"text","text":""} · {"type":"tool_use","id":"toolu_…","name":"get_weather","input":{},"caller":{"type":"direct"}} · {"type":"thinking","thinking":"","signature":""}
content_block_delta {type, index, delta} — delta types below
content_block_stop {type, index}
message_delta {type, delta: {stop_reason, stop_sequence, stop_details, container}, usage} usage: {input_tokens: 40, cache_creation_input_tokens: 0, cache_read_input_tokens: 0, output_tokens: 39, output_tokens_details: {thinking_tokens: 32}}
message_stop {type}
ping {type: "ping"} observed right after the first content_block_start in all 3 streams
error {type: "error", error: {type, message}} documented: {"type":"overloaded_error","message":"Overloaded"}

# content_block_delta.delta types

delta.type Payload Applies to Notes
text_delta text text
input_json_delta partial_json tool_use, server_tool_use concatenate, parse at content_block_stop; first fragment often "". Live fragments: "", {"city", : "Pari, s"}. With tools[].eager_input_streaming: true (fine-grained; replaces the fine-grained-tool-streaming-2025-05-14 header) fragments are unbuffered/unvalidated — guard the parse, return is_error tool_result on bad JSON.
thinking_delta thinking thinking empty string when display: "omitted"
signature_delta signature thinking once, right before the block's content_block_stop
citations_delta citation text one citation object per event

# Observed sequences (Haiku 4.5, 2026-09-18)

text
# (b) "Reply with OK.", max_tokens 16
message_start → content_block_start(0,text) → ping → content_block_delta(text_delta "OK") → content_block_delta(text_delta ".")
→ content_block_stop(0) → message_delta(stop_reason end_turn, usage.output_tokens 5) → message_stop

# (m2) forced tool_choice {type: tool, name: get_weather}
message_start → content_block_start(0,tool_use input {}) → ping → input_json_delta×4 ("", {"city", : "Pari, s"})
→ content_block_stop(0) → message_delta(stop_reason tool_use, output_tokens 33) → message_stop

# (m3) thinking {type: enabled, budget_tokens: 1024}
message_start → content_block_start(0,thinking) → ping → thinking_delta×8 → signature_delta → content_block_stop(0)
→ content_block_start(1,text) → text_delta("OK.") → content_block_stop(1)
→ message_delta(stop_reason end_turn, usage.output_tokens 39, output_tokens_details.thinking_tokens 32) → message_stop

Documented sequence with web search: text block → server_tool_use (input_json_delta) → web_search_tool_result (start/stop, no deltas) → text block(s) → message_delta with usage.server_tool_use.web_search_requests.

# Streaming + thinking / structured outputs / tools

  • Thinking: thinking_delta then one signature_delta; interleaved thinking may put thinking blocks between tool calls. display: "updates" (beta) streams only progress updates.
  • Structured outputs (output_config.format): the JSON arrives as ordinary text_delta fragments of one text block; parse at content_block_stop/message_stop (SDK: messages.parse is non-streaming; TS stream() + finalMessage()).
  • Tool use: stop_reason: tool_use arrives in message_delta; tool_use.input is {} in content_block_start by design.

# SDK helpers

SDK Helper Accumulated message
Python 1.7.0 with client.messages.stream(...) as s: for text in s.text_stream / for event in s (adds synthetic text, content_block, message events) s.get_final_message(), s.get_final_text(), s.until_done(), s.current_message_snapshot, s.request_id
TypeScript 0.126.0 client.messages.stream(params).on('text' | 'streamEvent' | 'contentBlock' | 'message' | 'finalMessage' | 'error') / for await await stream.finalMessage(), stream.finalText(), stream.controller.abort()
low-level (both) client.messages.create(..., stream=True) → raw event iterator, no accumulation —
Go / Java / C# / Ruby / PHP NewStreaming + message.Accumulate(ev) / createStreaming + MessageAccumulator / CreateStreaming(...).Aggregate() / .stream(...).accumulated_message / createStream + MessageAccumulator::forMessages()

# Long requests & recovery

Use streaming whenever max_tokens is large: SDKs refuse non-streaming calls expected to exceed 10 minutes ("Streaming is required for operations that may take longer than 10 minutes"). Recovery after a dropped stream: Claude ≤4.5 — resend the partial assistant text as an assistant prefill; Claude 4.6+ — put the partial text in a user message asking to continue (prefill is rejected). Tool-use and thinking blocks cannot be partially recovered.

Examples: examples/anthropic/streaming/{raw_sse.sh, sdk_stream.py, sdk_stream.ts, manual_sse_parser.py} (all LIVE_VERIFIED).