Anthropic streaming (stream: true) — SSE event reference + observed sequences
Status: DOCUMENTED + LIVE_VERIFIED 2026-09-18 (3 streams on claude-haiku-4-5-20251001: text, forced tool_use, extended thinking; raws tmp-live/anthropic-core/{b_stream,m2_stream_tool_use,m3_stream_thinking}.json). citations_delta, error event and beta blocks DOCUMENTED only.
Sources: https://platform.claude.com/docs/en/build-with-claude/streaming · https://platform.claude.com/docs/en/api/messages (Raw*Event types) · https://platform.claude.com/docs/en/agents-and-tools/tool-use/fine-grained-tool-streaming · https://platform.claude.com/docs/en/api/versioning
Machine-readable: generated/fragments/streaming-events/anthropic-messages.json
Last verified: 2026-09-18
Wire format
content-type: text/event-stream. Since version 2023-06-01 every frame is a named event: event: <type>\n data: <json>\n\n, where data.type == event name; deltas are incremental; there is no data: [DONE]. Rate-limit and request-id headers arrive with the 200 as usual.
Event flow (documented ordering rules)
message_start— exactly once, first.messagehascontent: [],stop_reason: null, initialusage.- Zero or more content blocks, each
content_block_start→content_block_delta* →content_block_stop, with increasingindex. Server-tool result blocks and betafallbackblocks arrive as a start/stop pair with no deltas. - One or more
message_delta—delta{stop_reason, stop_sequence, stop_details, container}and cumulativeusage. message_stop— exactly once, last.pingmay appear anywhere;errormay appear after the 200; unknown event types must be ignored (versioning policy).
Event catalogue
| Event | data |
Live example |
|---|---|---|
message_start |
{type, message: Message} |
usage: {input_tokens: 11, cache_creation_input_tokens: 0, cache_read_input_tokens: 0, cache_creation: {…}, output_tokens: 1, service_tier: "standard", inference_geo: "not_available"} |
content_block_start |
{type, index, content_block} |
{"type":"text","text":""} · {"type":"tool_use","id":"toolu_…","name":"get_weather","input":{},"caller":{"type":"direct"}} · {"type":"thinking","thinking":"","signature":""} |
content_block_delta |
{type, index, delta} — delta types below |
|
content_block_stop |
{type, index} |
|
message_delta |
{type, delta: {stop_reason, stop_sequence, stop_details, container}, usage} |
usage: {input_tokens: 40, cache_creation_input_tokens: 0, cache_read_input_tokens: 0, output_tokens: 39, output_tokens_details: {thinking_tokens: 32}} |
message_stop |
{type} |
|
ping |
{type: "ping"} |
observed right after the first content_block_start in all 3 streams |
error |
{type: "error", error: {type, message}} |
documented: {"type":"overloaded_error","message":"Overloaded"} |
content_block_delta.delta types
delta.type |
Payload | Applies to | Notes |
|---|---|---|---|
text_delta |
text |
text | |
input_json_delta |
partial_json |
tool_use, server_tool_use | concatenate, parse at content_block_stop; first fragment often "". Live fragments: "", {"city", : "Pari, s"}. With tools[].eager_input_streaming: true (fine-grained; replaces the fine-grained-tool-streaming-2025-05-14 header) fragments are unbuffered/unvalidated — guard the parse, return is_error tool_result on bad JSON. |
thinking_delta |
thinking |
thinking | empty string when display: "omitted" |
signature_delta |
signature |
thinking | once, right before the block's content_block_stop |
citations_delta |
citation |
text | one citation object per event |
Observed sequences (Haiku 4.5, 2026-09-18)
# (b) "Reply with OK.", max_tokens 16
message_start → content_block_start(0,text) → ping → content_block_delta(text_delta "OK") → content_block_delta(text_delta ".")
→ content_block_stop(0) → message_delta(stop_reason end_turn, usage.output_tokens 5) → message_stop
# (m2) forced tool_choice {type: tool, name: get_weather}
message_start → content_block_start(0,tool_use input {}) → ping → input_json_delta×4 ("", {"city", : "Pari, s"})
→ content_block_stop(0) → message_delta(stop_reason tool_use, output_tokens 33) → message_stop
# (m3) thinking {type: enabled, budget_tokens: 1024}
message_start → content_block_start(0,thinking) → ping → thinking_delta×8 → signature_delta → content_block_stop(0)
→ content_block_start(1,text) → text_delta("OK.") → content_block_stop(1)
→ message_delta(stop_reason end_turn, usage.output_tokens 39, output_tokens_details.thinking_tokens 32) → message_stopDocumented sequence with web search: text block → server_tool_use (input_json_delta) → web_search_tool_result (start/stop, no deltas) → text block(s) → message_delta with usage.server_tool_use.web_search_requests.
Streaming + thinking / structured outputs / tools
- Thinking:
thinking_deltathen onesignature_delta; interleaved thinking may put thinking blocks between tool calls.display: "updates"(beta) streams only progress updates. - Structured outputs (
output_config.format): the JSON arrives as ordinarytext_deltafragments of one text block; parse atcontent_block_stop/message_stop(SDK:messages.parseis non-streaming; TSstream()+finalMessage()). - Tool use:
stop_reason: tool_usearrives inmessage_delta;tool_use.inputis{}incontent_block_startby design.
SDK helpers
| SDK | Helper | Accumulated message |
|---|---|---|
| Python 1.7.0 | with client.messages.stream(...) as s: for text in s.text_stream / for event in s (adds synthetic text, content_block, message events) |
s.get_final_message(), s.get_final_text(), s.until_done(), s.current_message_snapshot, s.request_id |
| TypeScript 0.126.0 | client.messages.stream(params).on('text' | 'streamEvent' | 'contentBlock' | 'message' | 'finalMessage' | 'error') / for await |
await stream.finalMessage(), stream.finalText(), stream.controller.abort() |
| low-level (both) | client.messages.create(..., stream=True) → raw event iterator, no accumulation |
— |
| Go / Java / C# / Ruby / PHP | NewStreaming + message.Accumulate(ev) / createStreaming + MessageAccumulator / CreateStreaming(...).Aggregate() / .stream(...).accumulated_message / createStream + MessageAccumulator::forMessages() |
Long requests & recovery
Use streaming whenever max_tokens is large: SDKs refuse non-streaming calls expected to exceed 10 minutes ("Streaming is required for operations that may take longer than 10 minutes"). Recovery after a dropped stream: Claude ≤4.5 — resend the partial assistant text as an assistant prefill; Claude 4.6+ — put the partial text in a user message asking to continue (prefill is rejected). Tool-use and thinking blocks cannot be partially recovered.
Examples: examples/anthropic/streaming/{raw_sse.sh, sdk_stream.py, sdk_stream.ts, manual_sse_parser.py} (all LIVE_VERIFIED).