SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
7.0 KB · 82 lines markdown
Rendered Raw Blame History
1# Anthropic streaming (`stream: true`) — SSE event reference + observed sequences23**Status:** `DOCUMENTED` + `LIVE_VERIFIED` 2026-09-18 (3 streams on `claude-haiku-4-5-20251001`: text, forced tool_use, extended thinking; raws `tmp-live/anthropic-core/{b_stream,m2_stream_tool_use,m3_stream_thinking}.json`). `citations_delta`, `error` event and beta blocks `DOCUMENTED` only.4**Sources:** https://platform.claude.com/docs/en/build-with-claude/streaming · https://platform.claude.com/docs/en/api/messages (Raw*Event types) · https://platform.claude.com/docs/en/agents-and-tools/tool-use/fine-grained-tool-streaming · https://platform.claude.com/docs/en/api/versioning5**Machine-readable:** `generated/fragments/streaming-events/anthropic-messages.json`6**Last verified:** 2026-09-1878## Wire format910`content-type: text/event-stream`. Since version `2023-06-01` every frame is a **named event**: `event: <type>\n data: <json>\n\n`, where `data.type == event name`; deltas are incremental; there is no `data: [DONE]`. Rate-limit and `request-id` headers arrive with the 200 as usual.1112## Event flow (documented ordering rules)13141. `message_start` — exactly once, first. `message` has `content: []`, `stop_reason: null`, initial `usage`.152. Zero or more content blocks, each `content_block_start` → `content_block_delta`* → `content_block_stop`, with increasing `index`. Server-tool result blocks and beta `fallback` blocks arrive as a start/stop pair with **no deltas**.163. One or more `message_delta` — `delta{stop_reason, stop_sequence, stop_details, container}` and **cumulative** `usage`.174. `message_stop` — exactly once, last.18`ping` may appear anywhere; `error` may appear after the 200; unknown event types must be ignored (versioning policy).1920## Event catalogue2122| Event | `data` | Live example |23|---|---|---|24| `message_start` | `{type, message: Message}` | `usage: {input_tokens: 11, cache_creation_input_tokens: 0, cache_read_input_tokens: 0, cache_creation: {…}, output_tokens: 1, service_tier: "standard", inference_geo: "not_available"}` |25| `content_block_start` | `{type, index, content_block}` | `{"type":"text","text":""}` · `{"type":"tool_use","id":"toolu_…","name":"get_weather","input":{},"caller":{"type":"direct"}}` · `{"type":"thinking","thinking":"","signature":""}` |26| `content_block_delta` | `{type, index, delta}` — delta types below | |27| `content_block_stop` | `{type, index}` | |28| `message_delta` | `{type, delta: {stop_reason, stop_sequence, stop_details, container}, usage}` | `usage: {input_tokens: 40, cache_creation_input_tokens: 0, cache_read_input_tokens: 0, output_tokens: 39, output_tokens_details: {thinking_tokens: 32}}` |29| `message_stop` | `{type}` | |30| `ping` | `{type: "ping"}` | observed right after the first `content_block_start` in all 3 streams |31| `error` | `{type: "error", error: {type, message}}` | documented: `{"type":"overloaded_error","message":"Overloaded"}` |3233### `content_block_delta.delta` types3435| `delta.type` | Payload | Applies to | Notes |36|---|---|---|---|37| `text_delta` | `text` | text | |38| `input_json_delta` | `partial_json` | tool_use, server_tool_use | concatenate, parse at `content_block_stop`; first fragment often `""`. Live fragments: `""`, `{"city"`, `: "Pari`, `s"}`. With `tools[].eager_input_streaming: true` (fine-grained; replaces the `fine-grained-tool-streaming-2025-05-14` header) fragments are unbuffered/unvalidated — guard the parse, return `is_error` tool_result on bad JSON. |39| `thinking_delta` | `thinking` | thinking | empty string when `display: "omitted"` |40| `signature_delta` | `signature` | thinking | once, right before the block's `content_block_stop` |41| `citations_delta` | `citation` | text | one citation object per event |4243## Observed sequences (Haiku 4.5, 2026-09-18)4445```46# (b) "Reply with OK.", max_tokens 1647message_start → content_block_start(0,text) → ping → content_block_delta(text_delta "OK") → content_block_delta(text_delta ".")48→ content_block_stop(0) → message_delta(stop_reason end_turn, usage.output_tokens 5) → message_stop4950# (m2) forced tool_choice {type: tool, name: get_weather}51message_start → content_block_start(0,tool_use input {}) → ping → input_json_delta×4 ("", {"city", : "Pari, s"})52→ content_block_stop(0) → message_delta(stop_reason tool_use, output_tokens 33) → message_stop5354# (m3) thinking {type: enabled, budget_tokens: 1024}55message_start → content_block_start(0,thinking) → ping → thinking_delta×8 → signature_delta → content_block_stop(0)56→ content_block_start(1,text) → text_delta("OK.") → content_block_stop(1)57→ message_delta(stop_reason end_turn, usage.output_tokens 39, output_tokens_details.thinking_tokens 32) → message_stop58```5960Documented sequence with web search: text block → `server_tool_use` (input_json_delta) → `web_search_tool_result` (start/stop, no deltas) → text block(s) → `message_delta` with `usage.server_tool_use.web_search_requests`.6162## Streaming + thinking / structured outputs / tools6364- Thinking: `thinking_delta` then one `signature_delta`; interleaved thinking may put thinking blocks between tool calls. `display: "updates"` (beta) streams only progress updates.65- Structured outputs (`output_config.format`): the JSON arrives as ordinary `text_delta` fragments of one text block; parse at `content_block_stop`/`message_stop` (SDK: `messages.parse` is non-streaming; TS `stream()` + `finalMessage()`).66- Tool use: `stop_reason: tool_use` arrives in `message_delta`; `tool_use.input` is `{}` in `content_block_start` by design.6768## SDK helpers6970| SDK | Helper | Accumulated message |71|---|---|---|72| Python 1.7.0 | `with client.messages.stream(...) as s: for text in s.text_stream` / `for event in s` (adds synthetic `text`, `content_block`, `message` events) | `s.get_final_message()`, `s.get_final_text()`, `s.until_done()`, `s.current_message_snapshot`, `s.request_id` |73| TypeScript 0.126.0 | `client.messages.stream(params).on('text' \| 'streamEvent' \| 'contentBlock' \| 'message' \| 'finalMessage' \| 'error')` / `for await` | `await stream.finalMessage()`, `stream.finalText()`, `stream.controller.abort()` |74| low-level (both) | `client.messages.create(..., stream=True)` → raw event iterator, no accumulation | — |75| Go / Java / C# / Ruby / PHP | `NewStreaming` + `message.Accumulate(ev)` / `createStreaming` + `MessageAccumulator` / `CreateStreaming(...).Aggregate()` / `.stream(...).accumulated_message` / `createStream` + `MessageAccumulator::forMessages()` | |7677## Long requests & recovery7879Use streaming whenever `max_tokens` is large: SDKs refuse non-streaming calls expected to exceed 10 minutes ("Streaming is required for operations that may take longer than 10 minutes"). Recovery after a dropped stream: Claude ≤4.5 — resend the partial assistant text as an assistant prefill; Claude 4.6+ — put the partial text in a **user** message asking to continue (prefill is rejected). Tool-use and thinking blocks cannot be partially recovered.8081Examples: `examples/anthropic/streaming/{raw_sse.sh, sdk_stream.py, sdk_stream.ts, manual_sse_parser.py}` (all `LIVE_VERIFIED`).82