SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
11.2 KB

# Anthropic — Thinking (extended, adaptive, interleaved, preserved)

Status: DOCUMENTED · LIVE_VERIFIED (extended thinking on claude-haiku-4-5-20251001; adaptive on claude-sonnet-5; interleaved header on claude-sonnet-4-6; 2026-09-18). display: "updates" and block_binding are BETA. Sources: Thinking · Extended thinking · Steering thinking · Troubleshooting · Preserved thinking · Context windows · Messages API · Beta Messages API Last verified: 2026-09-18

# 1. Modes

Mode Request Behaviour
Extended (manual) thinking: {"type": "enabled", "budget_tokens": N[, "display"]} Claude always thinks against a target budget. Only mode on 4.5-and-earlier models; deprecated on Opus/Sonnet 4.6; 400 on 4.7+, Sonnet 5, Opus 5, Fable/Mythos ("thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.)
Adaptive thinking: {"type": "adaptive"[, "display"]} Claude decides whether/how deeply to think; depth steered by output_config.effort. Interleaves between tool calls automatically. 400 on 4.5 models (adaptive thinking is not supported on this model).
Disabled thinking: {"type": "disabled"} Off. 400 on Fable/Mythos (always on); 400 on Opus 5 combined with effort xhigh/max.
Omitted parameter — Always-on on Fable 5.x/Mythos; on by default on Opus 5 / Sonnet 5; off on Opus 4.6–4.8, Sonnet 4.6 and all 4.5 models.

# Per-model table (docs + Models API capabilities.thinking.types + live)

Model Types accepted Default Rejected (400) display default Prior-turn thinking kept Live
Fable 5.1 / Mythos 5.1 adaptive always on enabled, disabled omitted all —
Fable 5 / Mythos 5 adaptive always on enabled, disabled omitted all —
Mythos Preview adaptive, extended always on disabled omitted all —
Opus 5 adaptive on enabled; disabled at xhigh/max omitted all disabled accepted (fast_04)
Opus 4.8 / 4.7 adaptive off enabled omitted all —
Sonnet 5 adaptive on enabled omitted all adaptive ✔, disabled ✔, enabled ✘ 400
Opus 4.6 / Sonnet 4.6 adaptive, extended (deprecated) off none summarized all Sonnet 4.6 enabled + interleaved header ✔
Opus 4.5 extended off adaptive summarized all —
Sonnet 4.5 extended off adaptive summarized last turn —
Haiku 4.5 extended off adaptive summarized last turn enabled ✔, adaptive ✘ 400

# 2. Budget rules (manual mode)

  • budget_tokens ≥ 1024 (400 thinking.enabled.budget_tokens: Input should be greater than or equal to 1024) and < max_tokens (400 max_tokens must be greater than thinking.budget_tokens) — except with interleaved thinking where the budget spans the whole turn and may exceed max_tokens.
  • The budget is a target, not a cap; max_tokens is the hard ceiling (thinking + text). Live: budget 1024 on "2+2" used 30 thinking tokens.
  • max_tokens: 0 (cache pre-warm) is incompatible with enabled thinking.
  • Budgets > 32k → use Batches (timeouts). SDKs require streaming when max_tokens > 21,333.
  • On Opus 4.5 effort composes with budget_tokens.
  • Adaptive mode: no budget; bound cost with max_tokens (hard) and effort (soft). stop_reason: "max_tokens" → raise max_tokens or lower effort.

# 3. Response shape

json
{"type": "thinking", "thinking": "The user is asking for 2+2, which equals 4…", "signature": "EtQCCpoBCBEYAipAGx2p…"}
{"type": "text", "text": "4"}
  • thinking text is a summary (never raw chain of thought). display: "omitted" → thinking: "" with the same signature (faster TTFT when streaming, same billing). display: "updates" (beta thinking-display-updates-2026-08-18) → reasoning blocks empty, progress-update blocks (Fable 5.x) carry text; without the header: 400 thinking.adaptive.display: Input should be 'summarized', 'omitted'.
  • redacted_thinking: {"type": "redacted_thinking", "data": "<opaque>"} — safety-redacted; pass back unchanged.
  • usage.output_tokens_details.thinking_tokens reports billed reasoning tokens (live: 30 of 37 output tokens). Streaming: only on the final message_delta.
  • Adaptive mode may return no thinking block (live Sonnet 5 "2+2": text only, thinking_tokens: 0).
  • Fable 5.x may refuse attempts to extract reasoning: stop_details.category: "reasoning_extraction".

# 4. Signatures, preservation and tool use

Rules

  1. Within a tool-use turn you must pass every thinking/redacted_thinking block back, complete and unmodified, alongside the tool_use it accompanied. Recommended across turns; allowed to omit outside tool use.
  2. Manual mode requires the last assistant turn of a thinking request to start with a thinking block; adaptive mode relaxes this.
  3. tool_choice any/tool is incompatible with manual thinking (400 Thinking may not be enabled when tool_choice forces tool use.); allowed with adaptive except Fable 5.1 / Mythos 5.1 (reject forced tool use on every request).
  4. Toggling thinking mid-turn does not error: the API silently disables thinking for that request (check for thinking blocks).
  5. Signatures are opaque, cross-platform (API/Bedrock/Vertex), identical whatever display is; a block is readable only by the model that produced it or a newer one — older models silently drop it (unbilled).
  6. Prefix (conversation) check — Fable 5.1 / Mythos 5.1: a replayed block is valid only if system, tools and all earlier messages are unchanged. Enforced by default for accounts created ≥ 2026-08-31; opt in with thinking.block_binding.prefix_mismatch_behavior: "error" | "drop_block" under thinking-binding-controls-2026-08-01, which also adds input_transformations[] (type: thinking_dropped | thinking_mismatch_allowed, reason: prefix_binding_mismatch | model_binding_mismatch, path). Live (Sonnet 5): accepted, input_transformations: [].
  7. Preservation default: Opus 4.5+, Sonnet 4.6+, Fable/Mythos keep all prior-turn thinking (billed as input, cache-friendly); Haiku 4.5, Sonnet 4.5 and earlier keep only the last turn (older blocks stripped for free). Override with context editing clear_thinking_20251015.

Live round trip (Haiku 4.5, budget 1024): turn 1 → [thinking, tool_use], stop_reason: tool_use, 47 thinking tokens. Turn 2 with the thinking block + tool_result → [text], 200. Observed deviations from the docs: (a) the same turn 2 with the thinking text edited (signature unchanged) also returned 200 (docs: "Modified thinking blocks are rejected with a 400"); (b) turn 2 with the thinking block removed returned 200; (c) assistant prefill with thinking enabled returned 200 with thinking silently disabled (docs: prefill not allowed while thinking is on). Treat these as Haiku-specific graceful degradation; do not rely on them.

# 5. Interleaved thinking

  • Adaptive models: automatic, no header (Fable/Mythos, Opus 4.7+, Opus 5, Sonnet 5, Opus/Sonnet 4.6 in adaptive mode).
  • Manual mode: header interleaved-thinking-2025-05-14 on Opus 4.5, Sonnet 4.5, Claude 4; Sonnet 4.6 manual+header deprecated; Opus 4.6 manual mode never interleaves; Haiku 4.5 does not support it (header accepted and ignored — live: 200, ordinary [thinking, text]).
  • The header is accepted on any model on the Claude API and ignored where unsupported. Only for tools used through the Messages API.
  • Progress updates (Fable 5.x): extra thinking blocks with their own signature right before each tool_use; text visible only under display: "updates"/"summarized"; last block text This part of the response was interrupted before it finished. when the response stops early.

# 6. Streaming

Observed sequence (Haiku, enabled): message_start → content_block_start{thinking} → ping → 10× content_block_delta{thinking_delta} → content_block_delta{signature_delta} → content_block_stop → content_block_start{text} → text_delta → content_block_stop → message_delta{usage.output_tokens_details.thinking_tokens: 33} → message_stop. With display: "omitted": one thinking_delta with "", then signature_delta. Chunky delivery is expected. Use stream.get_final_message() / finalMessage() to reassemble.

# 7. Incompatibilities and errors (live text)

Request Result
temperature: 0.5 + thinking (Haiku) 400 temperature may only be set to 1 when thinking is enabled.
top_k: 5 + thinking (Haiku) 400 top_k must be unset when thinking is enabled.
top_p + thinking (≤4.6) allowed only in [0.95, 1]
temperature: 0.2 on Sonnet 5 (no thinking param) 400 temperature is deprecated for this model. — 4.7+/5.x/Fable reject non-default sampling on every request
budget_tokens: 512 400 … greater than or equal to 1024
budget_tokens ≥ max_tokens 400 max_tokens must be greater than thinking.budget_tokens
type: adaptive on Haiku 400 adaptive thinking is not supported on this model
type: enabled on Sonnet 5 400 "thinking.type.enabled" is not supported for this model…
tool_choice: any + enabled 400 Thinking may not be enabled when tool_choice forces tool use.
display: "updates" without header 400 thinking.adaptive.display: Input should be 'summarized', 'omitted'
block_binding without header 400 … block_binding: Extra inputs are not permitted (documented)
Modified thinking (4.6+/Fable) 400 "thinking blocks cannot be modified" (documented; Haiku accepted, see §4)
Invalid signature 400 Invalid signature in thinking block

# 8. Context window & pricing

  • Current-turn thinking counts toward max_tokens, is billed as output and occupies window space. Prior-turn thinking: input tokens on keep-all models; stripped (free) on last-turn models.
  • Input + max_tokens > window on 4.5+ → accepted; generation stops with stop_reason: "model_context_window_exceeded".
  • Max output: 128K on 4.6+/5.x/Fable; 64K on 4.5 models; batches beta output-300k-2026-03-24 → 300K on Opus 4.6–5 / Sonnet 4.6–5.
  • A specialized system prompt is injected when thinking is active (small input-token increase).

# 9. Migration (enabled → adaptive)

Remove budget_tokens, set thinking: {"type": "adaptive"}, steer with output_config.effort (default high == omitted). Drop interleaved-thinking-2025-05-14. Expect Claude to skip thinking on easy inputs at lower effort. First request after the switch restarts the prompt cache.

# 10. Examples & tests

examples/anthropic/thinking/ · tests/anthropic/test_thinking.py.