Anthropic — Thinking (extended, adaptive, interleaved, preserved)
Status: DOCUMENTED · LIVE_VERIFIED (extended thinking on claude-haiku-4-5-20251001; adaptive on claude-sonnet-5; interleaved header on claude-sonnet-4-6; 2026-09-18). display: "updates" and block_binding are BETA.
Sources: Thinking · Extended thinking · Steering thinking · Troubleshooting · Preserved thinking · Context windows · Messages API · Beta Messages API
Last verified: 2026-09-18
1. Modes
| Mode | Request | Behaviour |
|---|---|---|
| Extended (manual) | thinking: {"type": "enabled", "budget_tokens": N[, "display"]} |
Claude always thinks against a target budget. Only mode on 4.5-and-earlier models; deprecated on Opus/Sonnet 4.6; 400 on 4.7+, Sonnet 5, Opus 5, Fable/Mythos ("thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.) |
| Adaptive | thinking: {"type": "adaptive"[, "display"]} |
Claude decides whether/how deeply to think; depth steered by output_config.effort. Interleaves between tool calls automatically. 400 on 4.5 models (adaptive thinking is not supported on this model). |
| Disabled | thinking: {"type": "disabled"} |
Off. 400 on Fable/Mythos (always on); 400 on Opus 5 combined with effort xhigh/max. |
| Omitted parameter | — | Always-on on Fable 5.x/Mythos; on by default on Opus 5 / Sonnet 5; off on Opus 4.6–4.8, Sonnet 4.6 and all 4.5 models. |
Per-model table (docs + Models API capabilities.thinking.types + live)
| Model | Types accepted | Default | Rejected (400) | display default |
Prior-turn thinking kept | Live |
|---|---|---|---|---|---|---|
| Fable 5.1 / Mythos 5.1 | adaptive | always on | enabled, disabled |
omitted | all | — |
| Fable 5 / Mythos 5 | adaptive | always on | enabled, disabled |
omitted | all | — |
| Mythos Preview | adaptive, extended | always on | disabled |
omitted | all | — |
| Opus 5 | adaptive | on | enabled; disabled at xhigh/max |
omitted | all | disabled accepted (fast_04) |
| Opus 4.8 / 4.7 | adaptive | off | enabled |
omitted | all | — |
| Sonnet 5 | adaptive | on | enabled |
omitted | all | adaptive ✔, disabled ✔, enabled ✘ 400 |
| Opus 4.6 / Sonnet 4.6 | adaptive, extended (deprecated) | off | none | summarized | all | Sonnet 4.6 enabled + interleaved header ✔ |
| Opus 4.5 | extended | off | adaptive |
summarized | all | — |
| Sonnet 4.5 | extended | off | adaptive |
summarized | last turn | — |
| Haiku 4.5 | extended | off | adaptive |
summarized | last turn | enabled ✔, adaptive ✘ 400 |
2. Budget rules (manual mode)
budget_tokens≥ 1024 (400 thinking.enabled.budget_tokens: Input should be greater than or equal to 1024) and <max_tokens(400 max_tokens must be greater than thinking.budget_tokens) — except with interleaved thinking where the budget spans the whole turn and may exceedmax_tokens.- The budget is a target, not a cap;
max_tokensis the hard ceiling (thinking + text). Live: budget 1024 on "2+2" used 30 thinking tokens. max_tokens: 0(cache pre-warm) is incompatible with enabled thinking.- Budgets > 32k → use Batches (timeouts). SDKs require streaming when
max_tokens> 21,333. - On Opus 4.5 effort composes with
budget_tokens. - Adaptive mode: no budget; bound cost with
max_tokens(hard) andeffort(soft).stop_reason: "max_tokens"→ raisemax_tokensor lower effort.
3. Response shape
{"type": "thinking", "thinking": "The user is asking for 2+2, which equals 4…", "signature": "EtQCCpoBCBEYAipAGx2p…"}
{"type": "text", "text": "4"}thinkingtext is a summary (never raw chain of thought).display: "omitted"→thinking: ""with the same signature (faster TTFT when streaming, same billing).display: "updates"(betathinking-display-updates-2026-08-18) → reasoning blocks empty, progress-update blocks (Fable 5.x) carry text; without the header:400 thinking.adaptive.display: Input should be 'summarized', 'omitted'.redacted_thinking:{"type": "redacted_thinking", "data": "<opaque>"}— safety-redacted; pass back unchanged.usage.output_tokens_details.thinking_tokensreports billed reasoning tokens (live: 30 of 37 output tokens). Streaming: only on the finalmessage_delta.- Adaptive mode may return no thinking block (live Sonnet 5 "2+2": text only,
thinking_tokens: 0). - Fable 5.x may refuse attempts to extract reasoning:
stop_details.category: "reasoning_extraction".
4. Signatures, preservation and tool use
Rules
- Within a tool-use turn you must pass every
thinking/redacted_thinkingblock back, complete and unmodified, alongside thetool_useit accompanied. Recommended across turns; allowed to omit outside tool use. - Manual mode requires the last assistant turn of a thinking request to start with a thinking block; adaptive mode relaxes this.
tool_choiceany/toolis incompatible with manual thinking (400 Thinking may not be enabled when tool_choice forces tool use.); allowed with adaptive except Fable 5.1 / Mythos 5.1 (reject forced tool use on every request).- Toggling thinking mid-turn does not error: the API silently disables thinking for that request (check for thinking blocks).
- Signatures are opaque, cross-platform (API/Bedrock/Vertex), identical whatever
displayis; a block is readable only by the model that produced it or a newer one — older models silently drop it (unbilled). - Prefix (conversation) check — Fable 5.1 / Mythos 5.1: a replayed block is valid only if
system,toolsand all earlier messages are unchanged. Enforced by default for accounts created ≥ 2026-08-31; opt in withthinking.block_binding.prefix_mismatch_behavior: "error" | "drop_block"underthinking-binding-controls-2026-08-01, which also addsinput_transformations[](type: thinking_dropped | thinking_mismatch_allowed,reason: prefix_binding_mismatch | model_binding_mismatch,path). Live (Sonnet 5): accepted,input_transformations: []. - Preservation default: Opus 4.5+, Sonnet 4.6+, Fable/Mythos keep all prior-turn thinking (billed as input, cache-friendly); Haiku 4.5, Sonnet 4.5 and earlier keep only the last turn (older blocks stripped for free). Override with context editing
clear_thinking_20251015.
Live round trip (Haiku 4.5, budget 1024): turn 1 → [thinking, tool_use], stop_reason: tool_use, 47 thinking tokens. Turn 2 with the thinking block + tool_result → [text], 200. Observed deviations from the docs: (a) the same turn 2 with the thinking text edited (signature unchanged) also returned 200 (docs: "Modified thinking blocks are rejected with a 400"); (b) turn 2 with the thinking block removed returned 200; (c) assistant prefill with thinking enabled returned 200 with thinking silently disabled (docs: prefill not allowed while thinking is on). Treat these as Haiku-specific graceful degradation; do not rely on them.
5. Interleaved thinking
- Adaptive models: automatic, no header (Fable/Mythos, Opus 4.7+, Opus 5, Sonnet 5, Opus/Sonnet 4.6 in adaptive mode).
- Manual mode: header
interleaved-thinking-2025-05-14on Opus 4.5, Sonnet 4.5, Claude 4; Sonnet 4.6 manual+header deprecated; Opus 4.6 manual mode never interleaves; Haiku 4.5 does not support it (header accepted and ignored — live: 200, ordinary[thinking, text]). - The header is accepted on any model on the Claude API and ignored where unsupported. Only for tools used through the Messages API.
- Progress updates (Fable 5.x): extra
thinkingblocks with their own signature right before eachtool_use; text visible only underdisplay: "updates"/"summarized"; last block textThis part of the response was interrupted before it finished.when the response stops early.
6. Streaming
Observed sequence (Haiku, enabled): message_start → content_block_start{thinking} → ping → 10× content_block_delta{thinking_delta} → content_block_delta{signature_delta} → content_block_stop → content_block_start{text} → text_delta → content_block_stop → message_delta{usage.output_tokens_details.thinking_tokens: 33} → message_stop.
With display: "omitted": one thinking_delta with "", then signature_delta. Chunky delivery is expected. Use stream.get_final_message() / finalMessage() to reassemble.
7. Incompatibilities and errors (live text)
| Request | Result |
|---|---|
temperature: 0.5 + thinking (Haiku) |
400 temperature may only be set to 1 when thinking is enabled. |
top_k: 5 + thinking (Haiku) |
400 top_k must be unset when thinking is enabled. |
top_p + thinking (≤4.6) |
allowed only in [0.95, 1] |
temperature: 0.2 on Sonnet 5 (no thinking param) |
400 temperature is deprecated for this model. — 4.7+/5.x/Fable reject non-default sampling on every request |
budget_tokens: 512 |
400 … greater than or equal to 1024 |
budget_tokens ≥ max_tokens |
400 max_tokens must be greater than thinking.budget_tokens |
type: adaptive on Haiku |
400 adaptive thinking is not supported on this model |
type: enabled on Sonnet 5 |
400 "thinking.type.enabled" is not supported for this model… |
tool_choice: any + enabled |
400 Thinking may not be enabled when tool_choice forces tool use. |
display: "updates" without header |
400 thinking.adaptive.display: Input should be 'summarized', 'omitted' |
block_binding without header |
400 … block_binding: Extra inputs are not permitted (documented) |
| Modified thinking (4.6+/Fable) | 400 "thinking blocks cannot be modified" (documented; Haiku accepted, see §4) |
| Invalid signature | 400 Invalid signature in thinking block |
8. Context window & pricing
- Current-turn thinking counts toward
max_tokens, is billed as output and occupies window space. Prior-turn thinking: input tokens on keep-all models; stripped (free) on last-turn models. - Input +
max_tokens> window on 4.5+ → accepted; generation stops withstop_reason: "model_context_window_exceeded". - Max output: 128K on 4.6+/5.x/Fable; 64K on 4.5 models; batches beta
output-300k-2026-03-24→ 300K on Opus 4.6–5 / Sonnet 4.6–5. - A specialized system prompt is injected when thinking is active (small input-token increase).
9. Migration (enabled → adaptive)
Remove budget_tokens, set thinking: {"type": "adaptive"}, steer with output_config.effort (default high == omitted). Drop interleaved-thinking-2025-05-14. Expect Claude to skip thinking on easy inputs at lower effort. First request after the switch restarts the prompt cache.
10. Examples & tests
examples/anthropic/thinking/ · tests/anthropic/test_thinking.py.