# Anthropic — Thinking (extended, adaptive, interleaved, preserved) **Status:** `DOCUMENTED` · `LIVE_VERIFIED` (extended thinking on claude-haiku-4-5-20251001; adaptive on claude-sonnet-5; interleaved header on claude-sonnet-4-6; 2026-09-18). `display: "updates"` and `block_binding` are `BETA`. **Sources:** [Thinking](https://platform.claude.com/docs/en/build-with-claude/thinking) · [Extended thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) · [Steering thinking](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost) · [Troubleshooting](https://platform.claude.com/docs/en/build-with-claude/thinking-troubleshooting) · [Preserved thinking](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking) · [Context windows](https://platform.claude.com/docs/en/build-with-claude/context-windows) · [Messages API](https://platform.claude.com/docs/en/api/messages/create) · [Beta Messages API](https://platform.claude.com/docs/en/api/beta/messages) **Last verified:** 2026-09-18 ## 1. Modes | Mode | Request | Behaviour | |---|---|---| | Extended (manual) | `thinking: {"type": "enabled", "budget_tokens": N[, "display"]}` | Claude always thinks against a target budget. Only mode on 4.5-and-earlier models; **deprecated** on Opus/Sonnet 4.6; **400** on 4.7+, Sonnet 5, Opus 5, Fable/Mythos (`"thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.`) | | Adaptive | `thinking: {"type": "adaptive"[, "display"]}` | Claude decides whether/how deeply to think; depth steered by `output_config.effort`. Interleaves between tool calls automatically. 400 on 4.5 models (`adaptive thinking is not supported on this model`). | | Disabled | `thinking: {"type": "disabled"}` | Off. 400 on Fable/Mythos (always on); 400 on Opus 5 combined with effort `xhigh`/`max`. | | Omitted parameter | — | Always-on on Fable 5.x/Mythos; on by default on Opus 5 / Sonnet 5; off on Opus 4.6–4.8, Sonnet 4.6 and all 4.5 models. | ### Per-model table (docs + Models API `capabilities.thinking.types` + live) | Model | Types accepted | Default | Rejected (400) | `display` default | Prior-turn thinking kept | Live | |---|---|---|---|---|---|---| | Fable 5.1 / Mythos 5.1 | adaptive | always on | `enabled`, `disabled` | omitted | all | — | | Fable 5 / Mythos 5 | adaptive | always on | `enabled`, `disabled` | omitted | all | — | | Mythos Preview | adaptive, extended | always on | `disabled` | omitted | all | — | | Opus 5 | adaptive | on | `enabled`; `disabled` at xhigh/max | omitted | all | `disabled` accepted (fast_04) | | Opus 4.8 / 4.7 | adaptive | off | `enabled` | omitted | all | — | | Sonnet 5 | adaptive | on | `enabled` | omitted | all | adaptive ✔, disabled ✔, enabled ✘ 400 | | Opus 4.6 / Sonnet 4.6 | adaptive, extended (deprecated) | off | none | summarized | all | Sonnet 4.6 enabled + interleaved header ✔ | | Opus 4.5 | extended | off | `adaptive` | summarized | all | — | | Sonnet 4.5 | extended | off | `adaptive` | summarized | last turn | — | | Haiku 4.5 | extended | off | `adaptive` | summarized | last turn | enabled ✔, adaptive ✘ 400 | ## 2. Budget rules (manual mode) * `budget_tokens` ≥ **1024** (`400 thinking.enabled.budget_tokens: Input should be greater than or equal to 1024`) and **< `max_tokens`** (`400 max_tokens must be greater than thinking.budget_tokens`) — except with interleaved thinking where the budget spans the whole turn and may exceed `max_tokens`. * The budget is a target, not a cap; `max_tokens` is the hard ceiling (thinking + text). Live: budget 1024 on "2+2" used 30 thinking tokens. * `max_tokens: 0` (cache pre-warm) is incompatible with enabled thinking. * Budgets > 32k → use Batches (timeouts). SDKs require streaming when `max_tokens` > 21,333. * On Opus 4.5 effort composes with `budget_tokens`. * Adaptive mode: no budget; bound cost with `max_tokens` (hard) and `effort` (soft). `stop_reason: "max_tokens"` → raise `max_tokens` or lower effort. ## 3. Response shape ```json {"type": "thinking", "thinking": "The user is asking for 2+2, which equals 4…", "signature": "EtQCCpoBCBEYAipAGx2p…"} {"type": "text", "text": "4"} ``` * `thinking` text is a **summary** (never raw chain of thought). `display: "omitted"` → `thinking: ""` with the same signature (faster TTFT when streaming, same billing). `display: "updates"` (beta `thinking-display-updates-2026-08-18`) → reasoning blocks empty, progress-update blocks (Fable 5.x) carry text; without the header: `400 thinking.adaptive.display: Input should be 'summarized', 'omitted'`. * `redacted_thinking`: `{"type": "redacted_thinking", "data": ""}` — safety-redacted; pass back unchanged. * `usage.output_tokens_details.thinking_tokens` reports billed reasoning tokens (live: 30 of 37 output tokens). Streaming: only on the final `message_delta`. * Adaptive mode may return **no** thinking block (live Sonnet 5 "2+2": text only, `thinking_tokens: 0`). * Fable 5.x may refuse attempts to extract reasoning: `stop_details.category: "reasoning_extraction"`. ## 4. Signatures, preservation and tool use **Rules** 1. Within a tool-use turn you **must** pass every `thinking`/`redacted_thinking` block back, complete and unmodified, alongside the `tool_use` it accompanied. Recommended across turns; allowed to omit outside tool use. 2. Manual mode requires the last assistant turn of a thinking request to start with a thinking block; adaptive mode relaxes this. 3. `tool_choice` `any`/`tool` is incompatible with manual thinking (`400 Thinking may not be enabled when tool_choice forces tool use.`); allowed with adaptive except Fable 5.1 / Mythos 5.1 (reject forced tool use on every request). 4. Toggling thinking mid-turn does not error: the API silently disables thinking for that request (check for thinking blocks). 5. Signatures are opaque, cross-platform (API/Bedrock/Vertex), identical whatever `display` is; a block is readable only by the model that produced it or a newer one — older models silently drop it (unbilled). 6. **Prefix (conversation) check** — Fable 5.1 / Mythos 5.1: a replayed block is valid only if `system`, `tools` and all earlier messages are unchanged. Enforced by default for accounts created ≥ 2026-08-31; opt in with `thinking.block_binding.prefix_mismatch_behavior: "error" | "drop_block"` under `thinking-binding-controls-2026-08-01`, which also adds `input_transformations[]` (`type: thinking_dropped | thinking_mismatch_allowed`, `reason: prefix_binding_mismatch | model_binding_mismatch`, `path`). Live (Sonnet 5): accepted, `input_transformations: []`. 7. Preservation default: Opus 4.5+, Sonnet 4.6+, Fable/Mythos keep **all** prior-turn thinking (billed as input, cache-friendly); Haiku 4.5, Sonnet 4.5 and earlier keep only the **last turn** (older blocks stripped for free). Override with context editing `clear_thinking_20251015`. **Live round trip (Haiku 4.5, budget 1024):** turn 1 → `[thinking, tool_use]`, `stop_reason: tool_use`, 47 thinking tokens. Turn 2 with the thinking block + `tool_result` → `[text]`, 200. **Observed deviations from the docs:** (a) the same turn 2 with the thinking *text edited* (signature unchanged) also returned 200 (docs: "Modified thinking blocks are rejected with a 400"); (b) turn 2 with the thinking block *removed* returned 200; (c) assistant **prefill** with thinking enabled returned 200 with thinking silently disabled (docs: prefill not allowed while thinking is on). Treat these as Haiku-specific graceful degradation; do not rely on them. ## 5. Interleaved thinking * Adaptive models: automatic, no header (Fable/Mythos, Opus 4.7+, Opus 5, Sonnet 5, Opus/Sonnet 4.6 in adaptive mode). * Manual mode: header `interleaved-thinking-2025-05-14` on Opus 4.5, Sonnet 4.5, Claude 4; Sonnet 4.6 manual+header deprecated; Opus 4.6 manual mode never interleaves; **Haiku 4.5 does not support it** (header accepted and ignored — live: 200, ordinary `[thinking, text]`). * The header is accepted on any model on the Claude API and ignored where unsupported. Only for tools used through the Messages API. * Progress updates (Fable 5.x): extra `thinking` blocks with their own signature right before each `tool_use`; text visible only under `display: "updates"`/`"summarized"`; last block text `This part of the response was interrupted before it finished.` when the response stops early. ## 6. Streaming Observed sequence (Haiku, `enabled`): `message_start` → `content_block_start{thinking}` → `ping` → 10× `content_block_delta{thinking_delta}` → `content_block_delta{signature_delta}` → `content_block_stop` → `content_block_start{text}` → `text_delta` → `content_block_stop` → `message_delta{usage.output_tokens_details.thinking_tokens: 33}` → `message_stop`. With `display: "omitted"`: one `thinking_delta` with `""`, then `signature_delta`. Chunky delivery is expected. Use `stream.get_final_message()` / `finalMessage()` to reassemble. ## 7. Incompatibilities and errors (live text) | Request | Result | |---|---| | `temperature: 0.5` + thinking (Haiku) | `400 temperature may only be set to 1 when thinking is enabled.` | | `top_k: 5` + thinking (Haiku) | `400 top_k must be unset when thinking is enabled.` | | `top_p` + thinking (≤4.6) | allowed only in [0.95, 1] | | `temperature: 0.2` on Sonnet 5 (no thinking param) | `400 temperature is deprecated for this model.` — 4.7+/5.x/Fable reject non-default sampling on every request | | `budget_tokens: 512` | `400 … greater than or equal to 1024` | | `budget_tokens ≥ max_tokens` | `400 max_tokens must be greater than thinking.budget_tokens` | | `type: adaptive` on Haiku | `400 adaptive thinking is not supported on this model` | | `type: enabled` on Sonnet 5 | `400 "thinking.type.enabled" is not supported for this model…` | | `tool_choice: any` + enabled | `400 Thinking may not be enabled when tool_choice forces tool use.` | | `display: "updates"` without header | `400 thinking.adaptive.display: Input should be 'summarized', 'omitted'` | | `block_binding` without header | `400 … block_binding: Extra inputs are not permitted` (documented) | | Modified thinking (4.6+/Fable) | 400 "thinking blocks cannot be modified" (documented; Haiku accepted, see §4) | | Invalid signature | `400 Invalid signature in thinking block` | ## 8. Context window & pricing * Current-turn thinking counts toward `max_tokens`, is billed as output and occupies window space. Prior-turn thinking: input tokens on keep-all models; stripped (free) on last-turn models. * Input + `max_tokens` > window on 4.5+ → accepted; generation stops with `stop_reason: "model_context_window_exceeded"`. * Max output: 128K on 4.6+/5.x/Fable; 64K on 4.5 models; batches beta `output-300k-2026-03-24` → 300K on Opus 4.6–5 / Sonnet 4.6–5. * A specialized system prompt is injected when thinking is active (small input-token increase). ## 9. Migration (enabled → adaptive) Remove `budget_tokens`, set `thinking: {"type": "adaptive"}`, steer with `output_config.effort` (default `high` == omitted). Drop `interleaved-thinking-2025-05-14`. Expect Claude to skip thinking on easy inputs at lower effort. First request after the switch restarts the prompt cache. ## 10. Examples & tests `examples/anthropic/thinking/` · `tests/anthropic/test_thinking.py`.