SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
6.7 KB

# Anthropic — Effort (output_config.effort) and task budgets

Status: DOCUMENTED · LIVE_VERIFIED (claude-sonnet-5: low/xhigh/max accepted; invalid value 400; claude-haiku-4-5 400) · per-message effort BETA (mid-conversation-output-config-2026-07-01) · task budgets BETA (task-budgets-2026-03-13). Sources: Effort · Steering thinking · Task budgets · Mid-conversation system messages · Beta Messages API · Models API capabilities.effort Last verified: 2026-09-18

# Location and values

  • Exact location: output_config.effort (top level of the request body). No beta header. The value is a string: low | medium | high | xhigh | max. adaptive is not an effort value.
  • Default high; setting high explicitly is identical to omitting it (and does not invalidate the prompt cache).
  • Invalid value → 400 output_config.effort: Input should be 'low', 'medium', 'high', 'xhigh' or 'max' (live).
  • Unsupported model → 400 This model does not support the effort parameter. (live, Haiku 4.5).

# Per-model support (docs + Models API capabilities.effort.* + live)

Model low medium high xhigh max Default Per-message effort (beta) Notes
Fable 5.1 / Mythos 5.1 ✔ ✔ ✔ ✔ ✔ high ✔ prefer per-message changes; fewer progress updates at higher effort
Fable 5 / Mythos 5 ✔ ✔ ✔ ✔ ✔ high ✘ (400 output_config.effort requires a model that supports per-turn effort) effort is the primary quality/latency lever
Opus 5 ✔ ✔ ✔ ✔ ✔ high ✔ thinking: disabled + xhigh/max → 400
Opus 4.8 / 4.7 ✔ ✔ ✔ ✔ ✔ high ✘ start at xhigh for coding/agentic work; 64k max_tokens at xhigh/max
Sonnet 5 ✔ ✔ ✔ ✔ ✔ high ✘ (live 400) live: low/xhigh/max 200; disabled thinking + max accepted (restriction is Opus 5+)
Sonnet 4.6 / Opus 4.6 ✔ ✔ ✔ ✘ ✔ high ✘ Sonnet 4.6: set medium explicitly for latency
Opus 4.5 ✔ ✔ ✔ ✘ ✘ high ✘ only extended-thinking model with effort; composes with budget_tokens
Sonnet 4.5 / Haiku 4.5 ✘ ✘ ✘ ✘ ✘ — ✘ 400

# What effort does

Level Thinking behaviour (adaptive) Typical use
max always thinks, unconstrained depth frontier problems; may overthink structured tasks
xhigh always thinks deeply, extended exploration >30-minute agentic/coding runs, token budgets in the millions
high (default) almost always thinks complex reasoning, agentic tasks
medium moderate; may skip thinking on simple queries balanced agentic work
low minimizes thinking; terser and fewer tool calls subagents, latency-sensitive chat

Effort affects all output tokens (text, tool-call arguments, thinking) and works with thinking disabled too. It is a behavioural signal, not a token budget; max_tokens remains the hard cap. On Opus 5 effort does not reliably shorten visible responses — prompt for length instead.

# Interactions

With Behaviour
Thinking thinking decides whether thinking blocks exist; effort decides how much work. Opus 5+: cannot disable thinking at xhigh/max. Opus 4.5: effort + budget_tokens both apply.
Prompt caching The resolved effort is rendered into the prompt: changing the top-level value between requests invalidates message breakpoints (and tool/system breakpoints on some models). Explicit default = no change.
Per-message effort (beta) {"role": "system", "content": [], "output_config": {"effort": "low"}} inside messages + header mid-conversation-output-config-2026-07-01; takes effect from the next user turn, keeps the cached prefix; allowed anywhere in messages. Fable 5.1 / Mythos 5.1 / Opus 5 only (Claude API + Vertex). Live on Sonnet 5: 400 as documented.
Tools Lower effort → fewer, combined tool calls, no preamble. Higher → more calls, plans and summaries.
Task budgets Effort tunes depth per step; task budget bounds total work across the loop.
Managed Agents effort can be set inside the agent's model object.

# Task budgets (beta task-budgets-2026-03-13)

json
"output_config": {"task_budget": {"type": "tokens", "total": 200000, "remaining": 150000}}
  • Advisory countdown injected server-side for one agentic turn (a user message without tool_result starts a new turn). Includes thinking, tool calls, tool results, output. Not a hard cap; max_tokens still truncates per request.
  • Documented minimum total 20,000 (the beta reference schema says minimum: 1024 — discrepancy); values too small cause refusal-like behaviour.
  • remaining only when you compact client-side; a changed value invalidates the cache; not allowed together with the compaction parameter or a signed compaction block (400).
  • Supported: Fable 5.1, Mythos 5.1, Opus 5, Fable 5, Mythos 5, Opus 4.8, Opus 4.7. Documented as unsupported on Sonnet 5 / 4.6 / Opus 4.6 / Haiku — live: Sonnet 5 accepted task_budget (200, no error) and the response gained usage.iterations; unknown whether the countdown was applied. LIVE_DISCOVERED.
  • The response never reports the remaining budget; sum usage.output_tokens + tool-result tokens client-side.

# Live log (2026-09-18)

Call Model Result
effort low / xhigh / max ("Reply with OK.") claude-sonnet-5 200, OK, 4 output tokens, thinking_tokens: 0
adaptive + effort low vs max ("What is 2+2?", example adaptive_effort.ts) claude-sonnet-5 low → text only, no thinking block; max → [thinking, text], 49 thinking tokens — effort visibly changes whether Claude thinks
effort ultra claude-sonnet-5 400 Input should be 'low', 'medium', 'high', 'xhigh' or 'max'
effort low claude-haiku-4-5-20251001 400 This model does not support the effort parameter.
effort max + thinking: disabled claude-sonnet-5 200
per-message effort + beta header claude-sonnet-5 400 output_config.effort requires a model that supports per-turn effort; this model does not
task_budget total 20000 + beta header claude-sonnet-5 200, usage.iterations: [{type: "message", …}]

Examples: examples/anthropic/thinking/ (adaptive + effort). Tests: tests/anthropic/test_thinking.py.