SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
10.3 KB

# Anthropic — Context management: windows, 1M, context editing, compaction

Status: context windows DOCUMENTED/LIVE_VERIFIED · context editing BETA (context-management-2025-06-27) LIVE_VERIFIED (Haiku) · threshold compaction BETA (compact-2026-01-12) LIVE_VERIFIED (config accepted on Sonnet 5; no compaction triggered — needs ≥ 50k tokens) · on-demand compaction BETA (compact-2026-09-04) LIVE_VERIFIED (Sonnet 5 round trip). Sources: Context windows · Context editing · Compaction · Task budgets · Beta Messages API · Release notes Last verified: 2026-09-18

# 1. Context windows

Model family Window Max output Media per request Header Pricing
Fable 5.1/5, Mythos 5.1/5/Preview, Opus 5, Opus 4.8/4.7/4.6, Sonnet 5, Sonnet 4.6 1M (default) 128K (300K in Batches with output-300k-2026-03-24) 600 images / PDF pages none standard (no long-context premium)
Opus 4.5, Sonnet 4.5, Haiku 4.5 200K 64K 100 — —
  • context-1m-2025-08-07 is legacy: retired 2026-04-30 for Sonnet 4.5/Sonnet 4 (no effect; > 200k → error). Live: header accepted and ignored on Sonnet 5 and Haiku 4.5 (200).
  • Everything counts: system, messages (tool results, images, documents), tool definitions, output incl. thinking. Cached tokens still occupy the window (input + cache_read + cache_creation).
  • Overflow: input alone > window → 400 prompt is too long. Input + max_tokens > window on 4.5+ → accepted, stops with stop_reason: "model_context_window_exceeded" (older models: validation error unless model-context-window-exceeded-2025-08-26).
  • Context awareness (Sonnet 5, Sonnet 4.6, Sonnet 4.5, Haiku 4.5): the API injects the remaining-token budget after each tool call; not on Opus 4.7+/Fable (use task budgets there).
  • Thinking: current-turn thinking → max_tokens; prior-turn thinking kept (input tokens) on Opus 4.5+/Sonnet 4.6+/Fable, stripped on Haiku/Sonnet 4.5.

# 2. Context editing (context_management.edits, beta context-management-2025-06-27)

Server-side: your client keeps the full history; the API edits the prompt before the model sees it. Exact strategy strings (verified in SDK types BetaClearToolUses20250919Edit, BetaClearThinking20251015Edit, BetaCompact20260112Edit):

Strategy type Fields Defaults Notes
clear_tool_uses_20250919 trigger: {type: input_tokens|tool_uses, value ≥ 1}, keep: {type: tool_uses, value}, clear_at_least: {type: input_tokens, value}, exclude_tools: [names], clear_tool_inputs: bool | [names] trigger 100,000 input tokens; keep 3; clear results only clears oldest results first, replaces with placeholder text; breaks the cache at the clear point (use clear_at_least)
clear_thinking_20251015 keep: {type: thinking_turns, value ≥ 1} | {type: all} | "all" model-specific (all on Opus 4.5+/Sonnet 4.6+/Fable; last turn on Haiku & earlier) must be first in edits when combined; requires thinking enabled/adaptive (400 clear_thinking_20251015 strategy requires thinking to be enabled or adaptive, live)
compact_20260112 trigger: {type: input_tokens, value ≥ 50000}, pause_after_compaction: bool, instructions: string trigger 150,000 beta header compact-2026-01-12; 4.6+ models only (400 'claude-haiku-4-5-20251001' does not support the 'compact_20260112' context management strategy.); value 1000 → 400 trigger.value must be at least 50000

Response field:

json
"context_management": {"applied_edits": [
  {"type": "clear_tool_uses_20250919", "cleared_tool_uses": 2, "cleared_input_tokens": 174},
  {"type": "clear_thinking_20251015", "cleared_thinking_turns": 3, "cleared_input_tokens": 15000}]}

Streaming: in the final message_delta. POST /v1/messages/count_tokens accepts context_management and returns {"input_tokens": 838, "context_management": {"original_input_tokens": 928}} (live). Without the header: 400 context_management: Extra inputs are not permitted. Works with the memory tool (Claude is warned before clearing).

Live (Haiku): 3 fake tool round-trips, trigger input_tokens 1, keep tool_uses 1, clear_at_least 1 → applied_edits: [{clear_tool_uses_20250919, cleared_tool_uses: 2, cleared_input_tokens: 174}], 200.

# 3. Compaction

# Threshold compaction (compact_20260112)

  1. Input tokens reach trigger → the API summarizes, emits a compaction block at the start of the assistant response, continues.
  2. Pass the whole response back; everything before the last compaction block is ignored. pause_after_compaction: true → stop_reason: "compaction" so you can inject content first.
  3. Default prompt writes a <summary> for continuation; instructions replaces it entirely (tell the model not to call tools — otherwise content: null). On Fable 5.1 custom instructions summarize the visible conversation only.
  4. Streaming: content_block_start → one content_block_delta with the whole summary → content_block_stop.
  5. usage.iterations[] lists {type: "compaction"} and {type: "message"} entries; top-level input/output_tokens exclude the compaction iteration → sum iterations for billing. With the beta header every response carries iterations (live Sonnet 5: [{type: "message", …}] with no compaction).
  6. Same model summarizes; cache_control allowed on the block; keep a system breakpoint so only the summary is rewritten. Token counting applies existing blocks but never triggers new ones. Server tools: trigger checked at every sampling iteration.
  7. Images, documents, container_upload blocks and fetched URLs inside the summarized range are lost.

# On-demand compaction (compaction: {"type": "summarize"}, beta compact-2026-09-04)

  • Separate request (same system/tools/thinking/max_tokens as the conversation) → response = one signed compaction block (content, signature), stop_reason: "compaction", top-level usage 0, usage.iterations: [{type: "compaction", input_tokens: 90, output_tokens: 116}] (live). Can run in the background.
  • Continue: send the block first in messages (own assistant message or first block of the first message) in place of the summarized messages, header on every request, exactly one block. Live: [assistant{compaction}, user"What is my name…"] → "Ada — favourite color: teal." Summarized messages left in front → 400 messages.1.content.0: compaction block must be sent first, in place of the messages it summarizes; remove those messages (compaction_block_misplaced).
  • Keep-tail: leave recent turns out of the summarize request, then put the block before them; their thinking stays valid on preserved-thinking models if system/non-deferred tools are unchanged.
  • Rejected alongside: context_management, stop_sequences, output_config.format, tool_choice any/tool, task_budget.remaining, a last assistant turn with unresolved tool calls. Token counting ignores compaction. No summary cases return 200 with empty content and the summarizer's stop_reason (max_tokens, model_context_window_exceeded, refusal, tool_use, end_turn). 529 overloaded_error with error.details.error_code: compaction_unavailable is retryable.
  • Models: Fable 5.1/5, Mythos 5.1/5/Preview, Opus 5, 4.8, 4.7, 4.6, Sonnet 5, 4.6 (Claude API only). Models API with the header exposes capabilities.compaction.

# Client-side SDK compaction

compaction_control in TypeScript/Ruby tool_runner — deprecated; removed from the Python SDK v1.0. Prefer server-side.

# 4. Token budgets for long-running agents

  • max_tokens = hard per-request cap; effort = soft depth; task_budget (beta, Fable/Opus 4.7+) = advisory loop-wide countdown; compaction trigger = window guard.
  • Cost accounting with compaction: Σ iterations[].(input + cache_read + cache_creation + output).

# 5. Recipe — a 500-turn agent conversation that never blows the window

  1. Model: a 1M model (Sonnet 5 / Opus 5 / Fable 5.1). Keep thinking config and top-level effort constant for the whole session (cache).
  2. Headers: anthropic-beta: compact-2026-01-12,context-management-2025-06-27 (add thinking-binding-controls-2026-08-01 on Fable 5.1 and set prefix_mismatch_behavior: "drop_block" if you ever rewrite history).
  3. Prompt layout: tools → system (explicit cache_control breakpoint on the last system block) → messages, plus top-level automatic cache_control for the growing tail. Never edit earlier turns; add instructions with {"role": "system"} messages (Fable 5.x / Opus 4.8+) or in the newest user turn.
  4. Context editing: clear_tool_uses_20250919 with trigger 60000, keep 5, clear_at_least 15000, exclude_tools: ["memory"]; pair with the memory tool so Claude offloads facts before they are cleared.
  5. Compaction: compact_20260112 with trigger 150000 (default) and instructions that forbid tool calls and list what to retain (files touched, decisions, open tasks). Append every response verbatim; when stop_reason == "compaction" (if pause_after_compaction) re-add the last user message and continue.
  6. Thinking: pass all thinking/redacted_thinking blocks back unchanged; add clear_thinking_20251015 keep: {thinking_turns: 2} first in edits if thinking history grows faster than compaction reclaims it.
  7. Budgets: max_tokens 16k–64k per request; task_budget if the model supports it; monitor usage.iterations and context_management.applied_edits every turn; sum iterations for cost.
  8. Watch stop_reason for max_tokens, model_context_window_exceeded, compaction, pause_turn; usage.cache_read_input_tokens should stay ≈ the prefix size — a drop to 0 means a config change or an edit before a breakpoint.
  9. Between sessions: persist the last compaction block (or use on-demand compaction at session end) and resume from it.

Examples: examples/anthropic/context-management/ · tests: covered in tests/anthropic/test_thinking.py (context editing on Haiku).