Python 88.3%
TypeScript 7.6%
Shell 4.1%
1# Anthropic — Context management: windows, 1M, context editing, compaction23**Status:** context windows `DOCUMENTED`/`LIVE_VERIFIED` · context editing `BETA` (`context-management-2025-06-27`) `LIVE_VERIFIED` (Haiku) · threshold compaction `BETA` (`compact-2026-01-12`) `LIVE_VERIFIED` (config accepted on Sonnet 5; no compaction triggered — needs ≥ 50k tokens) · on-demand compaction `BETA` (`compact-2026-09-04`) `LIVE_VERIFIED` (Sonnet 5 round trip).4**Sources:** [Context windows](https://platform.claude.com/docs/en/build-with-claude/context-windows) · [Context editing](https://platform.claude.com/docs/en/build-with-claude/context-editing) · [Compaction](https://platform.claude.com/docs/en/build-with-claude/compaction) · [Task budgets](https://platform.claude.com/docs/en/build-with-claude/task-budgets) · [Beta Messages API](https://platform.claude.com/docs/en/api/beta/messages) · [Release notes](https://platform.claude.com/docs/en/release-notes/api)5**Last verified:** 2026-09-1867## 1. Context windows89| Model family | Window | Max output | Media per request | Header | Pricing |10|---|---|---|---|---|---|11| Fable 5.1/5, Mythos 5.1/5/Preview, Opus 5, Opus 4.8/4.7/4.6, Sonnet 5, Sonnet 4.6 | **1M** (default) | 128K (300K in Batches with `output-300k-2026-03-24`) | 600 images / PDF pages | none | standard (no long-context premium) |12| Opus 4.5, Sonnet 4.5, Haiku 4.5 | 200K | 64K | 100 | — | — |1314* `context-1m-2025-08-07` is **legacy**: retired 2026-04-30 for Sonnet 4.5/Sonnet 4 (no effect; > 200k → error). Live: header accepted and ignored on Sonnet 5 and Haiku 4.5 (200).15* Everything counts: system, messages (tool results, images, documents), tool definitions, output incl. thinking. Cached tokens still occupy the window (`input + cache_read + cache_creation`).16* Overflow: input alone > window → `400 prompt is too long`. Input + `max_tokens` > window on 4.5+ → accepted, stops with `stop_reason: "model_context_window_exceeded"` (older models: validation error unless `model-context-window-exceeded-2025-08-26`).17* **Context awareness** (Sonnet 5, Sonnet 4.6, Sonnet 4.5, Haiku 4.5): the API injects the remaining-token budget after each tool call; not on Opus 4.7+/Fable (use task budgets there).18* Thinking: current-turn thinking → `max_tokens`; prior-turn thinking kept (input tokens) on Opus 4.5+/Sonnet 4.6+/Fable, stripped on Haiku/Sonnet 4.5.1920## 2. Context editing (`context_management.edits`, beta `context-management-2025-06-27`)2122Server-side: your client keeps the full history; the API edits the prompt before the model sees it. Exact strategy strings (verified in SDK types `BetaClearToolUses20250919Edit`, `BetaClearThinking20251015Edit`, `BetaCompact20260112Edit`):2324| Strategy `type` | Fields | Defaults | Notes |25|---|---|---|---|26| `clear_tool_uses_20250919` | `trigger: {type: input_tokens\|tool_uses, value ≥ 1}`, `keep: {type: tool_uses, value}`, `clear_at_least: {type: input_tokens, value}`, `exclude_tools: [names]`, `clear_tool_inputs: bool \| [names]` | trigger 100,000 input tokens; keep 3; clear results only | clears oldest results first, replaces with placeholder text; breaks the cache at the clear point (use `clear_at_least`) |27| `clear_thinking_20251015` | `keep: {type: thinking_turns, value ≥ 1} \| {type: all} \| "all"` | model-specific (all on Opus 4.5+/Sonnet 4.6+/Fable; last turn on Haiku & earlier) | must be **first** in `edits` when combined; requires thinking enabled/adaptive (`400 clear_thinking_20251015 strategy requires thinking to be enabled or adaptive`, live) |28| `compact_20260112` | `trigger: {type: input_tokens, value ≥ 50000}`, `pause_after_compaction: bool`, `instructions: string` | trigger 150,000 | beta header `compact-2026-01-12`; 4.6+ models only (`400 'claude-haiku-4-5-20251001' does not support the 'compact_20260112' context management strategy.`); `value 1000 → 400 trigger.value must be at least 50000` |2930Response field:31```json32"context_management": {"applied_edits": [33 {"type": "clear_tool_uses_20250919", "cleared_tool_uses": 2, "cleared_input_tokens": 174},34 {"type": "clear_thinking_20251015", "cleared_thinking_turns": 3, "cleared_input_tokens": 15000}]}35```36Streaming: in the final `message_delta`. `POST /v1/messages/count_tokens` accepts `context_management` and returns `{"input_tokens": 838, "context_management": {"original_input_tokens": 928}}` (live). Without the header: `400 context_management: Extra inputs are not permitted`. Works with the memory tool (Claude is warned before clearing).3738Live (Haiku): 3 fake tool round-trips, `trigger input_tokens 1`, `keep tool_uses 1`, `clear_at_least 1` → `applied_edits: [{clear_tool_uses_20250919, cleared_tool_uses: 2, cleared_input_tokens: 174}]`, 200.3940## 3. Compaction4142### Threshold compaction (`compact_20260112`)431. Input tokens reach `trigger` → the API summarizes, emits a `compaction` block at the start of the assistant response, continues.442. Pass the whole response back; everything before the last `compaction` block is ignored. `pause_after_compaction: true` → `stop_reason: "compaction"` so you can inject content first.453. Default prompt writes a `<summary>` for continuation; `instructions` replaces it entirely (tell the model **not to call tools** — otherwise `content: null`). On Fable 5.1 custom instructions summarize the visible conversation only.464. Streaming: `content_block_start` → **one** `content_block_delta` with the whole summary → `content_block_stop`.475. `usage.iterations[]` lists `{type: "compaction"}` and `{type: "message"}` entries; top-level `input/output_tokens` **exclude** the compaction iteration → sum `iterations` for billing. With the beta header every response carries `iterations` (live Sonnet 5: `[{type: "message", …}]` with no compaction).486. Same model summarizes; `cache_control` allowed on the block; keep a system breakpoint so only the summary is rewritten. Token counting applies existing blocks but never triggers new ones. Server tools: trigger checked at every sampling iteration.497. Images, documents, `container_upload` blocks and fetched URLs inside the summarized range are lost.5051### On-demand compaction (`compaction: {"type": "summarize"}`, beta `compact-2026-09-04`)52* Separate request (same `system`/`tools`/thinking/`max_tokens` as the conversation) → response = one **signed** `compaction` block (`content`, `signature`), `stop_reason: "compaction"`, top-level usage 0, `usage.iterations: [{type: "compaction", input_tokens: 90, output_tokens: 116}]` (live). Can run in the background.53* Continue: send the block **first** in `messages` (own assistant message or first block of the first message) in place of the summarized messages, header on every request, exactly one block. Live: `[assistant{compaction}, user"What is my name…"]` → "Ada — favourite color: teal." Summarized messages left in front → `400 messages.1.content.0: compaction block must be sent first, in place of the messages it summarizes; remove those messages` (`compaction_block_misplaced`).54* Keep-tail: leave recent turns out of the summarize request, then put the block before them; their thinking stays valid on preserved-thinking models if `system`/non-deferred `tools` are unchanged.55* Rejected alongside: `context_management`, `stop_sequences`, `output_config.format`, `tool_choice any/tool`, `task_budget.remaining`, a last assistant turn with unresolved tool calls. Token counting ignores `compaction`. No summary cases return 200 with empty `content` and the summarizer's `stop_reason` (`max_tokens`, `model_context_window_exceeded`, `refusal`, `tool_use`, `end_turn`). 529 `overloaded_error` with `error.details.error_code: compaction_unavailable` is retryable.56* Models: Fable 5.1/5, Mythos 5.1/5/Preview, Opus 5, 4.8, 4.7, 4.6, Sonnet 5, 4.6 (Claude API only). Models API with the header exposes `capabilities.compaction`.5758### Client-side SDK compaction59`compaction_control` in TypeScript/Ruby `tool_runner` — **deprecated**; removed from the Python SDK v1.0. Prefer server-side.6061## 4. Token budgets for long-running agents62* `max_tokens` = hard per-request cap; `effort` = soft depth; `task_budget` (beta, Fable/Opus 4.7+) = advisory loop-wide countdown; compaction trigger = window guard.63* Cost accounting with compaction: `Σ iterations[].(input + cache_read + cache_creation + output)`.6465## 5. Recipe — a 500-turn agent conversation that never blows the window661. **Model**: a 1M model (Sonnet 5 / Opus 5 / Fable 5.1). Keep thinking config and top-level effort **constant** for the whole session (cache).672. **Headers**: `anthropic-beta: compact-2026-01-12,context-management-2025-06-27` (add `thinking-binding-controls-2026-08-01` on Fable 5.1 and set `prefix_mismatch_behavior: "drop_block"` if you ever rewrite history).683. **Prompt layout**: tools → system (explicit `cache_control` breakpoint on the last system block) → messages, plus top-level automatic `cache_control` for the growing tail. Never edit earlier turns; add instructions with `{"role": "system"}` messages (Fable 5.x / Opus 4.8+) or in the newest user turn.694. **Context editing**: `clear_tool_uses_20250919` with `trigger 60000`, `keep 5`, `clear_at_least 15000`, `exclude_tools: ["memory"]`; pair with the memory tool so Claude offloads facts before they are cleared.705. **Compaction**: `compact_20260112` with `trigger 150000` (default) and `instructions` that forbid tool calls and list what to retain (files touched, decisions, open tasks). Append every response verbatim; when `stop_reason == "compaction"` (if `pause_after_compaction`) re-add the last user message and continue.716. **Thinking**: pass all `thinking`/`redacted_thinking` blocks back unchanged; add `clear_thinking_20251015` `keep: {thinking_turns: 2}` **first** in `edits` if thinking history grows faster than compaction reclaims it.727. **Budgets**: `max_tokens` 16k–64k per request; `task_budget` if the model supports it; monitor `usage.iterations` and `context_management.applied_edits` every turn; sum iterations for cost.738. **Watch** `stop_reason` for `max_tokens`, `model_context_window_exceeded`, `compaction`, `pause_turn`; `usage.cache_read_input_tokens` should stay ≈ the prefix size — a drop to 0 means a config change or an edit before a breakpoint.749. **Between sessions**: persist the last compaction block (or use on-demand `compaction` at session end) and resume from it.7576Examples: `examples/anthropic/context-management/` · tests: covered in `tests/anthropic/test_thinking.py` (context editing on Haiku).77