The agentic tool loop (client tools, server tools, pause_turn, mixed turns, programmatic calls)
Status: DOCUMENTED · LIVE_VERIFIED 2026-09-18 (examples/shared/tool-loop/anthropic_tool_loop.py and .ts ran: 2 turns, parallel get_weather + get_time, end_turn; e1/e2 code execution, g1/g2 programmatic pause/resume).
Sources: How tool use works · Handle tool calls · Parallel tool use · Server tools · Stop reasons · Tool runner (SDK).
Last verified: 2026-09-18.
State machine
send(messages, tools) ──► stop_reason?
├─ end_turn / max_tokens / stop_sequence / refusal / model_context_window_exceeded → done (handle truncation)
├─ pause_turn → server-side loop hit its cap (default 10 iterations/request): append assistant content AS-IS, same tools, send again
└─ tool_use → for every tool_use block (parallel calls!): run it, build tool_result{tool_use_id, content, is_error?, toolset_name?}
├─ any server_tool_use / mcp_tool_use WITHOUT a matching *_tool_result in this response, or any tool_use with
│ caller.type != "direct" → user message must contain ONLY tool_result blocks; keep the same tools array;
│ programmatic: also send top-level container = response.container.id
└─ otherwise text may follow the tool_result blocksRules that produce 400s if broken (all seen live): tool_result not first / missing for an id (a9); tool_result for a srvtoolu_ id (k24 "unexpected tool_use_id"); dropping the pending server tool from tools (docs: …but no `web_fetch` tool was provided); text after results while a server call is pending (docs: was found without a corresponding `web_search_tool_result` block).
Parallel calls
Claude 4+ emits several tool_use blocks per turn by default (a3: get_weather + get_time). Run them concurrently or sequentially — your choice — but return all results in one user message; separate user messages per result "teach" Claude to stop parallelising. If you skip a call, still answer it with is_error:true ("Not executed: …"). Computer/browser toolsets are stricter: run in order, halt at first failure, exact halt text (see computer-use). Limit to one call with tool_choice.disable_parallel_tool_use: true (a4). Fable 5.1 issues fewer parallel calls in long agent loops; add a batching instruction.
Error handling
| Situation | What to do |
|---|---|
| Tool threw | tool_result with is_error:true and an instructive message ("Rate limit exceeded. Retry after 60 seconds.") — a8 shows Claude relaying it |
| Invalid / missing params | return an error result; Claude retries 2–3× — or use strict:true |
| Fine-grained streaming produced invalid JSON | is_error:true, content {"INVALID_JSON": "<raw>"} serialized |
| Server tool error | already inside the *_tool_result (error_code), 200 response; nothing to send |
max_tokens mid tool call |
partial JSON; retry with a higher max_tokens |
| Prompt injection in results | keep untrusted content inside tool_result; Claude is trained to distrust instructions there (it may ask for confirmation) |
Server tools inside the loop
- Results normally arrive in the same response (
server_tool_use→web_search_tool_resultetc., paired bytool_use_id, not position). - Mixed turn (server tool + client tool in one parallel group):
stop_reason: tool_use, theserver_tool_usehas no result yet; after yourtool_result-only reply the next response starts with that result block, then new content. pause_turn: never leaves a clienttool_usewaiting; just re-send. Can repeat — cap retries.- Programmatic calls (
caller.type == code_execution_20260120): the code is paused inside the container; yourtool_resultresumes it. Timeout ≈ 4 minutes →TimeoutErrorin stderr; idle containers reclaimed after ≈ 5 minutes; hard limit 30 days. - With prompt caching enabled, the API auto-places a 5-minute breakpoint on each server tool result between iterations (shows up as
cache_creation.ephemeral_5m_input_tokens).
Reference implementations
examples/shared/tool-loop/anthropic_tool_loop.py(dependency-free, logs viascripts/live.py) andanthropic_tool_loop.ts(@anthropic-ai/sdk0.126,client.beta.messageswhenbetasare needed).- Anthropic SDKs also provide
client.beta.messages.tool_runner()/toolRunner()which implement this loop (and memory helpers, client-side compaction).