SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
5.0 KB

# The agentic tool loop (client tools, server tools, pause_turn, mixed turns, programmatic calls)

Status: DOCUMENTED · LIVE_VERIFIED 2026-09-18 (examples/shared/tool-loop/anthropic_tool_loop.py and .ts ran: 2 turns, parallel get_weather + get_time, end_turn; e1/e2 code execution, g1/g2 programmatic pause/resume). Sources: How tool use works · Handle tool calls · Parallel tool use · Server tools · Stop reasons · Tool runner (SDK). Last verified: 2026-09-18.

# State machine

text
send(messages, tools) ──► stop_reason?
   ├─ end_turn / max_tokens / stop_sequence / refusal / model_context_window_exceeded → done (handle truncation)
   ├─ pause_turn  → server-side loop hit its cap (default 10 iterations/request): append assistant content AS-IS, same tools, send again
   └─ tool_use    → for every tool_use block (parallel calls!): run it, build tool_result{tool_use_id, content, is_error?, toolset_name?}
                    ├─ any server_tool_use / mcp_tool_use WITHOUT a matching *_tool_result in this response, or any tool_use with
                    │  caller.type != "direct"  → user message must contain ONLY tool_result blocks; keep the same tools array;
                    │  programmatic: also send top-level container = response.container.id
                    └─ otherwise text may follow the tool_result blocks

Rules that produce 400s if broken (all seen live): tool_result not first / missing for an id (a9); tool_result for a srvtoolu_ id (k24 "unexpected tool_use_id"); dropping the pending server tool from tools (docs: …but no `web_fetch` tool was provided); text after results while a server call is pending (docs: was found without a corresponding `web_search_tool_result` block).

# Parallel calls

Claude 4+ emits several tool_use blocks per turn by default (a3: get_weather + get_time). Run them concurrently or sequentially — your choice — but return all results in one user message; separate user messages per result "teach" Claude to stop parallelising. If you skip a call, still answer it with is_error:true ("Not executed: …"). Computer/browser toolsets are stricter: run in order, halt at first failure, exact halt text (see computer-use). Limit to one call with tool_choice.disable_parallel_tool_use: true (a4). Fable 5.1 issues fewer parallel calls in long agent loops; add a batching instruction.

# Error handling

Situation What to do
Tool threw tool_result with is_error:true and an instructive message ("Rate limit exceeded. Retry after 60 seconds.") — a8 shows Claude relaying it
Invalid / missing params return an error result; Claude retries 2–3× — or use strict:true
Fine-grained streaming produced invalid JSON is_error:true, content {"INVALID_JSON": "<raw>"} serialized
Server tool error already inside the *_tool_result (error_code), 200 response; nothing to send
max_tokens mid tool call partial JSON; retry with a higher max_tokens
Prompt injection in results keep untrusted content inside tool_result; Claude is trained to distrust instructions there (it may ask for confirmation)

# Server tools inside the loop

  • Results normally arrive in the same response (server_tool_use → web_search_tool_result etc., paired by tool_use_id, not position).
  • Mixed turn (server tool + client tool in one parallel group): stop_reason: tool_use, the server_tool_use has no result yet; after your tool_result-only reply the next response starts with that result block, then new content.
  • pause_turn: never leaves a client tool_use waiting; just re-send. Can repeat — cap retries.
  • Programmatic calls (caller.type == code_execution_20260120): the code is paused inside the container; your tool_result resumes it. Timeout ≈ 4 minutes → TimeoutError in stderr; idle containers reclaimed after ≈ 5 minutes; hard limit 30 days.
  • With prompt caching enabled, the API auto-places a 5-minute breakpoint on each server tool result between iterations (shows up as cache_creation.ephemeral_5m_input_tokens).

# Reference implementations

  • examples/shared/tool-loop/anthropic_tool_loop.py (dependency-free, logs via scripts/live.py) and anthropic_tool_loop.ts (@anthropic-ai/sdk 0.126, client.beta.messages when betas are needed).
  • Anthropic SDKs also provide client.beta.messages.tool_runner() / toolRunner() which implement this loop (and memory helpers, client-side compaction).