SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
6.1 KB

# The agentic tool loop on the Responses API — patterns, approvals, parallel calls, error handling

Status: DOCUMENTED · LIVE_VERIFIED (generic loop executed 2026-09-18 in Python and TypeScript with gpt-5.4-nano: 2 turns, function call → output → answer; MCP approval and local shell / apply_patch branches verified in the per-tool probes) Sources: https://developers.openai.com/api/docs/guides/function-calling · https://developers.openai.com/api/docs/guides/tools · https://developers.openai.com/api/docs/guides/tools-connectors-mcp · https://developers.openai.com/api/docs/guides/tools-programmatic-tool-calling · https://developers.openai.com/api/docs/guides/async-tool-calling · https://developers.openai.com/api/docs/guides/tools-tool-search Last verified: 2026-09-18 · code: examples/shared/tool-loop/openai_tool_loop.py, openai_tool_loop.ts

# 1. Which items require your action

Output item You must send back Who executes
function_call function_call_output {call_id, output} you
custom_tool_call custom_tool_call_output {call_id, output} you
mcp_approval_request mcp_approval_response {approval_request_id, approve, reason?} OpenAI runs the MCP call after approval
shell_call with environment null/local shell_call_output {call_id, output:[{stdout, stderr, outcome}]} you
apply_patch_call apply_patch_call_output {call_id, status: completed|failed, output} you
computer_call computer_call_output {call_id, output:{type: computer_screenshot, image_url|file_id}, acknowledged_safety_checks} you
tool_search_call with execution: client tool_search_output {call_id, execution: client, tools[]} you
local_shell_call (legacy) local_shell_call_output {id, output} you
web_search_call, file_search_call, code_interpreter_call, image_generation_call, mcp_list_tools, mcp_call, hosted shell_call/shell_call_output, hosted tool_search_* nothing OpenAI (hosted)

Terminate when a response contains none of the actionable items (typically ends with a message). Bound the number of turns and total cost.

# 2. State: previous_response_id vs replaying items

  • previous_response_id: resp.id + input: [outputs…] — simplest; the server keeps reasoning items, mcp_list_tools, loaded tools (tool search) in context. Verified in every live probe.
  • Stateless (store: false / ZDR): replay all prior output items including reasoning items (reasoning models require it; encrypted_content by default with programmatic tool calling) before your outputs.
  • Always resend tools (and instructions) on continuation requests; keep tool lists stable to profit from prompt caching (use tool_choice: allowed_tools to restrict instead of editing tools).

# 3. Parallel calls

parallel_tool_calls defaults to true: a turn may contain several function_call/custom_tool_call items — execute all of them (concurrently if independent) and return all outputs in one request, matching call_ids. Built-in tools are never batched with function calls. parallel_tool_calls: false guarantees ≤ 1 call per turn (verified with allowed_tools). Async tools (async: true, GPT-6 Astra+) let the model continue while a job runs; pair with a wait tool and task_handles; do not mix async tools with parallel calls in multi-agent mode.

# 4. Approvals and human-in-the-loop

  • MCP: require_approval defaults to always; the model pauses at mcp_approval_request; approve/deny with a reason (denial verified: the model reports the rejection to the user). Narrow with allowed_tools/read_only and require_approval: {never: {tool_names: […]}} for trusted read-only tools; deep research needs never.
  • Your own tools: gate write/high-impact actions in your handler regardless of the caller (caller.type may be program when programmatic tool calling generated the call).
  • Computer use: acknowledge pending_safety_checks explicitly and confirm at the point of risk.

# 5. Error handling

Situation Recommended handling (docs + live)
Tool raises / not found return a string describing the error in function_call_output — never leave a call_id unanswered (400 on the next turn)
Invalid strict schema HTTP 400 invalid_function_parameters at request time — validate schemas in CI
status: incomplete (incomplete_details.reason: max_output_tokens) before any tool call raise max_output_tokens (hosted searches with reasoning needed ≥ 300 tokens live)
Model answers without using an expected hosted tool force with tool_choice: {type: <tool>} (live: code_interpreter, web_search) or "required"
mcp_call.error (mcp_protocol_error / mcp_tool_execution_error / http_error) surfaced in the item; the model usually explains — retry or fall back
Shell/apply_patch failures outcome: {type: timeout} + partial output, non-zero exit_code, status: failed + message — the model re-plans
Local execution sandbox, allow-lists, timeouts (timeout_ms is a hint), path validation (no .., no absolute paths)

# 6. Programmatic tool calling

Add {type: programmatic_tool_calling} and set allowed_callers on eligible tools (function, custom, mcp, apply_patch, shell, code_interpreter). The model writes JavaScript (isolated V8, no network) that calls tools; nested function_call items arrive with caller: {type: "program", caller_id} and you answer them like direct calls. Use output_schema on functions so the program can parse results. Not live-tested (no eligible cheap model identified; check model pages).

# 7. Reference implementation (verified)

examples/shared/tool-loop/openai_tool_loop.py / .ts: bounded loop (max_turns), handlers map for functions/custom tools, MCP approval callback (deny by default), local shell stub with an allow-list, apply_patch dry-run with path validation, previous_response_id chaining, every request logged to reports/live-requests.jsonl. Live run: turn 0 ['reasoning','function_call'], turn 1 ['message'] → "Clear skies, 18°C in Paris."