The agentic tool loop on the Responses API — patterns, approvals, parallel calls, error handling
Status: DOCUMENTED · LIVE_VERIFIED (generic loop executed 2026-09-18 in Python and TypeScript with gpt-5.4-nano: 2 turns, function call → output → answer; MCP approval and local shell / apply_patch branches verified in the per-tool probes)
Sources: https://developers.openai.com/api/docs/guides/function-calling · https://developers.openai.com/api/docs/guides/tools · https://developers.openai.com/api/docs/guides/tools-connectors-mcp · https://developers.openai.com/api/docs/guides/tools-programmatic-tool-calling · https://developers.openai.com/api/docs/guides/async-tool-calling · https://developers.openai.com/api/docs/guides/tools-tool-search
Last verified: 2026-09-18 · code: examples/shared/tool-loop/openai_tool_loop.py, openai_tool_loop.ts
1. Which items require your action
| Output item | You must send back | Who executes |
|---|---|---|
function_call |
function_call_output {call_id, output} |
you |
custom_tool_call |
custom_tool_call_output {call_id, output} |
you |
mcp_approval_request |
mcp_approval_response {approval_request_id, approve, reason?} |
OpenAI runs the MCP call after approval |
shell_call with environment null/local |
shell_call_output {call_id, output:[{stdout, stderr, outcome}]} |
you |
apply_patch_call |
apply_patch_call_output {call_id, status: completed|failed, output} |
you |
computer_call |
computer_call_output {call_id, output:{type: computer_screenshot, image_url|file_id}, acknowledged_safety_checks} |
you |
tool_search_call with execution: client |
tool_search_output {call_id, execution: client, tools[]} |
you |
local_shell_call (legacy) |
local_shell_call_output {id, output} |
you |
web_search_call, file_search_call, code_interpreter_call, image_generation_call, mcp_list_tools, mcp_call, hosted shell_call/shell_call_output, hosted tool_search_* |
nothing | OpenAI (hosted) |
Terminate when a response contains none of the actionable items (typically ends with a message). Bound the number of turns and total cost.
2. State: previous_response_id vs replaying items
previous_response_id: resp.id+input: [outputs…]— simplest; the server keeps reasoning items,mcp_list_tools, loaded tools (tool search) in context. Verified in every live probe.- Stateless (
store: false/ ZDR): replay all prior output items includingreasoningitems (reasoning models require it;encrypted_contentby default with programmatic tool calling) before your outputs. - Always resend
tools(andinstructions) on continuation requests; keep tool lists stable to profit from prompt caching (usetool_choice: allowed_toolsto restrict instead of editingtools).
3. Parallel calls
parallel_tool_calls defaults to true: a turn may contain several function_call/custom_tool_call items — execute all of them (concurrently if independent) and return all outputs in one request, matching call_ids. Built-in tools are never batched with function calls. parallel_tool_calls: false guarantees ≤ 1 call per turn (verified with allowed_tools). Async tools (async: true, GPT-6 Astra+) let the model continue while a job runs; pair with a wait tool and task_handles; do not mix async tools with parallel calls in multi-agent mode.
4. Approvals and human-in-the-loop
- MCP:
require_approvaldefaults toalways; the model pauses atmcp_approval_request; approve/deny with a reason (denial verified: the model reports the rejection to the user). Narrow withallowed_tools/read_onlyandrequire_approval: {never: {tool_names: […]}}for trusted read-only tools; deep research needsnever. - Your own tools: gate write/high-impact actions in your handler regardless of the caller (
caller.typemay beprogramwhen programmatic tool calling generated the call). - Computer use: acknowledge
pending_safety_checksexplicitly and confirm at the point of risk.
5. Error handling
| Situation | Recommended handling (docs + live) |
|---|---|
| Tool raises / not found | return a string describing the error in function_call_output — never leave a call_id unanswered (400 on the next turn) |
| Invalid strict schema | HTTP 400 invalid_function_parameters at request time — validate schemas in CI |
status: incomplete (incomplete_details.reason: max_output_tokens) before any tool call |
raise max_output_tokens (hosted searches with reasoning needed ≥ 300 tokens live) |
| Model answers without using an expected hosted tool | force with tool_choice: {type: <tool>} (live: code_interpreter, web_search) or "required" |
mcp_call.error (mcp_protocol_error / mcp_tool_execution_error / http_error) |
surfaced in the item; the model usually explains — retry or fall back |
| Shell/apply_patch failures | outcome: {type: timeout} + partial output, non-zero exit_code, status: failed + message — the model re-plans |
| Local execution | sandbox, allow-lists, timeouts (timeout_ms is a hint), path validation (no .., no absolute paths) |
6. Programmatic tool calling
Add {type: programmatic_tool_calling} and set allowed_callers on eligible tools (function, custom, mcp, apply_patch, shell, code_interpreter). The model writes JavaScript (isolated V8, no network) that calls tools; nested function_call items arrive with caller: {type: "program", caller_id} and you answer them like direct calls. Use output_schema on functions so the program can parse results. Not live-tested (no eligible cheap model identified; check model pages).
7. Reference implementation (verified)
examples/shared/tool-loop/openai_tool_loop.py / .ts: bounded loop (max_turns), handlers map for functions/custom tools, MCP approval callback (deny by default), local shell stub with an allow-list, apply_patch dry-run with path validation, previous_response_id chaining, every request logged to reports/live-requests.jsonl. Live run: turn 0 ['reasoning','function_call'], turn 1 ['message'] → "Clear skies, 18°C in Paris."