SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
24.9 KB

# Tool execution — OpenAI Responses · Anthropic Messages · xAI Responses · Gemini generateContent / Interactions

Status: shapes and statuses from generated/tools.json (69 records: OpenAI 19, Anthropic 27, xAI 12, Gemini 11), generated/parameters.json (tool entries), generated/compatibility/{model-tool-matrix,gemini-tool-model-matrix}.json; live observations quoted from docs/openai/tool-loop.md, docs/tools/openai/.md, docs/tools/anthropic/.md, docs/tools/xai/.md, docs/xai/tool-loop.md, docs/tools/gemini/.md, docs/gemini/tool-loop.md (2026-09-18/19). Gemini googleSearch and computerUse are ACCOUNT_RESTRICTED on this free-tier key; xAI tool_search is alpha-only (403). Sources: https://developers.openai.com/api/docs/guides/function-calling · …/guides/tools · https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview · https://docs.x.ai/developers/tools/overview · https://docs.x.ai/developers/model-capabilities/text/function-calling · https://ai.google.dev/gemini-api/docs/function-calling · https://ai.google.dev/gemini-api/docs/tools Last verified: 2026-09-18

# 1. Taxonomy

Category OpenAI (tools[].type) Anthropic (tools[].type) xAI (tools[].type, Responses) Gemini (tools[] keys) Who executes
Your own tools function, custom (free-form / grammar), grouped in namespace custom tool (name + input_schema) function (also the only tool type on Chat Completions and /v1/messages) functionDeclarations[] (Interactions {type: function}) you
Vendor-defined, you execute computer, computer_use_preview, apply_patch, shell (environment.type: local) bash_20250124, text_editor_20250728, memory_20250818, computer_toolset_20260801, computer_20251124, browser_toolset_20260801 shell (environment: {type: local}) computerUse (predefined functionCalls: click, type, scroll, navigate, open_app…; PREVIEW) you
Vendor-executed (hosted / server) web_search, file_search, code_interpreter, image_generation, shell (container), tool_search (server), programmatic_tool_calling web_search_*, web_fetch_*, code_execution_*, tool_search_*, advisor_20260301 web_search, x_search, code_interpreter (= code_execution), file_search (= collections_search), image_generation, implicit attachment search, view_image/view_x_video sub-tools, tool_search (alpha) googleSearch (+ searchTypes.imageSearch), googleMaps, urlContext, codeExecution, fileSearch; legacy googleSearchRetrieval vendor, inside the request
Remote MCP mcp {server_url|connector_id|tunnel_id} mcp_servers[] + mcp_toolset (beta) mcp {server_url, server_label, allowed_tools, authorization, headers} mcpServers[] {streamableHttpTransport} (schema only, UNVERIFIED) / Interactions {type: mcp_server}; SDK-side MCP (mcpToTool()) vendor calls the server (Gemini SDK-side: your process)
Skills skill_reference in shell.environment.skills[], containers, Agents container.skills[], Managed Agents skills[] shell.environment.skills[] {name, description, path} (Skills API 404 for this team) .agents/skills/<name>/SKILL.md in custom-agent environments vendor sandbox / local

# 2. Definition mapping (client tools)

Field OpenAI tools[type=function] Anthropic custom tool xAI tools[type=function] Gemini functionDeclarations[] Notes
name name name (^[a-zA-Z0-9_-]{1,128}$) name name (≤128, [A-Za-z0-9_.:-], start with letter/underscore)
description description description description description (required per REST reference)
schema parameters (JSON Schema) input_schema (JSON Schema) parameters (JSON Schema; root must be an object) parameters (OpenAPI 3.03 subset) or parametersJsonSchema (full JSON Schema)
strict strict: true (omit → try strict, fall back) strict: true (≤20 strict tools) always strict (strict accepted, ignored) toolConfig.functionCallingConfig.mode: VALIDATED (request-level)
examples — input_examples[] — —
result schema output_schema (programmatic callers) — — response / responseJsonSchema
lazy load defer_loading: true (+ tool_search) defer_loading: true (+ tool search tool) defer_loading: true (+ tool_search, alpha 403) — (best practice 10–20 active declarations; max 512)
callers allowed_callers: ["direct","programmatic"] allowed_callers: ["direct","code_execution_20260120"] — —
async async: true (GPT-6 Astra) — — behavior: NON_BLOCKING (Live API only)
cache breakpoint per-part prompt_cache_breakpoint (5.6+) cache_control on the tool automatic inside cachedContents.tools
limits 2,000 per agent 128 per agent ≤350 tools (Chat) 512 declarations (SDK)

# 3. The loop — identical task on all four providers

Task: the model must call get_weather(city); you run it and return the result.

Request 1

json
// OpenAI — POST /v1/responses
{"model": "gpt-5.4-nano", "input": "Weather in Paris?",
 "tools": [{"type": "function", "name": "get_weather", "description": "Current weather for a city",
            "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"], "additionalProperties": false}, "strict": true}],
 "tool_choice": "auto", "parallel_tool_calls": true}
json
// Anthropic — POST /v1/messages
{"model": "claude-haiku-4-5-20251001", "max_tokens": 256,
 "messages": [{"role": "user", "content": "Weather in Paris?"}],
 "tools": [{"name": "get_weather", "description": "Current weather for a city",
            "input_schema": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"], "additionalProperties": false}, "strict": true}],
 "tool_choice": {"type": "auto"}}
json
// xAI — POST /v1/responses (same body as OpenAI; strict is implicit)
{"model": "grok-4.3", "input": "Weather in Paris?", "reasoning": {"effort": "low"},
 "tools": [{"type": "function", "name": "get_weather", "description": "Current weather for a city",
            "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}],
 "tool_choice": "auto", "parallel_tool_calls": true}
json
// Gemini — POST /v1beta/models/gemini-3.5-flash-lite:generateContent
{"contents": [{"role": "user", "parts": [{"text": "Weather in Paris?"}]}],
 "tools": [{"functionDeclarations": [{"name": "get_weather", "description": "Current weather for a city",
            "parameters": {"type": "OBJECT", "properties": {"city": {"type": "STRING"}}, "required": ["city"]}}]}],
 "toolConfig": {"functionCallingConfig": {"mode": "AUTO"}}}

Response 1 (shapes observed live)

json
// OpenAI output[] item
{"type": "function_call", "id": "fc_…", "call_id": "call_…", "name": "get_weather", "arguments": "{\"city\":\"Paris\"}", "status": "completed"}
json
// Anthropic content[] block, stop_reason: "tool_use"
{"type": "tool_use", "id": "toolu_…", "name": "get_weather", "input": {"city": "Paris"}, "caller": {"type": "direct"}}
json
// xAI output[] item (preceded by a reasoning item)
{"type": "function_call", "id": "fc_…", "call_id": "call-<uuid>-0", "name": "get_weather", "arguments": "{\"city\":\"Paris\"}", "status": "completed"}
json
// Gemini candidates[0].content.parts[] — the thoughtSignature MUST be echoed back on Gemini 3
{"functionCall": {"name": "get_weather", "args": {"city": "Paris"}, "id": "call_…"}, "thoughtSignature": "<base64>"}

Request 2 — return the result

json
// OpenAI (chained)
{"model": "gpt-5.4-nano", "previous_response_id": "resp_…", "tools": [ …same tools… ],
 "input": [{"type": "function_call_output", "call_id": "call_…", "output": "{\"temp_c\": 18, \"sky\": \"clear\"}"}]}
json
// Anthropic (replay)
{"model": "claude-haiku-4-5-20251001", "max_tokens": 256, "tools": [ …same tools… ],
 "messages": [{"role": "user", "content": "Weather in Paris?"},
   {"role": "assistant", "content": [{"type": "tool_use", "id": "toolu_…", "name": "get_weather", "input": {"city": "Paris"}}]},
   {"role": "user", "content": [{"type": "tool_result", "tool_use_id": "toolu_…", "content": "{\"temp_c\": 18, \"sky\": \"clear\"}"}]}]}
json
// xAI (chained — identical to OpenAI; a fresh max_turns budget starts here)
{"model": "grok-4.3", "previous_response_id": "resp_…", "tools": [ …same tools… ],
 "input": [{"type": "function_call_output", "call_id": "call-<uuid>-0", "output": "{\"temp_c\": 18, \"sky\": \"clear\"}"}]}
json
// Gemini (replay; all functionResponse parts of the step in ONE user Content, after ALL functionCall parts)
{"contents": [{"role": "user", "parts": [{"text": "Weather in Paris?"}]},
   {"role": "model", "parts": [{"functionCall": {"name": "get_weather", "args": {"city": "Paris"}, "id": "call_…"}, "thoughtSignature": "<base64>"}]},
   {"role": "user", "parts": [{"functionResponse": {"name": "get_weather", "id": "call_…", "response": {"temp_c": 18, "sky": "clear"}}}]}],
 "tools": [ …same tools… ]}

Differences that break naive ports: arguments is a JSON string on OpenAI/xAI, an object on Anthropic (input) and Gemini (args); the result is a top-level item (OpenAI/xAI), a tool_result block first in a user turn (Anthropic) or a functionResponse part whose response is a JSON object (Gemini); Anthropic signals with stop_reason: tool_use, OpenAI/xAI with a function_call item and status: completed, Gemini with a functionCall part and finishReason: STOP; Gemini 3 rejects a missing/edited thoughtSignature (400 or finishReason: MISSING_THOUGHT_SIGNATURE); tool errors are a string (OpenAI/xAI output), is_error: true (Anthropic) or an {error: …} object in response (Gemini); every call_id / tool_use_id / functionCall.id must be answered.

# 4. Tool choice and parallelism

Intent OpenAI tool_choice Anthropic tool_choice xAI tool_choice Gemini toolConfig.functionCallingConfig (Interactions generation_config.tool_choice)
let the model decide "auto" {"type": "auto"} "auto" (Chat/Responses); /v1/messages {type: auto} {"mode": "AUTO"} ("auto")
forbid tools "none" {"type": "none"} "none" {"mode": "NONE"} ("none")
must call some tool "required" {"type": "any"} "required" (/v1/messages {type: any}) {"mode": "ANY"} (constrained decoding) ("any")
must call tool X {"type": "function", "name": "X"} (also hosted types) {"type": "tool", "name": "X"} {"type": "function", "name": "X"} (Chat: {type: function, function:{name}}); server tools cannot be forced {"mode": "ANY", "allowedFunctionNames": ["X"]}; built-in tools cannot be forced (allowed_tools: {mode, tools[]})
restrict to a subset {"type": "allowed_tools", "mode": "auto|required", "tools": […]} — (remove from tools or defer_loading) — allowedFunctionNames[] (ANY mode) / Interactions allowed_tools
schema-validated decoding strict per tool strict per tool always {"mode": "VALIDATED"} (default when built-ins or structured output are combined) ("validated")
at most one call parallel_tool_calls: false tool_choice.disable_parallel_tool_use: true parallel_tool_calls: false — (no switch; the model may emit several functionCall parts)
forced choice + reasoning allowed any/tool ⟂ manual thinking; 400 on Fable 5.1 / Mythos 5.1 allowed (reasoning always on) allowed; Live setup.toolConfig → close 1007

# 5. Hosted / server tools side by side

Task OpenAI Anthropic xAI Gemini Portable JSON?
Search the web {"type":"web_search","search_context_size":"medium","filters":{"allowed_domains":["example.com"]}} → web_search_call + url_citation; $10/1k {"type":"web_search_20260318","name":"web_search","max_uses":5,"allowed_domains":["example.com"]} → server_tool_use + web_search_tool_result + web_search_result_location; $10/1k {"type":"web_search","allowed_domains":["example.com"],"enable_image_search":true} → web_search_call {action: search|open_page|find_in_page} + url_citation + inline [[N]](url); $5/1k; search_context_size → 400 {"googleSearch":{"timeRangeFilter":{…},"searchTypes":{"webSearch":{}}}} → groundingMetadata {webSearchQueries, searchEntryPoint (must render), groundingChunks[].web, groundingSupports}; 5,000 free queries/month then $14/1k (3.x) concept yes; shapes no
Search social / places — — {"type":"x_search","allowed_x_handles":["xai"],"from_date":"2026-09-01","enable_video_understanding":true} → custom_tool_call (x_keyword_search/x_semantic_search); $5/1k → per-post/profile from 2026-09-21 {"googleMaps":{"enableWidget":true}} + toolConfig.retrievalConfig.latLng → groundingChunks[].maps {placeId, uri, reviewSnippets}; $14/1k queries (3.x) no
Fetch a URL — {"type":"web_fetch_20260318","name":"web_fetch","max_uses":3,"citations":{"enabled":true}} (URLs already in context) — (web search open_page) {"urlContext":{}} (≤20 public URLs named in the prompt) → urlContextMetadata.urlMetadata[]; free tool, content = input tokens no
Run Python {"type":"code_interpreter","container":{"type":"auto","memory_limit":"1g"}} → code_interpreter_call {code, outputs}; container files; per-session billing {"type":"code_execution_20260521","name":"code_execution"} → server_tool_use + bash_code_execution_tool_result {stdout, stderr, return_code, content[{file_id}]}; container reuse; container-hours {"type":"code_interpreter"} → code_interpreter_call {code, outputs[{type: logs, logs: "<JSON stdout/stderr/exit_code>"}]} (needs include: ["code_interpreter_call.outputs"]); no container object; $5/1k calls {"codeExecution":{}} → parts executableCode {language: PYTHON, code} + codeExecutionResult {outcome, output} (+ PNG inlineData); 30 s; no fee concept yes
Search your documents {"type":"file_search","vector_store_ids":["vs_…"],"max_num_results":5} (Vector Stores API) pass search_result / document blocks with citations.enabled (your retrieval) {"type":"file_search","vector_store_ids":["<collection_id>"],"max_num_results":5} (Collections API; collections_search alias) → file_search_call {queries, results[{file_id, filename, score, text}]}; $2.50/1k + $0.10/GiB/day {"fileSearch":{"fileSearchStoreNames":["fileSearchStores/my-store"],"metadataFilter":"author = \"X\""}} (File Search stores; indexing $0.15/1M once) → groundingChunks[].retrievedContext OpenAI ↔ xAI literal; Gemini concept
Generate an image {"type":"image_generation","quality":"low","size":"1024x1024"} → image_generation_call — {"type":"image_generation"} → image_generation_call at Imagine per-image rates not a tool: generationConfig.responseModalities: ["TEXT","IMAGE"] on an image model OpenAI ↔ xAI
Call an MCP server {"type":"mcp","server_label":"deepwiki","server_url":"https://mcp.deepwiki.com/mcp","require_approval":"never","allowed_tools":["read_wiki_structure"]} → mcp_list_tools, mcp_call, mcp_approval_request/response header anthropic-beta: mcp-client-2025-11-20; mcp_servers:[{type:url,url,name}] + tools:[{type:mcp_toolset, mcp_server_name, configs}] → mcp_tool_use/mcp_tool_result; no approvals {"type":"mcp","server_label":"deepwiki","server_url":"https://mcp.deepwiki.com/mcp","allowed_tools":["read_wiki_structure"]} → mcp_call {name, server_label, arguments, output, error}; no approvals (require_approval ignored); tokens only Interactions {"type":"mcp_server","name":"deepwiki","url":"https://mcp.deepwiki.com/mcp","allowed_tools":[…]} (documented; server names without -); generateContent mcpServers[] UNVERIFIED; SDK mcpToTool() runs calls client-side OpenAI ↔ xAI literal
Find a tool among many {"type":"tool_search"} + defer_loading:true → tool_search_call/tool_search_output {"type":"tool_search_tool_regex_20251119","name":"tool_search_tool_regex"} + defer_loading → tool_search_tool_result {"type":"tool_search"} + defer_loading:true → 403 alpha users only — concept
Let the model script tool calls {"type":"programmatic_tool_calling"} + allowed_callers:["programmatic"] → program/program_output (JavaScript) code_execution_20260120+ + allowed_callers:["code_execution_20260120"] (Python; reply with tool_result + container) — — (compositional function calling across turns) OpenAI ↔ Anthropic concept
Shell on your machine {"type":"shell","environment":{"type":"local"}} → shell_call {action.commands[]} → shell_call_output {output[{stdout,stderr,outcome}]} {"type":"bash_20250124","name":"bash"} → tool_use {command|restart} → tool_result {"type":"shell","environment":{"type":"local","skills":[…]}} → shell_call → shell_call_output {stdout, stderr, outcome} — (Antigravity agent sandbox only) OpenAI ↔ xAI literal
Edit files {"type":"apply_patch"} → apply_patch_call {operation, diff} → apply_patch_call_output {"type":"text_editor_20250728","name":"str_replace_based_edit_tool"} → tool_use {command: view|str_replace|create|insert} — — (Antigravity built-ins write_to_file, replace_file_content, view_file…) concept
Drive a computer {"type":"computer"} → computer_call {action|actions[], pending_safety_checks} → computer_call_output {computer_screenshot, acknowledged_safety_checks} {"type":"computer_toolset_20260801"} → member tool_use {name: screenshot|click|type|zoom…} → tool_result {toolset_name: "computer", content:[image]} — {"computerUse":{"environment":"ENVIRONMENT_BROWSER","enablePromptInjectionDetection":true}} → functionCall {name: click, args:{x,y (0–999), intent}} (+ safety_decision) → functionResponse with screenshot inlineData (+ safety_acknowledgement) concept
Persistent notes — (Agents API sessions) {"type":"memory_20250818","name":"memory"} — — no
Ask a stronger model — {"type":"advisor_20260301","name":"advisor","model":"claude-opus-5"} (beta) — — no

# 6. Server-tool loop semantics that differ

Topic OpenAI Anthropic xAI Gemini
Where hosted results land items in the same output[]; one call completes the request (max_tool_calls caps built-in calls) *_tool_result blocks in the same content[]; ~10 iterations → stop_reason: pause_turn → resend as-is items in the same output[]; the server loops inside one request up to max_turns (default = server cap; docs 1–2 quick, 3–5 balanced, 10+ deep); a client-side function/shell call ends the request and the follow-up gets a fresh budget parts in the same candidates[].content (executableCode/codeExecutionResult, groundingMetadata, urlContextMetadata); Gemini 3 exposes built-in calls as toolCall/toolResponse parts when toolConfig.includeServerSideToolInvocations: true (echo them back)
Mixed server + client calls in one turn hosted calls never batched with function calls possible (stop_reason: tool_use with an unanswered server_tool_use) possible: server tools run first, a client call ends the turn Gemini 3 only ("tool combination", Preview; mode VALIDATED); Live API allows googleSearch + functions only
Approvals MCP require_approval (default always) → mcp_approval_request item none in Messages (Managed Agents permission_policy) none (require_approval silently accepted) computer use safety_decision: require_confirmation → safety_acknowledgement: true; no MCP approvals
Result replay keep hosted items in context (auto with previous_response_id) replay encrypted_content/encrypted_index unchanged (400 otherwise) keep items / use previous_response_id; replay reasoning.encrypted_content for cache hits echo executableCode/codeExecutionResult parts with their id and thoughtSignature; groundingChunks accumulate across stream chunks
Errors mcp_call.error {type}; code_interpreter_call.status: failed *_tool_result_error {error_code} inside a 200 file_search_call.status: failed (collection not indexed → 404 on search); mcp_call.error string; tool type unknown → 422 bare string codeExecutionResult.outcome: OUTCOME_FAILED|OUTCOME_DEADLINE_EXCEEDED; urlRetrievalStatus: URL_RETRIEVAL_STATUS_ERROR|PAYWALL|UNSAFE; finishReason: MALFORMED_FUNCTION_CALL|UNEXPECTED_TOOL_CALL|TOO_MANY_TOOL_CALLS
Billing counters tool_usage {web_search: {num_requests}} (undocumented); container sessions usage.server_tool_use {web_search_requests, web_fetch_requests}; container hours usage.server_side_tool_usage_details {web_search_calls, x_search_calls, code_interpreter_calls, file_search_calls, document_search_calls, mcp_calls, image_generation_calls}, num_server_side_tools_used, cost_in_usd_ticks; only successful executions billed usageMetadata.toolUsePromptTokenCount (+ per-modality details); Interactions usage.total_tool_use_tokens, grounding_tool_count[]; search queries counted against the 5,000/month quota

# 7. Model coverage (from generated/compatibility/*matrix*.json and tools.json)

Tool family OpenAI models (current) Anthropic models (current) xAI models Gemini models
function / custom tools all text models (38 listed) all 13 active Claude models all 7 Grok text models (+ voice session, /v1/messages) 21 listed (all 3.x / 2.5 text models, Live models, Gemma unknown)
web search 43 models all 13 all 7 (Responses only) 26 listed incl. gemini-3.1-flash-image (web + image search); Live models
code execution 36 models all 13 all 7 17 (all 3.x / 2.5 text models; not Live)
MCP 42 models all 13 (beta header) all 7 (+ voice) mcpServers schema lists 135 (UNVERIFIED); SDK-side MCP 83
file search / RAG vector stores on Responses models — all 7 (file_search/collections_search) 15 (fileSearch; 3.1 Pro 'AI Studio only')
tool search 11 (gpt-5.4+, 5.5, 5.6, 6) 12 7 listed, all 403 (alpha) —
computer use 11 toolset GA 7; beta 11 — 9 (3.8/3.7/3.6/3.5 Flash, 3.5 Flash-Lite, 3 Flash Preview, 2.5 CU legacy, robotics) — PREVIEW
shell / editor / skills 14–15 all 13 (bash, editor); skills with code execution shell (local) all 7; skills 404 — (Antigravity only)
programmatic tool calling not enumerated (docs: GPT-5.4+/6) 12 (not Haiku 4.5) — —
image generation tool / modality 31 — all 7 (image_generation) image models only (responseModalities)
URL context / web fetch — all 13 (web_fetch_*) — 16 (urlContext)
Maps / X search — — all 7 (x_search) 15 (googleMaps, not Live)

# 8. Portability checklist

  1. Keep tool definitions provider-neutral (name, description, JSON Schema with additionalProperties: false, all fields required); map parameters ↔ input_schema ↔ parameters/parametersJsonSchema; Gemini's OpenAPI-subset parameters uses uppercase type (OBJECT, STRING) — use parametersJsonSchema to keep one schema.
  2. Normalise the call record to {id, name, args(object)} — parse arguments on OpenAI/xAI; keep Gemini's thoughtSignature alongside the call.
  3. Emit results as function_call_output items (OpenAI/xAI), a leading user turn of tool_result blocks (Anthropic), or one user Content of functionResponse parts after all calls of the step (Gemini); represent errors as text (OpenAI/xAI), is_error: true (Anthropic) or an error object in response (Gemini).
  4. Loop until no tool call: OpenAI/xAI → no function_call item in output[]; Anthropic → stop_reason != tool_use && != pause_turn; Gemini → no functionCall part (check finishReason for MALFORMED_FUNCTION_CALL / MISSING_THOUGHT_SIGNATURE).
  5. Keep tools identical across turns on all four (cache prefix; Gemini Interactions and xAI previous_response_id do not need them resent but accept changes); restrict with allowed_tools (OpenAI) / allowedFunctionNames (Gemini) rather than editing the list.
  6. Bound the loop (max_turns on xAI is server-side; use your own counter elsewhere) and honour approvals/safety checks — reference implementations: examples/shared/tool-loop/{openai,anthropic,gemini}_tool_loop.{py,ts}, examples/xai/responses/tool_loop.py, examples/xai/tools/function-calling/loop.ts (all LIVE_VERIFIED).

Related: features · state-management · streaming · docs/openai/tool-loop.md · docs/tools/anthropic/tool-use-loop.md · docs/xai/tool-loop.md · docs/gemini/tool-loop.md · docs/tools/{openai,anthropic,xai,gemini}/index.md.