SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
6.8 KB

# Untrusted tool outputs

Status: DOCUMENTED Sources: https://platform.claude.com/docs/en/agents-and-tools/tool-use/handle-tool-calls · …/bash-tool · …/browser-use-tool · OpenAI OpenAPI spec (function_call_output, mcp_call, web_search_call) · https://developers.openai.com/api/docs/guides/tools-connectors-mcp · https://platform.claude.com/docs/en/build-with-claude/streaming · xAI: https://docs.x.ai/developers/tools/function-calling (function_call_output, chat role: tool), https://docs.x.ai/developers/tools/tool-usage-details (mcp_call.output, code_interpreter_call.outputs[].logs, web_search_call, custom_tool_call for x_search), https://docs.x.ai/developers/tools/overview (shell_call_output {stdout, stderr, outcome}) · Gemini: https://ai.google.dev/gemini-api/docs/function-calling (functionResponse {name, id, response: object}, "all functionResponses of a turn in one Content", error objects), https://ai.google.dev/api/generate-content (#FunctionResponse parts[] multimodal results, scheduling/willContinue (Live), #CodeExecutionResult outcome, #UrlContextMetadata, #GroundingMetadata), https://ai.google.dev/gemini-api/docs/thought-signatures Last verified: 2026-09-19

# Principle

Anything a tool returns — API JSON, a web page, an X post, a file's text, a shell's stdout, an MCP result, a screenshot — is untrusted input to the model and untrusted input to your code. It can carry prompt injection (for the model) and classic injection payloads (for your parsers, templates, SQL, shells).

# Feeding results back safely

OpenAI Responses Anthropic Messages xAI Responses / Chat Gemini generateContent
Result item {type:"function_call_output", call_id, output} (string) {type:"tool_result", tool_use_id, content, is_error?} in a user turn Responses: {type:"function_call_output", call_id, output: string | parts[]} (+ previous_response_id or replayed output[]); Chat: {role:"tool", tool_call_id, content}; local shell: shell_call_output {call_id, output:[{stdout, stderr, outcome:{type:"exit", exit_code} | {type:"timeout"}}], max_output_length} {functionResponse:{name, id?, response: OBJECT, parts?: [inlineData…]}} in a user Content — response must be a JSON object (wrap scalars: {"result": …})
Signalling failure no flag → machine-readable status in output ([TOOL ERROR] prefix in this atlas) is_error: true no flag → same as OpenAI ([TOOL ERROR] prefix); shell: exit_code/timeout outcome no flag → {"error": "…"} object as the response (documented pattern; the shared adapter does this on is_error)
Multiple results one item per call_id before the next model turn all tool_result blocks in one user message one function_call_output per call_id in the same input all functionResponse parts of a turn in ONE user Content, same order as the calls; echo the model turn verbatim first (thoughtSignature) or get 400 "missing a thought_signature" / finishReason: MISSING_THOUGHT_SIGNATURE
Size bound it max_content_tokens on web_fetch bound it; server-tool outputs (code_interpreter_call.outputs, web_search_call.action.sources, file_search_call.results) are only returned with include[] — request only what you need bound it; server-tool content is billed as toolUsePromptTokenCount even when retrieval failed

Typed server-tool results to inspect rather than trust: Anthropic web_search_tool_result / web_fetch_tool_result / code_execution_tool_result (error codes); OpenAI web_search_call annotations; xAI web_search_call.action, custom_tool_call.input (x_search sub-tools), mcp_call {output, error}, code_interpreter_call.outputs[].logs (a JSON string with stdout/stderr/exit_code/command_timed_out — parse, don't regex); Gemini codeExecutionResult {outcome, output}, groundingMetadata.groundingChunks[] (redirect URIs), urlContextMetadata.urlMetadata[].urlRetrievalStatus, executableCode.code (code the sandbox ran — model-authored).

# Rules

  1. Truncate and summarise tool output before returning it (head + tail, byte cap). A hostile page that is 2 MB of "ignore previous instructions" is also a cost attack.
  2. Never execute or render tool output without validation. Model → tool arguments → your code: validate against the tool's JSON Schema (schema-validation.md). Tool → model → user: sanitise HTML/markdown, block auto-loading resources, escape before templating.
  3. Do not leak stack traces or secrets in is_error messages. Return a short, sanitised reason (the reference tool loop in examples/shared/provider-abstraction/ does f"Tool error: {type(e).__name__}: {e}" — review what e can contain in your tools; strip paths, hostnames, tokens).
  4. Label provenance in the content: {"source":"web","url":…,"fetched_at":…,"content":…} so the model can weight trust and cite. Anthropic server tools already return typed result blocks (web_search_tool_result, web_fetch_tool_result, code_execution_tool_result, each with typed error codes such as url_not_in_prior_context, max_uses_exceeded).
  5. Treat tool descriptions as untrusted too (MCP): a server can change a tool's description to inject instructions between list and call (OpenAI: "MCP servers may update tool behavior unexpectedly"). Pin/review descriptions; alert on change.
  6. Streaming: partial tool arguments (input_json_delta, function_call_arguments.delta) are previews — never act before the done/content_block_stop event and a successful parse (docs/architecture/streaming-patterns.md).
  7. Idempotency of tool execution: the model may call the same tool twice (parallel calls, retries after a dropped stream). Make side-effecting tools idempotent on a key derived from call_id/tool_use.id.
  8. Images/screenshots returned to the model are injection vectors (Anthropic computer-use docs). Do not pass user-supplied images to tool-equipped models without the same care as text.

# Checklist

  • Every tool output capped in size and JSON-encoded with provenance (Gemini: as an object in functionResponse.response).
  • Failures reported with is_error: true (Anthropic) / explicit status text (OpenAI, xAI) / {"error": …} object (Gemini) / outcome (xAI shell), no stack traces or secrets.
  • Gemini model turns echoed verbatim (thought signatures) — never rebuilt from your own summaries; all parallel calls answered in one Content.
  • Model-produced arguments validated by code before the tool runs.
  • Tool output rendered to end users is sanitised/escaped.
  • MCP tool lists and descriptions pinned; changes reviewed.
  • Side-effecting tools idempotent per call id; parallel calls handled.
  • Act only on final (non-delta) tool arguments.