Untrusted tool outputs
Status: DOCUMENTED
Sources: https://platform.claude.com/docs/en/agents-and-tools/tool-use/handle-tool-calls · …/bash-tool · …/browser-use-tool · OpenAI OpenAPI spec (function_call_output, mcp_call, web_search_call) · https://developers.openai.com/api/docs/guides/tools-connectors-mcp · https://platform.claude.com/docs/en/build-with-claude/streaming · xAI: https://docs.x.ai/developers/tools/function-calling (function_call_output, chat role: tool), https://docs.x.ai/developers/tools/tool-usage-details (mcp_call.output, code_interpreter_call.outputs[].logs, web_search_call, custom_tool_call for x_search), https://docs.x.ai/developers/tools/overview (shell_call_output {stdout, stderr, outcome}) · Gemini: https://ai.google.dev/gemini-api/docs/function-calling (functionResponse {name, id, response: object}, "all functionResponses of a turn in one Content", error objects), https://ai.google.dev/api/generate-content (#FunctionResponse parts[] multimodal results, scheduling/willContinue (Live), #CodeExecutionResult outcome, #UrlContextMetadata, #GroundingMetadata), https://ai.google.dev/gemini-api/docs/thought-signatures
Last verified: 2026-09-19
Principle
Anything a tool returns — API JSON, a web page, an X post, a file's text, a shell's stdout, an MCP result, a screenshot — is untrusted input to the model and untrusted input to your code. It can carry prompt injection (for the model) and classic injection payloads (for your parsers, templates, SQL, shells).
Feeding results back safely
| OpenAI Responses | Anthropic Messages | xAI Responses / Chat | Gemini generateContent | |
|---|---|---|---|---|
| Result item | {type:"function_call_output", call_id, output} (string) |
{type:"tool_result", tool_use_id, content, is_error?} in a user turn |
Responses: {type:"function_call_output", call_id, output: string | parts[]} (+ previous_response_id or replayed output[]); Chat: {role:"tool", tool_call_id, content}; local shell: shell_call_output {call_id, output:[{stdout, stderr, outcome:{type:"exit", exit_code} | {type:"timeout"}}], max_output_length} |
{functionResponse:{name, id?, response: OBJECT, parts?: [inlineData…]}} in a user Content — response must be a JSON object (wrap scalars: {"result": …}) |
| Signalling failure | no flag → machine-readable status in output ([TOOL ERROR] prefix in this atlas) |
is_error: true |
no flag → same as OpenAI ([TOOL ERROR] prefix); shell: exit_code/timeout outcome |
no flag → {"error": "…"} object as the response (documented pattern; the shared adapter does this on is_error) |
| Multiple results | one item per call_id before the next model turn |
all tool_result blocks in one user message |
one function_call_output per call_id in the same input |
all functionResponse parts of a turn in ONE user Content, same order as the calls; echo the model turn verbatim first (thoughtSignature) or get 400 "missing a thought_signature" / finishReason: MISSING_THOUGHT_SIGNATURE |
| Size | bound it | max_content_tokens on web_fetch |
bound it; server-tool outputs (code_interpreter_call.outputs, web_search_call.action.sources, file_search_call.results) are only returned with include[] — request only what you need |
bound it; server-tool content is billed as toolUsePromptTokenCount even when retrieval failed |
Typed server-tool results to inspect rather than trust: Anthropic web_search_tool_result / web_fetch_tool_result / code_execution_tool_result (error codes); OpenAI web_search_call annotations; xAI web_search_call.action, custom_tool_call.input (x_search sub-tools), mcp_call {output, error}, code_interpreter_call.outputs[].logs (a JSON string with stdout/stderr/exit_code/command_timed_out — parse, don't regex); Gemini codeExecutionResult {outcome, output}, groundingMetadata.groundingChunks[] (redirect URIs), urlContextMetadata.urlMetadata[].urlRetrievalStatus, executableCode.code (code the sandbox ran — model-authored).
Rules
- Truncate and summarise tool output before returning it (head + tail, byte cap). A hostile page that is 2 MB of "ignore previous instructions" is also a cost attack.
- Never execute or render tool output without validation. Model → tool arguments → your code: validate against the tool's JSON Schema (
schema-validation.md). Tool → model → user: sanitise HTML/markdown, block auto-loading resources, escape before templating. - Do not leak stack traces or secrets in
is_errormessages. Return a short, sanitised reason (the reference tool loop inexamples/shared/provider-abstraction/doesf"Tool error: {type(e).__name__}: {e}"— review whatecan contain in your tools; strip paths, hostnames, tokens). - Label provenance in the content:
{"source":"web","url":…,"fetched_at":…,"content":…}so the model can weight trust and cite. Anthropic server tools already return typed result blocks (web_search_tool_result,web_fetch_tool_result,code_execution_tool_result, each with typed error codes such asurl_not_in_prior_context,max_uses_exceeded). - Treat tool descriptions as untrusted too (MCP): a server can change a tool's description to inject instructions between list and call (OpenAI: "MCP servers may update tool behavior unexpectedly"). Pin/review descriptions; alert on change.
- Streaming: partial tool arguments (
input_json_delta,function_call_arguments.delta) are previews — never act before thedone/content_block_stopevent and a successful parse (docs/architecture/streaming-patterns.md). - Idempotency of tool execution: the model may call the same tool twice (parallel calls, retries after a dropped stream). Make side-effecting tools idempotent on a key derived from
call_id/tool_use.id. - Images/screenshots returned to the model are injection vectors (Anthropic computer-use docs). Do not pass user-supplied images to tool-equipped models without the same care as text.
Checklist
- Every tool output capped in size and JSON-encoded with provenance (Gemini: as an object in
functionResponse.response). - Failures reported with
is_error: true(Anthropic) / explicit status text (OpenAI, xAI) /{"error": …}object (Gemini) /outcome(xAI shell), no stack traces or secrets. - Gemini model turns echoed verbatim (thought signatures) — never rebuilt from your own summaries; all parallel calls answered in one Content.
- Model-produced arguments validated by code before the tool runs.
- Tool output rendered to end users is sanitised/escaped.
- MCP tool lists and descriptions pinned; changes reviewed.
- Side-effecting tools idempotent per call id; parallel calls handled.
- Act only on final (non-delta) tool arguments.