Code execution and sandboxing
Status: DOCUMENTED (xAI code_interpreter and Gemini codeExecution LIVE_VERIFIED with print(2+2)-class prompts on 2026-09-18/19)
Sources: https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool · …/programmatic-tool-calling · OpenAI OpenAPI spec (code_interpreter, /v1/containers*) · https://developers.openai.com/api/docs/guides/tools-computer-use · https://platform.claude.com/docs/en/agents-and-tools/tool-use/bash-tool · xAI: https://docs.x.ai/developers/tools/code-execution ("Python with NumPy/Pandas/Matplotlib/SciPy; no network, no persistent filesystem, per-request stateless; time/memory limits"), https://docs.x.ai/developers/tools/streaming-and-sync (include: code_interpreter_call.outputs), https://docs.x.ai/developers/tools/overview (shell tool: shell_call / shell_call_output, env local), sources/xai/openapi/openapi.json (CodeInterpreterCall, CodeInterpreterOutput) · Gemini: https://ai.google.dev/gemini-api/docs/code-execution (Python only, 30 s runtime, up to 5 automatic regenerations, fixed library list, no pip, no network, matplotlib images only, not in Live API), https://ai.google.dev/api/generate-content (#ExecutableCode #CodeExecutionResult #Outcome), https://ai.google.dev/gemini-api/docs/agent-environment (managed-agent sandboxes)
Last verified: 2026-09-19
Hosted sandboxes (provider-run)
Anthropic code_execution_* |
OpenAI code_interpreter |
xAI code_interpreter (alias code_execution) |
Gemini codeExecution |
|
|---|---|---|---|---|
| Where code runs | Anthropic's container | OpenAI's container (/v1/containers) |
xAI's sandbox, Responses API only (gRPC/xai-sdk code_execution()); Chat Completions has no tools |
Google's sandbox; generateContent/streamGenerateContent/Interactions ({"type":"code_execution"})/Batch; not Live API |
| Language / libraries | Python + bash, pre-installed libs | Python, container packages | Python with NumPy/Pandas/Matplotlib/SciPy | Python only (other languages generated, not run), fixed list (numpy, pandas, scipy, sklearn, matplotlib, opencv, tensorflow, sympy, PyPDF2, python-docx, reportlab…) — no pip |
| Network | none | assume none | none | none |
| State | new container per request unless container.id reused; files via container_upload |
container persists for a while; file_ids |
stateless per request, no persistent filesystem; attach data with input_file (Files API, agentic attachment search $10/1k) or input_image |
stateless; inputs via inlineData / fileData (Files API, 48 h); model retries failing code up to 5× automatically |
| Limits | container-hour billing, Files 500 MB | usage-based | time/memory limits (unspecified); billing $5 per 1k successful calls + tokens (failed attempts free) | 30 s per execution; no fee (tokens only: code + output + thinking billed as output, re-read context as "intermediate" input); free tier: free |
| Output shape | code_execution_tool_result blocks |
code_interpreter_call outputs |
code_interpreter_call {code, outputs:[{type:"logs", logs:"<JSON string: stdout, stderr, exit_code, command_timed_out>"}, {type:"image", url}]} — outputs[] only with include:["code_interpreter_call.outputs"]; logs is a JSON-encoded string (parse it, don't regex it) |
parts executableCode {language: PYTHON, code, id} + codeExecutionResult {outcome: OUTCOME_OK | OUTCOME_FAILED | OUTCOME_DEADLINE_EXCEEDED, output, id} (+ inlineData images, thoughtSignature — echo whole parts in history) |
| Streaming | content_block_* |
response.code_interpreter_call_code.delta |
response.code_interpreter_call.in_progress → …call_code.delta → …call_code.done → …call.interpreting → …call.completed (observed) |
parts arrive as produced in streamGenerateContent |
| Combos | programmatic tool calling (allowed_callers) |
— | with web_search/x_search/files in the same request (bounded by max_turns) |
with googleSearch, urlContext, function calling (Gemini 3), structured outputs; not with fileSearch |
| What it protects you from | code runs off your infrastructure | same | same | same |
| What it does not protect you from | the output is untrusted; produced files may be malicious; uploaded files carry injection text | same | same — logs.stdout is model-influenced text, generated images are untrusted bytes |
same — codeExecutionResult.output is untrusted; OUTCOME_FAILED text may contain injected instructions; matplotlib PNGs are untrusted bytes |
Hosted execution is the safe default for "run this analysis" use cases: the code never touches your machines. Gemini's 30-second / 5-retry loop and xAI's max_turns are also cost bounds — keep maxOutputTokens/max_output_tokens and max_turns small while iterating.
Self-hosted execution (Anthropic bash_20250124, text_editor_*, OpenAI shell / apply_patch, xAI shell, custom tools)
xAI's shell tool (tools:[{type:"shell", env:"local"}]) makes the model emit shell_call items that you run and answer with shell_call_output {call_id, output:[{stdout, stderr, outcome:{type:"exit", exit_code} | {type:"timeout"}}], max_output_length} — exactly the self-hosted case below; there is no server-side filtering. Gemini has no shell tool (only codeExecution in Google's sandbox and managed-agent environments).
Here the model's commands run on your infrastructure. Minimum bar:
- Isolation: ephemeral container/microVM per session (gVisor, Firecracker, Docker with
--cap-drop ALL --security-opt no-new-privileges, read-only rootfs, non-root user, PID limit). Destroy after use. - No secrets inside: no cloud credentials, no API keys, no SSH agent, no mounted home. The sandbox should be able to leak everything it can see and still cause no harm.
- No network by default; if needed, an egress allowlist through a proxy (block RFC 1918, link-local
169.254.0.0/16, cloud metadata endpoints, localhost). - Resource limits: CPU, memory, disk quota, wall-clock timeout per command and per session; kill process groups.
- Filesystem scope: a working directory the model may edit (
text_editorpaths canonicalised and confined — seecommand-and-path-injection.md); everything else read-only or absent. - Command policy: run commands via
execvearrays where possible; if a shell is unavoidable (bashtool), run the whole command inside the sandbox where damage is bounded — do not try to "filter dangerous commands" with regexes. - Outputs: cap stdout/stderr bytes fed back (
untrusted-tool-outputs.md); report failures withis_error: true(Anthropic bash tool docs) so the model re-plans. - Artifacts: scan files produced before serving them to users (content type sniffing, no HTML rendering from the sandbox origin, size caps).
- Human gate for anything that leaves the sandbox: pushing commits, deploying, sending — the sandbox proposes, a human (or a strict allowlist) approves.
Both
- Uploaded documents (
container_upload,input_file) are untrusted text for the model (indirect injection) — the sandbox does not change that. - Log commands and code with request ids; retain long enough for incident response.
- Prefer hosted execution unless you need local resources; if you need both, keep two separate tool sets and never give the local shell tool in the same request as untrusted web content.
Checklist
- Hosted sandbox chosen where possible; container ids stored per user/session, never shared across tenants (xAI/Gemini sandboxes are stateless — nothing to store, nothing to leak between tenants).
- Self-hosted (Anthropic bash, OpenAI shell, xAI shell): ephemeral, unprivileged, no secrets, no network (or allowlisted egress), resource limits, confined filesystem;
shell_call_output.outcomereports timeouts honestly. - Command/file outputs capped and returned with
is_errorsemantics (xAI: parse thelogsJSON string; Gemini: branch oncodeExecutionResult.outcome). - Artifacts scanned before delivery; never rendered as HTML from a trusted origin.
-
allowed_callersreviewed for every tool exposed to programmatic calling. - Human approval before sandbox results cause external effects.