# Code execution and sandboxing **Status:** DOCUMENTED (xAI `code_interpreter` and Gemini `codeExecution` LIVE_VERIFIED with `print(2+2)`-class prompts on 2026-09-18/19) **Sources:** https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool · …/programmatic-tool-calling · OpenAI OpenAPI spec (`code_interpreter`, `/v1/containers*`) · https://developers.openai.com/api/docs/guides/tools-computer-use · https://platform.claude.com/docs/en/agents-and-tools/tool-use/bash-tool · xAI: https://docs.x.ai/developers/tools/code-execution ("Python with NumPy/Pandas/Matplotlib/SciPy; no network, no persistent filesystem, per-request stateless; time/memory limits"), https://docs.x.ai/developers/tools/streaming-and-sync (`include: code_interpreter_call.outputs`), https://docs.x.ai/developers/tools/overview (`shell` tool: `shell_call` / `shell_call_output`, env local), `sources/xai/openapi/openapi.json` (`CodeInterpreterCall`, `CodeInterpreterOutput`) · Gemini: https://ai.google.dev/gemini-api/docs/code-execution (Python only, 30 s runtime, up to 5 automatic regenerations, fixed library list, no pip, no network, matplotlib images only, not in Live API), https://ai.google.dev/api/generate-content (#ExecutableCode #CodeExecutionResult #Outcome), https://ai.google.dev/gemini-api/docs/agent-environment (managed-agent sandboxes) **Last verified:** 2026-09-19 ## Hosted sandboxes (provider-run) | | Anthropic `code_execution_*` | OpenAI `code_interpreter` | xAI `code_interpreter` (alias `code_execution`) | Gemini `codeExecution` | |---|---|---|---|---| | Where code runs | Anthropic's container | OpenAI's container (`/v1/containers`) | xAI's sandbox, Responses API only (gRPC/xai-sdk `code_execution()`); Chat Completions has no tools | Google's sandbox; `generateContent`/`streamGenerateContent`/Interactions (`{"type":"code_execution"}`)/Batch; **not** Live API | | Language / libraries | Python + bash, pre-installed libs | Python, container packages | Python with NumPy/Pandas/Matplotlib/SciPy | **Python only** (other languages generated, not run), fixed list (numpy, pandas, scipy, sklearn, matplotlib, opencv, tensorflow, sympy, PyPDF2, python-docx, reportlab…) — **no pip** | | Network | **none** | assume none | **none** | **none** | | State | new container per request unless `container.id` reused; files via `container_upload` | container persists for a while; `file_ids` | **stateless per request, no persistent filesystem**; attach data with `input_file` (Files API, agentic attachment search $10/1k) or `input_image` | stateless; inputs via `inlineData` / `fileData` (Files API, 48 h); model retries failing code up to **5×** automatically | | Limits | container-hour billing, Files 500 MB | usage-based | time/memory limits (unspecified); billing **$5 per 1k successful calls** + tokens (failed attempts free) | **30 s per execution**; no fee (tokens only: code + output + thinking billed as output, re-read context as "intermediate" input); free tier: free | | Output shape | `code_execution_tool_result` blocks | `code_interpreter_call` outputs | `code_interpreter_call {code, outputs:[{type:"logs", logs:""}, {type:"image", url}]}` — `outputs[]` only with `include:["code_interpreter_call.outputs"]`; `logs` is a **JSON-encoded string** (parse it, don't regex it) | parts `executableCode {language: PYTHON, code, id}` + `codeExecutionResult {outcome: OUTCOME_OK \| OUTCOME_FAILED \| OUTCOME_DEADLINE_EXCEEDED, output, id}` (+ `inlineData` images, `thoughtSignature` — echo whole parts in history) | | Streaming | `content_block_*` | `response.code_interpreter_call_code.delta` | `response.code_interpreter_call.in_progress → …call_code.delta → …call_code.done → …call.interpreting → …call.completed` (observed) | parts arrive as produced in `streamGenerateContent` | | Combos | programmatic tool calling (`allowed_callers`) | — | with `web_search`/`x_search`/files in the same request (bounded by `max_turns`) | with `googleSearch`, `urlContext`, function calling (Gemini 3), structured outputs; **not** with `fileSearch` | | What it protects you from | code runs off your infrastructure | same | same | same | | What it does **not** protect you from | the *output* is untrusted; produced files may be malicious; uploaded files carry injection text | same | same — `logs.stdout` is model-influenced text, generated images are untrusted bytes | same — `codeExecutionResult.output` is untrusted; `OUTCOME_FAILED` text may contain injected instructions; matplotlib PNGs are untrusted bytes | Hosted execution is the *safe default* for "run this analysis" use cases: the code never touches your machines. Gemini's 30-second / 5-retry loop and xAI's `max_turns` are also **cost bounds** — keep `maxOutputTokens`/`max_output_tokens` and `max_turns` small while iterating. ## Self-hosted execution (Anthropic `bash_20250124`, `text_editor_*`, OpenAI `shell` / `apply_patch`, **xAI `shell`**, custom tools) xAI's `shell` tool (`tools:[{type:"shell", env:"local"}]`) makes the model emit `shell_call` items that **you** run and answer with `shell_call_output {call_id, output:[{stdout, stderr, outcome:{type:"exit", exit_code} | {type:"timeout"}}], max_output_length}` — exactly the self-hosted case below; there is no server-side filtering. Gemini has no shell tool (only `codeExecution` in Google's sandbox and managed-agent environments). Here the model's commands run on **your** infrastructure. Minimum bar: 1. **Isolation**: ephemeral container/microVM per session (gVisor, Firecracker, Docker with `--cap-drop ALL --security-opt no-new-privileges`, read-only rootfs, non-root user, PID limit). Destroy after use. 2. **No secrets inside**: no cloud credentials, no API keys, no SSH agent, no mounted home. The sandbox should be able to leak everything it can see and still cause no harm. 3. **No network by default**; if needed, an egress allowlist through a proxy (block RFC 1918, link-local `169.254.0.0/16`, cloud metadata endpoints, localhost). 4. **Resource limits**: CPU, memory, disk quota, wall-clock timeout per command and per session; kill process groups. 5. **Filesystem scope**: a working directory the model may edit (`text_editor` paths canonicalised and confined — see `command-and-path-injection.md`); everything else read-only or absent. 6. **Command policy**: run commands via `execve` arrays where possible; if a shell is unavoidable (`bash` tool), run the *whole* command inside the sandbox where damage is bounded — do not try to "filter dangerous commands" with regexes. 7. **Outputs**: cap stdout/stderr bytes fed back (`untrusted-tool-outputs.md`); report failures with `is_error: true` (Anthropic bash tool docs) so the model re-plans. 8. **Artifacts**: scan files produced before serving them to users (content type sniffing, no HTML rendering from the sandbox origin, size caps). 9. **Human gate** for anything that leaves the sandbox: pushing commits, deploying, sending — the sandbox proposes, a human (or a strict allowlist) approves. ## Both - Uploaded documents (`container_upload`, `input_file`) are untrusted text for the model (indirect injection) — the sandbox does not change that. - Log commands and code with request ids; retain long enough for incident response. - Prefer hosted execution unless you need local resources; if you need both, keep two separate tool sets and never give the local shell tool in the same request as untrusted web content. ## Checklist - [ ] Hosted sandbox chosen where possible; container ids stored per user/session, never shared across tenants (xAI/Gemini sandboxes are stateless — nothing to store, nothing to leak between tenants). - [ ] Self-hosted (Anthropic bash, OpenAI shell, **xAI shell**): ephemeral, unprivileged, no secrets, no network (or allowlisted egress), resource limits, confined filesystem; `shell_call_output.outcome` reports timeouts honestly. - [ ] Command/file outputs capped and returned with `is_error` semantics (xAI: parse the `logs` JSON string; Gemini: branch on `codeExecutionResult.outcome`). - [ ] Artifacts scanned before delivery; never rendered as HTML from a trusted origin. - [ ] `allowed_callers` reviewed for every tool exposed to programmatic calling. - [ ] Human approval before sandbox results cause external effects.