SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
8.2 KB

# Code execution and sandboxing

Status: DOCUMENTED (xAI code_interpreter and Gemini codeExecution LIVE_VERIFIED with print(2+2)-class prompts on 2026-09-18/19) Sources: https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool · …/programmatic-tool-calling · OpenAI OpenAPI spec (code_interpreter, /v1/containers*) · https://developers.openai.com/api/docs/guides/tools-computer-use · https://platform.claude.com/docs/en/agents-and-tools/tool-use/bash-tool · xAI: https://docs.x.ai/developers/tools/code-execution ("Python with NumPy/Pandas/Matplotlib/SciPy; no network, no persistent filesystem, per-request stateless; time/memory limits"), https://docs.x.ai/developers/tools/streaming-and-sync (include: code_interpreter_call.outputs), https://docs.x.ai/developers/tools/overview (shell tool: shell_call / shell_call_output, env local), sources/xai/openapi/openapi.json (CodeInterpreterCall, CodeInterpreterOutput) · Gemini: https://ai.google.dev/gemini-api/docs/code-execution (Python only, 30 s runtime, up to 5 automatic regenerations, fixed library list, no pip, no network, matplotlib images only, not in Live API), https://ai.google.dev/api/generate-content (#ExecutableCode #CodeExecutionResult #Outcome), https://ai.google.dev/gemini-api/docs/agent-environment (managed-agent sandboxes) Last verified: 2026-09-19

# Hosted sandboxes (provider-run)

Anthropic code_execution_* OpenAI code_interpreter xAI code_interpreter (alias code_execution) Gemini codeExecution
Where code runs Anthropic's container OpenAI's container (/v1/containers) xAI's sandbox, Responses API only (gRPC/xai-sdk code_execution()); Chat Completions has no tools Google's sandbox; generateContent/streamGenerateContent/Interactions ({"type":"code_execution"})/Batch; not Live API
Language / libraries Python + bash, pre-installed libs Python, container packages Python with NumPy/Pandas/Matplotlib/SciPy Python only (other languages generated, not run), fixed list (numpy, pandas, scipy, sklearn, matplotlib, opencv, tensorflow, sympy, PyPDF2, python-docx, reportlab…) — no pip
Network none assume none none none
State new container per request unless container.id reused; files via container_upload container persists for a while; file_ids stateless per request, no persistent filesystem; attach data with input_file (Files API, agentic attachment search $10/1k) or input_image stateless; inputs via inlineData / fileData (Files API, 48 h); model retries failing code up to 5× automatically
Limits container-hour billing, Files 500 MB usage-based time/memory limits (unspecified); billing $5 per 1k successful calls + tokens (failed attempts free) 30 s per execution; no fee (tokens only: code + output + thinking billed as output, re-read context as "intermediate" input); free tier: free
Output shape code_execution_tool_result blocks code_interpreter_call outputs code_interpreter_call {code, outputs:[{type:"logs", logs:"<JSON string: stdout, stderr, exit_code, command_timed_out>"}, {type:"image", url}]} — outputs[] only with include:["code_interpreter_call.outputs"]; logs is a JSON-encoded string (parse it, don't regex it) parts executableCode {language: PYTHON, code, id} + codeExecutionResult {outcome: OUTCOME_OK | OUTCOME_FAILED | OUTCOME_DEADLINE_EXCEEDED, output, id} (+ inlineData images, thoughtSignature — echo whole parts in history)
Streaming content_block_* response.code_interpreter_call_code.delta response.code_interpreter_call.in_progress → …call_code.delta → …call_code.done → …call.interpreting → …call.completed (observed) parts arrive as produced in streamGenerateContent
Combos programmatic tool calling (allowed_callers) — with web_search/x_search/files in the same request (bounded by max_turns) with googleSearch, urlContext, function calling (Gemini 3), structured outputs; not with fileSearch
What it protects you from code runs off your infrastructure same same same
What it does not protect you from the output is untrusted; produced files may be malicious; uploaded files carry injection text same same — logs.stdout is model-influenced text, generated images are untrusted bytes same — codeExecutionResult.output is untrusted; OUTCOME_FAILED text may contain injected instructions; matplotlib PNGs are untrusted bytes

Hosted execution is the safe default for "run this analysis" use cases: the code never touches your machines. Gemini's 30-second / 5-retry loop and xAI's max_turns are also cost bounds — keep maxOutputTokens/max_output_tokens and max_turns small while iterating.

# Self-hosted execution (Anthropic bash_20250124, text_editor_*, OpenAI shell / apply_patch, xAI shell, custom tools)

xAI's shell tool (tools:[{type:"shell", env:"local"}]) makes the model emit shell_call items that you run and answer with shell_call_output {call_id, output:[{stdout, stderr, outcome:{type:"exit", exit_code} | {type:"timeout"}}], max_output_length} — exactly the self-hosted case below; there is no server-side filtering. Gemini has no shell tool (only codeExecution in Google's sandbox and managed-agent environments).

Here the model's commands run on your infrastructure. Minimum bar:

  1. Isolation: ephemeral container/microVM per session (gVisor, Firecracker, Docker with --cap-drop ALL --security-opt no-new-privileges, read-only rootfs, non-root user, PID limit). Destroy after use.
  2. No secrets inside: no cloud credentials, no API keys, no SSH agent, no mounted home. The sandbox should be able to leak everything it can see and still cause no harm.
  3. No network by default; if needed, an egress allowlist through a proxy (block RFC 1918, link-local 169.254.0.0/16, cloud metadata endpoints, localhost).
  4. Resource limits: CPU, memory, disk quota, wall-clock timeout per command and per session; kill process groups.
  5. Filesystem scope: a working directory the model may edit (text_editor paths canonicalised and confined — see command-and-path-injection.md); everything else read-only or absent.
  6. Command policy: run commands via execve arrays where possible; if a shell is unavoidable (bash tool), run the whole command inside the sandbox where damage is bounded — do not try to "filter dangerous commands" with regexes.
  7. Outputs: cap stdout/stderr bytes fed back (untrusted-tool-outputs.md); report failures with is_error: true (Anthropic bash tool docs) so the model re-plans.
  8. Artifacts: scan files produced before serving them to users (content type sniffing, no HTML rendering from the sandbox origin, size caps).
  9. Human gate for anything that leaves the sandbox: pushing commits, deploying, sending — the sandbox proposes, a human (or a strict allowlist) approves.

# Both

  • Uploaded documents (container_upload, input_file) are untrusted text for the model (indirect injection) — the sandbox does not change that.
  • Log commands and code with request ids; retain long enough for incident response.
  • Prefer hosted execution unless you need local resources; if you need both, keep two separate tool sets and never give the local shell tool in the same request as untrusted web content.

# Checklist

  • Hosted sandbox chosen where possible; container ids stored per user/session, never shared across tenants (xAI/Gemini sandboxes are stateless — nothing to store, nothing to leak between tenants).
  • Self-hosted (Anthropic bash, OpenAI shell, xAI shell): ephemeral, unprivileged, no secrets, no network (or allowlisted egress), resource limits, confined filesystem; shell_call_output.outcome reports timeouts honestly.
  • Command/file outputs capped and returned with is_error semantics (xAI: parse the logs JSON string; Gemini: branch on codeExecutionResult.outcome).
  • Artifacts scanned before delivery; never rendered as HTML from a trusted origin.
  • allowed_callers reviewed for every tool exposed to programmatic calling.
  • Human approval before sandbox results cause external effects.