Security guide — building on the OpenAI, Anthropic, xAI and Gemini APIs
Status: DOCUMENTED (every mitigation cites the exact parameter/header/setting in the official docs, spec or discovery document; nothing here was "tested for security" live)
Sources: listed per page; primary: https://developers.openai.com/api/docs/guides/safety-best-practices, …/production-best-practices, …/tools-connectors-mcp#risks-and-safety, …/webhooks · https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-fetch-tool, …/computer-use-tool, …/mcp-connector, https://platform.claude.com/docs/en/manage-claude/api-and-data-retention · xAI: https://docs.x.ai/developers/faq/security, …/management-api-guide (ACLs, per-key limits, rotation), …/tools/{web-search,x-search,remote-mcp,code-execution}, …/rest-api-reference/inference/other (GET /v1/api-key) · Gemini: https://ai.google.dev/gemini-api/docs/api-key (auth keys, restrictions, leaked-key blocking), https://ai.google.dev/gemini-api/terms (Unpaid Services data use), …/docs/webhooks, …/docs/function-calling, …/docs/url-context, …/docs/code-execution, …/docs/computer-use
Last verified: 2026-09-19
Threat model in one line: an LLM application is a confused deputy — it holds your credentials and tool powers while reading text written by strangers (users, web pages, X posts, documents, tool outputs, MCP servers). Every page below names the provider knobs that shrink the blast radius, and ends with a checklist.
| Page | Covers | Key provider controls (OpenAI · Anthropic · xAI · Gemini) |
|---|---|---|
| api-key-handling | storage, rotation, expiry, key types, scopes | OpenAI project keys / service accounts / governance / IP allowlist · Anthropic workspace keys, expiration, Admin keys · xAI team keys with ACLs + per-key qps/qpm/tpm + expireTime, GET /v1/api-key introspection, separate Management keys, 400 (not 401) on a bad key · Gemini x-goog-api-key header vs ?key= URL leak, standard vs auth (service-account-bound) keys, Cloud-console restrictions, auto-blocking of leaked keys, free-tier data-use terms |
| least-privilege-and-admin-credentials | runtime vs admin keys, projects/workspaces/teams, spend limits | OpenAI Admin API keys, project limits · Anthropic Admin scopes, workspace limits · xAI Management API (SCOPE_TEAM / SCOPE_ORGANIZATION, api-key:endpoint:* / api-key:model:* ACLs, spending limits, audit GET /audit/teams/{id}/events) · Gemini Google Cloud IAM + API key restrictions, per-project quotas and spend caps, Batch/Interactions on store=true |
| prompt-injection | direct & indirect injection, defence in depth | length limits, safety_identifier, Moderation API, Anthropic classifiers, URL-in-context rule · xAI safety_identifier, max_turns, no approval step for MCP · Gemini safetySettings, labels.safety_identifier, computerUse.enablePromptInjectionDetection, URL_RETRIEVAL_STATUS_UNSAFE |
| untrusted-tool-outputs | tool results, web content, files as data not instructions | Anthropic is_error, OpenAI function_call_output · xAI function_call_output / chat role: tool (no error flag) · Gemini functionResponse.response object ({"error": …} pattern), codeExecutionResult.outcome |
| tool-and-mcp-security | remote MCP servers, connectors, approvals | OpenAI require_approval, allowed_tools, authorization · Anthropic mcp_servers[].tool_configuration, https-only · xAI mcp tool: allowed_tools, authorization/headers forwarded by xAI, no approval round-trip (require_approval silently accepted/ignored) · Gemini SDK-side MCP (mcpToTool, automatic_function_calling) vs server-side tools[].mcpServers (UNVERIFIED) / Interactions mcp_server{allowed_tools} |
| web-browsing-and-fetch | web search / fetch exfiltration | Anthropic web_fetch/web_search domain lists + URL-in-context rule · OpenAI web_search · xAI web_search allowed_domains/excluded_domains (≤ 5, exclusive), x_search allowed_x_handles/excluded_x_handles + date range, enable_image_understanding · Gemini googleSearch (no domain filter; Search Suggestions display obligation; 30-day storage) and urlContext (public URLs only, 20/request, 34 MB) |
| computer-use | GUI agents | OpenAI pending_safety_checks → acknowledged_safety_checks · Anthropic VM isolation, classifiers · Gemini computerUse safety_decision {decision: require_confirmation | blocked} → safety_acknowledgement: true after a human confirms, disabledSafetyPolicies, excludedPredefinedFunctions · xAI: no computer-use tool (only shell local tool — your sandbox) |
| code-execution-and-sandboxing | hosted and self-hosted code execution | Anthropic code_execution_* container (no internet) · OpenAI code_interpreter · xAI code_interpreter/code_execution (Python, no network, no persistent FS, stateless per request; logs is a JSON string) and the local shell tool · Gemini codeExecution (Python only, 30 s, fixed library list, no pip/network, images only as artifacts) |
| file-uploads-and-ssrf | user files, URLs, SSRF | size limits · xAI Files API (team-scoped, permanent unless expires_after, public CDN URLs via /public-url, input_file/file_url in Responses, attachment search $10/1k) · Gemini Files API (48 h auto-delete, 2 GB, fileData.fileUri also accepts public HTTPS/YouTube URLs the provider fetches, files:register for GCS) |
| command-and-path-injection | shell / filesystem tools | Anthropic bash/text_editor, OpenAI shell/apply_patch · xAI shell (shell_call / shell_call_output executed on your machine) · Gemini: no shell tool (code runs in Google's sandbox) |
| schema-validation | validating model output before acting | OpenAI text.format json_schema strict · Anthropic output_config.format · xAI response_format/text.format json_schema (strict implicit, additionalProperties defaults false) · Gemini responseMimeType+responseJsonSchema / responseFormat (values not validated, unsupported keywords ignored, MAX_TOKENS truncation) |
| secret-leakage | keys/PII in prompts, logs, stored responses | OpenAI store:false, ZDR · Anthropic ZDR · xAI 30-day retention, team-wide ZDR (x-zero-data-retention, /v1/me.zdr_status), store/previous_response_id disabled under ZDR, metadata rejected · Gemini free tier = "Unpaid Services" (prompts/responses may be used to improve products, human review), paid tier 55-day abuse logs, store, no full ZDR on the Developer API (Vertex AI for that) |
| domain-allowlists | egress control for every network-capable tool | per-tool allowlists for all four providers + network policy on your sandbox |
| webhook-verification | inbound webhooks | OpenAI webhook-id/-timestamp/-signature (Standard Webhooks) · Anthropic inference-hook webhook-id · Gemini static webhooks (same Standard Webhooks headers, signing secret shown once, rotate_secret) and dynamic webhook_config webhooks signed with JWT/JWKS (Webhook-Signature, https://generativelanguage.googleapis.com/.well-known/jwks.json, RS256) · xAI: no webhooks (poll deferred completions / batches) |
Cross-cutting principles
- Least privilege per key: one project/workspace/team key per deployment, read-only Admin/Management keys for monitoring, expiry enforced (OpenAI governance, Anthropic expiration, xAI
expireTime, Gemini auth keys bound to a service account with API restrictions). - Untrusted by default: user text, documents, web pages, X posts, screenshots, tool outputs and MCP tool descriptions are data. Only your system prompt is instructions — and even it cannot be relied on to stop injection.
- Human in the loop at the point of risk: approvals for MCP tools with side effects (OpenAI has a server-side flow; Anthropic/xAI/Gemini do not — gate client-side), computer-use safety checks (OpenAI
pending_safety_checks, Geminisafety_decision), irreversible actions. - Constrain, then validate: constrain output with schemas and tool allowlists; validate again with your own code before executing anything — on all four providers
strictguarantees shape, not meaning. - Minimise data:
store: false/ ZDR when eligible; know what cannot be turned off (Gemini grounding storage, Gemini free-tier data use, xAI stateful features under ZDR); no secrets in prompts; masked logs. - Observe:
safety_identifier(OpenAI, xAI),metadata.user_id(Anthropic),labels.safety_identifier(Gemini), request ids (x-request-id,request-id, Gemini bodyresponseId), webhook ids — so abuse can be traced and revoked per end-user without rotating your key.
Related engineering docs: docs/architecture/resilience.md (what to retry — retrying auth errors is also a security smell; xAI answers a bad key with 400), docs/architecture/streaming-patterns.md (bounded buffers), docs/architecture/multi-provider-abstraction.md (state and reasoning artefacts that must never cross providers).