SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
8.2 KB

# Domain allowlists — egress control for every network-capable tool (4 providers)

Status: DOCUMENTED (xAI web_search.allowed_domains echoed live 2026-09-19; Gemini urlContext public-URL-only rule LIVE_VERIFIED via URL_RETRIEVAL_STATUS_ERROR on an unreachable host) Sources: https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-fetch-tool · …/web-search-tool ("allowed_domains or blocked_domains, not both → 400") · …/computer-use-tool · https://developers.openai.com/api/docs/guides/tools-connectors-mcp · https://developers.openai.com/api/docs/guides/error-codes (401 "IP not authorized") · https://platform.claude.com/docs/en/api/ip-addresses · xAI: https://docs.x.ai/developers/tools/web-search (allowed_domains / excluded_domains, ≤ 5, exclusive; enable_image_understanding), https://docs.x.ai/developers/tools/x-search (allowed_x_handles / excluded_x_handles, from_date/to_date), https://docs.x.ai/developers/tools/remote-mcp (allowed_tools), sources/xai/openapi/openapi.json (WebSearchFilters: "domains without protocol or subdomains … maximum of 5 … cannot be set together with excluded_domains") · Gemini: https://ai.google.dev/gemini-api/docs/url-context (public URLs only — no localhost/private networks/tunnels, 20 URLs, 34 MB, moderation → URL_RETRIEVAL_STATUS_UNSAFE), https://ai.google.dev/gemini-api/docs/google-search (timeRangeFilter, searchTypes — no domain filter), https://ai.google.dev/gemini-api/docs/api-key (application restrictions: IP / referrer), https://ai.google.dev/gemini-api/docs/computer-use (disabledSafetyPolicies, sandbox) Last verified: 2026-09-19

# Principle

Every tool that can reach the network is an exfiltration channel and an injection source. An allowlist bounds both. Allowlists beat blocklists: a blocklist must anticipate every attacker domain; an allowlist only has to know your business.

# Where allowlists exist

Surface Mechanism Notes
Anthropic web_search_* allowed_domains xor blocked_domains (400 if both) bare domains, optional path (example.com/docs); subdomains included per docs; max_uses caps calls
Anthropic web_fetch_* allowed_domains / blocked_domains, max_uses, max_content_tokens plus the structural rule that only URLs already in context can be fetched
Anthropic computer use / browser your network layer (VM egress) — the docs list a domain allowlist as precaution #3 implement with a forward proxy or firewall in the sandbox network
Anthropic code execution no network at all inside the container nothing to allowlist
Anthropic MCP connector mcp_servers[].url must be https://…; tool_configuration.allowed_tools the server list is your allowlist
OpenAI web_search tool-level filters per version (see docs/tools/) hosted; still restrict what you can
OpenAI MCP / connectors server_url / connector_id you choose; allowed_tools; require_approval the server list is the allowlist; Secure MCP Tunnels for private servers
OpenAI computer use your sandbox network (execute_in_sandbox boundary) same as Anthropic
xAI web_search allowed_domains xor excluded_domains — bare domains, ≤ 5 each, no protocol/subdomain; also accepted nested as filters.allowed_domains (OpenAI shape) server-side browsing sub-tools (browse_page, open_page, open_page_with_find, search_images) all obey the list; search_context_size → 400; enable_image_understanding adds view_image (image tokens)
xAI x_search allowed_x_handles xor excluded_x_handles (≤ 10 spec / ≤ 20 guide) + from_date/to_date the "domain" is X itself — restrict accounts; enable_video_understanding fetches post videos
xAI mcp server_url you choose + allowed_tools (empty = all) no approval flow → the allowlist is the only gate; HTTPS
xAI code_interpreter no network inside the sandbox nothing to allowlist
xAI shell (local) your machine — your firewall/proxy same rules as self-hosted tools
Gemini googleSearch no domain filter — only timeRangeFilter{startTime,endTime} and searchTypes{webSearch, imageSearch}; results come through Google redirect URIs (vertexaisearch.cloud.google.com/grounding-api-redirect/…) cannot be scoped; treat grounded text as untrusted; you must display Search Suggestions; grounding data is stored 30 days
Gemini urlContext no allowlist parameter, but a structural rule: public URLs only (localhost, 127.0.0.1, private ranges, ngrok/pinggy tunnels, logins, paywalls fail), ≤ 20 URLs/request, 34 MB/URL, content moderation (URL_RETRIEVAL_STATUS_UNSAFE), nested links not followed the URLs come from the prompt text — an injected URL in user content will be fetched; pre-filter prompts against your own allowlist
Gemini codeExecution no network, no pip nothing to allowlist
Gemini computerUse your browser sandbox network; disabledSafetyPolicies are preferences, not filters domain allowlist at the proxy, like the others
Gemini SDK-side MCP the servers you connect from your process your network policy applies
Your own fetch/HTTP tools code-level allowlist + DNS pinning (file-uploads-and-ssrf.md) apply after redirects
Inbound to the API OpenAI project/org IP allowlist (401 "IP not authorized"); Gemini Cloud-console application restrictions (IP addresses / HTTP referrers / app ids) on the key; xAI team mTLS (mtls.api.x.ai, client cert required) and per-key ACLs; Anthropic publishes egress IP ranges for its outbound calls different direction, same idea

# Designing the list

  1. One source of truth: a versioned allowed_domains.json per application, loaded into (a) Anthropic tool parameters, (b) your proxy/firewall rules, (c) your own fetcher. Drift between layers is how exfil slips through.
  2. Exact hosts over wildcards where possible (docs.example.com rather than example.com) — user-generated-content hosts (*.github.io, pastebin, *.s3.amazonaws.com, URL shorteners, translation proxies) are attacker-controllable and defeat the purpose.
  3. Deny by default for new domains; add via review. Log every blocked attempt (Anthropic returns url_not_allowed) — a spike is an injection attempt signal.
  4. Path prefixes when the provider supports them (Anthropic: example.com/docs).
  5. Combine with quantity limits: max_uses, max_content_tokens, step limits — even an allowed domain can host an injection page.
  6. Network enforcement for local tools: forward proxy with explicit CONNECT allowlist + DNS resolver that refuses private ranges; no direct egress from the sandbox. Block cloud metadata (169.254.169.254, fd00:ec2::254), localhost and RFC 1918 unconditionally.
  7. Review cadence: expire entries; re-justify quarterly; remove domains no longer needed.

# Checklist

  • Single allowlist artifact shared by tool params, proxy and code (Anthropic allowed_domains, xAI allowed_domains ≤ 5 — split into several requests if you need more, your proxy for everything else).
  • Anthropic web tools: allowed_domains set (never both lists); max_uses and max_content_tokens set.
  • xAI: allowed_domains (web) / allowed_x_handles (X) set and exclusive with the exclusion lists; max_turns bounded; shell tool only inside your sandbox.
  • Gemini: accept that googleSearch cannot be domain-scoped (or don't enable it); pre-filter URLs in prompts before urlContext; check urlContextMetadata.urlRetrievalStatus and treat UNSAFE/ERROR as signals.
  • MCP: explicit server inventory; allowed_tools on all providers (xAI: mandatory, empty = all).
  • Sandboxes (computer use, shell): egress via allowlisting proxy; metadata/private ranges blocked.
  • Own fetchers: allowlist applied after each redirect with DNS pinning.
  • Blocked attempts logged and alerted (Anthropic url_not_allowed, Gemini URL_RETRIEVAL_STATUS_*, xAI web_search_call.status: failed).
  • Inbound: OpenAI IP allowlist, Gemini key application restrictions, xAI mTLS/ACLs configured for production egress.