# Domain allowlists — egress control for every network-capable tool (4 providers) **Status:** DOCUMENTED (xAI `web_search.allowed_domains` echoed live 2026-09-19; Gemini `urlContext` public-URL-only rule LIVE_VERIFIED via `URL_RETRIEVAL_STATUS_ERROR` on an unreachable host) **Sources:** https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-fetch-tool · …/web-search-tool ("`allowed_domains` or `blocked_domains`, not both → 400") · …/computer-use-tool · https://developers.openai.com/api/docs/guides/tools-connectors-mcp · https://developers.openai.com/api/docs/guides/error-codes (401 "IP not authorized") · https://platform.claude.com/docs/en/api/ip-addresses · xAI: https://docs.x.ai/developers/tools/web-search (`allowed_domains` / `excluded_domains`, ≤ 5, exclusive; `enable_image_understanding`), https://docs.x.ai/developers/tools/x-search (`allowed_x_handles` / `excluded_x_handles`, `from_date`/`to_date`), https://docs.x.ai/developers/tools/remote-mcp (`allowed_tools`), `sources/xai/openapi/openapi.json` (`WebSearchFilters`: "domains without protocol or subdomains … maximum of 5 … cannot be set together with excluded_domains") · Gemini: https://ai.google.dev/gemini-api/docs/url-context (public URLs only — no localhost/private networks/tunnels, 20 URLs, 34 MB, moderation → `URL_RETRIEVAL_STATUS_UNSAFE`), https://ai.google.dev/gemini-api/docs/google-search (`timeRangeFilter`, `searchTypes` — no domain filter), https://ai.google.dev/gemini-api/docs/api-key (application restrictions: IP / referrer), https://ai.google.dev/gemini-api/docs/computer-use (`disabledSafetyPolicies`, sandbox) **Last verified:** 2026-09-19 ## Principle Every tool that can *reach the network* is an exfiltration channel and an injection source. An allowlist bounds both. Allowlists beat blocklists: a blocklist must anticipate every attacker domain; an allowlist only has to know your business. ## Where allowlists exist | Surface | Mechanism | Notes | |---|---|---| | Anthropic `web_search_*` | `allowed_domains` **xor** `blocked_domains` (400 if both) | bare domains, optional path (`example.com/docs`); subdomains included per docs; `max_uses` caps calls | | Anthropic `web_fetch_*` | `allowed_domains` / `blocked_domains`, `max_uses`, `max_content_tokens` | plus the structural rule that only URLs already in context can be fetched | | Anthropic computer use / browser | **your** network layer (VM egress) — the docs list a domain allowlist as precaution #3 | implement with a forward proxy or firewall in the sandbox network | | Anthropic code execution | no network at all inside the container | nothing to allowlist | | Anthropic MCP connector | `mcp_servers[].url` must be `https://…`; `tool_configuration.allowed_tools` | the *server list* is your allowlist | | OpenAI `web_search` | tool-level filters per version (see `docs/tools/`) | hosted; still restrict what you can | | OpenAI MCP / connectors | `server_url` / `connector_id` you choose; `allowed_tools`; `require_approval` | the server list is the allowlist; Secure MCP Tunnels for private servers | | OpenAI computer use | your sandbox network (`execute_in_sandbox` boundary) | same as Anthropic | | **xAI `web_search`** | `allowed_domains` **xor** `excluded_domains` — bare domains, **≤ 5 each**, no protocol/subdomain; also accepted nested as `filters.allowed_domains` (OpenAI shape) | server-side browsing sub-tools (`browse_page`, `open_page`, `open_page_with_find`, `search_images`) all obey the list; `search_context_size` → 400; `enable_image_understanding` adds `view_image` (image tokens) | | **xAI `x_search`** | `allowed_x_handles` **xor** `excluded_x_handles` (≤ 10 spec / ≤ 20 guide) + `from_date`/`to_date` | the "domain" is X itself — restrict *accounts*; `enable_video_understanding` fetches post videos | | **xAI `mcp`** | `server_url` you choose + `allowed_tools` (empty = all) | no approval flow → the allowlist is the only gate; HTTPS | | **xAI `code_interpreter`** | no network inside the sandbox | nothing to allowlist | | **xAI `shell`** (local) | **your** machine — your firewall/proxy | same rules as self-hosted tools | | **Gemini `googleSearch`** | **no domain filter** — only `timeRangeFilter{startTime,endTime}` and `searchTypes{webSearch, imageSearch}`; results come through Google redirect URIs (`vertexaisearch.cloud.google.com/grounding-api-redirect/…`) | cannot be scoped; treat grounded text as untrusted; you must display Search Suggestions; grounding data is stored 30 days | | **Gemini `urlContext`** | no allowlist parameter, but a structural rule: **public URLs only** (localhost, 127.0.0.1, private ranges, ngrok/pinggy tunnels, logins, paywalls fail), ≤ 20 URLs/request, 34 MB/URL, content moderation (`URL_RETRIEVAL_STATUS_UNSAFE`), nested links not followed | the URLs come from the **prompt text** — an injected URL in user content will be fetched; pre-filter prompts against your own allowlist | | **Gemini `codeExecution`** | no network, no pip | nothing to allowlist | | **Gemini `computerUse`** | your browser sandbox network; `disabledSafetyPolicies` are preferences, not filters | domain allowlist at the proxy, like the others | | **Gemini SDK-side MCP** | the servers you connect from your process | your network policy applies | | Your own fetch/HTTP tools | code-level allowlist + DNS pinning (`file-uploads-and-ssrf.md`) | apply after redirects | | **Inbound** to the API | OpenAI project/org **IP allowlist** (401 "IP not authorized"); **Gemini** Cloud-console **application restrictions** (IP addresses / HTTP referrers / app ids) on the key; **xAI** team **mTLS** (`mtls.api.x.ai`, client cert required) and per-key ACLs; Anthropic publishes egress IP ranges for *its* outbound calls | different direction, same idea | ## Designing the list 1. **One source of truth**: a versioned `allowed_domains.json` per application, loaded into (a) Anthropic tool parameters, (b) your proxy/firewall rules, (c) your own fetcher. Drift between layers is how exfil slips through. 2. **Exact hosts over wildcards** where possible (`docs.example.com` rather than `example.com`) — user-generated-content hosts (`*.github.io`, `pastebin`, `*.s3.amazonaws.com`, URL shorteners, translation proxies) are attacker-controllable and defeat the purpose. 3. **Deny by default for new domains**; add via review. Log every blocked attempt (Anthropic returns `url_not_allowed`) — a spike is an injection attempt signal. 4. **Path prefixes** when the provider supports them (Anthropic: `example.com/docs`). 5. **Combine with quantity limits**: `max_uses`, `max_content_tokens`, step limits — even an allowed domain can host an injection page. 6. **Network enforcement for local tools**: forward proxy with explicit CONNECT allowlist + DNS resolver that refuses private ranges; no direct egress from the sandbox. Block cloud metadata (`169.254.169.254`, `fd00:ec2::254`), localhost and RFC 1918 unconditionally. 7. **Review cadence**: expire entries; re-justify quarterly; remove domains no longer needed. ## Checklist - [ ] Single allowlist artifact shared by tool params, proxy and code (Anthropic `allowed_domains`, xAI `allowed_domains` ≤ 5 — split into several requests if you need more, your proxy for everything else). - [ ] Anthropic web tools: `allowed_domains` set (never both lists); `max_uses` and `max_content_tokens` set. - [ ] xAI: `allowed_domains` (web) / `allowed_x_handles` (X) set and exclusive with the exclusion lists; `max_turns` bounded; `shell` tool only inside your sandbox. - [ ] Gemini: accept that `googleSearch` cannot be domain-scoped (or don't enable it); pre-filter URLs in prompts before `urlContext`; check `urlContextMetadata.urlRetrievalStatus` and treat `UNSAFE`/`ERROR` as signals. - [ ] MCP: explicit server inventory; `allowed_tools` on all providers (xAI: mandatory, empty = all). - [ ] Sandboxes (computer use, shell): egress via allowlisting proxy; metadata/private ranges blocked. - [ ] Own fetchers: allowlist applied after each redirect with DNS pinning. - [ ] Blocked attempts logged and alerted (Anthropic `url_not_allowed`, Gemini `URL_RETRIEVAL_STATUS_*`, xAI `web_search_call.status: failed`). - [ ] Inbound: OpenAI IP allowlist, Gemini key application restrictions, xAI mTLS/ACLs configured for production egress.