Python 88.3%
TypeScript 7.6%
Shell 4.1%
1# Domain allowlists — egress control for every network-capable tool (4 providers)23**Status:** DOCUMENTED (xAI `web_search.allowed_domains` echoed live 2026-09-19; Gemini `urlContext` public-URL-only rule LIVE_VERIFIED via `URL_RETRIEVAL_STATUS_ERROR` on an unreachable host)4**Sources:** https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-fetch-tool · …/web-search-tool ("`allowed_domains` or `blocked_domains`, not both → 400") · …/computer-use-tool · https://developers.openai.com/api/docs/guides/tools-connectors-mcp · https://developers.openai.com/api/docs/guides/error-codes (401 "IP not authorized") · https://platform.claude.com/docs/en/api/ip-addresses · xAI: https://docs.x.ai/developers/tools/web-search (`allowed_domains` / `excluded_domains`, ≤ 5, exclusive; `enable_image_understanding`), https://docs.x.ai/developers/tools/x-search (`allowed_x_handles` / `excluded_x_handles`, `from_date`/`to_date`), https://docs.x.ai/developers/tools/remote-mcp (`allowed_tools`), `sources/xai/openapi/openapi.json` (`WebSearchFilters`: "domains without protocol or subdomains … maximum of 5 … cannot be set together with excluded_domains") · Gemini: https://ai.google.dev/gemini-api/docs/url-context (public URLs only — no localhost/private networks/tunnels, 20 URLs, 34 MB, moderation → `URL_RETRIEVAL_STATUS_UNSAFE`), https://ai.google.dev/gemini-api/docs/google-search (`timeRangeFilter`, `searchTypes` — no domain filter), https://ai.google.dev/gemini-api/docs/api-key (application restrictions: IP / referrer), https://ai.google.dev/gemini-api/docs/computer-use (`disabledSafetyPolicies`, sandbox)5**Last verified:** 2026-09-1967## Principle89Every tool that can *reach the network* is an exfiltration channel and an injection source. An allowlist bounds both. Allowlists beat blocklists: a blocklist must anticipate every attacker domain; an allowlist only has to know your business.1011## Where allowlists exist1213| Surface | Mechanism | Notes |14|---|---|---|15| Anthropic `web_search_*` | `allowed_domains` **xor** `blocked_domains` (400 if both) | bare domains, optional path (`example.com/docs`); subdomains included per docs; `max_uses` caps calls |16| Anthropic `web_fetch_*` | `allowed_domains` / `blocked_domains`, `max_uses`, `max_content_tokens` | plus the structural rule that only URLs already in context can be fetched |17| Anthropic computer use / browser | **your** network layer (VM egress) — the docs list a domain allowlist as precaution #3 | implement with a forward proxy or firewall in the sandbox network |18| Anthropic code execution | no network at all inside the container | nothing to allowlist |19| Anthropic MCP connector | `mcp_servers[].url` must be `https://…`; `tool_configuration.allowed_tools` | the *server list* is your allowlist |20| OpenAI `web_search` | tool-level filters per version (see `docs/tools/`) | hosted; still restrict what you can |21| OpenAI MCP / connectors | `server_url` / `connector_id` you choose; `allowed_tools`; `require_approval` | the server list is the allowlist; Secure MCP Tunnels for private servers |22| OpenAI computer use | your sandbox network (`execute_in_sandbox` boundary) | same as Anthropic |23| **xAI `web_search`** | `allowed_domains` **xor** `excluded_domains` — bare domains, **≤ 5 each**, no protocol/subdomain; also accepted nested as `filters.allowed_domains` (OpenAI shape) | server-side browsing sub-tools (`browse_page`, `open_page`, `open_page_with_find`, `search_images`) all obey the list; `search_context_size` → 400; `enable_image_understanding` adds `view_image` (image tokens) |24| **xAI `x_search`** | `allowed_x_handles` **xor** `excluded_x_handles` (≤ 10 spec / ≤ 20 guide) + `from_date`/`to_date` | the "domain" is X itself — restrict *accounts*; `enable_video_understanding` fetches post videos |25| **xAI `mcp`** | `server_url` you choose + `allowed_tools` (empty = all) | no approval flow → the allowlist is the only gate; HTTPS |26| **xAI `code_interpreter`** | no network inside the sandbox | nothing to allowlist |27| **xAI `shell`** (local) | **your** machine — your firewall/proxy | same rules as self-hosted tools |28| **Gemini `googleSearch`** | **no domain filter** — only `timeRangeFilter{startTime,endTime}` and `searchTypes{webSearch, imageSearch}`; results come through Google redirect URIs (`vertexaisearch.cloud.google.com/grounding-api-redirect/…`) | cannot be scoped; treat grounded text as untrusted; you must display Search Suggestions; grounding data is stored 30 days |29| **Gemini `urlContext`** | no allowlist parameter, but a structural rule: **public URLs only** (localhost, 127.0.0.1, private ranges, ngrok/pinggy tunnels, logins, paywalls fail), ≤ 20 URLs/request, 34 MB/URL, content moderation (`URL_RETRIEVAL_STATUS_UNSAFE`), nested links not followed | the URLs come from the **prompt text** — an injected URL in user content will be fetched; pre-filter prompts against your own allowlist |30| **Gemini `codeExecution`** | no network, no pip | nothing to allowlist |31| **Gemini `computerUse`** | your browser sandbox network; `disabledSafetyPolicies` are preferences, not filters | domain allowlist at the proxy, like the others |32| **Gemini SDK-side MCP** | the servers you connect from your process | your network policy applies |33| Your own fetch/HTTP tools | code-level allowlist + DNS pinning (`file-uploads-and-ssrf.md`) | apply after redirects |34| **Inbound** to the API | OpenAI project/org **IP allowlist** (401 "IP not authorized"); **Gemini** Cloud-console **application restrictions** (IP addresses / HTTP referrers / app ids) on the key; **xAI** team **mTLS** (`mtls.api.x.ai`, client cert required) and per-key ACLs; Anthropic publishes egress IP ranges for *its* outbound calls | different direction, same idea |3536## Designing the list37381. **One source of truth**: a versioned `allowed_domains.json` per application, loaded into (a) Anthropic tool parameters, (b) your proxy/firewall rules, (c) your own fetcher. Drift between layers is how exfil slips through.392. **Exact hosts over wildcards** where possible (`docs.example.com` rather than `example.com`) — user-generated-content hosts (`*.github.io`, `pastebin`, `*.s3.amazonaws.com`, URL shorteners, translation proxies) are attacker-controllable and defeat the purpose.403. **Deny by default for new domains**; add via review. Log every blocked attempt (Anthropic returns `url_not_allowed`) — a spike is an injection attempt signal.414. **Path prefixes** when the provider supports them (Anthropic: `example.com/docs`).425. **Combine with quantity limits**: `max_uses`, `max_content_tokens`, step limits — even an allowed domain can host an injection page.436. **Network enforcement for local tools**: forward proxy with explicit CONNECT allowlist + DNS resolver that refuses private ranges; no direct egress from the sandbox. Block cloud metadata (`169.254.169.254`, `fd00:ec2::254`), localhost and RFC 1918 unconditionally.447. **Review cadence**: expire entries; re-justify quarterly; remove domains no longer needed.4546## Checklist4748- [ ] Single allowlist artifact shared by tool params, proxy and code (Anthropic `allowed_domains`, xAI `allowed_domains` ≤ 5 — split into several requests if you need more, your proxy for everything else).49- [ ] Anthropic web tools: `allowed_domains` set (never both lists); `max_uses` and `max_content_tokens` set.50- [ ] xAI: `allowed_domains` (web) / `allowed_x_handles` (X) set and exclusive with the exclusion lists; `max_turns` bounded; `shell` tool only inside your sandbox.51- [ ] Gemini: accept that `googleSearch` cannot be domain-scoped (or don't enable it); pre-filter URLs in prompts before `urlContext`; check `urlContextMetadata.urlRetrievalStatus` and treat `UNSAFE`/`ERROR` as signals.52- [ ] MCP: explicit server inventory; `allowed_tools` on all providers (xAI: mandatory, empty = all).53- [ ] Sandboxes (computer use, shell): egress via allowlisting proxy; metadata/private ranges blocked.54- [ ] Own fetchers: allowlist applied after each redirect with DNS pinning.55- [ ] Blocked attempts logged and alerted (Anthropic `url_not_allowed`, Gemini `URL_RETRIEVAL_STATUS_*`, xAI `web_search_call.status: failed`).56- [ ] Inbound: OpenAI IP allowlist, Gemini key application restrictions, xAI mTLS/ACLs configured for production egress.57