SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
183.1 KB

# Feature × Provider matrix — OpenAI · Anthropic · xAI · Gemini

Status: synthesis of generated/*.json (endpoints, parameters, tools, models, headers, pricing, streaming-events, webhook-events, deprecations) and the domain pages under docs/; every status shown is the status recorded in those files (LIVE_VERIFIED means called successfully with this atlas's keys on 2026-09-18; xAI probes ran on a Tier 0 team, Gemini probes on a free-tier key — hence ACCOUNT_RESTRICTED on paid-only Gemini features and on xAI's Management/Skills/Embeddings). Machine-readable twin: generated/compatibility/cross-provider-feature-matrix.json (same rows, openai/anthropic/xai/gemini objects per record, providers_supporting[], provider_count). Sources: the ref column of the JSON twin names, per row, the generated file or docs page each cell was taken from. Canonical vendor pages: https://developers.openai.com/api/docs · https://platform.claude.com/docs/en · https://docs.x.ai/developers · https://ai.google.dev/gemini-api/docs. Last verified: 2026-09-18

Legend — Portable = the same task can be expressed with an equivalent parameter on every provider that offers it (mapping in the per-topic pages; a feature offered by one provider only is never portable); “— not offered” = no documented surface. Statuses use the atlas vocabulary (DOCUMENTED, LIVE_VERIFIED, LIVE_DISCOVERED, BETA, PREVIEW, GA, LEGACY, DEPRECATED, RETIRED, ACCOUNT_RESTRICTED, UNVERIFIED, FAILED_VERIFICATION, DOCUMENTATION_INCOMPLETE).

# Contents

  1. Core generation (12 rows)
  2. Structured output & sampling (10 rows)
  3. Tools — client side (13 rows)
  4. Tools — server side / hosted (13 rows)
  5. Multimodal input & media APIs (13 rows)
  6. Files & batch (3 rows)
  7. Prompt caching (2 rows)
  8. Reasoning / thinking (6 rows)
  9. Context management (6 rows)
  10. Service tiers, limits, safety (8 rows)
  11. Auth, versioning, SDKs, platform (14 rows)
  12. Managed agents platforms (15 rows)
  13. Model customisation & evaluation (5 rows)
  14. Legacy / retired surfaces (5 rows)

Totals: 125 features — on all four providers: 45 · on three: 34 · on two: 16 · on one: 26 · on none (legacy/absent everywhere): 4 · marked portable: 86.

Coverage Count Features
All four 45 Primary text/multimodal generation endpoint; System / developer prompt; Message roles; Multi-turn conversation; Token counting; Model listing / catalogue; Structured outputs (JSON Schema constrained); Strict tool arguments; Sampling parameters (temperature / top_p / top_k); Stop sequences; Refusal / safety signalling in-band; Function / custom tools (you execute); Tool choice; Parallel tool calls; Web search; Code execution (sandboxed Python); Remote MCP servers; Citations / grounding metadata; Image input; PDF / document input; Files API; Batch processing (async, discounted); Prompt / context caching; Reasoning control; Reasoning visibility; Reasoning replay across turns; Context window; Max output tokens; Service tiers / processing modes; End-user identifier for abuse detection; Rate-limit tiers; Overload / capacity error; Error envelope; Authentication; Official SDKs; Webhooks (platform events); Usage & cost reporting; Cloud availability; Zero data retention / data-use terms; Managed agent harness; Create an agent session / run; Send input / events to a session; Session event stream; Client-side agent framework / CLI; Retired / deprecated beta headers, parameters and SDKs
anthropic + xai + gemini 2 Task-wide token budget (advisory); Session budgets
openai + anthropic + gemini 6 Computer use; Private-network MCP (tunnels) / agent credentials; Server-side compaction (in-flight); Hosted sandbox; Credential vaults / secrets; Artifacts / deliverables
openai + anthropic + xai 12 Deferred tool loading / tool search; Shell execution; Skills (packaged instructions + files); Stand-alone compaction request; Rate-limit headers; Request correlation; Administration API; Audit / compliance data access; Spend limits; Data residency; Self-hosted execution; Multi-agent / subagents
openai + xai + gemini 14 Legacy / secondary chat endpoint; Response storage, retrieval, deletion; Background / deferred (async) execution; JSON mode (valid JSON, no schema); File search / managed RAG; Image generation as a tool / modality inside a text call; Text-to-speech; Speech-to-text / transcription; Realtime speech-to-speech (WebSocket / WebRTC / SIP); Image generation API; Video generation API; Multipart / resumable upload; Pro / extended-compute mode; OpenAI-compatibility layer
anthropic + gemini 3 Web fetch / URL context (retrieve specific URLs); API version header / path version; Scheduled runs
anthropic + xai 1 Anthropic-compatible Messages endpoint
openai + anthropic 6 Programmatic tool calling (model writes code that calls tools); File editing tool; Cache diagnostics; Change effort mid-conversation without breaking the cache; Mid-conversation tool add/remove (cache-preserving); Beta opt-in header
openai + gemini 5 Configurable safety thresholds; Async tools (model continues while a tool runs); Container management API; Audio / video input (understanding); Embeddings
xai + gemini 1 Social / vertical search (X posts, Google Maps)
OpenAI only 15 Legacy completion (prompt-in / text-out) endpoint; Output verbosity control; Log probabilities; Free-form / grammar-constrained custom tools; Tool namespaces; Built-in SaaS connectors; Full-duplex voice with backend delegation (Live) / voice agents; Moderation endpoint; Content provenance; Embeddable chat UI & workspace agents; Fine-tuning / tuning; Evals; Graders (standalone); Stored completions / distillation; Reusable prompt templates
Anthropic only 8 Assistant prefill (continue a partial assistant turn); Server-side model fallback on refusal; Memory tool (client-side persistent memory); Browser use toolset; Advisor (consult a stronger model mid-response); Server-side context editing (clear old tool results / thinking); Outcome grading (define outcome + rubric); Server-side memory stores
xAI only 2 Per-request dollar cost in the response; Retired model ids keep resolving (redirect aliases)
Gemini only 1 Music generation
None 4 Idempotency key; Retired agent / thread APIs; Retired media models / endpoints; Retired text models

# Core generation

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Primary text/multimodal generation endpoint
4/4
Responses API: input (string or Item[]), output[] items
Endpoint: POST /v1/responses
Params: model, input, instructions, tools, text, reasoning
Status: DOCUMENTED · LIVE_VERIFIED
stored by default (store:true, 30 days)
Messages API: messages[] of content blocks, content[] blocks out
Endpoint: POST /v1/messages
Params: model, messages, system, max_tokens, tools, output_config, thinking
Status: DOCUMENTED · LIVE_VERIFIED
stateless; max_tokens required (400 if missing)
Responses API clone: same input items / output[] items as OpenAI (message, reasoning, function_call, web_search_call, code_interpreter_call, mcp_call…); xAI extras max_turns, top_k, min_p, reasoning_effort; usage.cost_in_usd_ticks
Endpoint: POST /v1/responses
Params: model, input, instructions, tools, text, reasoning, max_turns, store
Status: DOCUMENTED · LIVE_VERIFIED
stored by default (30 days); background → 400; metadata → 400; every Grok call bills reasoning tokens
generateContent: contents[] {role: user|model, parts[]} → candidates[].content.parts[]; generationConfig holds every knob; Google calls it 'legacy' since June 2026 in favour of the Interactions API (POST /v1beta/interactions, snake_case, model or agent, input, steps[] out)
Endpoint: POST /v1beta/models/{model}:generateContent
Params: contents, systemInstruction, generationConfig, tools, toolConfig, safetySettings, cachedContent
Status: DOCUMENTED · LIVE_VERIFIED
/v1 stable twin exists for generateContent; Interactions /v1beta BETA + LIVE_VERIFIED, /v1 GA but UNVERIFIED
yes Four envelopes, two shapes: OpenAI and xAI share the items model (xAI reimplements the Responses API); Anthropic uses content blocks; Gemini uses parts inside contents[] with generationConfig. Required fields differ: Anthropic needs max_tokens; Gemini needs the last turn to be user; xAI rejects background/metadata.
Ref: docs/openai/responses.md · docs/anthropic/messages-api.md · docs/xai/responses.md · docs/gemini/generate-content.md · docs/gemini/interactions-api.md
Legacy / secondary chat endpoint
3/4
Chat Completions (messages[], choices[]), still GA and maintained
Endpoint: POST /v1/chat/completions
Params: messages, response_format, reasoning_effort, web_search_options
Status: DOCUMENTED · LIVE_VERIFIED
no hosted tools except search models; list/retrieve/update/delete of stored completions
— not offered
none — Messages is the only chat surface
Chat Completions, OpenAI-compatible; xAI calls it the legacy predecessor of /v1/responses (new features land on Responses first); only function tools; deferred:true + GET /v1/chat/deferred-completion/{id}; reasoning_content field
Endpoint: POST /v1/chat/completions
Params: messages, reasoning_effort, response_format, deferred, prompt_cache_key, max_completion_tokens
Status: DOCUMENTED · LEGACY · LIVE_VERIFIED
file parts → 400 (use Responses); server tools → 422; Live Search params → 410
OpenAI-compatibility layer POST /v1beta/openai/chat/completions (Bearer key required; extra_body.google.thinking_config, cached_content; unknown params silently ignored); native generateContent itself is now labelled legacy vs Interactions
Endpoint: POST /v1beta/openai/chat/completions
Params: messages, reasoning_effort, response_format, extra_body.google
Status: DOCUMENTED · BETA · LIVE_VERIFIED
no Responses / Assistants / audio routes on the compat layer
yes OpenAI, xAI and Gemini all expose an OpenAI-shaped chat/completions; Anthropic has none. Only OpenAI's is a first-class, fully featured surface.
Ref: docs/comparisons/responses-vs-chat-completions.md · docs/xai/chat-completions.md · docs/gemini/openai-compatibility.md
Anthropic-compatible Messages endpoint
2/4
— not offered the native surface
Endpoint: POST /v1/messages
Params: messages, system, max_tokens, tools
Status: DOCUMENTED · LIVE_VERIFIED
POST /v1/messages accepting Anthropic shapes (system, messages, tools[{name,description,input_schema}], tool_choice {auto|any|tool}, thinking blocks with empty signature) — fully deprecated by xAI ('migrate to Responses or gRPC'); Bearer auth, anthropic-version ignored; /v1/complete twin RETIRED (400)
Endpoint: POST /v1/messages
Params: messages, system, max_tokens, tools, tool_choice
Status: DOCUMENTED · DEPRECATED · LIVE_VERIFIED
top_k 400, stop_sequences 400 on reasoning models, document blocks 422, no count_tokens (404), no batches, cache_control ignored
— not offered yes xAI is the only third party exposing the Anthropic Messages wire format, and is retiring it.
Ref: docs/xai/messages-compat.md · parameters.json (xai POST /v1/messages)
Legacy completion (prompt-in / text-out) endpoint
1/4
gpt-3.5-turbo-instruct, davinci-002, babbage-002 only; shutdowns 2026-09-28
Endpoint: POST /v1/completions
Params: prompt, suffix, max_tokens
Status: DOCUMENTED · LEGACY · LIVE_VERIFIED
documented as Legacy but every live call returns 400 'has been deprecated'
Endpoint: POST /v1/complete
Params: prompt, max_tokens_to_sample
Status: DOCUMENTED · LEGACY · DEPRECATED · FAILED_VERIFICATION
removed from Python SDK v1
POST /v1/completions (OpenAI shape) and POST /v1/complete (Anthropic shape) both answer 400 'Raw sampling is not supported for reasoning models' for every live model, incl. the non-reasoning grok-4.20
Endpoint: POST /v1/completions
Params: prompt, max_tokens
Status: DOCUMENTED · LEGACY · RETIRED
gRPC Sample service still lists raw sampling
PaLM-era :generateText / :generateMessage remain in the v1beta discovery document but return 404 / 501; only models/aqa:generateAnswer still answers
Endpoint: POST /v1beta/models/{model}:generateText
Status: DOCUMENTED · LEGACY · FAILED_VERIFICATION
legacy-palm family: 7 endpoints
no Every provider still documents a prompt-completion route; only OpenAI's answers, and it shuts down 2026-09-28.
Ref: docs/openai/completions-legacy.md · docs/anthropic/text-completions-legacy.md · docs/xai/legacy-completions.md · docs/gemini/legacy-palm-methods.md
System / developer prompt
4/4
instructions (per request, not carried by previous_response_id) or developer/system role message items
Endpoint: POST /v1/responses
Params: instructions, input[](message).role
Status: DOCUMENTED · LIVE_VERIFIED
instructions participate in the cache prefix
top-level system (string or text blocks with cache_control/citations); mid-conversation role:system messages for tool changes/effort (beta)
Endpoint: POST /v1/messages
Params: system, system[].cache_control, messages[].output_config.effort
Status: DOCUMENTED · LIVE_VERIFIED
schema lists a system role but docs say use the top-level field
instructions (Responses) or a system / developer role message; Chat: system role; /v1/messages: system string or text blocks
Endpoint: POST /v1/responses
Params: instructions, input[](message).role
Status: DOCUMENTED · LIVE_VERIFIED
instructions + previous_response_id → 400 (the previous system prompt is reused)
systemInstruction (Content, text parts only, role ignored) — not a turn; Interactions: system_instruction string (must be resent when chaining)
Endpoint: POST /v1beta/models/{model}:generateContent
Params: systemInstruction
Status: DOCUMENTED · LIVE_VERIFIED
counted in promptTokenCount; cacheable inside cachedContents
yes All four keep the system prompt outside the turn list. OpenAI/xAI accept it either as a field or as a role item; Anthropic and Gemini only as a field.
Ref: parameters.json (instructions / system / systemInstruction)
Message roles
4/4
user, assistant, system, developer (+ phase: commentary|final_answer on assistant)
Endpoint: POST /v1/responses
Params: input[](message).role, input[](message).phase
Status: DOCUMENTED · LIVE_VERIFIED
user, assistant; consecutive same-role turns are merged; ≤100,000 messages
Endpoint: POST /v1/messages
Params: messages[].role
Status: DOCUMENTED · LIVE_VERIFIED
system role reserved for beta mid-conversation blocks
user, assistant, system, developer (Responses); Chat adds tool (tool_call_id) and legacy function; no ordering constraint
Endpoint: POST /v1/responses
Params: input[](message).role, messages[].role
Status: DOCUMENTED · LIVE_VERIFIED
developer accepted live on Chat although not in the spec
user, model only (omitted role = user); last turn must be user (400 'Requests ending with a model turn are not supported'); function results go in a user Content of functionResponse parts
Endpoint: POST /v1beta/models/{model}:generateContent
Params: contents[].role
Status: DOCUMENTED · LIVE_VERIFIED
Interactions steps use role: user|model too
yes The assistant role is spelled model on Gemini; Anthropic enforces alternation by merging; Gemini enforces a trailing user turn; OpenAI/xAI accept any order.
Ref: docs/anthropic/messages-api.md §1 · docs/gemini/generate-content.md
Assistant prefill (continue a partial assistant turn)
1/4
no prefill semantics; an assistant item is history only
Endpoint: POST /v1/responses
Status: DOCUMENTED
end messages with an assistant turn → model continues it. Deprecated: 400 on Claude 4.6+ / Fable / Mythos, incompatible with thinking and structured outputs
Endpoint: POST /v1/messages
Params: messages[].content[] (assistant prefill)
Status: DOCUMENTED · DEPRECATED · LIVE_VERIFIED
live: Haiku 4.5 OK, Sonnet 5 → 400
assistant history items / assistant role accepted, but no documented continuation semantics
Endpoint: POST /v1/responses
Status: DOCUMENTED
impossible: a request ending with a model turn is rejected (400)
Endpoint: POST /v1beta/models/{model}:generateContent
Status: DOCUMENTED · LIVE_VERIFIED
no Only Anthropic ever offered prefill, and only its 4.5 models still honour it; Gemini rejects the pattern outright.
Ref: docs/anthropic/messages-api.md · deprecations.json · docs/gemini/generate-content.md
Multi-turn conversation
4/4
three modes: manual replay of output[] items, previous_response_id (server keeps chain), conversation (Conversations API)
Endpoint: POST /v1/responses
Params: previous_response_id, conversation, store
Status: DOCUMENTED · LIVE_VERIFIED
previous_response_id and conversation are mutually exclusive
manual replay only: resend full messages[] each call (incl. thinking/tool_use/tool_result blocks)
Endpoint: POST /v1/messages
Params: messages
Status: DOCUMENTED · LIVE_VERIFIED
server-side state exists only in Managed Agents sessions
manual replay or previous_response_id (server rehydrates the whole agentic trajectory incl. reasoning and tool outputs; follow-ups may change tools/model); Chat: replay incl. reasoning_content
Endpoint: POST /v1/responses
Params: previous_response_id, store, input
Status: DOCUMENTED · LIVE_VERIFIED
no Conversations API; x-grok-conv-id / prompt_cache_key are cache-routing keys, not stored conversations
generateContent: manual replay of contents[] (echo thoughtSignature on Gemini 3 function calls); Interactions: previous_interaction_id chains stored interactions (store:true default; only history is carried — resend tools/system/config)
Endpoint: POST /v1beta/models/{model}:generateContent
Params: contents, previous_interaction_id, store
Status: DOCUMENTED · LIVE_VERIFIED
chaining on an in_progress interaction → 400
yes Portable pattern = manual replay everywhere; OpenAI, xAI and Gemini (Interactions) add server-side chaining by id.
Ref: docs/comparisons/state-management.md
Response storage, retrieval, deletion
3/4
store:true default → GET/DELETE /v1/responses/{id}, GET …/input_items; 30-day retention
Endpoint: GET /v1/responses/{response_id}
Params: store, include
Status: DOCUMENTED · LIVE_VERIFIED
store:false → 404 on retrieve, reasoning returned as encrypted_content
no response store; nothing to retrieve after the HTTP response
Status: DOCUMENTED
Batches results are retrievable 29 days
store:true default → GET/DELETE /v1/responses/{id}, GET …/input_items (limit 1–100, order, after); 30-day retention; ZDR teams cannot store
Endpoint: GET /v1/responses/{response_id}
Params: store
Status: DOCUMENTED · LIVE_VERIFIED · LIVE_DISCOVERED
LIVE_DISCOVERED: GET still 200 for a store:false id
Interactions: store:true default → GET /v1beta/interactions/{id} (full timeline incl. user_input), DELETE, POST …/cancel; retention 55 days paid (AI Studio 7/14/28/55), 1 day free; store:false → no id, stateless. generateContent stores nothing (per-request store logging flag LIVE_DISCOVERED)
Endpoint: GET /v1beta/interactions/{id}
Params: store
Status: DOCUMENTED · BETA · LIVE_VERIFIED
GET /v1beta/interactions (list) → 404
yes Three stateful-by-default surfaces (OpenAI Responses, xAI Responses, Gemini Interactions) vs Anthropic's stateless Messages.
Ref: docs/openai/responses.md §1 · docs/xai/responses.md · docs/gemini/interactions-api.md
Background / deferred (async) execution
3/4
background:true → status:queued, poll GET, POST …/cancel, resumable stream ?stream=true&starting_after=N
Endpoint: POST /v1/responses
Params: background
Status: DOCUMENTED · LIVE_VERIFIED
not available over WebSocket; not in EU region
no per-request background mode; use Message Batches (≤24 h) or Managed Agents sessions
Endpoint: POST /v1/messages/batches
Status: DOCUMENTED
Chat Completions only: deferred:true → {request_id}; GET /v1/chat/deferred-completion/{request_id} → 202 while pending, 200 when done (docs: retrievable once within 24 h; live: second GET also 200)
Endpoint: GET /v1/chat/deferred-completion/{request_id}
Params: deferred
Status: DOCUMENTED · LIVE_VERIFIED
Responses background → 400 'Argument not supported'
Interactions background:true (requires store) → queued/in_progress; poll GET /v1beta/interactions/{id} or resume the stream with ?stream=true&last_event_id=; POST …/cancel; mandatory for Deep Research (≤60 min); webhook_config for completion callbacks
Endpoint: POST /v1beta/interactions
Params: background, stream, webhook_config
Status: DOCUMENTED · BETA · LIVE_VERIFIED
Veo/Batch use long-running Operations instead
yes OpenAI, xAI (Chat only) and Gemini (Interactions) offer per-request async; Anthropic only batches.
Ref: docs/openai/responses.md §5 · docs/xai/deferred-completions.md · docs/gemini/interactions-api.md
Token counting
4/4
count input tokens of a Responses payload (free)
Endpoint: POST /v1/responses/input_tokens
Params: model, input, tools, instructions
Status: DOCUMENTED · LIVE_VERIFIED
supports images/files
count tokens of a Messages payload (free; separate RPM bucket 5k/10k/20k)
Endpoint: POST /v1/messages/count_tokens
Params: model, messages, system, tools, thinking, output_config.format
Status: DOCUMENTED · LIVE_VERIFIED
server tools other than advisor → 400; mcp_servers rejected
POST /v1/tokenize-text {model, text} → token_ids[{token_id, string_token, token_bytes}] (tokenizer, not a request counter); free
Endpoint: POST /v1/tokenize-text
Params: model, text
Status: DOCUMENTED · LIVE_VERIFIED
no per-request counter; /v1/messages/count_tokens → 404
:countTokens with {contents} or {generateContentRequest:{model, contents, systemInstruction, tools, cachedContent, generationConfig}} → totalTokens, promptTokensDetails[], cachedContentTokenCount; free
Endpoint: POST /v1beta/models/{model}:countTokens
Params: contents, generateContentRequest
Status: DOCUMENTED · LIVE_VERIFIED
SDKs refuse system_instruction/tools here although REST accepts them
yes Free everywhere; xAI tokenises raw text only (no tools/images), the other three count a full request.
Ref: docs/anthropic/token-counting.md · docs/openai/multimodal-input.md · docs/xai/index.md · docs/gemini/token-counting.md
Model listing / catalogue
4/4
list/retrieve (shutdown_date field on deprecated ids); some aliases (e.g. gpt-5.6) 404 on GET but work on POST
Endpoint: GET /v1/models
Status: DOCUMENTED · LIVE_VERIFIED
136 ids live
list/retrieve with capabilities block (thinking, effort, structured_outputs, context_management…)
Endpoint: GET /v1/models
Params: limit, before_id, after_id
Status: DOCUMENTED · LIVE_VERIFIED
11 ids live; invite-only Mythos ids → 404
GET /v1/models (OpenAI shape, 12 ids) plus typed catalogues GET /v1/language-models, /v1/image-generation-models, /v1/video-generation-models, /v1/embedding-models with live price ticks, aliases, input_modalities, fingerprint; retired slugs redirect (grok-3 → grok-4.3 object)
Endpoint: GET /v1/language-models
Status: DOCUMENTED · LIVE_VERIFIED
voice models absent from the catalogue; /v1/embedding-models → {models: []} for this team
GET /v1beta/models (58 ids: inputTokenLimit, outputTokenLimit, supportedGenerationMethods, thinking, temperature/topP/topK defaults, version) and GET /v1beta/models/{model}; /v1/models lists only 22 stable ids
Endpoint: GET /v1beta/models
Params: pageSize, pageToken
Status: DOCUMENTED · LIVE_VERIFIED
agents (deep-research, antigravity) appear as models; shut-down previews still listed
yes xAI publishes prices in the catalogue, Anthropic publishes capability flags, Gemini publishes token limits and generation methods, OpenAI publishes shutdown dates.
Ref: sources/*/models-api-raw.json

# Structured output & sampling

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Structured outputs (JSON Schema constrained)
4/4
text.format {type:json_schema, name, schema, strict, description} (flat)
Endpoint: POST /v1/responses
Params: text.format, text.format(json_schema).strict
Status: DOCUMENTED · LIVE_VERIFIED
Chat: response_format.json_schema{…} wrapper; refusal → refusal content part
output_config.format {type:json_schema, schema} (GA, no beta header; legacy output_format → 400)
Endpoint: POST /v1/messages
Params: output_config.format, output_config.format.schema
Status: DOCUMENTED · LIVE_VERIFIED
grammar compiled once (24 h cache), first call slower; stop_reason: refusal possible
text.format {type:json_schema, name, schema, strict, description} (Responses) / response_format {type:json_schema, json_schema:{name, strict, schema}} (Chat); Draft 2020-12 preferred; additionalProperties defaults false; formats date/time/email/uuid/uri enforced; pattern ECMA subset
Endpoint: POST /v1/responses
Params: text.format, response_format
Status: DOCUMENTED · LIVE_VERIFIED
strict accepted and ignored (always strict); all Grok 4 models
generationConfig.responseMimeType: application/json + responseJsonSchema (JSON Schema) or legacy responseSchema (OpenAPI subset, propertyOrdering); new canonical responseFormat.text {mimeType: APPLICATION_JSON, schema}; enum mode text/x.enum; XML/YAML mime types accepted; Interactions response_format {type:text, mime_type, schema}
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.responseMimeType, generationConfig.responseJsonSchema, generationConfig.responseSchema, generationConfig.responseFormat
Status: DOCUMENTED · LIVE_VERIFIED
values are syntactically valid but not semantically validated; maxOutputTokens can truncate the JSON; SO + tools = Gemini 3 preview
yes Same task on all four. Schema dialects differ: OpenAI/xAI require additionalProperties:false semantics (xAI defaults it), Anthropic rejects numeric/string constraints, Gemini ignores unsupported keywords silently and offers an OpenAPI-style alternative.
Ref: docs/openai/structured-outputs.md · docs/anthropic/structured-outputs.md · docs/xai/structured-outputs.md · docs/gemini/structured-outputs.md
Strict tool arguments
4/4
tools[type=function].strict:true (Responses omits → tries strict then falls back)
Endpoint: POST /v1/responses
Params: tools[type=function].strict
Status: DOCUMENTED · LIVE_VERIFIED
tools[].strict:true grammar-constrained tool_use.input; ≤20 strict tools/request
Endpoint: POST /v1/messages
Params: tools[].strict
Status: DOCUMENTED · LIVE_VERIFIED
not on toolsets / mcp_toolset / programmatic callers
tool parameters are always strictly enforced ('strict flag implicitly true'); explicit strict accepted and ignored
Endpoint: POST /v1/responses
Params: tools[].strict
Status: DOCUMENTED · LIVE_VERIFIED
toolConfig.functionCallingConfig.mode: VALIDATED (schema-validated constrained decoding; default when built-ins or structured output are combined) or ANY (forced, constrained)
Endpoint: POST /v1beta/models/{model}:generateContent
Params: toolConfig.functionCallingConfig.mode
Status: DOCUMENTED · LIVE_VERIFIED
no per-tool flag; ANY may reject very large/deep schemas
yes Per-tool flag on OpenAI/Anthropic, always-on on xAI, a request-level mode on Gemini.
Ref: tools.json · docs/tools/gemini/function-calling.md
JSON mode (valid JSON, no schema)
3/4
text.format {type:json_object}; prompt must mention JSON (400 otherwise)
Endpoint: POST /v1/responses
Params: text.format(json_object)
Status: DOCUMENTED · LIVE_VERIFIED
legacy
— not offered text.format {type:json_object} / response_format {type:json_object}
Endpoint: POST /v1/responses
Params: text.format(json_object)
Status: DOCUMENTED · LIVE_VERIFIED
responseMimeType: application/json without a schema
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.responseMimeType
Status: DOCUMENTED · LIVE_VERIFIED
yes Anthropic only offers schema-constrained output.
Ref: docs/openai/structured-outputs.md · docs/gemini/structured-outputs.md
Output verbosity control
1/4
text.verbosity: low|medium|high
Endpoint: POST /v1/responses
Params: text.verbosity
Status: DOCUMENTED · LIVE_VERIFIED
no direct knob; output_config.effort shapes length indirectly
Status: DOCUMENTED
no verbosity parameter (text.format only)
Status: DOCUMENTED
no verbosity parameter; thinkingLevel and maxOutputTokens only
Status: DOCUMENTED
no OpenAI-only knob.
Ref: parameters.json
Sampling parameters (temperature / top_p / top_k)
4/4
temperature 0–2, top_p; rejected by reasoning models
Endpoint: POST /v1/responses
Params: temperature, top_p
Status: DOCUMENTED · LIVE_VERIFIED
no top_k
temperature 0–1, top_p, top_k — DEPRECATED: 400 when non-default on Claude 4.7+ / Fable / Mythos; Python SDK v1 removed the kwargs
Endpoint: POST /v1/messages
Params: temperature, top_p, top_k
Status: DOCUMENTED · DEPRECATED · LIVE_VERIFIED
temperature 0–2, top_p, plus xAI-specific top_k (≥1) and min_p (0–1); accepted on reasoning models (echo 0.7/0.95); Chat seed → system_fingerprint
Endpoint: POST /v1/responses
Params: temperature, top_p, top_k, min_p, seed
Status: DOCUMENTED · LIVE_VERIFIED
presence_penalty/frequency_penalty/stop → 400 on reasoning models
generationConfig.temperature 0–2 (default 1.0), topP (0.95), topK (64), seed; DEPRECATED guidance since 2026-07-21: keep defaults on Gemini 3.x (temperature < 1 can cause looping); presencePenalty/frequencyPenalty → 400 'not enabled'
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.temperature, generationConfig.topP, generationConfig.topK, generationConfig.seed
Status: DOCUMENTED · DEPRECATED · LIVE_VERIFIED
candidateCount > 1 → 400 on 3.x
yes Sampling knobs are shrinking on all frontier lines: rejected by OpenAI reasoning models and Claude 4.7+, deprecated-by-guidance on Gemini 3.x, still fully accepted on Grok.
Ref: deprecations.json (anthropic, gemini api_features) · parameters.json
Stop sequences
4/4
Chat Completions stop (≤4); not a Responses parameter
Endpoint: POST /v1/chat/completions
Params: stop
Status: DOCUMENTED
stop_sequences[] → stop_reason: stop_sequence + stop_sequence
Endpoint: POST /v1/messages
Params: stop_sequences
Status: DOCUMENTED · LIVE_VERIFIED
Chat stop (≤4) and /v1/messages stop_sequences — 400 on reasoning models; not on Responses
Endpoint: POST /v1/chat/completions
Params: stop, stop_sequences
Status: DOCUMENTED · LIVE_VERIFIED
usable only on grok-4.20-0309-non-reasoning
generationConfig.stopSequences[] (≤5; 6 → 400); Interactions generation_config.stop_sequences
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.stopSequences
Status: DOCUMENTED · LIVE_VERIFIED
yes Absent from both Responses APIs; Gemini allows 5, the others 4.
Ref: parameters.json
Log probabilities
1/4
top_logprobs + include: ["message.output_text.logprobs"]
Endpoint: POST /v1/responses
Params: top_logprobs, include
Status: DOCUMENTED
— not offered logprobs/top_logprobs (0–8) accepted but silently ignored on grok-4.20 and newer (DEPRECATED); include: message.output_text.logprobs ignored
Endpoint: POST /v1/chat/completions
Params: logprobs, top_logprobs
Status: DOCUMENTED · DEPRECATED · LIVE_VERIFIED
responseLogprobs / logprobs (0–20) in the schema but 400 Logprobs is not enabled for this model on every current model
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.responseLogprobs, generationConfig.logprobs
Status: DOCUMENTED · FAILED_VERIFICATION
no Only OpenAI still returns logprobs; xAI and Gemini keep the parameters as dead compatibility fields.
Ref: parameters.json
Refusal / safety signalling in-band
4/4
output[].content[] {type: refusal} (structured outputs) / status: incomplete + incomplete_details.reason: content_filter; HTTP 403 misalignment_policy_violation, cyber_policy
Endpoint: POST /v1/responses
Status: DOCUMENTED
HTTP 200 with stop_reason: refusal + stop_details {category, explanation}; beta fallbacks re-runs on another model
Endpoint: POST /v1/messages
Params: fallbacks, fallback_credit_token
Status: DOCUMENTED · BETA
categories: cyber, bio, frontier_llm, reasoning_extraction, general_harms
Chat message.refusal field; respect_moderation flag on image/video results; usage-guideline violations are billed (+ $0.05 fee when caught pre-generation on Responses); no error code catalogue for refusals
Endpoint: POST /v1/chat/completions
Status: DOCUMENTED
HTTP 200 with no candidates + promptFeedback.blockReason (SAFETY|OTHER|BLOCKLIST|PROHIBITED_CONTENT|IMAGE_SAFETY) for prompt blocks; candidates[].finishReason SAFETY|RECITATION|SPII|IMAGE_SAFETY|… (21 values) + safetyRatings[] for output blocks; thresholds via safetySettings[]
Endpoint: POST /v1beta/models/{model}:generateContent
Params: safetySettings, candidates[].finishReason, promptFeedback.blockReason
Status: DOCUMENTED · LIVE_VERIFIED
Interactions mirror finishReason as snake_case error codes
yes All four signal in-band with different shapes; only Gemini lets the caller tune thresholds; only xAI charges for violating requests.
Ref: docs/anthropic/stop-reasons.md · docs/openai/safety.md · docs/gemini/safety.md · docs/xai/pricing.md §5
Server-side model fallback on refusal
1/4
— not offered fallbacks parameter (+ fallback_credit_token billing credit)
Endpoint: POST /v1/messages
Params: fallbacks, fallback_credit_token
Status: DOCUMENTED · BETA
beta server-side-fallback-2026-07-01, fallback-credit-2026-07-01
— not offered — not offered no Anthropic-only.
Ref: generated/fragments/headers/anthropic-beta-headers.json
Configurable safety thresholds
2/4
inline moderation {model, policy} on Responses/Chat (policy selection, not thresholds)
Endpoint: POST /v1/responses
Params: moderation
Status: DOCUMENTED
— not offered — not offered safetySettings[] {category: HARM_CATEGORY_HARASSMENT|HATE_SPEECH|SEXUALLY_EXPLICIT|DANGEROUS_CONTENT|JAILBREAK (CIVIC_INTEGRITY deprecated → enableEnhancedCivicAnswers), threshold: OFF|BLOCK_NONE|BLOCK_ONLY_HIGH|BLOCK_MEDIUM_AND_ABOVE|BLOCK_LOW_AND_ABOVE}; default Off on 2.5/3.x; safetyRatings[] returned when a threshold is set
Endpoint: POST /v1beta/models/{model}:generateContent
Params: safetySettings
Status: DOCUMENTED · LIVE_VERIFIED
Interactions: custom safety settings not supported
no Only Gemini exposes per-category blocking thresholds; OpenAI selects a moderation policy.
Ref: docs/gemini/safety.md · docs/openai/moderation.md

# Tools — client side

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Function / custom tools (you execute)
4/4
tools[type=function] {name, description, parameters, strict, output_schema, async, defer_loading, allowed_callers} → function_call item → you reply function_call_output {call_id, output}
Endpoint: POST /v1/responses
Params: tools[type=function], input[](function_call_output)
Status: DOCUMENTED · LIVE_VERIFIED
38 compatible models listed
tools[] {name, description, input_schema, strict, input_examples, cache_control, defer_loading, allowed_callers} → tool_use block → you reply user message of tool_result {tool_use_id, content, is_error}
Endpoint: POST /v1/messages
Params: tools, messages[].content[] (tool_result)
Status: DOCUMENTED · LIVE_VERIFIED
13 compatible models
Responses tools[type=function] {name, description, parameters} → function_call {call_id: call-…, name, arguments} → function_call_output; Chat tools[{type:function, function:{…}}] → message.tool_calls[] → {role: tool, tool_call_id}; ≤350 tools; parameters always strict; a client-side call ends the agentic request (fresh max_turns budget on the follow-up)
Endpoint: POST /v1/responses
Params: tools[type=function], input[](function_call_output), parallel_tool_calls
Status: DOCUMENTED · LIVE_VERIFIED
all 7 Grok text models; also inside the voice session and /v1/messages
tools[].functionDeclarations[] {name, description, parameters | parametersJsonSchema, response | responseJsonSchema, behavior} → parts[].functionCall {name, args, id} (+ mandatory thoughtSignature on Gemini 3) → you reply a user Content of functionResponse {name, id, response, parts[] (multimodal)}; max 512 declarations
Endpoint: POST /v1beta/models/{model}:generateContent
Params: tools[].functionDeclarations, contents[].parts[].functionResponse, toolConfig.functionCallingConfig
Status: DOCUMENTED · LIVE_VERIFIED
Interactions: {type: function, name, description, parameters} + function_call/function_result steps; Live: toolCall/toolResponse messages
yes Same loop on all four. Arguments are a JSON string on OpenAI/xAI and an object on Anthropic (input) and Gemini (args); results are a top-level item (OpenAI/xAI), a leading tool_result block in a user turn (Anthropic) or functionResponse parts in a user Content (Gemini). Gemini 3 additionally requires echoing the thoughtSignature of the first call of each step (400 / MISSING_THOUGHT_SIGNATURE).
Ref: docs/comparisons/tool-execution.md · docs/tools/xai/function-calling.md · docs/tools/gemini/function-calling.md
Free-form / grammar-constrained custom tools
1/4
tools[type=custom] {format: {type:text} | {type:grammar, syntax: lark|regex, definition}}
Endpoint: POST /v1/responses
Params: tools[type=custom].format
Status: DOCUMENTED · LIVE_VERIFIED
— not offered no custom tool type in the deserializer (422 'unknown variant'); no grammar mode
Endpoint: POST /v1/responses
Status: DOCUMENTED
— not offered no OpenAI-only.
Ref: tools.json · docs/xai/structured-outputs.md
Tool choice
4/4
tool_choice: auto|none|required string, {type:function,name}, hosted {type:web_search|mcp|shell|apply_patch|…}, {type:allowed_tools, mode, tools[]}
Endpoint: POST /v1/responses
Params: tool_choice, tool_choice(allowed_tools)
Status: DOCUMENTED · LIVE_VERIFIED
tool_choice {type: auto|any|tool|none, name?, disable_parallel_tool_use?}
Endpoint: POST /v1/messages
Params: tool_choice.type, tool_choice.disable_parallel_tool_use
Status: DOCUMENTED · LIVE_VERIFIED
any/tool → 400 on Fable 5.1 / Mythos 5.1 and with manual thinking
tool_choice: auto|none|required or {type:function, name} (Responses) / {type:function, function:{name}} (Chat); /v1/messages: {type: auto|any|tool} (disable_parallel_tool_use → 400); no allowed_tools, no forcing of server tools
Endpoint: POST /v1/responses
Params: tool_choice, tool_choice.type, tool_choice.name
Status: DOCUMENTED · LIVE_VERIFIED
toolConfig.functionCallingConfig {mode: AUTO|ANY|NONE|VALIDATED, allowedFunctionNames[]} (ANY + names = forced subset); Interactions generation_config.tool_choice: auto|any|none|validated or {allowed_tools:{mode, tools[]}}; built-in tools cannot be forced
Endpoint: POST /v1beta/models/{model}:generateContent
Params: toolConfig.functionCallingConfig.mode, toolConfig.functionCallingConfig.allowedFunctionNames
Status: DOCUMENTED · LIVE_VERIFIED
Live setup.toolConfig → close 1007
yes required ≈ any ≈ ANY; only OpenAI and Gemini (allowedFunctionNames / Interactions allowed_tools) can restrict to a subset; only OpenAI can force a hosted tool.
Ref: parameters.json
Parallel tool calls
4/4
parallel_tool_calls (default true); hosted tools never batched with functions
Endpoint: POST /v1/responses
Params: parallel_tool_calls
Status: DOCUMENTED · LIVE_VERIFIED
default on (Claude 4+); tool_choice.disable_parallel_tool_use:true to force ≤1; all results in one user message
Endpoint: POST /v1/messages
Params: tool_choice.disable_parallel_tool_use
Status: DOCUMENTED · LIVE_VERIFIED
parallel_tool_calls (default true; false = at most one call) on Chat and Responses
Endpoint: POST /v1/responses
Params: parallel_tool_calls
Status: DOCUMENTED · LIVE_VERIFIED
/v1/messages disable_parallel_tool_use → 400
several functionCall parts in one model Content; answer all of them in one user Content (interleaving FC1,FR1,FC2 → 400); only the first call carries the thought signature; no on/off switch
Endpoint: POST /v1beta/models/{model}:generateContent
Params: contents[].parts[].functionCall
Status: DOCUMENTED · LIVE_VERIFIED
compositional_function_calling chains calls across turns
yes Default-on everywhere; Gemini has no switch to disable it.
Ref: docs/openai/tool-loop.md · docs/tools/anthropic/tool-use-loop.md · docs/tools/gemini/function-calling.md
Tool namespaces
1/4
tools[type=namespace] {name, description, tools[]}; calls carry namespace
Endpoint: POST /v1/responses
Params: tools[type=namespace]
Status: DOCUMENTED · LIVE_VERIFIED
no namespaces; toolsets (computer_toolset_20260801, mcp_toolset) group Anthropic-defined members only
Status: DOCUMENTED
— not offered — not offered no OpenAI-only.
Ref: tools.json
Deferred tool loading / tool search
3/4
tools[type=tool_search] {execution: server|client} + defer_loading:true on function/custom/mcp; tool_search_call/tool_search_output items; additional_tools input item
Endpoint: POST /v1/responses
Params: tools[type=tool_search], tools[type=function].defer_loading
Status: DOCUMENTED · LIVE_VERIFIED
gpt-5.4+ only; gpt-5.4-nano lacks it
tool_search_tool_regex_20251119 / tool_search_tool_bm25_20251119 (server) + defer_loading:true (≤10,000 tools); tool_reference blocks expanded server-side; client-side search via tool_reference in tool_result
Endpoint: POST /v1/messages
Params: tools[].defer_loading, tools[].type
Status: DOCUMENTED · LIVE_VERIFIED
GA no header; not cacheable together with cache_control
tools[type=tool_search] {execution} + defer_loading:true on function/mcp tools → tool_search_call {arguments:{query, limit}} / tool_search_output {tools[]} — alpha: 403 'only available for alpha users'
Endpoint: POST /v1/responses
Params: tools[type=tool_search], tools[].defer_loading
Status: DOCUMENTED · ACCOUNT_RESTRICTED
listed as compatible with all 7 text models
— not offered yes OpenAI-shaped on xAI (gated), two algorithms on Anthropic; Gemini has no deferred loading (best practice: keep 10–20 active declarations).
Ref: docs/tools/openai/tool-search.md · docs/tools/anthropic/tool-search.md · tools.json (xai tool_search)
Programmatic tool calling (model writes code that calls tools)
2/4
tools[type=programmatic_tool_calling] + allowed_callers:["programmatic"]; program/program_output items; nested calls caller.type: program; JavaScript in isolated V8
Endpoint: POST /v1/responses
Params: tools[type=programmatic_tool_calling], tools[type=function].allowed_callers
Status: DOCUMENTED · UNVERIFIED
no compatible model list in tools.json
code_execution_20260120+ + allowed_callers:["code_execution_20260120"]; Python await tool({...}) in the sandbox; caller {type, tool_id}; reply requires top-level container
Endpoint: POST /v1/messages
Params: tools[].allowed_callers, container
Status: DOCUMENTED · LIVE_VERIFIED
not Haiku 4.5 (400)
— not offered not a generateContent feature; nearest: compositional function calling (chained calls across turns) and code execution over functionResponse data; Antigravity agents script tools inside their sandbox
Endpoint: POST /v1beta/models/{model}:generateContent
Status: DOCUMENTED
yes OpenAI and Anthropic only (JS vs Python).
Ref: docs/tools/anthropic/programmatic-tool-calling.md · docs/openai/tool-loop.md §6 · docs/tools/gemini/function-calling.md
Async tools (model continues while a tool runs)
2/4
tools[type=function].async:true + wait tool / task handles (GPT-6 Astra+)
Endpoint: POST /v1/responses
Params: tools[type=function].async, input[](function_call).async
Status: DOCUMENTED
— not offered — not offered Live API only: functionDeclarations[].behavior: NON_BLOCKING (default on gemini-3.8-live) + toolResponse.functionResponses[].scheduling: INTERRUPT|WHEN_IDLE|SILENT, willContinue; generateContent → 400
Endpoint: WSS BidiGenerateContent
Params: setup.tools[].functionDeclarations[].behavior, toolResponse.functionResponses[].scheduling
Status: DOCUMENTED · LIVE_VERIFIED
gemini-3.8-live-extended-thinking: async only
yes Different scopes: OpenAI in text Responses, Gemini in the voice Live session.
Ref: parameters.json · docs/gemini/live-api.md
Shell execution
3/4
tools[type=shell] hosted (environment.type: container_auto|container_reference) or local (type: local — you run commands); legacy local_shell → 400
Endpoint: POST /v1/responses
Params: tools[type=shell].environment
Status: DOCUMENTED · LIVE_VERIFIED
gpt-5.2+, codex, gpt-6-astra
bash_20250124 client tool (you run a persistent bash) or server bash_code_execution sub-tool of code_execution_20250825+
Endpoint: POST /v1/messages
Params: tools[].type
Status: DOCUMENTED · LIVE_VERIFIED
all current models
tools[type=shell] {environment:{type: local, skills:[{name, description, path}]}} → shell_call → you reply shell_call_output {stdout, stderr, outcome}; local only (no hosted container variant); hosted Python via code_interpreter
Endpoint: POST /v1/responses
Params: tools[type=shell].environment, input[](shell_call_output)
Status: DOCUMENTED · LIVE_VERIFIED
accepted live, not invoked; Grok Build CLI runs its own sandbox
no shell tool on generateContent; the Antigravity managed agent runs bash/python/node code_execution and file tools inside its Linux sandbox (Interactions API only, PREVIEW)
Endpoint: POST /v1beta/interactions
Params: agent_config(antigravity)
Status: DOCUMENTED · PREVIEW
gemini-3.1-pro-preview-customtools is tuned for bash-style custom tools you define yourself
yes OpenAI has hosted + local, Anthropic hosted (code execution) + local (bash), xAI local only, Gemini only inside its managed agent.
Ref: docs/tools/openai/shell.md · docs/tools/anthropic/bash.md · docs/xai/skills-api.md · docs/gemini/interactions-api.md
File editing tool
2/4
tools[type=apply_patch] → apply_patch_call {operation: create_file|update_file|delete_file, diff}; you reply apply_patch_call_output
Endpoint: POST /v1/responses
Params: tools[type=apply_patch]
Status: DOCUMENTED · LIVE_VERIFIED
undocumented SSE events response.apply_patch_call_operation_diff.* (LIVE_DISCOVERED)
text_editor_20250728 (name: str_replace_based_edit_tool, max_characters) → commands view/str_replace/create/insert; server variant text_editor_code_execution
Endpoint: POST /v1/messages
Params: tools[].max_characters
Status: DOCUMENTED · LIVE_VERIFIED
older 20250429/20250124 → 400 on current models
— not offered no editor tool on generateContent; Antigravity 09-2026 built-ins write_to_file, replace_file_content, view_file, list_dir, find_by_name, grep_search (agent sandbox only)
Endpoint: POST /v1beta/interactions
Status: DOCUMENTED · PREVIEW
05-2026 tool names deprecated → 2026-10-05
yes Diff-based (OpenAI) vs command-based (Anthropic); Gemini only inside Antigravity; none on xAI.
Ref: docs/tools/openai/apply-patch.md · docs/tools/anthropic/text-editor.md · docs/gemini/interactions-api.md
Memory tool (client-side persistent memory)
1/4
none in Responses; Agents API sessions persist items but no memory tool
Status: DOCUMENTED
memory_20250818 client tool (you store files under /memories)
Endpoint: POST /v1/messages
Params: tools[].type
Status: DOCUMENTED · LIVE_VERIFIED
Managed Agents add server-side memory stores
— not offered — not offered no Anthropic-only.
Ref: docs/tools/anthropic/memory.md
Computer use
3/4
tools[type=computer] (current, gpt-5.4+/gpt-6) or computer_use_preview (+ computer-use-preview model, RETIRED); computer_call {action|actions[], pending_safety_checks} → computer_call_output {computer_screenshot, acknowledged_safety_checks}
Endpoint: POST /v1/responses
Params: tools[type=computer], tools[type=computer_use_preview]
Status: DOCUMENTED · PREVIEW · ACCOUNT_RESTRICTED
computer UNVERIFIED live
computer_toolset_20260801 (GA, no header, Fable/Mythos/Opus 5/Sonnet 5/Opus 4.8) or beta computer_20251124 (computer-use-2025-11-24) / computer_20250124; member tools screenshot/zoom/click…; batch actions
Endpoint: POST /v1/messages
Params: tools[].display_width_px, tools[].enable_zoom, tools[].configs
Status: DOCUMENTED · LIVE_VERIFIED · BETA
~4,500-token toolset definition
— not offered tools[].computerUse {environment: ENVIRONMENT_BROWSER|MOBILE|DESKTOP, excludedPredefinedFunctions[], enablePromptInjectionDetection, disabledSafetyPolicies[]} → predefined functionCalls (click, type, scroll, navigate, open_app…; coordinates 0–999) → you reply functionResponse with a screenshot; safety_decision: require_confirmation → safety_acknowledgement; Interactions {type: computer_use, environment: browser}
Endpoint: POST /v1beta/models/{model}:generateContent
Params: tools[].computerUse, tools[].computerUse.environment
Status: DOCUMENTED · PREVIEW · ACCOUNT_RESTRICTED
gemini-3.8-flash recommended; legacy gemini-2.5-computer-use-preview-10-2025 browser-only; no free tier
yes Three client-executed screenshot loops (OpenAI, Anthropic, Gemini); none on xAI. Gemini normalises coordinates to 0–999 and adds mobile/desktop environments.
Ref: docs/tools/openai/computer-use.md · docs/tools/anthropic/computer-use.md · docs/tools/gemini/computer-use.md
Browser use toolset
1/4
— not offered browser_toolset_20260801 (client toolset, browser_state result blocks)
Endpoint: POST /v1/messages
Params: tools[].type
Status: DOCUMENTED · LIVE_VERIFIED
≈6,600-token definition; 7 models
— not offered browser is an environment of the computer-use tool, not a separate toolset
Endpoint: POST /v1beta/models/{model}:generateContent
Status: DOCUMENTED
no Anthropic-only as a dedicated toolset.
Ref: tools.json

# Tools — server side / hosted

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Web search
4/4
tools[type=web_search] {search_context_size, user_location, filters.allowed_domains, external_web_access, return_token_budget, search_content_types}; web_search_call item + url_citation annotations; $10/1k calls
Endpoint: POST /v1/responses
Params: tools[type=web_search]
Status: DOCUMENTED · LIVE_VERIFIED
Chat: search models + web_search_options; preview variants LEGACY
web_search_20260318|20260209|20250305 {max_uses, allowed_domains XOR blocked_domains, user_location}; server_tool_use + web_search_tool_result blocks, web_search_result_location citations; $10/1k searches; dynamic filtering via code execution (20260209+)
Endpoint: POST /v1/messages
Params: tools[].max_uses, tools[].allowed_domains, tools[].blocked_domains, tools[].user_location
Status: DOCUMENTED · LIVE_VERIFIED
13 models; pause_turn after 10 iterations
tools[type=web_search] {allowed_domains ≤5 XOR excluded_domains ≤5, enable_image_understanding, enable_image_search} → web_search_call {action: search|open_page|find_in_page} + url_citation annotations + inline [[N]](url) (disable via include: no_inline_citations); $5/1k successful calls; search_context_size → 400
Endpoint: POST /v1/responses
Params: tools[type=web_search], tools[].allowed_domains, tools[].excluded_domains, tools[].enable_image_search
Status: DOCUMENTED · LIVE_VERIFIED
Responses only; Chat Live Search → 410; usage server_side_tool_usage_details.web_search_calls
tools[{googleSearch: {timeRangeFilter?, searchTypes?{webSearch, imageSearch}}}] → groundingMetadata {webSearchQueries, searchEntryPoint (must be displayed — ToS), groundingChunks[].web, groundingSupports}; Gemini 3.x 5,000 free queries/month then $14/1k queries, 2.5: 1,500 RPD free then $35/1k grounded prompts; legacy googleSearchRetrieval (dynamic retrieval) DEPRECATED
Endpoint: POST /v1beta/models/{model}:generateContent
Params: tools[].googleSearch, tools[].googleSearch.searchTypes, tools[].googleSearchRetrieval
Status: DOCUMENTED · ACCOUNT_RESTRICTED
429 limit: 0 on this free-tier key; Interactions {type: google_search}; the only tool allowed with functions in the Live API
yes Same task on all four with four price models ($10 / $10 / $5 per 1k calls; Gemini per query with a free monthly quota). Domain filters exist on OpenAI, Anthropic and xAI; Gemini offers time-range and image-search filters instead.
Ref: docs/tools/openai/web-search.md · docs/tools/anthropic/web-search.md · docs/tools/xai/web-search.md · docs/tools/gemini/google-search-grounding.md
Social / vertical search (X posts, Google Maps)
2/4
— not offered — not offered tools[type=x_search] {allowed_x_handles ≤10 XOR excluded_x_handles, from_date, to_date, enable_image_understanding, enable_video_understanding} → live item custom_tool_call named x_keyword_search/x_semantic_search (docs: x_search_call) + url_citation to x.com posts; $5/1k calls until 2026-09-21, then $5/1k posts + $10/1k profiles fetched
Endpoint: POST /v1/responses
Params: tools[type=x_search], tools[].allowed_x_handles, tools[].from_date, tools[].to_date
Status: DOCUMENTED · LIVE_VERIFIED · LIVE_DISCOVERED
also in the voice session
tools[{googleMaps: {enableWidget?}}] + toolConfig.retrievalConfig {latLng, languageCode} → groundingChunks[].maps {uri, title, placeId, text, placeAnswerSources}; GA; text only; not in Live; Gemini 3.x $14/1k queries after 5,000/month (tools table: $25/1k grounded prompts, 1,500 RPD free)
Endpoint: POST /v1beta/models/{model}:generateContent
Params: tools[].googleMaps, toolConfig.retrievalConfig.latLng
Status: DOCUMENTED · GA · LIVE_VERIFIED
pricing inconsistency flagged (DOCUMENTATION_INCOMPLETE)
no Provider-specific data sources; no equivalent on OpenAI/Anthropic.
Ref: docs/tools/xai/x-search.md · docs/tools/gemini/google-maps-grounding.md
Web fetch / URL context (retrieve specific URLs)
2/4
no dedicated fetch tool; web search may open pages; input_file.file_url fetches PDFs only
Status: DOCUMENTED
web_fetch_20260318|20260309|20260209|20250910 {max_uses, allowed_domains, citations, max_content_tokens, use_cache, url_sources}; only URLs already in context; free (tokens only)
Endpoint: POST /v1/messages
Params: tools[].max_content_tokens, tools[].url_sources, tools[].use_cache
Status: DOCUMENTED · LIVE_VERIFIED
url_not_in_prior_context error
no fetch tool; web_search_call.action.type: open_page|find_in_page shows the search tool opening pages; input_file.file_url attaches a remote file (attachment search)
Endpoint: POST /v1/responses
Status: DOCUMENTED
tools[{urlContext: {}}] (no options): ≤20 public URLs per request, ≤34 MB each (HTML/JSON/text/CSV/RTF/PNG/JPEG/PDF; no YouTube/Workspace/paywalls) → urlContextMetadata.urlMetadata[] {retrievedUrl, urlRetrievalStatus}; free tool, content billed as input (toolUsePromptTokenCount)
Endpoint: POST /v1beta/models/{model}:generateContent
Params: tools[].urlContext
Status: DOCUMENTED · GA · LIVE_VERIFIED
Interactions {type: url_context}; combinable with search/code exec/functions
yes Anthropic and Gemini offer a fetch tool; Anthropic restricts to URLs already in context, Gemini to any public URL you name.
Ref: docs/tools/anthropic/web-fetch.md · docs/tools/gemini/url-context.md
File search / managed RAG
3/4
Vector Stores API (create, files, file_batches, search) + tools[type=file_search] {vector_store_ids, max_num_results, filters, ranking_options}; $2.50/1k calls + $0.10/GB/day
Endpoint: POST /v1/vector_stores · POST /v1/responses
Params: tools[type=file_search]
Status: DOCUMENTED · LIVE_VERIFIED
16 endpoints
no vector store; patterns: search_result blocks (your RAG, citable), document blocks from Files API, code execution over uploaded files
Endpoint: POST /v1/messages
Params: messages[].content[] {type:'search_result'}, messages[].content[] {type:'document'}
Status: DOCUMENTED · LIVE_VERIFIED
Collections API (/v1/collections, documents added from Files ids, index_configuration.model_name: grok-embedding-small, chunk_configuration, POST /v1/documents/search {query, retrieval_mode: hybrid|semantic|keyword}) + tools[type=file_search|collections_search] {vector_store_ids (= collection ids), max_num_results, filters, ranking_options} → file_search_call {queries, results[{file_id, filename, score, text}]}; $2.50/1k calls + $0.10/GiB/day; implicit attachment_search over input_file parts $10/1k
Endpoint: POST /v1/responses
Params: tools[type=file_search], tools[].vector_store_ids
Status: DOCUMENTED · LIVE_VERIFIED · ACCOUNT_RESTRICTED
13 collection endpoints; documented on management-api.x.ai but working on api.x.ai (LIVE_DISCOVERED); search 404 while indexing
File Search stores (/v1beta/fileSearchStores, :uploadToFileSearchStore resumable ≤100 MB, :importFile, documents, chunkingConfig, customMetadata[]; embedding model fixed at creation) + tools[{fileSearch: {fileSearchStoreNames[], metadataFilter (AIP-160), topK}}] → groundingChunks[].retrievedContext {title, text, pageNumber, customMetadata}; indexing $0.15/1M tokens once, storage + queries free; store quota 1 GB (free) … 1 TB (Tier 3)
Endpoint: POST /v1beta/models/{model}:generateContent
Params: tools[].fileSearch, tools[].fileSearch.fileSearchStoreNames, tools[].fileSearch.metadataFilter
Status: DOCUMENTED · PREVIEW · LIVE_VERIFIED
12 endpoints; not combinable with Search/URL context; not in Live
yes Three hosted stores (OpenAI vector stores, xAI collections, Gemini file-search stores) with different billing (per call / per call + storage / per indexed token); Anthropic supplies citation plumbing instead of storage.
Ref: docs/openai/vector-stores.md · docs/anthropic/citations.md · docs/xai/collections.md · docs/tools/gemini/file-search.md
Code execution (sandboxed Python)
4/4
tools[type=code_interpreter] {container: auto|cntr_id, file_ids, memory_limit 1g–64g, network_policy}; code_interpreter_call {code, outputs}; $0.03–$1.92 per 20-min session (per-minute since 2026-06-02)
Endpoint: POST /v1/responses
Params: tools[type=code_interpreter].container
Status: DOCUMENTED · LIVE_VERIFIED
36 models
code_execution_20260521|20260120|20250825 (bash + text editor sub-tools, Python 3.11, 5 GiB RAM, no internet); container {id, skills} reuse (30-day state); 1,550 free container-hours/org/month then $0.05/h; free with web_search/fetch 20260209+
Endpoint: POST /v1/messages
Params: tools[].type, container
Status: DOCUMENTED · LIVE_VERIFIED
13 models; usage counter code_execution_requests missing live
tools[type=code_interpreter] (alias code_execution) → code_interpreter_call {code, outputs[{type: logs, logs: <JSON string stdout/stderr/exit_code>} | {type: image, url}]} (outputs only with include: code_interpreter_call.outputs); Python + NumPy/Pandas/Matplotlib/SciPy, no network; $5/1k calls + tokens; no container object
Endpoint: POST /v1/responses
Params: tools[type=code_interpreter], include
Status: DOCUMENTED · LIVE_VERIFIED
all 7 text models; gRPC rejects the code_interpreter alias
tools[{codeExecution: {}}] → parts executableCode {language: PYTHON, code} + codeExecutionResult {outcome, output} (+ inlineData PNG for matplotlib); Python ≥3.10, fixed library set, 30 s per run, ≤5 retries, no pip/network; no fee — code and results billed as output then input tokens
Endpoint: POST /v1beta/models/{model}:generateContent
Params: tools[].codeExecution
Status: DOCUMENTED · LIVE_VERIFIED
12 models; Interactions {type: code_execution}; not in Live
yes Four sandboxes, four price models: per container-session (OpenAI), per container-hour with a free tier (Anthropic), per call (xAI), tokens only (Gemini). Only OpenAI/Anthropic expose a reusable container object.
Ref: docs/tools/openai/code-interpreter.md · docs/tools/anthropic/code-execution.md · docs/tools/xai/code-execution.md · docs/tools/gemini/code-execution.md
Container management API
2/4
CRUD containers and container files, download outputs
Endpoint: GET/POST/DELETE /v1/containers[/{id}/files]
Status: DOCUMENTED · LIVE_VERIFIED
9 endpoints; auto containers expire 20 min after last activity
no container endpoints; container id returned on the message (container.id, expires_at) and reused via the container param; outputs via Files API
Endpoint: POST /v1/messages · GET /v1/files/{id}/content
Params: container
Status: DOCUMENTED · LIVE_VERIFIED
no container object; container param accepted for OpenAI compatibility, not needed
Endpoint: POST /v1/responses
Status: DOCUMENTED
Environments API for Interactions agents (POST/GET/DELETE /v1beta/environments, files GET …/files/{path}, PUT /upload/v1beta/environments/{env}/files/{path}; sources repository/GCS/inline; network allowlist; idle 15 min, deleted after 7 days) — not usable by codeExecution
Endpoint: POST /v1beta/environments
Params: environment, sources, network
Status: DOCUMENTED · BETA · PREVIEW · LIVE_VERIFIED
7 endpoints; sandbox compute unbilled during preview
yes OpenAI containers and Gemini environments are both first-class resources but serve different tools (code interpreter vs managed agents).
Ref: docs/openai/containers.md · docs/tools/anthropic/code-execution.md · docs/gemini/interactions-api.md
Image generation as a tool / modality inside a text call
3/4
tools[type=image_generation] {model, quality, size, background, input_fidelity, partial_images, moderation}; image_generation_call item; streaming partials
Endpoint: POST /v1/responses
Params: tools[type=image_generation]
Status: DOCUMENTED · LIVE_VERIFIED
31 models
— not offered tools[type=image_generation] → image_generation_call billed at Imagine per-image rates (no call fee); accepted live, not invoked
Endpoint: POST /v1/responses
Params: tools[type=image_generation]
Status: DOCUMENTED · LIVE_VERIFIED
all 7 text models listed
not a tool: generationConfig.responseModalities: [TEXT, IMAGE] + imageConfig {aspectRatio, imageSize 512|1K|2K|4K} on image models (gemini-3.1-flash-image, -lite-image, gemini-3-pro-image); Interactions response_format {type: image}
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.responseModalities, generationConfig.imageConfig
Status: DOCUMENTED · ACCOUNT_RESTRICTED
no free tier (limit: 0 here); text models ignore IMAGE modality
yes Tool item on OpenAI/xAI, output modality on Gemini; Claude outputs text only.
Ref: tools.json · docs/gemini/image-generation.md · docs/xai/images.md
Remote MCP servers
4/4
tools[type=mcp] {server_label, server_url|connector_id|tunnel_id, authorization, headers, allowed_tools, require_approval (default always), defer_loading, allowed_callers}; mcp_list_tools, mcp_call, mcp_approval_request/response items; no beta header
Endpoint: POST /v1/responses
Params: tools[type=mcp]
Status: DOCUMENTED · LIVE_VERIFIED
42 models; also Realtime and Agents API
mcp_servers[] {type:url, url, name, authorization_token} (≤20) + tools[type=mcp_toolset] {mcp_server_name, default_config, configs}; mcp_tool_use/mcp_tool_result blocks; beta header mcp-client-2025-11-20
Endpoint: POST /v1/messages
Params: mcp_servers, tools[].mcp_server_name, tools[].default_config, tools[].configs
Status: DOCUMENTED · BETA · LIVE_VERIFIED
no approvals in Messages; the error text advertises mcp-client-2026-09-15 which is rejected (inconsistency)
tools[type=mcp] {server_url, server_label (required), server_description, allowed_tools[], authorization, headers} → mcp_call {server_label, name, arguments, output, error}; tokens only; no approval round-trip (require_approval silently accepted); also inside the voice session
Endpoint: POST /v1/responses
Params: tools[type=mcp], tools[].server_url, tools[].server_label, tools[].allowed_tools
Status: DOCUMENTED · LIVE_VERIFIED
LIVE_VERIFIED with DeepWiki; connector_id unsupported
server-side tools[].mcpServers[] {name, streamableHttpTransport {url, headers, timeout}} in the discovery schema (UNVERIFIED, no guide) and Interactions {type: mcp_server, name, url, headers, allowed_tools} (documented; not on Gemini 3 per the overview); SDK-side MCP (mcpToTool(), Python ClientSession in tools) runs the calls in your process (BETA)
Endpoint: POST /v1beta/interactions
Params: tools[].mcpServers, tools[](mcp_server)
Status: DOCUMENTED · DOCUMENTATION_INCOMPLETE · UNVERIFIED
Streamable HTTP only; server names must not contain '-'
yes OpenAI GA with approvals; xAI GA without approvals; Anthropic beta without approvals; Gemini server-side MCP is documented but unverified — its verified path is SDK-side execution.
Ref: docs/tools/openai/mcp-and-connectors.md · docs/tools/anthropic/mcp-connector.md · docs/tools/xai/mcp.md · docs/tools/gemini/mcp.md
Built-in SaaS connectors
1/4
connector_id: dropbox, gmail, googlecalendar, googledrive, microsoftteams, outlookcalendar, outlookemail, sharepoint (+ OAuth authorization) — deprecated for models after 2025-09-01
Endpoint: POST /v1/responses
Params: tools[type=mcp].connector_id
Status: DOCUMENTED · DEPRECATED
— not offered connector_id documented as unsupported
Endpoint: POST /v1/responses
Status: DOCUMENTED
— not offered no OpenAI-only (deprecated).
Ref: docs/tools/openai/mcp-and-connectors.md
Private-network MCP (tunnels) / agent credentials
3/4
Secure MCP Tunnel: outbound tunnel-client, tools[type=mcp].tunnel_id (tunnel_[a-z0-9]{32})
Endpoint: POST /v1/responses
Params: tools[type=mcp].tunnel_id
Status: DOCUMENTED
MCP tunnels API (/v1/tunnels, certificates, tokens) + tunnel agent; header mcp-tunnels-2026-06-22; WIF bearer with workspace:manage_tunnels
Endpoint: GET/POST /v1/tunnels
Status: DOCUMENTED · BETA · PREVIEW · ACCOUNT_RESTRICTED
10 endpoints; older /v1/organizations/tunnels deprecated
— not offered no tunnels; Credentials API (/v1beta/credentials: bearer_token, oauth2, environment_variable with injection_location and trusted_domains; secrets write-only) referenced from environment network allowlists
Endpoint: POST /v1beta/credentials
Status: DOCUMENTED · BETA · PREVIEW
5 endpoints, not tested
no Tunnels on OpenAI/Anthropic; Gemini solves the secret-injection half with credentials; nothing on xAI.
Ref: endpoints.json (managed-agents tunnels, gemini credentials) · docs/tools/openai/mcp-and-connectors.md
Skills (packaged instructions + files)
3/4
Skills API (POST /v1/skills, versions, content zip) referenced from tools[type=shell].environment.skills[], containers, Agents sessions; ≤500 files, ≤25 MB
Endpoint: POST /v1/skills
Params: tools[type=shell].environment.skills
Status: DOCUMENTED · LIVE_VERIFIED · FAILED_VERIFICATION
version-by-number endpoints returned 404 live; lists returned empty
Skills API (/v1/skills, versions, content; GA no header) + Anthropic skills (pptx, xlsx, docx, pdf); used via container.skills[] with code execution and on Managed Agents
Endpoint: POST /v1/skills
Params: container.skills
Status: DOCUMENTED · LIVE_VERIFIED
9 endpoints all LIVE_VERIFIED
Skills API GET/POST /v1/skills, GET/DELETE /v1/skills/{id}, GET …/content (zip; SKILL.md frontmatter name, description, when-to-use, paths, allowed-tools) — all 404 for this team (ACCOUNT_RESTRICTED); reach inference only via shell.environment.skills[]; Grok Build CLI also loads skills locally
Endpoint: POST /v1/skills
Params: tools[type=shell].environment.skills
Status: DOCUMENTED · ACCOUNT_RESTRICTED
5 endpoints, OpenAPI only
no Skills API; custom Interactions agents mount .agents/skills/<name>/SKILL.md from their environment sources
Endpoint: POST /v1beta/interactions
Status: DOCUMENTED · PREVIEW
yes Same SKILL.md bundle model on OpenAI, Anthropic and xAI (xAI gated); Gemini uses environment files.
Ref: docs/openai/skills-api.md · docs/anthropic/skills-api.md · docs/xai/skills-api.md · docs/gemini/interactions-api.md
Advisor (consult a stronger model mid-response)
1/4
— not offered advisor_20260301 {model, max_tokens, caching} server tool, beta advisor-tool-2026-03-01
Endpoint: POST /v1/messages
Params: tools[].model, tools[].max_tokens, tools[].caching
Status: DOCUMENTED · BETA · LIVE_VERIFIED
11 models
— not offered — not offered no Anthropic-only.
Ref: tools.json
Citations / grounding metadata
4/4
output_text.annotations[]: url_citation, file_citation, container_file_citation, file_path (produced by web/file search, code interpreter)
Endpoint: POST /v1/responses
Params: include
Status: DOCUMENTED · LIVE_VERIFIED
no citations for plain documents
citations {enabled:true} on document/search_result blocks → char_location, page_location, content_block_location, search_result_location, web_search_result_location; citations_delta when streaming
Endpoint: POST /v1/messages
Params: messages[].content[].citations, tools[].citations
Status: DOCUMENTED · LIVE_VERIFIED
all-or-nothing across documents
output_text.annotations[] {type: url_citation, url, start_index, end_index, title} (live indices 0/0, title = URL) + inline markdown [[N]](url); collections citations collections://<cid>/files/<fid>; include: web_search_call.action.sources
Endpoint: POST /v1/responses
Params: include
Status: DOCUMENTED · LIVE_VERIFIED
usage.num_sources_used
candidates[].groundingMetadata {groundingChunks[] (web|maps|retrievedContext), groundingSupports[] {segment, groundingChunkIndices, confidenceScores}, webSearchQueries, searchEntryPoint}, urlContextMetadata, citationMetadata (recitation); Interactions text_annotation_delta
Endpoint: POST /v1beta/models/{model}:generateContent
Params: candidates[].groundingMetadata, candidates[].urlContextMetadata
Status: DOCUMENTED · LIVE_VERIFIED
streaming chunks carry only new grounding chunks — accumulate
yes Anthropic cites any document you pass; OpenAI, xAI and Gemini cite only their own tool results (Gemini with segment-level supports and confidence scores).
Ref: docs/anthropic/citations.md · docs/tools/gemini/google-search-grounding.md

# Multimodal input & media APIs

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Image input
4/4
input_image {image_url|file_id, detail: low|high|auto|original}; patch/tile token formulas; ≤1,500 images, ≤512 MB
Endpoint: POST /v1/responses
Params: input[](message).content[](input_image)
Status: DOCUMENTED · LIVE_VERIFIED
image {source: base64|url|file, transformations}; ≤100 (200k models) / 600 (1M models) images; 10 MB each; tokens = ⌈w/28⌉×⌈h/28⌉ (2576 px hi-res on 4.7+)
Endpoint: POST /v1/messages
Params: messages[].content[] {type:'image'}
Status: DOCUMENTED · LIVE_VERIFIED
Responses input_image {image_url (https or data URL), detail} (≥512 px; file_id variant → 400); Chat image_url {url, detail}; image tokens priced as text input (image_input = input price); all Grok 4 text models accept images
Endpoint: POST /v1/responses
Params: input[](message).content[](input_image), messages[].content[](image_url)
Status: DOCUMENTED · LIVE_VERIFIED
usage.prompt_tokens_details.image_tokens
parts[].inlineData {mimeType, data} or fileData {fileUri} (Files API); mediaResolution LOW/MEDIUM/HIGH globally or per part (Gemini 3 adds ULTRA_HIGH, 2,240 tokens); docs 280/560/1,120/2,240 tokens by level (observed 256/529/1,089/2,209)
Endpoint: POST /v1beta/models/{model}:generateContent
Params: contents[].parts[].inlineData, contents[].parts[].fileData, generationConfig.mediaResolution
Status: DOCUMENTED · LIVE_VERIFIED
promptTokensDetails[] {modality: IMAGE}
yes Universal. Token accounting differs on every provider (patches/tiles, 28-px grid, text-price tokens, resolution levels).
Ref: docs/openai/multimodal-input.md · docs/anthropic/vision-and-documents.md · docs/xai/chat-completions.md · docs/gemini/multimodal-input.md
PDF / document input
4/4
input_file {file_id|file_url|file_data+filename, detail}; text + page images in context; <50 MB combined; office formats via file_id
Endpoint: POST /v1/responses
Params: input[](message).content[](input_file)
Status: DOCUMENTED
document {source: base64 pdf|url|file|text|content, title, context, citations}; ≤600 pages/request, 32 MB body
Endpoint: POST /v1/messages
Params: messages[].content[] {type:'document'}
Status: DOCUMENTED · LIVE_VERIFIED
citable
Responses input_file {file_id|file_url|file_data} → routed through the implicit attachment search tool ($10/1k calls); Chat file parts → 400; /v1/messages document → 422
Endpoint: POST /v1/responses
Params: input[](message).content[](input_file)
Status: DOCUMENTED · LIVE_VERIFIED
Files API 50 MB (spec) / 512 MB (guide)
inlineData {mimeType: application/pdf} or fileData (Files API, 2 GB); PDF pages billed at the image token rate (DOCUMENT modality, 560 tokens/page in countTokens vs IMAGE 520 observed in generateContent); pdf_input true on all 3.x text models
Endpoint: POST /v1beta/models/{model}:generateContent
Params: contents[].parts[].inlineData, contents[].parts[].fileData
Status: DOCUMENTED · LIVE_VERIFIED
no citations for documents; URL context handles remote PDFs
yes Universal; only Anthropic makes documents citable, only xAI bills a per-call fee for attachments.
Ref: docs/anthropic/vision-and-documents.md · docs/xai/files.md · docs/gemini/multimodal-input.md
Audio / video input (understanding)
2/4
Chat Completions input_audio {data, format} on gpt-audio-1.5, gpt-4o-audio-preview; Responses input_audio part in schema but UNVERIFIED; no video input
Endpoint: POST /v1/chat/completions
Params: messages[](user).content[](input_audio)
Status: DOCUMENTED · LIVE_DISCOVERED
— not offered no audio/video parts on text models (audio only in the voice session and STT; grok-imagine-video accepts video/audio as generation inputs; view_x_video sub-tool understands X videos)
Status: DOCUMENTED
native: inlineData/fileData audio (WAV/MP3/AIFF/AAC/OGG/FLAC; ≈25–32 tokens/s) and video (≤1 fps frames + audio; videoMetadata {startOffset, endOffset, fps}; YouTube URLs via fileData) on every 3.x text model; gemini-embedding-2 embeds audio/video too
Endpoint: POST /v1beta/models/{model}:generateContent
Params: contents[].parts[].inlineData, contents[].parts[].videoMetadata
Status: DOCUMENTED · LIVE_VERIFIED
audio_input price rows on 2.5/3.1-lite/3-flash; 3.5+ single price
yes Gemini is the only provider with native audio and video understanding in the text API; OpenAI has audio-in on dedicated chat models.
Ref: docs/openai/multimodal-input.md §4 · docs/gemini/multimodal-input.md
Text-to-speech
3/4
POST /v1/audio/speech (gpt-4o-mini-tts, tts-1, tts-1-hd; SSE speech.audio.delta) and Chat modalities:[text,audio]
Endpoint: POST /v1/audio/speech
Params: model, input, voice, instructions
Status: DOCUMENTED · LIVE_VERIFIED
custom voices ACCOUNT_RESTRICTED
— not offered POST /v1/tts {text ≤60,000 chars, voice_id, language (required), output_format {codec mp3|wav|pcm|mulaw|alaw, sample_rate, bit_rate}, speed, with_timestamps} + wss://api.x.ai/v1/tts streaming (text.delta → audio.delta); 28 built-in voices (GET /v1/tts/voices) + custom voices (/v1/custom-voices, Enterprise); $15 / 1M characters
Endpoint: POST /v1/tts
Params: text, voice_id, language, output_format
Status: DOCUMENTED · LIVE_VERIFIED
17 voice endpoints; voice models absent from GET /v1/models
generateContent on TTS models (gemini-3.1-flash-tts-preview, gemini-2.5-flash|pro-preview-tts) with responseModalities: [AUDIO] + speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName (30 voices) or multiSpeakerVoiceConfig (≤2 speakers); raw PCM 24 kHz out; $1 text in / $20 audio out per 1M (≈ $0.03/min); free tier (10 req/day observed)
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.responseModalities, generationConfig.speechConfig
Status: DOCUMENTED · PREVIEW · LIVE_VERIFIED
streaming TTS on 3.1 only; multi-speaker may return finishReason: OTHER
yes Dedicated endpoint (OpenAI, xAI) vs a generation modality (Gemini); none on Anthropic.
Ref: docs/openai/audio.md · docs/xai/voice.md · docs/gemini/speech-generation.md
Speech-to-text / transcription
3/4
POST /v1/audio/transcriptions (gpt-transcribe, gpt-4o-transcribe(-diarize), whisper-1; streaming transcript.text.delta), POST /v1/audio/translations (whisper-1)
Endpoint: POST /v1/audio/transcriptions
Params: model, file, response_format, stream
Status: DOCUMENTED · LIVE_VERIFIED
whisper-1 & gpt-4o-transcribe shutdown 2027-02-26
— not offered POST /v1/stt (multipart file ≤500 MB or url; language, diarize, keyterm, vad_threshold, model: grok-voice-transcribe-2.0|1.0) $0.10/h; wss://api.x.ai/v1/stt streaming (binary frames, interim_results, endpointing, smart_turn → transcript.partial/done) $0.20/h
Endpoint: POST /v1/stt
Params: file, url, language, diarize, model
Status: DOCUMENTED · LIVE_VERIFIED
25 languages; default model contradicts between release notes (1.0) and model page (2.0)
gemini-3.5-transcribe via generateContent + generationConfig.audioTranscriptionConfig {languageCodes, customVocabulary, mode VERBATIM|SMART, diarization, wordTimestamp} → parts[].audioTranscription {text, speakerLabel, words[]} (LIVE_DISCOVERED shape); ≤1 h audio; ≈ $0.005/min blended; gemini-3.5-transcribe-live over the Live WebSocket (10-min sessions); gemini-3.5-live-translate-preview speech-to-speech translation
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.audioTranscriptionConfig
Status: DOCUMENTED · LIVE_VERIFIED
free tier available
yes OpenAI/xAI dedicated endpoints (per minute / per hour); Gemini a dedicated model behind the generic endpoint; none on Anthropic.
Ref: docs/openai/audio.md · docs/xai/voice.md · docs/gemini/transcription.md
Realtime speech-to-speech (WebSocket / WebRTC / SIP)
3/4
Realtime API GA: WebSocket wss://api.openai.com/v1/realtime?model=, WebRTC POST /v1/realtime/calls, SIP, ephemeral POST /v1/realtime/client_secrets; transcription & translation sessions; 60-min sessions
Endpoint: WS /v1/realtime
Params: session.update, response.create
Status: DOCUMENTED · LIVE_VERIFIED
legacy /v1/realtime/sessions → 404
— not offered wss://api.x.ai/v1/realtime?model=grok-voice-latest (= grok-voice-think-fast-2.0) with the OpenAI Realtime event vocabulary (session.update, input_audio_buffer.*, conversation.item.create, response.create → response.output_audio.delta, response.done; 39 server / 9 client events); tools in-session (function, web_search, x_search, file_search, mcp); reasoning.effort high|none; ephemeral POST /v1/realtime/client_secrets (≤3600 s); SIP via /v2/phone-numbers, /v1/realtime/calls/{id}/refer|hangup, webhook realtime.call.incoming; $0.08/min + $0.004 per text item; 120-min sessions; concurrent sessions 10–200 by tier
Endpoint: WS wss://api.x.ai/v1/realtime
Params: session.update, conversation.item.create, response.create
Status: DOCUMENTED · LIVE_VERIFIED
text turn LIVE_VERIFIED; audio UNVERIFIED; ping undocumented
Live API wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent (setup → setupComplete; clientContent/realtimeInput {audio|video|text|activityStart|activityEnd}/toolResponse → serverContent {modelTurn, turnComplete, interrupted, inputTranscription, outputTranscription}, toolCall, goAway, sessionResumptionUpdate); models gemini-3.8-live (native audio, interleaved thinking, GA 2026-09-15), -extended-thinking, gemini-3.1-flash-live-preview, 2.5 native audio; ephemeral tokens POST /v1beta/auth_tokens on the …Constrained method; 16 kHz PCM in / 24 kHz out; ≈10-min connections, 15-min audio sessions (unlimited with contextWindowCompression), resumption handles 2 h; audio $3 in / $12 out per 1M (≈ $0.005 / $0.018 per min); free tier
Endpoint: WSS BidiGenerateContent
Params: setup, realtimeInput, toolResponse, setup.sessionResumption, setup.contextWindowCompression
Status: DOCUMENTED · PREVIEW · LIVE_VERIFIED
32 message types recorded; responseModalities: [TEXT] → close 1007; only googleSearch + functions as tools
yes Three voice stacks: xAI copies OpenAI's Realtime protocol (so client code ports almost verbatim), Gemini's Live API is a distinct message family; Anthropic has none. Only OpenAI and xAI offer WebRTC/SIP.
Ref: docs/openai/realtime.md · docs/xai/voice.md · docs/gemini/live-api.md · docs/gemini/live-events.md
Full-duplex voice with backend delegation (Live) / voice agents
1/4
Live API: POST /v1/live/sessions (WebRTC), WS /v1/live/sessions, fork/attach/SIP; gpt-live-1 $0.05/min; delegation to Responses or client
Endpoint: POST /v1/live/sessions
Params: session.delegation, session.store
Status: DOCUMENTED · LIVE_DISCOVERED
— not offered no separate delegation product; the realtime session itself runs server tools (web/x/file search, MCP) with reasoning.effort
Endpoint: WS wss://api.x.ai/v1/realtime
Status: DOCUMENTED
no delegation layer; gemini-3.8-live-extended-thinking runs background reasoning and speaks fillers (interactionStatus: IN_PROGRESS); proactivity.proactiveAudio always on for 3.8
Endpoint: WSS BidiGenerateContent
Status: DOCUMENTED
no OpenAI-only as a product; xAI and Gemini fold agentic behaviour into the voice session itself.
Ref: docs/openai/live.md · docs/xai/voice.md · docs/gemini/live-api.md
Image generation API
3/4
POST /v1/images/generations, /edits (gpt-image-2, gpt-image-2.5-flare/sunburst; DALL·E RETIRED, gpt-image-1.x DEPRECATED); streaming partials; arbitrary sizes ≤4K
Endpoint: POST /v1/images/generations
Params: prompt, size, quality, background, output_format, stream
Status: DOCUMENTED · LIVE_VERIFIED
/variations RETIRED (404)
— not offered POST /v1/images/generations {model, prompt, n 1–10, response_format url|b64_json, aspect_ratio (incl. 21:9, 5:2, auto), resolution 1k|1.5k|2k, quality low|medium|auto, storage_options} and POST /v1/images/edits (JSON only: image {url|file_id} or images[] 2–5); grok-imagine-image $0.02, grok-imagine-image-2.0 $0.04–$0.08, grok-imagine-image-quality $0.05 (DEPRECATED → 2026-11-02); sync ~5–10 s; always JPEG
Endpoint: POST /v1/images/generations
Params: prompt, n, aspect_ratio, resolution, quality, response_format
Status: DOCUMENTED · LIVE_VERIFIED
no size/mask/seed; grok-2-image RETIRED; not on us.api.x.ai
generateContent on image models with responseModalities + imageConfig {aspectRatio (14 ratios), imageSize 512|1K|2K|4K}; editing = pass the image as input and iterate; SynthID watermark; gemini-3.1-flash-image $0.045–$0.151, -lite-image $0.034, gemini-3-pro-image $0.134/$0.24; gemini-2.5-flash-image DEPRECATED → 2026-10-02; Imagen (:predict) RETIRED 2026-08-17; OpenAI-compat POST /v1beta/openai/images/generations subset
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.responseModalities, generationConfig.imageConfig.aspectRatio, generationConfig.imageConfig.imageSize
Status: DOCUMENTED · ACCOUNT_RESTRICTED
no free tier (limit: 0); batch 50 %
yes OpenAI and xAI have dedicated image endpoints (token-priced vs per-image); Gemini generates images through the text endpoint; Claude produces no images.
Ref: docs/openai/images.md · docs/xai/images.md · docs/gemini/image-generation.md
Video generation API
3/4
/v1/videos (Sora 2) — DEPRECATED, shutdown 2026-09-24; $0.10–$0.70/s
Endpoint: POST /v1/videos
Params: prompt, seconds, size
Status: DOCUMENTED · DEPRECATED · LIVE_VERIFIED
— not offered POST /v1/videos/generations {model, prompt, duration 1–15, aspect_ratio, resolution 480p|720p|1080p, generate_audio, image, reference_images[], reference_audios[] (1.5), last_frame (1.5)} → {request_id}; GET /v1/videos/{request_id} 202 pending → 200 {status: done, video:{url}}; POST /v1/videos/edits, /extensions; grok-imagine-video $0.05/s, grok-imagine-video-1.5 $0.08/s (1080p, reference-to-video)
Endpoint: POST /v1/videos/generations
Params: prompt, duration, aspect_ratio, resolution, image
Status: DOCUMENTED · LIVE_VERIFIED
4 endpoints; also via Batch (URLs expire 1 h)
Veo 3.1: POST /v1beta/models/veo-3.1-*:predictLongRunning {instances[{prompt, image, lastFrame, referenceImages[], video}], parameters {aspectRatio, resolution 720p|1080p|4k, durationSeconds 4|6|8, personGeneration, negativePrompt, seed}} → Operation; poll GET /v1beta/{name}; download files/{id}:download?alt=media; native audio; $0.40/$0.60 (Veo 3.1), $0.10–$0.30 (Fast), $0.05/$0.08 (Lite) per second; gemini-omni-1.1-flash video-out model via Interactions (≈ $0.10/s); OpenAI-compat POST /v1beta/openai/videos
Endpoint: POST /v1beta/models/{model}:predictLongRunning
Params: instances[].prompt, parameters.resolution, parameters.durationSeconds
Status: DOCUMENTED · PREVIEW
Veo 2.0/3.0 RETIRED 2026-06-30; no free tier; 400 on this key
yes Async polling everywhere; OpenAI's is shutting down, xAI's is the cheapest per second, Gemini's the only one with 4K.
Ref: docs/openai/video.md · docs/xai/videos.md · docs/gemini/video-generation.md
Music generation
1/4
— not offered — not offered — not offered lyria-3.5 (GA, $0.08/song), lyria-3-clip-preview ($0.04 / 30-s clip), lyria-3-pro-preview via generateContent (MP3) or Interactions response_format {type: audio} (WAV); Lyria RealTime WebSocket …BidiGenerateMusic (weightedPrompts, musicGenerationConfig {bpm, density, brightness, scale…}, playbackControl) → 48 kHz stereo PCM chunks (experimental, unpriced)
Endpoint: POST /v1beta/models/{model}:generateContent
Params: contents[].parts[].text, musicGenerationConfig, playbackControl
Status: DOCUMENTED · ACCOUNT_RESTRICTED
no free tier for Lyria 3.x; RealTime LIVE_VERIFIED
no Gemini-only.
Ref: docs/gemini/music-generation.md
Embeddings
2/4
text-embedding-3-small ($0.02/M), -3-large ($0.13/M), ada-002; dimensions, encoding_format; 8,192 tokens/input, 2,048 inputs
Endpoint: POST /v1/embeddings
Params: input, model, dimensions, encoding_format
Status: DOCUMENTED · LIVE_VERIFIED
— not offered POST /v1/embeddings {model, input, encoding_format, dimensions} documented (OpenAPI) but grok-embedding-small → 404 'does not exist or your team does not have access', GET /v1/embedding-models → {models: []}; no published price; used internally to index Collections
Endpoint: POST /v1/embeddings
Params: input, model, dimensions
Status: DOCUMENTED · ACCOUNT_RESTRICTED
gRPC Embedder.Embed exists
POST /v1beta/models/gemini-embedding-2:embedContent {content, outputDimensionality 128–3072, taskType (001 only), embedContentConfig} and :batchEmbedContents {requests[]}; multimodal (text, ≤6 images, ≤180 s audio, ≤120 s video, PDF ≤6 pages → one vector); $0.20 text / $0.45 image / $6.50 audio / $12 video per 1M, batch 50 %; free tier; gemini-embedding-001 DEPRECATED → 2028-05-14; OpenAI-compat /v1beta/openai/embeddings
Endpoint: POST /v1beta/models/{model}:embedContent
Params: content, outputDimensionality, taskType
Status: DOCUMENTED · LIVE_VERIFIED
:asyncBatchEmbedContent ACCOUNT_RESTRICTED (400 FAILED_PRECONDITION)
yes OpenAI (text) and Gemini (multimodal) ship embeddings; xAI's endpoint is documented but inaccessible; Anthropic has none.
Ref: docs/openai/embeddings.md · docs/gemini/embeddings.md · docs/models/xai-models.md
Moderation endpoint
1/4
POST /v1/moderations (omni-moderation-latest, text+image, free) and inline moderation {model, policy} on Responses/Chat
Endpoint: POST /v1/moderations
Params: input, model, moderation
Status: DOCUMENTED · LIVE_VERIFIED
text-moderation-* RETIRED
— not offered — not offered no moderation route among the 125 Gemini endpoints; safetySettings thresholds and safetyRatings instead
Status: DOCUMENTED
no OpenAI-only.
Ref: docs/openai/moderation.md · docs/gemini/safety.md
Content provenance
1/4
POST /v1/content_provenance_checks (C2PA verification)
Endpoint: POST /v1/content_provenance_checks
Status: DOCUMENTED · LIVE_VERIFIED
generated media from code execution carry C2PA credentials; no verification endpoint
Status: DOCUMENTED
no provenance endpoint; respect_moderation flag on image/video results
Status: DOCUMENTED
SynthID watermark on all generated images/video/audio (Nano Banana 2 Lite also C2PA); no verification endpoint on the Developer API
Status: DOCUMENTED
no Only OpenAI exposes a check endpoint.
Ref: endpoints.json · docs/gemini/image-generation.md

# Files & batch

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Files API
4/4
POST /v1/files (purposes user_data, batch, evals, assistants, vision, fine-tune), list/retrieve/delete/content; 512 MB/file; expires_after 1 h–30 d; batch files auto-expire 30 d
Endpoint: POST /v1/files
Params: file, purpose, expires_after
Status: DOCUMENTED · LIVE_VERIFIED
user_data content not downloadable (400)
POST /v1/files (any MIME; 500 MB/file; workspace-scoped), list (ids[]), metadata, content (tool-generated files only), delete; expires_in_seconds 1 h–90 d; GA no header
Endpoint: POST /v1/files
Params: file, expires_in_seconds
Status: DOCUMENTED · LIVE_VERIFIED
not on Bedrock/Vertex
POST /v1/files (multipart; purpose ignored; expires_after 1 h–30 d), list (filter AIP-160, sort_by), get, delete, GET …/content?format=original, public URLs (POST …/public-url, …/revoke; ≤50 MiB, ≤1,000 active/team), chunked upload files:initialize/:uploadChunks (UNVERIFIED); storage $0.025/GiB/day, downloads $0.20/GiB; used by Responses input_file, Collections, Imagine, Batch
Endpoint: POST /v1/files
Params: file, expires_after
Status: DOCUMENTED · LIVE_VERIFIED
10 endpoints; disabled under ZDR
POST /upload/v1beta/files (resumable X-Goog-Upload-* protocol, multipart or media), GET /v1beta/files[/{id}], DELETE, POST /v1beta/files:register {uris: gs://…} (GCS, 30 days), GET /v1beta/generatedFiles, download files/{id}:download?alt=media (generated files); 2 GB/file, 20 GB/project, 48 h TTL, free; used via fileData {fileUri}
Endpoint: POST /upload/v1beta/files
Params: file, displayName, mimeType
Status: DOCUMENTED · LIVE_VERIFIED
unknown file → 403 (never 404); not available on Vertex AI
yes All four are blob stores referenced by id; TTLs range from 48 h (Gemini, fixed) to 90 d (Anthropic); only xAI prices storage and only xAI mints public URLs.
Ref: docs/openai/files-and-uploads.md · docs/anthropic/files-api.md · docs/xai/files.md · docs/gemini/files.md
Multipart / resumable upload
3/4
Uploads API: ≤8 GB in ≤64 MB parts, 1 h TTL
Endpoint: POST /v1/uploads
Params: filename, purpose, bytes, mime_type
Status: DOCUMENTED · LIVE_VERIFIED
— not offered POST /v1/files:initialize + POST /v1/files:uploadChunks (documented, UNVERIFIED)
Endpoint: POST /v1/files:initialize
Status: DOCUMENTED · UNVERIFIED
Google resumable upload protocol on /upload/v1beta/files and /upload/v1beta/fileSearchStores/{s}:uploadToFileSearchStore (start → x-goog-upload-url → upload, finalize; 8 MiB chunk granularity); environments PUT /upload/v1beta/environments/{env}/files/{path} (≤2 GiB)
Endpoint: POST /upload/v1beta/files
Params: X-Goog-Upload-Protocol, X-Goog-Upload-Command
Status: DOCUMENTED · LIVE_VERIFIED
yes Three resumable schemes, all provider-specific.
Ref: docs/openai/files-and-uploads.md · docs/xai/files.md · docs/gemini/files.md
Batch processing (async, discounted)
4/4
JSONL file (custom_id, method, url, body) → POST /v1/batches {input_file_id, endpoint, completion_window:'24h'}; 50,000 requests / 200 MB; endpoints responses, chat, embeddings, completions, moderations, images, videos; 50 % off; output file 30 d
Endpoint: POST /v1/batches
Params: input_file_id, endpoint, completion_window, output_expires_after
Status: DOCUMENTED · LIVE_VERIFIED
cancel LIVE_DISCOVERED 409 on terminal
inline requests[] {custom_id, params} → POST /v1/messages/batches; 100,000 requests / 256 MB; 24 h expiry; results JSONL at results_url 29 d; 50 % off all token dims (stacks with caching); all Messages features incl. server tools; output-300k-2026-03-24 beta
Endpoint: POST /v1/messages/batches
Params: requests, requests[].custom_id, requests[].params
Status: DOCUMENTED · LIVE_VERIFIED
6 endpoints all LIVE_VERIFIED
POST /v1/batches {name} → POST /v1/batches/{id}/requests {batch_requests[{batch_request_id, batch_request: {chat_get_completion | responses | image_generation | image_edit | video_generation | video_extension}}]} (inline, ≤25 MB/request) or JSONL via Files (input_file_id, ≤50,000 lines / 200 MB); GET …/requests, GET …/results, POST …:cancel; 20 % off on grok-4.3 / 4.20 only (grok-4.6 / 4.5 / build unsupported; Imagine at standard rates); bypasses rate limits; not with priority; live turnaround 11 s
Endpoint: POST /v1/batches
Params: name, input_file_id, batch_requests
Status: DOCUMENTED · LIVE_VERIFIED
8 endpoints; DELETE → 405; responses bodies come back as chat_get_completion
POST /v1beta/models/{model}:batchGenerateContent {batch:{displayName, inputConfig:{requests:{requests[]} | fileName}, priority, webhookConfig}} (inline ≤20 MB or JSONL Files ≤2 GB) → Operation batches/{id}; :asyncBatchEmbedContent; GET /v1beta/batches[/{id}], :cancel, DELETE, PATCH :update*Batch; states PENDING→RUNNING→SUCCEEDED|FAILED|CANCELLED|EXPIRED (>48 h); 50 %; target 24 h; 100 concurrent jobs; per-model enqueued-token caps; results 6 weeks; OpenAI-compat /v1beta/openai/batches
Endpoint: POST /v1beta/models/{model}:batchGenerateContent
Params: batch.inputConfig.requests, batch.inputConfig.fileName, batch.displayName
Status: DOCUMENTED · ACCOUNT_RESTRICTED
9 endpoints; free tier → 400 FAILED_PRECONDITION; list LIVE_VERIFIED
yes Same 24-h / discounted idea on all four, with 50 % (OpenAI, Anthropic, Gemini) vs 20 % (xAI, three models only). Inline requests: Anthropic, xAI, Gemini; file-based: OpenAI, xAI, Gemini.
Ref: docs/openai/batch.md · docs/anthropic/message-batches.md · docs/xai/batches.md · docs/gemini/batch.md

# Prompt caching

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Prompt / context caching
4/4
automatic prefix caching (≥1,024 tokens; exact prefix; prompt_cache_key routing hint); GPT-5.6+: prompt_cache_options {mode: implicit|explicit, ttl:'30m', prewarm} and per-part prompt_cache_breakpoint {mode:explicit} (≤4); prompt_cache_retention: in_memory|24h (deprecated)
Endpoint: POST /v1/responses
Params: prompt_cache_key, prompt_cache_options, prompt_cache_retention, prompt_cache_breakpoint
Status: DOCUMENTED · LIVE_VERIFIED
reads 0.1× (5.6+) / model-specific; writes 1.25× only on GPT-5.6+
explicit cache_control {type: ephemeral, ttl: 5m|1h} on system/tools/content blocks (≤4 breakpoints) or top-level automatic; 20-block lookback; min 512 (Fable/Mythos/Opus 5) · 1,024 (Sonnet 5/4.6/4.5, Opus 4.8) · 2,048 (Opus 4.7) · 4,096 (Haiku 4.5, Opus 4.6/4.5)
Endpoint: POST /v1/messages
Params: cache_control, system[].cache_control, messages[].content[].cache_control, tools[].cache_control
Status: DOCUMENTED · LIVE_VERIFIED
writes 1.25× (5m) / 2× (1h); reads 0.1× (0.025× Fable 5.1)
automatic prefix cache, no markers (cache_control on /v1/messages ignored); sticky routing via prompt_cache_key (Responses, Chat) or header x-grok-conv-id (Chat/gRPC); include prior reasoning_content / encrypted reasoning or use previous_response_id to keep hits; no TTL or minimum documented, no write charge; cached reads 0.15×–0.25× (grok-4.6 $0.50, grok-4.3 $0.20); long-context tier doubles cached price too
Endpoint: POST /v1/responses
Params: prompt_cache_key, x-grok-conv-id
Status: DOCUMENTED · LIVE_VERIFIED
live: a fresh minimal request already shows ~192 cached tokens (hidden system prefix); cached tokens count toward TPM
implicit caching (automatic on 2.5+, min prompt 4,096 tokens on 3.x Flash/3.1 Pro, 2,048 on 2.5; hits in usageMetadata.cachedContentTokenCount) and explicit cachedContents (POST /v1beta/cachedContents {model, contents, systemInstruction, tools, ttl|expireTime} → use cachedContent: cachedContents/{id}; default TTL 1 h, PATCH TTL only; min 1,024 tokens live); cached reads 0.1× everywhere; explicit storage $0.50–$4.50 per 1M tokens per hour
Endpoint: POST /v1beta/models/{model}:generateContent
Params: cachedContent, generationConfig, ttl, expireTime
Status: DOCUMENTED · BETA · ACCOUNT_RESTRICTED
explicit caching v1beta-only, not via Interactions; free tier limit: 0; no write fee
yes Two philosophies: implicit (OpenAI default, xAI, Gemini implicit) vs explicit breakpoints (Anthropic, OpenAI 5.6+, Gemini cachedContents). Write fees exist only on Anthropic and GPT-5.6+; Gemini charges storage time instead; xAI charges nothing but publishes no TTL.
Ref: docs/comparisons/caching-and-reasoning.md · docs/xai/prompt-caching.md · docs/gemini/context-caching.md
Cache diagnostics
2/4
prompt_cache_options.comparison_response_id → prompt_cache_diagnostics {type: cache_hit|cache_miss{reason}} (GPT-5.6+)
Endpoint: POST /v1/responses
Params: prompt_cache_options.comparison_response_id
Status: DOCUMENTED
reasons: model_changed, tools_changed, text_format_changed, …
diagnostics.previous_message_id → response diagnostics cache-miss reasons; beta cache-diagnosis-2026-04-07
Endpoint: POST /v1/messages
Params: diagnostics.previous_message_id
Status: DOCUMENTED · BETA
no diagnostics; only usage.*cached_tokens counters (input_tokens_details.cached_tokens, prompt_tokens_details.cached_tokens, cache_read_input_tokens on /v1/messages)
Endpoint: POST /v1/responses
Status: DOCUMENTED
no diagnostics; usageMetadata.cachedContentTokenCount + cacheTokensDetails[] only
Endpoint: POST /v1beta/models/{model}:generateContent
Status: DOCUMENTED
yes OpenAI and Anthropic explain misses; xAI and Gemini only count hits.
Ref: parameters.json

# Reasoning / thinking

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Reasoning control
4/4
reasoning {effort: none|minimal|low|medium|high|xhigh|max, summary: auto|concise|detailed, context: auto|current_turn|all_turns, mode: standard|pro}; reasoning output items with encrypted_content; usage.output_tokens_details.reasoning_tokens
Endpoint: POST /v1/responses
Params: reasoning.effort, reasoning.summary, reasoning.context, reasoning.mode
Status: DOCUMENTED · LIVE_VERIFIED
Chat: reasoning_effort only; effort set per model (gpt-6-astra rejects none)
thinking {type: adaptive|enabled|disabled, budget_tokens, display: summarized|omitted|updates} + output_config.effort: low|medium|high|xhigh|max (default high); thinking/redacted_thinking blocks with signature; usage.output_tokens_details.thinking_tokens
Endpoint: POST /v1/messages
Params: thinking.type, thinking.budget_tokens, thinking.display, output_config.effort
Status: DOCUMENTED · LIVE_VERIFIED
enabled (manual budget) 400 on 4.7+; adaptive 400 on 4.5 models
Responses reasoning {effort: low|medium|high|xhigh, summary: auto|concise|detailed} (alias reasoning_effort); Chat reasoning_effort + message.reasoning_content (delta.reasoning_content when streaming); usage.completion_tokens_details.reasoning_tokens; reasoning is always on for grok-4.6/4.5/4.20-reasoning/build (reasoning_effort → 400 on 4.20-reasoning and build); grok-4.3 accepts none (LIVE_DISCOVERED, 0 reasoning tokens); grok-4.20-0309-non-reasoning has none; on grok-4.20-multi-agent-0309 effort selects 4 or 16 agents
Endpoint: POST /v1/responses
Params: reasoning.effort, reasoning.summary, reasoning_effort
Status: DOCUMENTED · LIVE_VERIFIED
max_output_tokens documented to include reasoning but not enforced live; 'Reply with OK.' costs 70–180 reasoning tokens
generationConfig.thinkingConfig {thinkingLevel: MINIMAL|LOW|MEDIUM|HIGH (Gemini 3), thinkingBudget: -1|0|N (2.5-era, LEGACY on 3.x; exclusive with level → 400), includeThoughts}; defaults high (3.1 Pro, 3 Flash) / medium (3.5–3.8 Flash) / minimal (Flash-Lite); MINIMAL → 400 on 3.8/3.7 Flash and 3.1 Pro; thought parts {text, thought: true}; usageMetadata.thoughtsTokenCount; Interactions generation_config.thinking_level + thinking_summaries: auto|none
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.thinkingConfig.thinkingLevel, generationConfig.thinkingConfig.thinkingBudget, generationConfig.thinkingConfig.includeThoughts
Status: DOCUMENTED · LIVE_VERIFIED
thinking cannot be disabled on 3.8/3.7 Flash, 3.1 Pro, 2.5 Pro; thinkingBudget: 0 works on 3.6/3.5 Flash; OpenAI-compat maps reasoning_effort minimal/low/medium/high
yes Four effort ladders overlap on low/medium/high: OpenAI adds none/minimal/xhigh/max, Anthropic xhigh/max, xAI xhigh, Gemini minimal. Off-switches: OpenAI none (most models), Anthropic disabled (not Fable/Mythos), xAI none on grok-4.3 only, Gemini thinkingBudget: 0 on two Flash models only. Only Anthropic and Gemini (2.5) expose a token budget.
Ref: docs/openai/reasoning.md · docs/anthropic/thinking.md · docs/xai/reasoning.md · docs/gemini/thinking.md
Reasoning visibility
4/4
summaries only (reasoning.summary); raw content[] empty for GPT models; encrypted replay item
Endpoint: POST /v1/responses
Params: reasoning.summary, include
Status: DOCUMENTED · LIVE_VERIFIED
summarized text by default (≤4.6) or omitted (4.7+); display: updates progress text (Fable 5.x beta); signature always present
Endpoint: POST /v1/messages
Params: thinking.display
Status: DOCUMENTED · BETA · LIVE_VERIFIED
Chat returns the full reasoning_content text by default; Responses returns reasoning items with summary[] (reasoning.summary accepted 'for compatibility'; the item is sometimes omitted); /v1/messages returns thinking blocks with an empty signature
Endpoint: POST /v1/responses
Params: reasoning.summary, include
Status: DOCUMENTED · LIVE_VERIFIED
includeThoughts: true → thought summaries as thought: true parts (not guaranteed on trivial prompts); thoughtSignature (opaque) on function-call and final parts; Interactions thought steps {signature, summary[]}
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.thinkingConfig.includeThoughts, contents[].parts[].thoughtSignature
Status: DOCUMENTED · LIVE_VERIFIED
billing is on full thoughts, not the summary
yes Only xAI (Chat Completions) exposes the raw chain of thought; the other three return summaries plus an opaque signature/encrypted blob.
Ref: docs/anthropic/thinking.md §3 · docs/xai/reasoning.md · docs/gemini/thinking.md
Reasoning replay across turns
4/4
replay reasoning items verbatim (auto with previous_response_id); encrypted_content decrypted in memory when store:false; reasoning.context: all_turns (GPT-5.6 default)
Endpoint: POST /v1/responses
Params: input[](reasoning).encrypted_content, reasoning.context
Status: DOCUMENTED · LIVE_VERIFIED
pass thinking blocks back unmodified within tool-use turns (400 if modified); prior-turn thinking kept for all turns on Opus 4.5+/Sonnet 4.6+/Fable, last turn on Haiku/Sonnet 4.5; clear_thinking_20251015 edit; beta thinking.block_binding
Endpoint: POST /v1/messages
Params: messages[].content[].signature, thinking.block_binding
Status: DOCUMENTED · BETA · LIVE_VERIFIED
include: ["reasoning.encrypted_content"] (xai-sdk use_encrypted_content=True) → replay the reasoning items with encrypted_content when store:false (ZDR) — also the 'top cause of cache misses' when omitted; previous_response_id rehydrates automatically; Chat: resend reasoning_content
Endpoint: POST /v1/responses
Params: include, input[](reasoning).encrypted_content, previous_response_id
Status: DOCUMENTED · LIVE_VERIFIED
mandatory on Gemini 3: echo thoughtSignature of the first functionCall of each step (400 'missing a thought_signature' / finishReason: MISSING_THOUGHT_SIGNATURE, even at minimal); text-part signatures recommended; tampering → 400 'Corrupted thought signature'; Interactions carry it in thought steps; 'thought preservation' across turns since 3.5 Flash
Endpoint: POST /v1beta/models/{model}:generateContent
Params: contents[].parts[].thoughtSignature
Status: DOCUMENTED · LIVE_VERIFIED
dummy values skip_thought_signature_validator bypass validation (verified)
yes All four carry hidden reasoning as an opaque blob; Anthropic and Gemini validate it (400 on edits / omissions), OpenAI and xAI encrypt it.
Ref: docs/anthropic/thinking.md §4 · docs/xai/reasoning.md · docs/gemini/thinking.md
Pro / extended-compute mode
3/4
reasoning.mode: pro (GPT-5.6) billed at standard rates; *-pro models (gpt-5.4-pro, gpt-5.5-pro)
Endpoint: POST /v1/responses
Params: reasoning.mode
Status: DOCUMENTED
effort: max is the top of the ladder; no separate mode
Params: output_config.effort
Status: DOCUMENTED
grok-4.20-multi-agent-0309 (BETA): one Responses call fans out to 4 (low/medium) or 16 (high/xhigh) agents; all agents' tokens billed at grok-4.20 rates; Responses-only (Chat → 400)
Endpoint: POST /v1/responses
Params: reasoning.effort
Status: DOCUMENTED · BETA · LIVE_VERIFIED
'Reply with OK.' cost 1,645 output tokens ≈ $0.0047
gemini-3.8-live-extended-thinking (Live) and Deep Research agents (deep-research-preview-04-2026, -max-; 60-min background runs, $1–7 per task estimated) rather than a mode; gemini-3.1-pro-preview is the deep-reasoning text model
Endpoint: POST /v1beta/interactions
Params: agent, agent_config(deep-research)
Status: DOCUMENTED · PREVIEW · LIVE_VERIFIED
no Different vehicles: a mode/model (OpenAI), a multi-agent model (xAI), an agent (Gemini), a ladder value (Anthropic).
Ref: parameters.json · docs/models/xai-models.md · docs/gemini/interactions-api.md
Change effort mid-conversation without breaking the cache
2/4
configuration_update input item {reasoning:{effort}} (gpt-6-astra)
Endpoint: POST /v1/responses
Params: input[](configuration_update)
Status: DOCUMENTED
messages[].output_config.effort on role:system entries; beta mid-conversation-output-config-2026-07-01 (Fable 5.1, Mythos 5.1, Opus 5)
Endpoint: POST /v1/messages
Params: messages[].output_config.effort
Status: DOCUMENTED · BETA
per-request reasoning.effort only (automatic cache; no documented invalidation rule)
Endpoint: POST /v1/responses
Status: DOCUMENTED
per-request thinkingLevel only; an explicit cachedContents prefix is unaffected by generationConfig changes
Endpoint: POST /v1beta/models/{model}:generateContent
Status: DOCUMENTED
yes OpenAI and Anthropic only.
Ref: parameters.json
Task-wide token budget (advisory)
3/4
none in Responses (max_tool_calls caps hosted tool calls; Agents API has session_budget_exceeded)
Endpoint: POST /v1/responses
Params: max_tool_calls
Status: DOCUMENTED
output_config.task_budget {type, total ≥20000, remaining}; beta task-budgets-2026-03-13 (Fable/Mythos/Opus 5/4.8/4.7)
Endpoint: POST /v1/messages
Params: output_config.task_budget
Status: DOCUMENTED · BETA
max_turns (agentic turns per Responses request; default = server cap) — a turn budget, not tokens
Endpoint: POST /v1/responses
Params: max_turns
Status: DOCUMENTED · LIVE_VERIFIED
Antigravity agent_config.max_total_tokens (Interactions, PREVIEW) → status incomplete when exhausted
Endpoint: POST /v1beta/interactions
Params: agent_config(antigravity).max_total_tokens
Status: DOCUMENTED · PREVIEW
no Three different budget units (tokens, turns, agent tokens); none on the OpenAI Responses API.
Ref: compatibility/anthropic-feature-model-matrix.json · docs/xai/responses.md · docs/gemini/interactions-api.md

# Context management

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Context window
4/4
1,050,000 (GPT-5.4/5.5/5.6/6 Astra; >272k input = long-context pricing 2×/1.5×), 400k (5.x mini/nano, GPT-5–5.3), 200k (o-series), 128k (gpt-4o)
Status: DOCUMENTED
1,000,000 default on Claude 4.6+ (no header, no premium); 200k on Opus 4.5 / Sonnet 4.5 / Haiku 4.5; context-1m-2025-08-07 header RETIRED
Status: DOCUMENTED · LIVE_VERIFIED
1,000,000 (grok-4.3, all grok-4.20 ids), 500,000 (grok-4.6, grok-4.5), 256,000 (grok-build-0.1); prompts ≥200,000 tokens switch the whole request to the 2× long-context tier
Status: DOCUMENTED · LIVE_VERIFIED
usage.context_details on Responses
1,048,576 on every Gemini 3.x / 2.5 text model (inputTokenLimit); 262,144 Gemma 4; 131,072 Live / image / agent models; Pro models bill >200k prompts at 2× input / 1.5× output, Flash models flat
Status: DOCUMENTED · LIVE_VERIFIED
yes ~1M on every flagship line; long-context premiums on OpenAI (>272k), xAI (≥200k, all tokens) and Gemini Pro (>200k); none on Anthropic or Gemini Flash.
Ref: models.json
Max output tokens
4/4
128,000 on GPT-5.x/6 (272,000 on gpt-5-pro); max_output_tokens ≥16
Endpoint: POST /v1/responses
Params: max_output_tokens
Status: DOCUMENTED · LIVE_VERIFIED
128,000 on Claude 4.6+ (64,000 on 4.5 models); max_tokens required; 300,000 in Batches with output-300k-2026-03-24
Endpoint: POST /v1/messages
Params: max_tokens
Status: DOCUMENTED · LIVE_VERIFIED
no documented limit (max_output: null on every Grok record; grok-4.6 page: 'No text output limit'); Responses max_output_tokens default 128,000 (docs) and not enforced on reasoning live; Chat max_completion_tokens caps visible output only
Endpoint: POST /v1/responses
Params: max_output_tokens, max_completion_tokens
Status: DOCUMENTED · LIVE_VERIFIED
max_tokens DEPRECATED alias on Chat
65,536 on every 3.x / 2.5 text model (outputTokenLimit); 32,768 Gemma 4 / image models; maxOutputTokens includes thinking tokens (hit → finishReason: MAX_TOKENS, possibly empty text)
Endpoint: POST /v1beta/models/{model}:generateContent
Params: generationConfig.maxOutputTokens
Status: DOCUMENTED · LIVE_VERIFIED
Interactions generation_config.max_output_tokens
yes 128k (OpenAI, Anthropic), 64k (Gemini), unpublished (xAI). Only Anthropic makes the cap mandatory.
Ref: models.json
Server-side compaction (in-flight)
3/4
context_management: [{type: compaction, compact_threshold ≥1000}] → compaction output item, SSE response.compaction.compacting
Endpoint: POST /v1/responses
Params: context_management[].compact_threshold
Status: DOCUMENTED
context_management.edits: [{type: compact_20260112, trigger ≥50000, pause_after_compaction, instructions}] → compaction block, stop_reason: compaction; beta compact-2026-01-12; 4.6+
Endpoint: POST /v1/messages
Params: context_management.edits[]
Status: DOCUMENTED · BETA · LIVE_VERIFIED
usage.iterations[] for billing
context_management[] accepted but 'parsed but not yet executed' (compat only); use the stand-alone compact endpoint
Endpoint: POST /v1/responses
Params: context_management
Status: DOCUMENTED
Live API only: setup.contextWindowCompression {triggerTokens, slidingWindow {targetTokens}} (unlimited session length); Antigravity agents compact around ~135k automatically; nothing on generateContent
Endpoint: WSS BidiGenerateContent
Params: setup.contextWindowCompression
Status: DOCUMENTED
yes OpenAI and Anthropic compact text conversations in-flight; Gemini only compresses Live sessions; xAI parses the field and ignores it.
Ref: docs/openai/responses.md §6 · docs/anthropic/context-management.md §3 · docs/gemini/live-api.md
Stand-alone compaction request
3/4
POST /v1/responses/compact {model, input|previous_response_id} → response.compaction with encrypted compaction item
Endpoint: POST /v1/responses/compact
Params: model, input, previous_response_id, instructions
Status: DOCUMENTED · LIVE_VERIFIED
compaction: {type: summarize} on Messages → single signed compaction block; beta compact-2026-09-04
Endpoint: POST /v1/messages
Params: compaction, compaction.type, compaction.instructions
Status: DOCUMENTED · BETA · LIVE_VERIFIED
block must be sent first
POST /v1/responses/compact {model, input} → {object: response.compaction, id: cmp_…, output:[{type: compaction, encrypted_content}], usage {…, dropped_message_count}}; put the output first in the next input; do not edit the blob; the pre-compaction conversation must still fit the window (May 2026 'Context Compaction API')
Endpoint: POST /v1/responses/compact
Params: model, input
Status: DOCUMENTED · LIVE_VERIFIED
48 parameter rows
— not offered yes OpenAI-shaped on xAI (same endpoint and item); Anthropic as a Messages parameter; none on Gemini.
Ref: docs/anthropic/context-management.md §3 · docs/xai/responses.md
Server-side context editing (clear old tool results / thinking)
1/4
no editing strategies; legacy truncation: auto drops oldest items
Endpoint: POST /v1/responses
Params: truncation
Status: DOCUMENTED · LEGACY · LIVE_VERIFIED
context_management.edits[]: clear_tool_uses_20250919, clear_thinking_20251015; beta context-management-2025-06-27; response context_management.applied_edits[]
Endpoint: POST /v1/messages
Params: context_management.edits[].type, context_management.edits[].trigger, context_management.edits[].keep
Status: DOCUMENTED · BETA · LIVE_VERIFIED
truncation accepted (disabled echoed) — 'not supported, compatibility only'
Endpoint: POST /v1/responses
Params: truncation
Status: DOCUMENTED
— not offered no Anthropic-only.
Ref: docs/anthropic/context-management.md §2
Mid-conversation tool add/remove (cache-preserving)
2/4
additional_tools developer item (adds tools mid-thread); tool_choice: allowed_tools to restrict
Endpoint: POST /v1/responses
Params: input[](additional_tools)
Status: DOCUMENTED
tool_addition/tool_removal blocks in role:system messages; beta mid-conversation-tool-changes-2026-07-01 (Fable 5, Mythos 5, Opus 4.8, Opus 5)
Endpoint: POST /v1/messages
Params: messages[].role
Status: DOCUMENTED · BETA
follow-ups via previous_response_id 'may change tools/model' — no dedicated item; cache effect undocumented
Endpoint: POST /v1/responses
Params: previous_response_id
Status: DOCUMENTED
resend the full tools[] each call (Interactions chaining also requires re-sending tools)
Endpoint: POST /v1beta/models/{model}:generateContent
Status: DOCUMENTED
yes OpenAI and Anthropic only.
Ref: anthropic-beta-headers.json

# Service tiers, limits, safety

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Service tiers / processing modes
4/4
service_tier: auto|default|flex|scale|priority|fast|ultrafast; flex = batch price, slower; fast = 2× (renamed from priority 2026-07-30); echoed in response
Endpoint: POST /v1/responses
Params: service_tier
Status: DOCUMENTED · LIVE_VERIFIED
service_tier: auto|standard_only (Priority Tier commitments, no longer sold) + speed: fast (beta fast-mode-2026-02-01, Opus 5 / 4.8 only, 2× price); usage.service_tier standard|priority|batch
Endpoint: POST /v1/messages
Params: service_tier, speed
Status: DOCUMENTED · BETA · PREVIEW · ACCOUNT_RESTRICTED
fast 429 'rate limit of 0' for our key
service_tier: default|priority (Chat + Responses): priority = 2× on every token type (after the cache discount), billed only when the response echoes service_tier: priority; not combinable with Batch
Endpoint: POST /v1/responses
Params: service_tier
Status: DOCUMENTED · LIVE_VERIFIED
live responses always echoed default
serviceTier: standard|flex|priority (service_tier on Interactions / OpenAI-compat): flex = 0.5× (1–15 min target, sheddable → 429), priority = 1.8× (0.3× rate limit, graceful downgrade to standard); echoed in usageMetadata.serviceTier and header X-Gemini-Service-Tier
Endpoint: POST /v1beta/models/{model}:generateContent
Params: serviceTier
Status: DOCUMENTED · LIVE_VERIFIED
flex accepted live; ten text models listed
yes Cheaper tier: OpenAI flex, Gemini flex (both 50 %). Faster tier: OpenAI fast 2×, Anthropic fast 2× (two models), xAI priority 2×, Gemini priority 1.8×.
Ref: docs/anthropic/service-tiers.md · docs/xai/pricing.md · docs/gemini/pricing.md · pricing.json
End-user identifier for abuse detection
4/4
safety_identifier (≤64 chars; replaces user) + prompt_cache_key; header OpenAI-Safety-Identifier on Realtime
Endpoint: POST /v1/responses
Params: safety_identifier, user
Status: DOCUMENTED · LIVE_VERIFIED
metadata.user_id (≤512 chars, no PII); beta anthropic-user-profile-id header
Endpoint: POST /v1/messages
Params: metadata.user_id, anthropic-user-profile-id
Status: DOCUMENTED · LIVE_VERIFIED
user and safety_identifier (both accepted on Chat and Responses); /v1/messages metadata.user_id
Endpoint: POST /v1/responses
Params: safety_identifier, user
Status: DOCUMENTED · LIVE_VERIFIED
labels {safety_identifier: …} (documented key; Cloud-label rules; accepted, not echoed) on generateContent and Interactions
Endpoint: POST /v1beta/models/{model}:generateContent
Params: labels
Status: DOCUMENTED · LIVE_VERIFIED
yes Same idea everywhere; the field name changes four times.
Ref: parameters.json
Rate-limit headers
3/4
x-ratelimit-{limit,remaining,reset}-{requests,tokens} (+ -project-tokens), Retry-After, retry-after-ms, x-should-retry
Status: DOCUMENTED · LIVE_VERIFIED
reset as Go durations (6m0s)
anthropic-ratelimit-{requests,tokens,input-tokens,output-tokens}-{limit,remaining,reset} (RFC 3339), retry-after, x-should-retry, anthropic-priority-*, anthropic-fast-*
Status: DOCUMENTED · LIVE_VERIFIED
not on GET /models or count_tokens
undocumented but observed: x-ratelimit-limit-requests (7,200 grok-4.6/4.5, 1,800 grok-4.3/4.20/build — per minute), x-ratelimit-remaining-requests, x-ratelimit-limit-tokens (50,000,000 / 10,000,000 = documented T0 TPM), x-ratelimit-remaining-tokens; no reset, no Retry-After; x-request-id, x-zero-data-retention
Status: DOCUMENTED · LIVE_DISCOVERED
absent on /v1/responses multi-agent calls and catalogue GETs
none — no x-ratelimit-* or Retry-After on 200/404/429; retry delay only inside the 429 message text ('Please retry in 54.2s') and google.rpc.RetryInfo; undocumented X-Gemini-Service-Tier header
Status: DOCUMENTED · LIVE_DISCOVERED
no request-id header either — use responseId in the body
yes Headers on three providers (xAI's undocumented); Gemini puts the information in the 429 body.
Ref: headers.json · rate-limits.json
Rate-limit tiers
4/4
Free, Tier 1–5 by cumulative spend ($5 → $1,000) with monthly usage caps $100 → $200k; per-model RPM/TPM/batch-queue tables; long-context tables >272k
Status: DOCUMENTED
Scale/Reserved Tier, Ultrafast preview above Tier 5
Start ($500/mo cap) / Build ($1,000) / Scale ($200k) / Custom; per model-class RPM/ITPM/OTPM (cache reads excluded from ITPM); Batches & count_tokens separate
Status: DOCUMENTED · LIVE_DISCOVERED
observed Scale-tier headers for our key
Tier 0–4 by cumulative spend since 2026-01-01 ($0 / $50 / $250 / $1,000 / $5,000; Enterprise on request); per-model RPS (= RPM/60) and TPM (prompt + completion + reasoning + cached): grok-4.6/4.5 150→500 RPS, 50M→100M TPM; grok-4.3/4.20/build 37→208 RPS, 10M→85M TPM; multi-agent 9→56 RPS; Imagine RPS-only (6→100 images, 10→158 videos); voice concurrent sessions 10→200; Batch bypasses limits; per-key qps/qpm/tpm caps via Management API
Status: DOCUMENTED · LIVE_DISCOVERED
console shows personalised limits
Free / Tier 1 (billing linked; $250 cap, $10 per rolling 10 min) / Tier 2 ($100 paid + 3 days; $2,000; $50) / Tier 3 ($1,000 + 30 days; $20k–100k+; $200); dimensions RPM / TPM / RPD (+ IPM, TPD) per project; per-model matrix published only in AI Studio; Pro and media models unavailable on Free (limit: 0); priority 0.3× limits; batch enqueued-token caps per model
Status: DOCUMENTED · LIVE_DISCOVERED
free-tier RPM observed 15/min on flash-lite
yes Spend-based tiers everywhere; xAI is the only provider publishing exact per-model RPS/TPM per tier in the docs, Gemini the only one with a genuinely free tier.
Ref: generated/rate-limits.json
Overload / capacity error
4/4
HTTP 503 server_is_overloaded (retryable, Retry-After)
Status: DOCUMENTED
HTTP 529 overloaded_error (also as SSE error event after 200); acceleration limits now 429
Status: DOCUMENTED
HTTP 429 (RPS/TPM/credits; gRPC RESOURCE_EXHAUSTED) and 5xx internal — no dedicated overload code; status.x.ai
Status: DOCUMENTED
no 429 triggered in this run
HTTP 503 UNAVAILABLE ('The model is overloaded. Please try again later.'), 429 RESOURCE_EXHAUSTED when Flex capacity is shed, 504 DEADLINE_EXCEEDED for long Flex/Deep Research requests
Status: DOCUMENTED
yes 503 (OpenAI, Gemini), 529 (Anthropic), 429 (xAI) for the same condition.
Ref: errors.json
Error envelope
4/4
{error: {message, type, param, code}}; types invalid_request_error, rate_limit_error, insufficient_quota, server_error…; codes e.g. model_not_found, context_length_exceeded, previous_response_not_found
Status: DOCUMENTED · LIVE_VERIFIED
empty-body 404 from Cloudflare on unknown URLs
{type: error, error: {type, message}, request_id}; types invalid_request_error, authentication_error, permission_error, not_found_error, request_too_large, rate_limit_error, api_error, overloaded_error, billing_error, timeout_error
Status: DOCUMENTED · LIVE_VERIFIED
no code field except error.details.error_code on some 429/529
{code: <kebab-case>, error: <message>} (invalid-argument, not-found, unauthenticated:no-credentials…); 422 = bare JSON string (serde message); Management API = gRPC-style {code: 16, message, details[]}; Realtime WS {type: error, error:{type, code, message}}; an invalid API key returns 400, not 401
Status: DOCUMENTED · LIVE_VERIFIED
12 error records; gRPC↔HTTP mapping 3→400, 16→401, 7→403, 5→404, 8→429
google.rpc Status: {error: {code: <http>, message, status: INVALID_ARGUMENT|FAILED_PRECONDITION|UNAUTHENTICATED|PERMISSION_DENIED|NOT_FOUND|ALREADY_EXISTS|RESOURCE_EXHAUSTED|INTERNAL|UNIMPLEMENTED|UNAVAILABLE|DEADLINE_EXCEEDED…, details[] (BadRequest.fieldViolations, QuotaFailure, RetryInfo, Help)}}; Interactions {error: {code: <snake_case>, message}}; soft failures at HTTP 200 (promptFeedback.blockReason, finishReason); Live WS close codes 1007/1008
Status: DOCUMENTED · LIVE_VERIFIED
18 error records; unknown File/Operation → 403, not 404
yes Four envelopes; only Anthropic and OpenAI carry a request id in the body/header pair; Gemini has no request-id header at all.
Ref: docs/errors/openai.md · docs/errors/anthropic.md · docs/xai/authentication-headers-errors.md · docs/errors/gemini.md
Idempotency key
0/4
not documented for api.openai.com (only Workspace Agents on api.chatgpt.com); dedupe via metadata/custom_id
Status: UNVERIFIED
not documented; SDKs retry on 409
Status: UNVERIFIED
not documented; public-url creation is idempotent by design; batch batch_request_id dedupes within a batch
Status: UNVERIFIED
not documented; batch key per JSONL line; seed for determinism
Status: UNVERIFIED
no No provider offers request idempotency keys on the model APIs.
Ref: headers.json
Per-request dollar cost in the response
1/4
token counts only (usage); costs via the Admin GET /v1/organization/costs report
Status: DOCUMENTED
token counts only; costs via /v1/organizations/cost_report
Status: DOCUMENTED
usage.cost_in_usd_ticks on every inference response (Chat, Responses, images, videos; 1 USD = 10^10 ticks; April 2026) — the effective price after cache, long-context, priority and regional multipliers; batch cost_breakdown (SDK/gRPC)
Endpoint: POST /v1/responses
Params: usage.cost_in_usd_ticks
Status: DOCUMENTED · LIVE_VERIFIED
catalogue endpoints return no usage/cost
token counts only (usageMetadata, Interactions usage); costs in Google Cloud Billing
Endpoint: POST /v1beta/models/{model}:generateContent
Status: DOCUMENTED
no xAI-only.
Ref: docs/xai/pricing.md · docs/xai/responses.md

# Auth, versioning, SDKs, platform

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Authentication
4/4
Authorization: Bearer <sk-proj-…|sk-…|service-account key|WIF access token>; optional OpenAI-Organization, OpenAI-Project; Admin keys sk-admin-…
Status: DOCUMENTED · LIVE_VERIFIED
WIF token exchange at auth.openai.com / mTLS
x-api-key: sk-ant-api03-… or Authorization: Bearer <key|sk-ant-oat01-… OAuth/WIF token>; anthropic-workspace-id; Admin keys sk-ant-admin01-…
Status: DOCUMENTED · LIVE_VERIFIED
WIF via POST /v1/oauth/token
Authorization: Bearer xai-… on REST, WebSocket and gRPC metadata; separate Management key for management-api.x.ai (inference key → 401 code 16); ephemeral xai-realtime… client secrets for browsers (POST /v1/realtime/client_secrets); per-key ACLs api-key:endpoint:*, api-key:model:* (new keys have no access by default); mTLS host mtls.api.x.ai (enterprise)
Status: DOCUMENTED · LIVE_VERIFIED
GET /v1/api-key, GET /v1/me introspection
x-goog-api-key: AIza… (recommended; ?key= discouraged); Authorization: Bearer <GEMINI_API_KEY> mandatory on /v1beta/openai/*; OAuth/ADC Bearer alternative (+ x-goog-user-project); ephemeral tokens POST /v1beta/auth_tokens for the Live API; standard keys rejected from September 2026 in favour of service-account-bound auth keys
Status: DOCUMENTED · LIVE_VERIFIED
limits are per project, not per key
yes Bearer everywhere except Gemini's native header; only OpenAI/Anthropic/xAI split admin vs inference keys; Gemini is the only one retiring a key type.
Ref: headers.json · docs/openai/authentication-and-keys.md · docs/anthropic/admin-api.md §1 · docs/xai/authentication-headers-errors.md · docs/gemini/authentication-headers-versions.md
API version header / path version
2/4
none (server answers openai-version: 2020-10-01)
Status: DOCUMENTED · LIVE_VERIFIED
anthropic-version: 2023-06-01 required on every request (400 otherwise)
Params: anthropic-version
Status: DOCUMENTED · LIVE_VERIFIED
no version header, no beta headers; features selected by body fields or base URL; Interactions-style Api-Revision does not exist
Status: DOCUMENTED · LIVE_VERIFIED
version in the URL path: /v1beta (86 methods, default for SDKs) vs /v1 (47; stable subset — no caching, tuning, Live, Files, agents); Interactions optional header Api-Revision: 2026-05-20
Params: Api-Revision
Status: DOCUMENTED · LIVE_VERIFIED
docs say every model is in both versions; live /v1/models lists 22 vs 58
no Header (Anthropic), path (Gemini), nothing (OpenAI, xAI).
Ref: headers.json · docs/gemini/authentication-headers-versions.md
Beta opt-in header
2/4
OpenAI-Beta: agents=v1, chatkit_beta=v1, workspace_agent_runs=v1, responses_multi_agent=v1, legacy assistants=v2, realtime=v1; plus ?beta=true surface with body openai-beta[]
Params: OpenAI-Beta
Status: DOCUMENTED
400 invalid_beta when missing
anthropic-beta: <feature>-<YYYY-MM-DD>[,…] (50 catalogued values; SDK betas=[…], client.beta.*); unknown → 400
Params: anthropic-beta, betas
Status: DOCUMENTED · LIVE_VERIFIED
see docs/faq.md for the current list
none; alpha features answer 403 ('only available for alpha users' — tool_search) or 404 (ACL)
Status: DOCUMENTED · LIVE_VERIFIED
none; gating by /v1beta path and -preview / -exp model ids (kind: preview 45 records)
Status: DOCUMENTED · LIVE_VERIFIED
no Anthropic gates parameters, OpenAI gates surfaces, Gemini gates by path/model id, xAI by account.
Ref: generated/fragments/headers/anthropic-beta-headers.json · headers.json · docs/faq.md Q6
Request correlation
3/4
response x-request-id; request X-Client-Request-Id (logged, not echoed); openai-processing-ms
Status: DOCUMENTED · LIVE_VERIFIED
response request-id (also request_id in error body); anthropic-organization-id
Status: DOCUMENTED · LIVE_VERIFIED
response x-request-id (= chat.completion.id), Server-Timing (Cloudflare cfEdge/cfOrigin), CF-RAY
Status: DOCUMENTED · LIVE_VERIFIED
absent on catalogue GETs
no request-id header; responseId in the JSON body; Server-Timing: gfet4t7; dur=…
Status: DOCUMENTED · LIVE_VERIFIED
yes Header on three providers, body field on Gemini.
Ref: headers.json
Official SDKs
4/4
Python openai 3.16.2, Node openai 7.18/7.19, .NET, Java 4.65 (beta label), Go v3 (beta), Ruby, CLI, Agents SDK (py/ts), Azure libraries; retries 2, timeout 600 s
Status: DOCUMENTED · LIVE_VERIFIED
Python anthropic 1.7.0, TS @anthropic-ai/sdk 0.126, Go, Java 2.63, Ruby, C# ≥10, PHP (beta), ant CLI 1.33; retries 2, timeout 10 min; cloud clients Bedrock/Vertex/AWS/Foundry
Status: DOCUMENTED · LIVE_VERIFIED
Python v1 removed sampling kwargs and completions
Python xai-sdk 1.19 (gRPC, client.chat.create().sample()/stream()/defer()/parse(), response.cost_usd, timeout 1620 s) — no official Node SDK: use openai with baseURL: https://api.x.ai/v1, @ai-sdk/xai, langchain-xai, or the Anthropic SDK against /v1/messages (deprecated); Grok Build CLI (grok, BETA); protos xai-org/xai-proto for buf curl; gRPC api.x.ai:443 (11 services / 38 RPCs) lacks Responses items, voice, skills
Status: DOCUMENTED · LIVE_VERIFIED
9 SDK records
Python google-genai 2.24 (genai.Client(); same SDK targets Vertex with vertexai=True), JS @google/genai 2.23, Go google.golang.org/genai, Java com.google.genai, C# Google.GenAI; Firebase AI Logic / Genkit / Vercel AI SDK; OpenAI SDK against /v1beta/openai; legacy google-generativeai, @google/generative-ai, Go/Dart/Swift/Android libs DEPRECATED 2025-11-30
Status: DOCUMENTED · LIVE_VERIFIED
8 SDK records; default api_version v1beta
yes Every provider has Python and Node coverage — xAI's Node path is the OpenAI SDK; xAI's own SDK is the only gRPC-first one.
Ref: sdks.json · docs/xai/sdks.md · docs/gemini/sdks.md
Webhooks (platform events)
4/4
/v1/webhook_endpoints CRUD + rotate_secret, test; /v1/webhook_event_types; events batch., response., fine_tuning., eval.run., video., realtime/live incoming calls, safety., agent.session.*; Standard Webhooks signature (webhook-id/-timestamp/-signature, whsec_)
Endpoint: GET /v1/webhook_endpoints
Status: DOCUMENTED · LIVE_VERIFIED
28 event types
Managed Agents only: endpoints registered in Console (no API), 44 event types (agent., session., deployment*., environment., vault*., memory_store.); same Standard Webhooks headers; ≤3 attempts, 5-min freshness; beta managed-agents-2026-04-01
Status: DOCUMENTED · BETA
no webhooks for Messages/Batches
only the SIP voice webhook realtime.call.incoming (Standard Webhooks headers, HMAC-SHA256) registered via POST /v2/phone-numbers {webhook:{url}}; no batch/response webhooks
Endpoint: POST /v2/phone-numbers
Status: DOCUMENTED
POST/GET /v1/webhooks, GET|PATCH|DELETE /v1/webhooks/{id}, POST …/rotate_secret (documented on /v1, BETA) + per-request webhook_config {uris[], user_metadata} / Batch webhookConfig; events interaction.completed|failed|cancelled|requires_action, batch.succeeded|failed; JWKS-signed
Endpoint: POST /v1/webhooks
Params: webhook_config
Status: DOCUMENTED · BETA
7 endpoints, not tested; :ping UNVERIFIED
yes OpenAI covers the widest event set; Gemini and Anthropic cover their agent/interaction resources; xAI only incoming calls.
Ref: generated/webhook-events.json · docs/openai/webhooks.md · docs/anthropic/managed-agents.md §17.2 · docs/gemini/interactions-api.md
Administration API
3/4
124 operations under /v1/organization/* and /v1/projects/*: admin keys, users, invites, projects, service accounts, project API keys, groups, roles, certificates (mTLS), data retention, spend limits/alerts, audit logs, usage & costs; Admin API key sk-admin-…
Endpoint: GET /v1/organization/*
Status: DOCUMENTED · ACCOUNT_RESTRICTED
all probes 403/401 with a project key
100 operations under /v1/organizations/*: me, users, invites, workspaces, api_keys, service accounts, federation (WIF), RBAC, spend limits, external keys, tunnels certs, usage_report & cost_report, analytics, Claude Code analytics; Admin key sk-ant-admin01-…
Endpoint: GET /v1/organizations/*
Status: DOCUMENTED · ACCOUNT_RESTRICTED · LIVE_VERIFIED
GET /v1/organizations/me works with a regular key
Management API https://management-api.x.ai (21 endpoints, separate Management key, camelCase): API keys (create with acls[], qps/qpm/tpm, expireTime; rotate; delete; propagation), team models/endpoints ACLs, billing (billing-info, invoices, payment methods, postpaid spending limits, prepaid balance/top-up, POST …/usage aggregated by key/model/IP/cluster), audit events; teams/members/ZDR toggle are console-only
Endpoint: GET /auth/teams/{teamId}/api-keys
Status: DOCUMENTED · ACCOUNT_RESTRICTED
inference key → 401 code 16
no admin API on the Developer API: keys/projects/quota live in AI Studio and Google Cloud IAM/Console; GET /v1beta/models is the only account-scoped read
Status: DOCUMENTED
yes Three admin surfaces (OpenAI, Anthropic, xAI — all inaccessible to this atlas's keys); Gemini delegates to Google Cloud.
Ref: docs/openai/admin-api.md · docs/anthropic/admin-api.md · docs/xai/management-api.md
Usage & cost reporting
4/4
GET /v1/organization/usage/{completions,embeddings,images,…} + GET /v1/organization/costs (buckets 1m/1h/1d, group_by)
Endpoint: GET /v1/organization/costs
Params: start_time, bucket_width, group_by
Status: DOCUMENTED · ACCOUNT_RESTRICTED
GET /v1/organizations/usage_report/messages, /usage_report/claude_code, /cost_report; analytics /analytics/{usage_report,cost_report,user_*}
Endpoint: GET /v1/organizations/cost_report
Params: starting_at, bucket_width, group_by
Status: DOCUMENTED · ACCOUNT_RESTRICTED
per-request usage.cost_in_usd_ticks (1 USD = 10^10 ticks) on every inference response + POST /v1/billing/teams/{team_id}/usage (Management API, aggregated by API key / model / IP / cluster / token type); batch cost_breakdown (SDK/gRPC)
Endpoint: POST /v1/billing/teams/{team_id}/usage
Status: DOCUMENTED · ACCOUNT_RESTRICTED · LIVE_VERIFIED
cost ticks LIVE_VERIFIED; billing endpoint restricted
no reporting endpoint; per-response usageMetadata (modality breakdown, thoughtsTokenCount, toolUsePromptTokenCount, cachedContentTokenCount, serviceTier) and Interactions usage {total_*_tokens, *_by_modality, grounding_tool_count}; billing dashboards in Google Cloud
Endpoint: POST /v1beta/models/{model}:generateContent
Params: usageMetadata
Status: DOCUMENTED · LIVE_VERIFIED
yes Org-level reports on OpenAI/Anthropic/xAI; xAI is the only one returning the dollar cost per call; Gemini only per-response token detail.
Ref: endpoints.json (admin, management) · docs/xai/pricing.md · docs/gemini/generate-content.md
Audit / compliance data access
3/4
GET /v1/organization/audit_logs (scope api.audit_logs.read); safety alerts/cases API
Endpoint: GET /v1/organization/audit_logs
Status: DOCUMENTED · ACCOUNT_RESTRICTED
Compliance API (36 endpoints, Claude Enterprise Compliance Access Key): activities feed, chats, projects, files, sessions, roles, groups; DELETE operations
Endpoint: GET /v1/compliance/activities
Status: DOCUMENTED · ACCOUNT_RESTRICTED
GET /audit/teams/{teamId}/events (Management API; administrative events only)
Endpoint: GET /audit/teams/{teamId}/events
Status: DOCUMENTED · ACCOUNT_RESTRICTED
none on the Developer API (Cloud Audit Logs on Vertex AI)
Status: DOCUMENTED
yes Admin-action logs on OpenAI/Anthropic/xAI; Anthropic additionally exports end-user content (Claude Enterprise).
Ref: docs/anthropic/compliance-and-iam.md · docs/xai/management-api.md
Spend limits
3/4
org/project spend_limit, spend_alerts Admin endpoints; 429 insufficient_quota codes when hit
Endpoint: POST /v1/organization/spend_limit
Status: DOCUMENTED
/v1/organizations/spend_limits, /spend_limits/effective, increase requests approve/deny; tier cap → 429 without retry-after; self-set limit → 400
Endpoint: POST /v1/organizations/spend_limits
Status: DOCUMENTED · ACCOUNT_RESTRICTED
GET/POST /v1/billing/teams/{team_id}/postpaid/spending-limits, prepaid balance / top-up (Management API); per-key qps/qpm/tpm caps
Endpoint: POST /v1/billing/teams/{team_id}/postpaid/spending-limits
Status: DOCUMENTED · ACCOUNT_RESTRICTED
tier spend caps ($250 / $2,000 / $20k+) and rolling 10-minute spend limits ($10 / $50 / $200) are enforced by Google, not configurable via API (Cloud Billing budgets instead)
Status: DOCUMENTED
yes API-settable on three providers; platform-imposed on Gemini.
Ref: endpoints.json (admin) · rate-limits.json (gemini)
Cloud availability
4/4
Azure OpenAI / Microsoft Foundry (AzureOpenAI clients, Azure libraries); regional hosts us.|eu.|au.|jp.|in.api.openai.com
Status: DOCUMENTED
Azure surface not catalogued in this atlas
Amazon Bedrock (Mantle bedrock-mantle.{region}.api.aws/anthropic/v1/messages and legacy InvokeModel), Google Cloud Vertex AI (rawPredict), Microsoft Foundry (/anthropic/v1/*), Claude Platform on AWS; feature gaps per platform (no Batches/Files/Skills/server tools on Bedrock/Vertex)
Status: DOCUMENTED · UNVERIFIED
24 cloud endpoint records
Google Cloud Vertex AI Model Garden (partner model) and Microsoft Foundry (Azure) — OpenAI-compatible chat + Responses, cloud billing; also OpenRouter, Vercel AI Gateway, Cloudflare, Cursor; regional first-party endpoints https://us.api.x.ai/v1 (US-pinned, grok-4.6, 1.1×) and https://eu-west-1.api.x.ai/v1 (undocumented, grok-4.3, LIVE_DISCOVERED); clusters us-east-1, us-west-2, us-central-1, eu-west-1, us-saltlake-2
Status: DOCUMENTED · LIVE_DISCOVERED
Gemini Developer API (generativelanguage.googleapis.com, global, API key) vs Vertex AI / Gemini Enterprise Agent Platform ({location}-aiplatform.googleapis.com, IAM only, ~40 regions, data residency, ZDR, CMEK, VPC-SC, provisioned throughput, supervised tuning, RAG/Agent Engine); one SDK, two backends; Developer-API-only: Interactions, Live, Files, File Search, agents, free tier
Status: DOCUMENTED
Vertex surface not catalogued in this atlas
yes Claude and Grok are resold on other clouds; OpenAI and Gemini have first-party cloud twins (Azure/Foundry, Vertex).
Ref: docs/anthropic/cloud-providers.md · docs/openai/data-residency-and-regions.md · docs/xai/sdks.md · docs/gemini/vertex-vs-gemini-api.md
Data residency
3/4
project region at creation; regional hosts us./eu./au./jp./in.api.openai.com; 10 % uplift for models ≥2026-03-05; EU lacks background:true
Status: DOCUMENTED · ACCOUNT_RESTRICTED
inference_geo: us|global per request (Claude 4.6+; 1.1× multiplier for us); workspace default_inference_geo; usage.inference_geo echo; Bedrock/Vertex regional endpoints +10 %
Endpoint: POST /v1/messages
Params: inference_geo
Status: DOCUMENTED · FAILED_VERIFICATION
400 on Haiku 4.5 live
base-URL choice: https://us.api.x.ai/v1 keeps request handling, inference, moderation and retained data in the US at 1.1× token prices (grok-4.6 only; no image/video/voice); global api.x.ai gives no region guarantee
Status: DOCUMENTED · LIVE_VERIFIED
eu-west-1.api.x.ai serves grok-4.3 at global prices (undocumented)
no residency option on the Developer API (global endpoint); residency, CMEK and ~40 regions are Vertex AI features
Status: DOCUMENTED
yes Regional pinning costs ~10 % on OpenAI, Anthropic and xAI alike; Gemini requires moving to Vertex.
Ref: docs/anthropic/regions-and-ips.md · docs/openai/data-residency-and-regions.md · docs/xai/pricing.md · docs/gemini/vertex-vs-gemini-api.md
Zero data retention / data-use terms
4/4
ZDR/Modified retention by approval; store:false keeps reasoning encrypted; Live recordings and Agents sessions excluded
Params: store
Status: DOCUMENTED
Messages stateless → ZDR-eligible on Claude API (per model flag zero_data_retention_eligible); not eligible: Files API, MCP connector, programmatic tool calling, Managed Agents, Fable 5.1 / Mythos 5.1
Status: DOCUMENTED
team-wide ZDR toggle (console): response header x-zero-data-retention: true|false, GET /v1/me → zdr_status: no_zdr|zdr; disables store, previous_response_id, Files, Collections, Batch, deferred completions, stored media; default retention 30 days, not used for training; encrypted reasoning replay keeps ZDR + caching
Status: DOCUMENTED · LIVE_VERIFIED
our team: no_zdr
Unpaid Services (free tier): prompts/responses may be used to improve Google products, with human review; Paid Services: not used for training, 55-day abuse logging; ZDR not achievable on the Developer API (grounding stores 30 days; Interactions state unless store:false) — Vertex AI offers ZDR; EEA/UK/CH end-user apps must use Paid Services
Params: store
Status: DOCUMENTED
yes ZDR is a program (OpenAI), a per-model eligibility (Anthropic), a team switch (xAI) and unavailable on Gemini's Developer API.
Ref: models.json (anthropic capabilities.zero_data_retention_eligible) · docs/xai/management-api.md · docs/gemini/authentication-headers-versions.md
OpenAI-compatibility layer
3/4
the native surface (Responses + Chat Completions)
Endpoint: POST /v1/chat/completions
Status: DOCUMENTED · LIVE_VERIFIED
none — Anthropic exposes only its own Messages format (partners such as xAI implement it)
Status: DOCUMENTED
native: the whole inference API is OpenAI-shaped (/v1/chat/completions, /v1/responses incl. items and SSE events, /v1/batches, /v1/files, /v1/images/*, /v1/realtime events); rejected/ignored OpenAI fields: background, metadata (Responses 400), logit_bias 400, stop/penalties on reasoning models 400, logprobs ignored, store ignored on Chat, non-function tools 422, images.edit() multipart unsupported; extras: reasoning_effort, reasoning_content, deferred, max_turns, top_k, min_p, cost_in_usd_ticks
Endpoint: POST /v1/responses
Status: DOCUMENTED · LIVE_VERIFIED
plus an Anthropic-compatible /v1/messages (deprecated)
/v1beta/openai/* (BETA): chat/completions (LIVE_VERIFIED), embeddings, models[/{id}], images/generations (subset), videos (Sora-style Veo), batches; Bearer key required; reasoning_effort → thinkingLevel mapping; extra_body.google {thinking_config, cached_content, safety_settings, tools[{google_search}]}; message.extra_content.google.thought_signature; unknown params silently ignored; no Responses, Assistants, audio, files, fine-tuning, moderations
Endpoint: POST /v1beta/openai/chat/completions
Params: extra_body.google.thinking_config, extra_body.google.cached_content, reasoning_effort
Status: DOCUMENTED · BETA · LIVE_VERIFIED
7 endpoints; google.rpc error envelope
yes xAI is OpenAI-shaped by design; Gemini offers a partial adapter; Anthropic none.
Ref: docs/xai/chat-completions.md · docs/xai/responses.md · docs/gemini/openai-compatibility.md

# Managed agents platforms

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Managed agent harness
4/4
Agents API (beta OpenAI-Beta: agents=v1): POST /v1/agents (model, instructions, reasoning, text, tools, multi_agent), sessions, turns, items, events, environments, templates, vaults; Codex harness; 34 + 9 endpoints
Endpoint: POST /v1/agents
Params: model, instructions, tools, multi_agent
Status: DOCUMENTED · BETA · LIVE_VERIFIED
gpt-5.4-nano rejected; gpt-5.6-luna accepted
Claude Managed Agents (beta managed-agents-2026-04-01): POST /v1/agents (name, model, system, tools, mcp_servers, skills, multiagent) with versions, /v1/sessions + events/stream, environments, deployments (cron), vaults, memory stores, dreams, tunnels, user profiles; 96 endpoints
Endpoint: POST /v1/agents
Params: name, model, system, tools, mcp_servers, skills, multiagent
Status: DOCUMENTED · BETA · LIVE_VERIFIED
Claude 4.5+ models
no agent resource: the Responses API agentic loop (server-side tools iterate inside one request, bounded by max_turns; stored responses + previous_response_id = the session); grok-4.20-multi-agent-0309 (BETA) as a built-in multi-agent model; Grok Build (grok-build-0.1 PREVIEW model + grok CLI BETA: TUI/headless/ACP, MCP, hooks, skills, subagents, sandbox) is a client-side harness
Endpoint: POST /v1/responses
Params: tools, max_turns, store, previous_response_id
Status: DOCUMENTED · LIVE_VERIFIED · BETA
Skills API 404 for this team
Interactions API agents (agent instead of model): Deep Research (deep-research-preview-04-2026, -max-, agent_config {collaborative_planning, thinking_summaries, visualization}, background only, ≤60 min), Antigravity coding agent (antigravity-preview-09-2026, Linux sandbox 4 vCPU/16 GB, agent_config {model, max_total_tokens}, hooks), custom managed agents POST /v1beta/agents {base_agent, agent_config, system_instruction, tools, base_environment} (≤1,000/project, no versioning); environments, credentials, triggers (cron), webhooks
Endpoint: POST /v1beta/interactions
Params: agent, agent_config, environment, background
Status: DOCUMENTED · BETA · PREVIEW · LIVE_VERIFIED
Deep Research LIVE_VERIFIED (create → in_progress → cancel); custom agents / triggers / credentials not tested; sandbox compute unbilled in preview
yes Four different shapes: agent → session → events (OpenAI, Anthropic), a stateful request loop (xAI), agent-as-model inside Interactions (Gemini).
Ref: docs/comparisons/agents-platforms.md · docs/xai/responses.md · docs/xai/grok-build.md · docs/gemini/interactions-api.md
Create an agent session / run
4/4
POST /v1/agents/sessions {agent|agent_id, environment (required: none|openai_hosted|self_hosted), input, stream, vault_ids} → 201 agent.session
Endpoint: POST /v1/agents/sessions
Params: agent_id, environment, input, stream
Status: DOCUMENTED · BETA · LIVE_VERIFIED
POST /v1/sessions {agent, environment_id (required), initial_events[], resources[], vault_ids[], budget, title} → 200 session
Endpoint: POST /v1/sessions
Params: agent, environment_id, initial_events, budget
Status: DOCUMENTED · BETA · LIVE_VERIFIED
POST /v1/responses {model, input, tools[], max_turns, store} → the stored response is the session; continue with previous_response_id
Endpoint: POST /v1/responses
Params: model, input, tools, max_turns
Status: DOCUMENTED · LIVE_VERIFIED
POST /v1beta/interactions {agent|model, input, background, store, environment, tools, webhook_config} → interaction {id, status, steps[]}; chain with previous_interaction_id
Endpoint: POST /v1beta/interactions
Params: agent, input, background, previous_interaction_id
Status: DOCUMENTED · BETA · LIVE_VERIFIED
/v1/interactions GA UNVERIFIED
yes See FAQ Q5 for the body-level comparison.
Ref: endpoints.json · docs/faq.md Q5
Send input / events to a session
4/4
POST /v1/agents/sessions/{id}/events with agent.session.input.message|tool_result|cancel → 202
Endpoint: POST /v1/agents/sessions/{session_id}/events
Status: DOCUMENTED · BETA · LIVE_VERIFIED
POST /v1/sessions/{id}/events with user.message|interrupt|tool_confirmation|custom_tool_result|tool_result|define_outcome, system.message
Endpoint: POST /v1/sessions/{session_id}/events
Status: DOCUMENTED · BETA · LIVE_VERIFIED
a new POST /v1/responses with previous_response_id (user message or function_call_output / shell_call_output items)
Endpoint: POST /v1/responses
Params: previous_response_id, input[](function_call_output)
Status: DOCUMENTED · LIVE_VERIFIED
a new POST /v1beta/interactions with previous_interaction_id and input (user text or function_result steps); POST …/cancel to stop
Endpoint: POST /v1beta/interactions
Params: previous_interaction_id, input
Status: DOCUMENTED · BETA · LIVE_VERIFIED
interaction.requires_action for function calls
yes Dedicated event endpoints (OpenAI, Anthropic) vs chained requests (xAI, Gemini).
Ref: streaming-events.json (client→server)
Session event stream
4/4
SSE via POST /sessions {stream:true} or GET …/events?stream=true; 31 event types agent.session.*, agent.output.*
Endpoint: GET /v1/agents/sessions/{session_id}/events
Status: DOCUMENTED · BETA · LIVE_VERIFIED
GET /v1/sessions/{id}/events/stream (+ per-thread /threads/{id}/stream); 37 event types agent.*, session.*, span.*, event_start/delta
Endpoint: GET /v1/sessions/{session_id}/events/stream
Status: DOCUMENTED · BETA · LIVE_VERIFIED
the ordinary Responses SSE stream (24 recorded events incl. response.code_interpreter_call.*, response.web_search_call.*, response.mcp_call.*) or wss://api.x.ai/v1/responses
Endpoint: POST /v1/responses
Params: stream
Status: DOCUMENTED · LIVE_VERIFIED
Interactions SSE (stream:true): interaction.created → interaction.status_update → step.start → step.delta {text|thought_signature|arguments_delta|function_result|*_call…} → step.stop → interaction.completed → done [DONE]; resumable with GET …?stream=true&last_event_id=; 15 recorded events (5 legacy names RETIRED 2026-06-08)
Endpoint: POST /v1beta/interactions
Params: stream, last_event_id
Status: DOCUMENTED · BETA · LIVE_VERIFIED
yes
Ref: streaming-events.json
Hosted sandbox
3/4
environment.type: openai_hosted {packages, setup_commands, network, env, skills, plugins, files, environment_template_id}; /workspace; artifacts from /workspace/outputs; ~1 h idle expiry; container rates
Endpoint: POST /v1/agents/sessions
Params: environment
Status: DOCUMENTED · BETA
POST /v1/environments {type: cloud, …} (Ubuntu 24.04, 8 GB RAM, 10 GB disk, /workspace, /mnt/session/{uploads,outputs}, /mnt/memory); state kept 30 days; $0.08/session-hour
Endpoint: POST /v1/environments
Status: DOCUMENTED · BETA · LIVE_VERIFIED
no hosted sandbox for agents (the code_interpreter tool runs Python without a container object; Grok Build sandboxes locally)
Endpoint: POST /v1/responses
Status: DOCUMENTED
POST /v1beta/environments {sources[] (repository ≤500 MB | gcs ≤2 GB | inline ≤1 MB/file), network (unrestricted|disabled|allowlist with credentials), from_environment}; Antigravity sandbox 4 vCPU / 16 GB (Python 3.12, Node 22); files GET …/files/{path}, PUT /upload/…/files/{path}; idle after 15 min, deleted after 7 days; compute unbilled during preview
Endpoint: POST /v1beta/environments
Params: sources, network
Status: DOCUMENTED · BETA · PREVIEW · LIVE_VERIFIED
live GET showed storage.tier: free, 1 GiB project limit
yes Three hosted sandboxes (OpenAI, Anthropic, Gemini); none on xAI.
Ref: docs/openai/agents-environments-and-vaults.md · docs/anthropic/managed-agents.md §8 · docs/gemini/interactions-api.md
Self-hosted execution
3/4
environment.type: self_hosted {workspace_directory} + codex exec-server --remote … executor with environment key CODEX_API_KEY; outbound only
Endpoint: POST /v1/agents/sessions
Params: environment
Status: DOCUMENTED · BETA
type: self_hosted environment = work queue (/environments/{id}/work, poll/ack/heartbeat/stop) + worker (ant beta:worker, SDK EnvironmentWorker) with environment key sk-ant-oat01-…
Endpoint: GET /v1/environments/{environment_id}/work/poll
Status: DOCUMENTED · BETA
client-side by construction: shell {environment: local} and function tools end the server loop and hand control back; Grok Build CLI runs everything locally
Endpoint: POST /v1/responses
Params: tools[type=shell].environment
Status: DOCUMENTED · LIVE_VERIFIED
no self-hosted worker protocol; function calls (interaction.requires_action) are the only client-side hook
Endpoint: POST /v1beta/interactions
Status: DOCUMENTED
yes Worker protocols on OpenAI/Anthropic; plain tool round-trips on xAI/Gemini.
Ref: endpoints.json (managed-agents)
Credential vaults / secrets
3/4
/v1/vaults, /vaults/{id}/credentials (rotate/delete); referenced by vault_ids[] and mcp.credential_id
Endpoint: POST /v1/vaults
Status: DOCUMENTED · BETA · LIVE_VERIFIED
9 endpoints
/v1/vaults, /vaults/{id}/credentials (mcp_oauth with refresh, environment_variable), mcp_oauth_validate; vault_credential.refresh_failed webhook
Endpoint: POST /v1/vaults
Status: DOCUMENTED · BETA · LIVE_VERIFIED
13 endpoints
pass authorization / headers on the mcp tool per request; no vault
Endpoint: POST /v1/responses
Params: tools[type=mcp].authorization
Status: DOCUMENTED
/v1beta/credentials (bearer_token, oauth2, environment_variable with injection_location, trusted_domains); write-only secrets; referenced from environment network allowlists
Endpoint: POST /v1beta/credentials
Status: DOCUMENTED · BETA · PREVIEW
5 endpoints, not tested
yes Three secret stores; xAI inlines credentials.
Ref: endpoints.json
Multi-agent / subagents
3/4
multi_agent {enabled, max_concurrent_subagents (6)}; /sessions/{id}/subagents[/{id}/items|turns]; items create_subagent_call, wait_for_subagents_call…; also Responses beta responses_multi_agent=v1
Endpoint: GET /v1/agents/sessions/{session_id}/subagents
Params: multi_agent
Status: DOCUMENTED · BETA · LIVE_VERIFIED
multiagent {type: coordinator, agents[] (≤20, incl. self and one advisor)}; threads (/sessions/{id}/threads, ≤25 concurrent); events agent.thread_message_*
Endpoint: GET /v1/sessions/{session_id}/threads
Params: multiagent
Status: DOCUMENTED · BETA
grok-4.20-multi-agent-0309 (BETA): the model itself fans out to 4 or 16 agents per reasoning.effort; no subagent resources; Grok Build CLI has client-side subagents
Endpoint: POST /v1/responses
Params: reasoning.effort
Status: DOCUMENTED · BETA · LIVE_VERIFIED
live output carried \confidence{80} markers
no sub-agents or delegation for custom agents (docs); Deep Research orchestrates internally
Endpoint: POST /v1beta/interactions
Status: DOCUMENTED
yes Orchestration APIs on OpenAI/Anthropic; a multi-agent model on xAI; internal only on Gemini.
Ref: docs/comparisons/agents-platforms.md · docs/models/xai-models.md
Session budgets
3/4
no documented budget parameter; turn error code session_budget_exceeded exists
Status: DOCUMENTED · UNVERIFIED
budget {type: limit, max_list_cost {amount (cents), currency: USD}} enforced between model requests; session.budget_reached webhook
Endpoint: POST /v1/sessions
Params: budget
Status: DOCUMENTED · BETA
max_turns per Responses request (turn budget)
Endpoint: POST /v1/responses
Params: max_turns
Status: DOCUMENTED · LIVE_VERIFIED
Antigravity agent_config.max_total_tokens → status: incomplete
Endpoint: POST /v1beta/interactions
Params: agent_config(antigravity).max_total_tokens
Status: DOCUMENTED · PREVIEW
no Money (Anthropic), turns (xAI), tokens (Gemini) — not interchangeable.
Ref: docs/anthropic/managed-agents.md §13 · docs/xai/responses.md · docs/gemini/interactions-api.md
Outcome grading (define outcome + rubric)
1/4
— not offered user.define_outcome event → grader loop (max_iterations ≤20), span.outcome_evaluation_* events, outcome_evaluations[] on the session
Endpoint: POST /v1/sessions/{session_id}/events
Status: DOCUMENTED · BETA
— not offered — not offered no Anthropic-only.
Ref: docs/anthropic/managed-agents.md §13
Server-side memory stores
1/4
— not offered /v1/memory_stores (+ memories, versions, redact); mounted at /mnt/memory; beta agent-memory-2026-07-22; Dreams (/v1/dreams, dreaming-2026-04-21) reorganise them
Endpoint: POST /v1/memory_stores
Status: DOCUMENTED · BETA · LIVE_VERIFIED
14 + 5 endpoints
— not offered — not offered no Anthropic-only.
Ref: endpoints.json
Scheduled runs
2/4
— not offered /v1/deployments (cron) → /v1/deployment_runs; pause/unpause/run now
Endpoint: POST /v1/deployments
Status: DOCUMENTED · BETA · LIVE_VERIFIED
— not offered /v1beta/triggers {schedule (cron), time_zone, display_name, max_consecutive_failures, execution_timeout_seconds, interaction}; PATCH {status: paused|active}; POST|GET …/executions
Endpoint: POST /v1beta/triggers
Params: schedule, time_zone, interaction
Status: DOCUMENTED · BETA · PREVIEW
7 endpoints, not tested
yes Anthropic deployments ≈ Gemini triggers.
Ref: endpoints.json
Artifacts / deliverables
3/4
/sessions/{id}/artifacts[/{id}/content] from /workspace/outputs (≤200 MiB each)
Endpoint: GET /v1/agents/sessions/{session_id}/artifacts
Status: DOCUMENTED · BETA · LIVE_VERIFIED
files written to /mnt/session/outputs → GET /v1/files?scope_id=<session> + /content
Endpoint: GET /v1/files
Status: DOCUMENTED · BETA
tool outputs (code_interpreter_call.outputs, image/video file_output) via include or the Files API; no artifact resource
Endpoint: POST /v1/responses
Params: include
Status: DOCUMENTED
environment files: GET /v1beta/environments/{env}/files/{path}?alt=media (file bytes or tar), legacy GET /v1beta/files/environment-{env}:download; Deep Research reports as model_output steps (text/image)
Endpoint: GET /v1beta/environments/{environment}/files/{path}
Status: DOCUMENTED · BETA · PREVIEW · LIVE_VERIFIED
partial 200 live
yes
Ref: docs/openai/agents-environments-and-vaults.md §2 · docs/gemini/interactions-api.md
Embeddable chat UI & workspace agents
1/4
ChatKit (OpenAI-Beta: chatkit_beta=v1, /v1/chatkit/sessions|threads) and Workspace Agents (api.chatgpt.com/v1/workspace_agents/{id}/trigger)
Endpoint: POST /v1/chatkit/sessions
Status: DOCUMENTED · BETA · LIVE_VERIFIED
Agent Builder shuts down 2026-11-30
— not offered Grok Apps / Grok Bot integrations are consumer products, not an embeddable API (docs/xai/grok-apps-and-integrations.md)
Status: DOCUMENTED
AI Studio and Firebase AI Logic are builder/SDK surfaces, not an embeddable chat API
Status: DOCUMENTED
no OpenAI-only.
Ref: docs/openai/chatkit.md · docs/openai/workspace-agents.md
Client-side agent framework / CLI
4/4
Agents SDK (openai-agents, @openai/agents) — loop over Responses API; handoffs, guardrails, tracing
Status: DOCUMENTED
Claude Agent SDK / ant CLI (ant beta:sessions, ant apply); SDK tool_runner helpers
Status: DOCUMENTED
Claude Agent SDK not catalogued in this atlas beyond the migration mapping
Grok Build CLI (curl -fsSL https://x.ai/cli/install.sh | bash; grok, grok -p … --output-format json, grok agent stdio (ACP); ~/.grok/config.toml api_backend = chat_completions|responses|messages; MCP servers, hooks, skills, plugins, subagents, Landlock/Seatbelt sandbox; reads CLAUDE.md/AGENTS.md) — BETA; xai-sdk tool_runner-style helpers absent
Status: DOCUMENTED · BETA
google-genai SDK automatic function calling (Python callables, maximum_remote_calls 10) and mcpToTool(); Genkit, Firebase AI Logic, Vercel AI SDK, LangGraph/CrewAI/LlamaIndex integrations; Antigravity is the hosted coding agent
Status: DOCUMENTED · BETA
yes
Ref: docs/openai/agents-sdk.md · docs/anthropic/cli.md · docs/xai/grok-build.md · docs/gemini/sdks.md

# Model customisation & evaluation

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Fine-tuning / tuning
1/4
/v1/fine_tuning/jobs (SFT, DPO, RFT, vision) on gpt-4.1*, gpt-4o*, o4-mini — DEPRECATED: no new jobs after 2027-01-06; our key 403 training_not_available
Endpoint: POST /v1/fine_tuning/jobs
Params: model, training_file, method
Status: DOCUMENTED · DEPRECATED · ACCOUNT_RESTRICTED
no GPT-5.x/6 fine-tuning
— not offered none offered by xAI (capability matrix row 'fine-tuning: none')
Status: DOCUMENTED
tunedModels.* (12 endpoints) still in the v1beta discovery document but RETIRED on the Developer API since the Gemini 1.5 Flash-001 deprecation (May 2025): create/list/get → 501 UNIMPLEMENTED; no tunable 3.x/2.x model; use Vertex AI supervised tuning
Endpoint: POST /v1beta/tunedModels
Status: DOCUMENTED · DEPRECATED · RETIRED
Gemma 4 tuning: not available
no Only OpenAI still accepts jobs, and only until 2027-01-06.
Ref: docs/openai/fine-tuning.md · docs/gemini/tuning.md · docs/models/xai-models.md
Evals
1/4
/v1/evals, runs, output items; graders — DEPRECATED: read-only 2026-10-31, shutdown 2026-11-30
Endpoint: POST /v1/evals
Status: DOCUMENTED · DEPRECATED · LIVE_VERIFIED
12 endpoints
— not offered — not offered no evals API on the Developer API (Vertex AI evaluation service)
Status: DOCUMENTED
no OpenAI-only (deprecated).
Ref: docs/openai/evals.md
Graders (standalone)
1/4
/v1/fine_tuning/alpha/graders/{run,validate}
Endpoint: POST /v1/fine_tuning/alpha/graders/run
Status: DOCUMENTED · DEPRECATED · BETA · LIVE_VERIFIED
— not offered — not offered — not offered no OpenAI-only.
Ref: docs/openai/graders.md
Stored completions / distillation
1/4
Chat store:true + metadata → list/retrieve/update/delete stored completions; distillation via SFT
Endpoint: GET /v1/chat/completions
Params: store, metadata
Status: DOCUMENTED · LIVE_VERIFIED
— not offered Chat store/metadata silently accepted; no stored-completion retrieval endpoints (Responses store instead)
Endpoint: POST /v1/chat/completions
Params: store
Status: DOCUMENTED · LIVE_DISCOVERED
— not offered no OpenAI-only.
Ref: docs/openai/chat-completions.md · parameters.json (xai store)
Reusable prompt templates
1/4
prompt {id, version, variables} on Responses — DEPRECATED, shutdown 2026-11-30
Endpoint: POST /v1/responses
Params: prompt
Status: DOCUMENTED · DEPRECATED
— not offered — not offered no prompt registry on the API (AI Studio saves prompts client-side); cachedContents reuse a prefix
Status: DOCUMENTED
no OpenAI-only (deprecated).
Ref: docs/openai/deprecations.md

# Legacy / retired surfaces

Feature OpenAI (how / endpoint / params / status) Anthropic xAI Gemini Portable Notes on differences
Retired agent / thread APIs
0/4
Assistants API RETIRED 2026-08-26 — 404 empty body; 23 operations mapped to Conversations/Responses/Agents
Endpoint: GET /v1/assistants
Status: RETIRED · DOCUMENTED
OpenAI-Beta: assistants=v2 historical
— not offered — not offered Interactions v1beta legacy schema (outputs → steps, old SSE names interaction.start, content.*) removed 2026-06-08; total_reasoning_tokens → total_thought_tokens
Endpoint: POST /v1beta/interactions
Status: RETIRED · DOCUMENTED
no
Ref: docs/openai/assistants-retired.md · docs/gemini/interactions-api.md
Retired media models / endpoints
0/4
DALL·E 2/3 RETIRED 2026-05-12; /v1/images/variations 404; Sora 2 + Videos API shut down 2026-09-24
Endpoint: POST /v1/images/variations
Status: RETIRED · DEPRECATED
— not offered grok-2-image(-1212) RETIRED (404); grok-imagine-image-pro → -quality (DEPRECATED, retires 2026-11-02 → grok-imagine-image-2.0 low); Live Search 410
Endpoint: POST /v1/images/generations
Status: RETIRED · DEPRECATED
Imagen 3/4 (:predict) RETIRED 2026-08-17; Veo 2.0/3.0 RETIRED 2026-06-30; gemini-2.5-flash-image DEPRECATED → 2026-10-02; gemini-omni-flash-preview → 2026-09-30; half-cascade Live models RETIRED 2025-12-09
Endpoint: POST /v1beta/models/{model}:predict
Status: RETIRED · DEPRECATED
no
Ref: docs/openai/images.md · docs/xai/deprecations-and-release-notes.md · docs/gemini/deprecations-and-changelog.md
Retired text models
0/4
gpt-3.5-turbo-instruct/babbage/davinci shut down 2026-09-28; o1/o3-mini/o4-mini/gpt-4/gpt-4-turbo 2026-10-23; gpt-5 2025 snapshots, o3 2026-12-11; codex ≤5.2, deep-research, computer-use-preview retired
Status: DEPRECATED · RETIRED
RETIRED on the Claude API: Opus 4.1 (2026-08-05), Sonnet 4 / Opus 4 (2026-06-15), Claude 3.x, 2.x, 1.x, Instant; some still served on Bedrock/Vertex
Endpoint: POST /v1/messages
Status: RETIRED
404 not_found_error
RETIRED 2026-05-15 with redirects: grok-3, grok-4-0709, grok-4-fast-*, grok-4-1-fast-* → grok-4.3 (billed at 4.3 rates; GET /v1/models/grok-3 returns the grok-4.3 object), grok-code-fast-1 → grok-build-0.1; LEGACY/UNVERIFIED: grok-2-*, grok-3-mini, grok-beta, grok-vision-beta, grok-4-latest
Endpoint: GET /v1/models/{model_id}
Status: RETIRED · LEGACY
no consolidated deprecations page; ~60-day notice observed
Gemini 2.0 Flash/-Lite RETIRED 2026-06-01; 2.5 previews 2025-11/2026-03; gemini-3-pro-preview 2026-03-09 (id repointed to 3.1 Pro), gemini-3.1-flash-lite-preview 2026-05-25; Gemini 2.5 Pro/Flash/Flash-Lite 'no longer available to new users' (404, undocumented); gemini-3.1-flash-lite DEPRECATED → 2027-05-07; gemini-embedding-001 → 2028-05-14; text-embedding-004 2026-01-14
Endpoint: POST /v1beta/models/{model}:generateContent
Status: RETIRED · DEPRECATED · ACCOUNT_RESTRICTED
shut-down ids still appear in GET /v1beta/models
no Retirement means 404 on Anthropic/Gemini, a dated shutdown on OpenAI, and a silent redirect on xAI.
Ref: deprecations.json · models.json
Retired / deprecated beta headers, parameters and SDKs
4/4
OpenAI-Beta: realtime=v1 (legacy /v1/realtime/sessions → 404), assistants=v2; prompt_cache_retention, truncation: auto, user (→ safety_identifier) legacy
Status: LEGACY · FAILED_VERIFICATION
RETIRED: context-1m-2025-08-07, computer-use-2024-10-22, max-tokens-3-5-sonnet-2024-07-15; DEPRECATED: mcp-client-2025-04-04; LEGACY (GA, header optional): prompt-caching, message-batches, pdfs, token-counting, files-api, skills, structured-outputs, extended-cache-ttl, code-execution-2025-05-22, interleaved/fine-grained streaming, effort-2025-11-24…
Status: RETIRED · DEPRECATED · LEGACY
no headers to retire; DEPRECATED params/features: max_tokens (→ max_completion_tokens), logprobs/top_logprobs (ignored ≥4.20), Anthropic-compatible /v1/messages, x_search per-call billing (→ 2026-09-21), zdr_status: pii_scrubbing, Management teamId (→ scope/scopeId); LEGACY: /v1/completions, /v1/complete, Live Search
Status: DEPRECATED · LEGACY · RETIRED
DEPRECATED: temperature/topP/topK guidance (2026-07-21), thinkingBudget (LEGACY on 3.x), HARM_CATEGORY_CIVIC_INTEGRITY (→ enableEnhancedCivicAnswers), googleSearchRetrieval, standard API keys (rejected from Sept 2026), legacy SDKs google-generativeai / @google/generative-ai (2025-11-30); LEGACY: PaLM methods, responseSchema/responseJsonSchema (→ responseFormat), mediaChunks, speechState
Status: DEPRECATED · LEGACY
no
Ref: generated/fragments/headers/anthropic-beta-headers.json · deprecations.json (xai, gemini api_features)
Retired model ids keep resolving (redirect aliases)
1/4
retired ids fail with model_not_found; shutdown_date exposed on GET /v1/models
Endpoint: GET /v1/models/{model}
Status: DOCUMENTED · LIVE_VERIFIED
retired ids → 404 not_found_error (some still served on Bedrock/Vertex)
Endpoint: POST /v1/messages
Status: DOCUMENTED
retired slugs redirect to their replacement and are billed at the replacement's price: grok-3, grok-4-0709, grok-4-fast-*, grok-4-1-fast-* → grok-4.3 (verified: GET /v1/models/grok-3 returns the grok-4.3 object), grok-code-fast-1 → grok-build-0.1, grok-imagine-image-pro → -quality → -2.0 low; response.model reveals the target
Endpoint: GET /v1/models/{model_id}
Status: DOCUMENTED · RETIRED · LIVE_VERIFIED
6 retired_redirect records in models.json
shut-down ids stay listed in GET /v1beta/models but generation fails (404); one documented repoint: gemini-3-pro-preview → gemini-3.1-pro-preview (2026-03-09); -latest aliases hot-swap targets
Endpoint: GET /v1beta/models
Status: DOCUMENTED · LIVE_DISCOVERED
no xAI-only behaviour (silent redirect); Gemini has one documented id repoint.
Ref: deprecations.json (xai) · docs/xai/deprecations-and-release-notes.md

# Reading the matrix programmatically

bash
# all features unique to xAI
jq '[.records[] | select(.providers_supporting == ["xai"]) | .feature]' generated/compatibility/cross-provider-feature-matrix.json
# features on all four providers that are marked portable
jq '[.records[] | select(.provider_count == 4 and .portable) | .feature]' generated/compatibility/cross-provider-feature-matrix.json
# every Gemini cell that is ACCOUNT_RESTRICTED (paid-tier feature probed with a free key)
jq '[.records[] | select(.gemini.status|index("ACCOUNT_RESTRICTED")) | {feature, endpoint: .gemini.endpoint}]' generated/compatibility/cross-provider-feature-matrix.json
# every beta-gated Anthropic cell
jq '[.records[] | select(.anthropic.status|index("BETA")) | {feature, endpoint: .anthropic.endpoint}]' generated/compatibility/cross-provider-feature-matrix.json

Related pages: index · models · state management · tool execution · streaming · agents platforms · pricing · caching and reasoning · realtime and media · endpoint catalogue · FAQ.