Feature × Provider matrix — OpenAI · Anthropic · xAI · Gemini
Status: synthesis of generated/*.json (endpoints, parameters, tools, models, headers, pricing, streaming-events, webhook-events, deprecations) and the domain pages under docs/; every status shown is the status recorded in those files (LIVE_VERIFIED means called successfully with this atlas's keys on 2026-09-18; xAI probes ran on a Tier 0 team, Gemini probes on a free-tier key — hence ACCOUNT_RESTRICTED on paid-only Gemini features and on xAI's Management/Skills/Embeddings). Machine-readable twin: generated/compatibility/cross-provider-feature-matrix.json (same rows, openai/anthropic/xai/gemini objects per record, providers_supporting[], provider_count).
Sources: the ref column of the JSON twin names, per row, the generated file or docs page each cell was taken from. Canonical vendor pages: https://developers.openai.com/api/docs · https://platform.claude.com/docs/en · https://docs.x.ai/developers · https://ai.google.dev/gemini-api/docs.
Last verified: 2026-09-18
Legend — Portable = the same task can be expressed with an equivalent parameter on every provider that offers it (mapping in the per-topic pages; a feature offered by one provider only is never portable); “— not offered” = no documented surface. Statuses use the atlas vocabulary (DOCUMENTED, LIVE_VERIFIED, LIVE_DISCOVERED, BETA, PREVIEW, GA, LEGACY, DEPRECATED, RETIRED, ACCOUNT_RESTRICTED, UNVERIFIED, FAILED_VERIFICATION, DOCUMENTATION_INCOMPLETE).
Contents
- Core generation (12 rows)
- Structured output & sampling (10 rows)
- Tools — client side (13 rows)
- Tools — server side / hosted (13 rows)
- Multimodal input & media APIs (13 rows)
- Files & batch (3 rows)
- Prompt caching (2 rows)
- Reasoning / thinking (6 rows)
- Context management (6 rows)
- Service tiers, limits, safety (8 rows)
- Auth, versioning, SDKs, platform (14 rows)
- Managed agents platforms (15 rows)
- Model customisation & evaluation (5 rows)
- Legacy / retired surfaces (5 rows)
Totals: 125 features — on all four providers: 45 · on three: 34 · on two: 16 · on one: 26 · on none (legacy/absent everywhere): 4 · marked portable: 86.
| Coverage | Count | Features |
|---|---|---|
| All four | 45 | Primary text/multimodal generation endpoint; System / developer prompt; Message roles; Multi-turn conversation; Token counting; Model listing / catalogue; Structured outputs (JSON Schema constrained); Strict tool arguments; Sampling parameters (temperature / top_p / top_k); Stop sequences; Refusal / safety signalling in-band; Function / custom tools (you execute); Tool choice; Parallel tool calls; Web search; Code execution (sandboxed Python); Remote MCP servers; Citations / grounding metadata; Image input; PDF / document input; Files API; Batch processing (async, discounted); Prompt / context caching; Reasoning control; Reasoning visibility; Reasoning replay across turns; Context window; Max output tokens; Service tiers / processing modes; End-user identifier for abuse detection; Rate-limit tiers; Overload / capacity error; Error envelope; Authentication; Official SDKs; Webhooks (platform events); Usage & cost reporting; Cloud availability; Zero data retention / data-use terms; Managed agent harness; Create an agent session / run; Send input / events to a session; Session event stream; Client-side agent framework / CLI; Retired / deprecated beta headers, parameters and SDKs |
| anthropic + xai + gemini | 2 | Task-wide token budget (advisory); Session budgets |
| openai + anthropic + gemini | 6 | Computer use; Private-network MCP (tunnels) / agent credentials; Server-side compaction (in-flight); Hosted sandbox; Credential vaults / secrets; Artifacts / deliverables |
| openai + anthropic + xai | 12 | Deferred tool loading / tool search; Shell execution; Skills (packaged instructions + files); Stand-alone compaction request; Rate-limit headers; Request correlation; Administration API; Audit / compliance data access; Spend limits; Data residency; Self-hosted execution; Multi-agent / subagents |
| openai + xai + gemini | 14 | Legacy / secondary chat endpoint; Response storage, retrieval, deletion; Background / deferred (async) execution; JSON mode (valid JSON, no schema); File search / managed RAG; Image generation as a tool / modality inside a text call; Text-to-speech; Speech-to-text / transcription; Realtime speech-to-speech (WebSocket / WebRTC / SIP); Image generation API; Video generation API; Multipart / resumable upload; Pro / extended-compute mode; OpenAI-compatibility layer |
| anthropic + gemini | 3 | Web fetch / URL context (retrieve specific URLs); API version header / path version; Scheduled runs |
| anthropic + xai | 1 | Anthropic-compatible Messages endpoint |
| openai + anthropic | 6 | Programmatic tool calling (model writes code that calls tools); File editing tool; Cache diagnostics; Change effort mid-conversation without breaking the cache; Mid-conversation tool add/remove (cache-preserving); Beta opt-in header |
| openai + gemini | 5 | Configurable safety thresholds; Async tools (model continues while a tool runs); Container management API; Audio / video input (understanding); Embeddings |
| xai + gemini | 1 | Social / vertical search (X posts, Google Maps) |
| OpenAI only | 15 | Legacy completion (prompt-in / text-out) endpoint; Output verbosity control; Log probabilities; Free-form / grammar-constrained custom tools; Tool namespaces; Built-in SaaS connectors; Full-duplex voice with backend delegation (Live) / voice agents; Moderation endpoint; Content provenance; Embeddable chat UI & workspace agents; Fine-tuning / tuning; Evals; Graders (standalone); Stored completions / distillation; Reusable prompt templates |
| Anthropic only | 8 | Assistant prefill (continue a partial assistant turn); Server-side model fallback on refusal; Memory tool (client-side persistent memory); Browser use toolset; Advisor (consult a stronger model mid-response); Server-side context editing (clear old tool results / thinking); Outcome grading (define outcome + rubric); Server-side memory stores |
| xAI only | 2 | Per-request dollar cost in the response; Retired model ids keep resolving (redirect aliases) |
| Gemini only | 1 | Music generation |
| None | 4 | Idempotency key; Retired agent / thread APIs; Retired media models / endpoints; Retired text models |
Core generation
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Primary text/multimodal generation endpoint 4/4 |
Responses API: input (string or Item[]), output[] itemsEndpoint: POST /v1/responsesParams: model, input, instructions, tools, text, reasoningStatus: DOCUMENTED · LIVE_VERIFIEDstored by default ( store:true, 30 days) |
Messages API: messages[] of content blocks, content[] blocks outEndpoint: POST /v1/messagesParams: model, messages, system, max_tokens, tools, output_config, thinkingStatus: DOCUMENTED · LIVE_VERIFIEDstateless; max_tokens required (400 if missing) |
Responses API clone: same input items / output[] items as OpenAI (message, reasoning, function_call, web_search_call, code_interpreter_call, mcp_call…); xAI extras max_turns, top_k, min_p, reasoning_effort; usage.cost_in_usd_ticksEndpoint: POST /v1/responsesParams: model, input, instructions, tools, text, reasoning, max_turns, storeStatus: DOCUMENTED · LIVE_VERIFIEDstored by default (30 days); background → 400; metadata → 400; every Grok call bills reasoning tokens |
generateContent: contents[] {role: user|model, parts[]} → candidates[].content.parts[]; generationConfig holds every knob; Google calls it 'legacy' since June 2026 in favour of the Interactions API (POST /v1beta/interactions, snake_case, model or agent, input, steps[] out)Endpoint: POST /v1beta/models/{model}:generateContentParams: contents, systemInstruction, generationConfig, tools, toolConfig, safetySettings, cachedContentStatus: DOCUMENTED · LIVE_VERIFIED/v1 stable twin exists for generateContent; Interactions /v1beta BETA + LIVE_VERIFIED, /v1 GA but UNVERIFIED |
yes | Four envelopes, two shapes: OpenAI and xAI share the items model (xAI reimplements the Responses API); Anthropic uses content blocks; Gemini uses parts inside contents[] with generationConfig. Required fields differ: Anthropic needs max_tokens; Gemini needs the last turn to be user; xAI rejects background/metadata.Ref: docs/openai/responses.md · docs/anthropic/messages-api.md · docs/xai/responses.md · docs/gemini/generate-content.md · docs/gemini/interactions-api.md |
| Legacy / secondary chat endpoint 3/4 |
Chat Completions (messages[], choices[]), still GA and maintainedEndpoint: POST /v1/chat/completionsParams: messages, response_format, reasoning_effort, web_search_optionsStatus: DOCUMENTED · LIVE_VERIFIEDno hosted tools except search models; list/retrieve/update/delete of stored completions |
— not offered none — Messages is the only chat surface |
Chat Completions, OpenAI-compatible; xAI calls it the legacy predecessor of /v1/responses (new features land on Responses first); only function tools; deferred:true + GET /v1/chat/deferred-completion/{id}; reasoning_content fieldEndpoint: POST /v1/chat/completionsParams: messages, reasoning_effort, response_format, deferred, prompt_cache_key, max_completion_tokensStatus: DOCUMENTED · LEGACY · LIVE_VERIFIEDfile parts → 400 (use Responses); server tools → 422; Live Search params → 410 |
OpenAI-compatibility layer POST /v1beta/openai/chat/completions (Bearer key required; extra_body.google.thinking_config, cached_content; unknown params silently ignored); native generateContent itself is now labelled legacy vs InteractionsEndpoint: POST /v1beta/openai/chat/completionsParams: messages, reasoning_effort, response_format, extra_body.googleStatus: DOCUMENTED · BETA · LIVE_VERIFIEDno Responses / Assistants / audio routes on the compat layer |
yes | OpenAI, xAI and Gemini all expose an OpenAI-shaped chat/completions; Anthropic has none. Only OpenAI's is a first-class, fully featured surface.Ref: docs/comparisons/responses-vs-chat-completions.md · docs/xai/chat-completions.md · docs/gemini/openai-compatibility.md |
| Anthropic-compatible Messages endpoint 2/4 |
— not offered | the native surface Endpoint: POST /v1/messagesParams: messages, system, max_tokens, toolsStatus: DOCUMENTED · LIVE_VERIFIED |
POST /v1/messages accepting Anthropic shapes (system, messages, tools[{name,description,input_schema}], tool_choice {auto|any|tool}, thinking blocks with empty signature) — fully deprecated by xAI ('migrate to Responses or gRPC'); Bearer auth, anthropic-version ignored; /v1/complete twin RETIRED (400)Endpoint: POST /v1/messagesParams: messages, system, max_tokens, tools, tool_choiceStatus: DOCUMENTED · DEPRECATED · LIVE_VERIFIEDtop_k 400, stop_sequences 400 on reasoning models, document blocks 422, no count_tokens (404), no batches, cache_control ignored |
— not offered | yes | xAI is the only third party exposing the Anthropic Messages wire format, and is retiring it. Ref: docs/xai/messages-compat.md · parameters.json (xai POST /v1/messages) |
| Legacy completion (prompt-in / text-out) endpoint 1/4 |
gpt-3.5-turbo-instruct, davinci-002, babbage-002 only; shutdowns 2026-09-28Endpoint: POST /v1/completionsParams: prompt, suffix, max_tokensStatus: DOCUMENTED · LEGACY · LIVE_VERIFIED |
documented as Legacy but every live call returns 400 'has been deprecated' Endpoint: POST /v1/completeParams: prompt, max_tokens_to_sampleStatus: DOCUMENTED · LEGACY · DEPRECATED · FAILED_VERIFICATIONremoved from Python SDK v1 |
POST /v1/completions (OpenAI shape) and POST /v1/complete (Anthropic shape) both answer 400 'Raw sampling is not supported for reasoning models' for every live model, incl. the non-reasoning grok-4.20Endpoint: POST /v1/completionsParams: prompt, max_tokensStatus: DOCUMENTED · LEGACY · RETIREDgRPC Sample service still lists raw sampling |
PaLM-era :generateText / :generateMessage remain in the v1beta discovery document but return 404 / 501; only models/aqa:generateAnswer still answersEndpoint: POST /v1beta/models/{model}:generateTextStatus: DOCUMENTED · LEGACY · FAILED_VERIFICATIONlegacy-palm family: 7 endpoints |
no | Every provider still documents a prompt-completion route; only OpenAI's answers, and it shuts down 2026-09-28. Ref: docs/openai/completions-legacy.md · docs/anthropic/text-completions-legacy.md · docs/xai/legacy-completions.md · docs/gemini/legacy-palm-methods.md |
| System / developer prompt 4/4 |
instructions (per request, not carried by previous_response_id) or developer/system role message itemsEndpoint: POST /v1/responsesParams: instructions, input[](message).roleStatus: DOCUMENTED · LIVE_VERIFIEDinstructions participate in the cache prefix |
top-level system (string or text blocks with cache_control/citations); mid-conversation role:system messages for tool changes/effort (beta)Endpoint: POST /v1/messagesParams: system, system[].cache_control, messages[].output_config.effortStatus: DOCUMENTED · LIVE_VERIFIEDschema lists a system role but docs say use the top-level field |
instructions (Responses) or a system / developer role message; Chat: system role; /v1/messages: system string or text blocksEndpoint: POST /v1/responsesParams: instructions, input[](message).roleStatus: DOCUMENTED · LIVE_VERIFIEDinstructions + previous_response_id → 400 (the previous system prompt is reused) |
systemInstruction (Content, text parts only, role ignored) — not a turn; Interactions: system_instruction string (must be resent when chaining)Endpoint: POST /v1beta/models/{model}:generateContentParams: systemInstructionStatus: DOCUMENTED · LIVE_VERIFIEDcounted in promptTokenCount; cacheable inside cachedContents |
yes | All four keep the system prompt outside the turn list. OpenAI/xAI accept it either as a field or as a role item; Anthropic and Gemini only as a field. Ref: parameters.json (instructions / system / systemInstruction) |
| Message roles 4/4 |
user, assistant, system, developer (+ phase: commentary|final_answer on assistant)Endpoint: POST /v1/responsesParams: input[](message).role, input[](message).phaseStatus: DOCUMENTED · LIVE_VERIFIED |
user, assistant; consecutive same-role turns are merged; ≤100,000 messagesEndpoint: POST /v1/messagesParams: messages[].roleStatus: DOCUMENTED · LIVE_VERIFIEDsystem role reserved for beta mid-conversation blocks |
user, assistant, system, developer (Responses); Chat adds tool (tool_call_id) and legacy function; no ordering constraintEndpoint: POST /v1/responsesParams: input[](message).role, messages[].roleStatus: DOCUMENTED · LIVE_VERIFIEDdeveloper accepted live on Chat although not in the spec |
user, model only (omitted role = user); last turn must be user (400 'Requests ending with a model turn are not supported'); function results go in a user Content of functionResponse partsEndpoint: POST /v1beta/models/{model}:generateContentParams: contents[].roleStatus: DOCUMENTED · LIVE_VERIFIEDInteractions steps use role: user|model too |
yes | The assistant role is spelled model on Gemini; Anthropic enforces alternation by merging; Gemini enforces a trailing user turn; OpenAI/xAI accept any order.Ref: docs/anthropic/messages-api.md §1 · docs/gemini/generate-content.md |
| Assistant prefill (continue a partial assistant turn) 1/4 |
no prefill semantics; an assistant item is history only Endpoint: POST /v1/responsesStatus: DOCUMENTED |
end messages with an assistant turn → model continues it. Deprecated: 400 on Claude 4.6+ / Fable / Mythos, incompatible with thinking and structured outputsEndpoint: POST /v1/messagesParams: messages[].content[] (assistant prefill)Status: DOCUMENTED · DEPRECATED · LIVE_VERIFIEDlive: Haiku 4.5 OK, Sonnet 5 → 400 |
assistant history items / assistant role accepted, but no documented continuation semanticsEndpoint: POST /v1/responsesStatus: DOCUMENTED |
impossible: a request ending with a model turn is rejected (400)Endpoint: POST /v1beta/models/{model}:generateContentStatus: DOCUMENTED · LIVE_VERIFIED |
no | Only Anthropic ever offered prefill, and only its 4.5 models still honour it; Gemini rejects the pattern outright. Ref: docs/anthropic/messages-api.md · deprecations.json · docs/gemini/generate-content.md |
| Multi-turn conversation 4/4 |
three modes: manual replay of output[] items, previous_response_id (server keeps chain), conversation (Conversations API)Endpoint: POST /v1/responsesParams: previous_response_id, conversation, storeStatus: DOCUMENTED · LIVE_VERIFIEDprevious_response_id and conversation are mutually exclusive |
manual replay only: resend full messages[] each call (incl. thinking/tool_use/tool_result blocks)Endpoint: POST /v1/messagesParams: messagesStatus: DOCUMENTED · LIVE_VERIFIEDserver-side state exists only in Managed Agents sessions |
manual replay or previous_response_id (server rehydrates the whole agentic trajectory incl. reasoning and tool outputs; follow-ups may change tools/model); Chat: replay incl. reasoning_contentEndpoint: POST /v1/responsesParams: previous_response_id, store, inputStatus: DOCUMENTED · LIVE_VERIFIEDno Conversations API; x-grok-conv-id / prompt_cache_key are cache-routing keys, not stored conversations |
generateContent: manual replay of contents[] (echo thoughtSignature on Gemini 3 function calls); Interactions: previous_interaction_id chains stored interactions (store:true default; only history is carried — resend tools/system/config)Endpoint: POST /v1beta/models/{model}:generateContentParams: contents, previous_interaction_id, storeStatus: DOCUMENTED · LIVE_VERIFIEDchaining on an in_progress interaction → 400 |
yes | Portable pattern = manual replay everywhere; OpenAI, xAI and Gemini (Interactions) add server-side chaining by id. Ref: docs/comparisons/state-management.md |
| Response storage, retrieval, deletion 3/4 |
store:true default → GET/DELETE /v1/responses/{id}, GET …/input_items; 30-day retentionEndpoint: GET /v1/responses/{response_id}Params: store, includeStatus: DOCUMENTED · LIVE_VERIFIEDstore:false → 404 on retrieve, reasoning returned as encrypted_content |
no response store; nothing to retrieve after the HTTP response Status: DOCUMENTEDBatches results are retrievable 29 days |
store:true default → GET/DELETE /v1/responses/{id}, GET …/input_items (limit 1–100, order, after); 30-day retention; ZDR teams cannot storeEndpoint: GET /v1/responses/{response_id}Params: storeStatus: DOCUMENTED · LIVE_VERIFIED · LIVE_DISCOVEREDLIVE_DISCOVERED: GET still 200 for a store:false id |
Interactions: store:true default → GET /v1beta/interactions/{id} (full timeline incl. user_input), DELETE, POST …/cancel; retention 55 days paid (AI Studio 7/14/28/55), 1 day free; store:false → no id, stateless. generateContent stores nothing (per-request store logging flag LIVE_DISCOVERED)Endpoint: GET /v1beta/interactions/{id}Params: storeStatus: DOCUMENTED · BETA · LIVE_VERIFIEDGET /v1beta/interactions (list) → 404 |
yes | Three stateful-by-default surfaces (OpenAI Responses, xAI Responses, Gemini Interactions) vs Anthropic's stateless Messages. Ref: docs/openai/responses.md §1 · docs/xai/responses.md · docs/gemini/interactions-api.md |
| Background / deferred (async) execution 3/4 |
background:true → status:queued, poll GET, POST …/cancel, resumable stream ?stream=true&starting_after=NEndpoint: POST /v1/responsesParams: backgroundStatus: DOCUMENTED · LIVE_VERIFIEDnot available over WebSocket; not in EU region |
no per-request background mode; use Message Batches (≤24 h) or Managed Agents sessions Endpoint: POST /v1/messages/batchesStatus: DOCUMENTED |
Chat Completions only: deferred:true → {request_id}; GET /v1/chat/deferred-completion/{request_id} → 202 while pending, 200 when done (docs: retrievable once within 24 h; live: second GET also 200)Endpoint: GET /v1/chat/deferred-completion/{request_id}Params: deferredStatus: DOCUMENTED · LIVE_VERIFIEDResponses background → 400 'Argument not supported' |
Interactions background:true (requires store) → queued/in_progress; poll GET /v1beta/interactions/{id} or resume the stream with ?stream=true&last_event_id=; POST …/cancel; mandatory for Deep Research (≤60 min); webhook_config for completion callbacksEndpoint: POST /v1beta/interactionsParams: background, stream, webhook_configStatus: DOCUMENTED · BETA · LIVE_VERIFIEDVeo/Batch use long-running Operations instead |
yes | OpenAI, xAI (Chat only) and Gemini (Interactions) offer per-request async; Anthropic only batches. Ref: docs/openai/responses.md §5 · docs/xai/deferred-completions.md · docs/gemini/interactions-api.md |
| Token counting 4/4 |
count input tokens of a Responses payload (free) Endpoint: POST /v1/responses/input_tokensParams: model, input, tools, instructionsStatus: DOCUMENTED · LIVE_VERIFIEDsupports images/files |
count tokens of a Messages payload (free; separate RPM bucket 5k/10k/20k) Endpoint: POST /v1/messages/count_tokensParams: model, messages, system, tools, thinking, output_config.formatStatus: DOCUMENTED · LIVE_VERIFIEDserver tools other than advisor → 400; mcp_servers rejected |
POST /v1/tokenize-text {model, text} → token_ids[{token_id, string_token, token_bytes}] (tokenizer, not a request counter); freeEndpoint: POST /v1/tokenize-textParams: model, textStatus: DOCUMENTED · LIVE_VERIFIEDno per-request counter; /v1/messages/count_tokens → 404 |
:countTokens with {contents} or {generateContentRequest:{model, contents, systemInstruction, tools, cachedContent, generationConfig}} → totalTokens, promptTokensDetails[], cachedContentTokenCount; freeEndpoint: POST /v1beta/models/{model}:countTokensParams: contents, generateContentRequestStatus: DOCUMENTED · LIVE_VERIFIEDSDKs refuse system_instruction/tools here although REST accepts them |
yes | Free everywhere; xAI tokenises raw text only (no tools/images), the other three count a full request. Ref: docs/anthropic/token-counting.md · docs/openai/multimodal-input.md · docs/xai/index.md · docs/gemini/token-counting.md |
| Model listing / catalogue 4/4 |
list/retrieve (shutdown_date field on deprecated ids); some aliases (e.g. gpt-5.6) 404 on GET but work on POSTEndpoint: GET /v1/modelsStatus: DOCUMENTED · LIVE_VERIFIED136 ids live |
list/retrieve with capabilities block (thinking, effort, structured_outputs, context_management…)Endpoint: GET /v1/modelsParams: limit, before_id, after_idStatus: DOCUMENTED · LIVE_VERIFIED11 ids live; invite-only Mythos ids → 404 |
GET /v1/models (OpenAI shape, 12 ids) plus typed catalogues GET /v1/language-models, /v1/image-generation-models, /v1/video-generation-models, /v1/embedding-models with live price ticks, aliases, input_modalities, fingerprint; retired slugs redirect (grok-3 → grok-4.3 object)Endpoint: GET /v1/language-modelsStatus: DOCUMENTED · LIVE_VERIFIEDvoice models absent from the catalogue; /v1/embedding-models → {models: []} for this team |
GET /v1beta/models (58 ids: inputTokenLimit, outputTokenLimit, supportedGenerationMethods, thinking, temperature/topP/topK defaults, version) and GET /v1beta/models/{model}; /v1/models lists only 22 stable idsEndpoint: GET /v1beta/modelsParams: pageSize, pageTokenStatus: DOCUMENTED · LIVE_VERIFIEDagents (deep-research, antigravity) appear as models; shut-down previews still listed |
yes | xAI publishes prices in the catalogue, Anthropic publishes capability flags, Gemini publishes token limits and generation methods, OpenAI publishes shutdown dates. Ref: sources/*/models-api-raw.json |
Structured output & sampling
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Structured outputs (JSON Schema constrained) 4/4 |
text.format {type:json_schema, name, schema, strict, description} (flat)Endpoint: POST /v1/responsesParams: text.format, text.format(json_schema).strictStatus: DOCUMENTED · LIVE_VERIFIEDChat: response_format.json_schema{…} wrapper; refusal → refusal content part |
output_config.format {type:json_schema, schema} (GA, no beta header; legacy output_format → 400)Endpoint: POST /v1/messagesParams: output_config.format, output_config.format.schemaStatus: DOCUMENTED · LIVE_VERIFIEDgrammar compiled once (24 h cache), first call slower; stop_reason: refusal possible |
text.format {type:json_schema, name, schema, strict, description} (Responses) / response_format {type:json_schema, json_schema:{name, strict, schema}} (Chat); Draft 2020-12 preferred; additionalProperties defaults false; formats date/time/email/uuid/uri enforced; pattern ECMA subsetEndpoint: POST /v1/responsesParams: text.format, response_formatStatus: DOCUMENTED · LIVE_VERIFIEDstrict accepted and ignored (always strict); all Grok 4 models |
generationConfig.responseMimeType: application/json + responseJsonSchema (JSON Schema) or legacy responseSchema (OpenAPI subset, propertyOrdering); new canonical responseFormat.text {mimeType: APPLICATION_JSON, schema}; enum mode text/x.enum; XML/YAML mime types accepted; Interactions response_format {type:text, mime_type, schema}Endpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.responseMimeType, generationConfig.responseJsonSchema, generationConfig.responseSchema, generationConfig.responseFormatStatus: DOCUMENTED · LIVE_VERIFIEDvalues are syntactically valid but not semantically validated; maxOutputTokens can truncate the JSON; SO + tools = Gemini 3 preview |
yes | Same task on all four. Schema dialects differ: OpenAI/xAI require additionalProperties:false semantics (xAI defaults it), Anthropic rejects numeric/string constraints, Gemini ignores unsupported keywords silently and offers an OpenAPI-style alternative.Ref: docs/openai/structured-outputs.md · docs/anthropic/structured-outputs.md · docs/xai/structured-outputs.md · docs/gemini/structured-outputs.md |
| Strict tool arguments 4/4 |
tools[type=function].strict:true (Responses omits → tries strict then falls back)Endpoint: POST /v1/responsesParams: tools[type=function].strictStatus: DOCUMENTED · LIVE_VERIFIED |
tools[].strict:true grammar-constrained tool_use.input; ≤20 strict tools/requestEndpoint: POST /v1/messagesParams: tools[].strictStatus: DOCUMENTED · LIVE_VERIFIEDnot on toolsets / mcp_toolset / programmatic callers |
tool parameters are always strictly enforced ('strict flag implicitly true'); explicit strict accepted and ignoredEndpoint: POST /v1/responsesParams: tools[].strictStatus: DOCUMENTED · LIVE_VERIFIED |
toolConfig.functionCallingConfig.mode: VALIDATED (schema-validated constrained decoding; default when built-ins or structured output are combined) or ANY (forced, constrained)Endpoint: POST /v1beta/models/{model}:generateContentParams: toolConfig.functionCallingConfig.modeStatus: DOCUMENTED · LIVE_VERIFIEDno per-tool flag; ANY may reject very large/deep schemas |
yes | Per-tool flag on OpenAI/Anthropic, always-on on xAI, a request-level mode on Gemini. Ref: tools.json · docs/tools/gemini/function-calling.md |
| JSON mode (valid JSON, no schema) 3/4 |
text.format {type:json_object}; prompt must mention JSON (400 otherwise)Endpoint: POST /v1/responsesParams: text.format(json_object)Status: DOCUMENTED · LIVE_VERIFIEDlegacy |
— not offered | text.format {type:json_object} / response_format {type:json_object}Endpoint: POST /v1/responsesParams: text.format(json_object)Status: DOCUMENTED · LIVE_VERIFIED |
responseMimeType: application/json without a schemaEndpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.responseMimeTypeStatus: DOCUMENTED · LIVE_VERIFIED |
yes | Anthropic only offers schema-constrained output. Ref: docs/openai/structured-outputs.md · docs/gemini/structured-outputs.md |
| Output verbosity control 1/4 |
text.verbosity: low|medium|highEndpoint: POST /v1/responsesParams: text.verbosityStatus: DOCUMENTED · LIVE_VERIFIED |
no direct knob; output_config.effort shapes length indirectlyStatus: DOCUMENTED |
no verbosity parameter (text.format only)Status: DOCUMENTED |
no verbosity parameter; thinkingLevel and maxOutputTokens onlyStatus: DOCUMENTED |
no | OpenAI-only knob. Ref: parameters.json |
| Sampling parameters (temperature / top_p / top_k) 4/4 |
temperature 0–2, top_p; rejected by reasoning modelsEndpoint: POST /v1/responsesParams: temperature, top_pStatus: DOCUMENTED · LIVE_VERIFIEDno top_k |
temperature 0–1, top_p, top_k — DEPRECATED: 400 when non-default on Claude 4.7+ / Fable / Mythos; Python SDK v1 removed the kwargsEndpoint: POST /v1/messagesParams: temperature, top_p, top_kStatus: DOCUMENTED · DEPRECATED · LIVE_VERIFIED |
temperature 0–2, top_p, plus xAI-specific top_k (≥1) and min_p (0–1); accepted on reasoning models (echo 0.7/0.95); Chat seed → system_fingerprintEndpoint: POST /v1/responsesParams: temperature, top_p, top_k, min_p, seedStatus: DOCUMENTED · LIVE_VERIFIEDpresence_penalty/frequency_penalty/stop → 400 on reasoning models |
generationConfig.temperature 0–2 (default 1.0), topP (0.95), topK (64), seed; DEPRECATED guidance since 2026-07-21: keep defaults on Gemini 3.x (temperature < 1 can cause looping); presencePenalty/frequencyPenalty → 400 'not enabled'Endpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.temperature, generationConfig.topP, generationConfig.topK, generationConfig.seedStatus: DOCUMENTED · DEPRECATED · LIVE_VERIFIEDcandidateCount > 1 → 400 on 3.x |
yes | Sampling knobs are shrinking on all frontier lines: rejected by OpenAI reasoning models and Claude 4.7+, deprecated-by-guidance on Gemini 3.x, still fully accepted on Grok. Ref: deprecations.json (anthropic, gemini api_features) · parameters.json |
| Stop sequences 4/4 |
Chat Completions stop (≤4); not a Responses parameterEndpoint: POST /v1/chat/completionsParams: stopStatus: DOCUMENTED |
stop_sequences[] → stop_reason: stop_sequence + stop_sequenceEndpoint: POST /v1/messagesParams: stop_sequencesStatus: DOCUMENTED · LIVE_VERIFIED |
Chat stop (≤4) and /v1/messages stop_sequences — 400 on reasoning models; not on ResponsesEndpoint: POST /v1/chat/completionsParams: stop, stop_sequencesStatus: DOCUMENTED · LIVE_VERIFIEDusable only on grok-4.20-0309-non-reasoning |
generationConfig.stopSequences[] (≤5; 6 → 400); Interactions generation_config.stop_sequencesEndpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.stopSequencesStatus: DOCUMENTED · LIVE_VERIFIED |
yes | Absent from both Responses APIs; Gemini allows 5, the others 4. Ref: parameters.json |
| Log probabilities 1/4 |
top_logprobs + include: ["message.output_text.logprobs"]Endpoint: POST /v1/responsesParams: top_logprobs, includeStatus: DOCUMENTED |
— not offered | logprobs/top_logprobs (0–8) accepted but silently ignored on grok-4.20 and newer (DEPRECATED); include: message.output_text.logprobs ignoredEndpoint: POST /v1/chat/completionsParams: logprobs, top_logprobsStatus: DOCUMENTED · DEPRECATED · LIVE_VERIFIED |
responseLogprobs / logprobs (0–20) in the schema but 400 Logprobs is not enabled for this model on every current modelEndpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.responseLogprobs, generationConfig.logprobsStatus: DOCUMENTED · FAILED_VERIFICATION |
no | Only OpenAI still returns logprobs; xAI and Gemini keep the parameters as dead compatibility fields. Ref: parameters.json |
| Refusal / safety signalling in-band 4/4 |
output[].content[] {type: refusal} (structured outputs) / status: incomplete + incomplete_details.reason: content_filter; HTTP 403 misalignment_policy_violation, cyber_policyEndpoint: POST /v1/responsesStatus: DOCUMENTED |
HTTP 200 with stop_reason: refusal + stop_details {category, explanation}; beta fallbacks re-runs on another modelEndpoint: POST /v1/messagesParams: fallbacks, fallback_credit_tokenStatus: DOCUMENTED · BETAcategories: cyber, bio, frontier_llm, reasoning_extraction, general_harms |
Chat message.refusal field; respect_moderation flag on image/video results; usage-guideline violations are billed (+ $0.05 fee when caught pre-generation on Responses); no error code catalogue for refusalsEndpoint: POST /v1/chat/completionsStatus: DOCUMENTED |
HTTP 200 with no candidates + promptFeedback.blockReason (SAFETY|OTHER|BLOCKLIST|PROHIBITED_CONTENT|IMAGE_SAFETY) for prompt blocks; candidates[].finishReason SAFETY|RECITATION|SPII|IMAGE_SAFETY|… (21 values) + safetyRatings[] for output blocks; thresholds via safetySettings[]Endpoint: POST /v1beta/models/{model}:generateContentParams: safetySettings, candidates[].finishReason, promptFeedback.blockReasonStatus: DOCUMENTED · LIVE_VERIFIEDInteractions mirror finishReason as snake_case error codes |
yes | All four signal in-band with different shapes; only Gemini lets the caller tune thresholds; only xAI charges for violating requests. Ref: docs/anthropic/stop-reasons.md · docs/openai/safety.md · docs/gemini/safety.md · docs/xai/pricing.md §5 |
| Server-side model fallback on refusal 1/4 |
— not offered | fallbacks parameter (+ fallback_credit_token billing credit)Endpoint: POST /v1/messagesParams: fallbacks, fallback_credit_tokenStatus: DOCUMENTED · BETAbeta server-side-fallback-2026-07-01, fallback-credit-2026-07-01 |
— not offered | — not offered | no | Anthropic-only. Ref: generated/fragments/headers/anthropic-beta-headers.json |
| Configurable safety thresholds 2/4 |
inline moderation {model, policy} on Responses/Chat (policy selection, not thresholds)Endpoint: POST /v1/responsesParams: moderationStatus: DOCUMENTED |
— not offered | — not offered | safetySettings[] {category: HARM_CATEGORY_HARASSMENT|HATE_SPEECH|SEXUALLY_EXPLICIT|DANGEROUS_CONTENT|JAILBREAK (CIVIC_INTEGRITY deprecated → enableEnhancedCivicAnswers), threshold: OFF|BLOCK_NONE|BLOCK_ONLY_HIGH|BLOCK_MEDIUM_AND_ABOVE|BLOCK_LOW_AND_ABOVE}; default Off on 2.5/3.x; safetyRatings[] returned when a threshold is setEndpoint: POST /v1beta/models/{model}:generateContentParams: safetySettingsStatus: DOCUMENTED · LIVE_VERIFIEDInteractions: custom safety settings not supported |
no | Only Gemini exposes per-category blocking thresholds; OpenAI selects a moderation policy. Ref: docs/gemini/safety.md · docs/openai/moderation.md |
Tools — client side
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Function / custom tools (you execute) 4/4 |
tools[type=function] {name, description, parameters, strict, output_schema, async, defer_loading, allowed_callers} → function_call item → you reply function_call_output {call_id, output}Endpoint: POST /v1/responsesParams: tools[type=function], input[](function_call_output)Status: DOCUMENTED · LIVE_VERIFIED38 compatible models listed |
tools[] {name, description, input_schema, strict, input_examples, cache_control, defer_loading, allowed_callers} → tool_use block → you reply user message of tool_result {tool_use_id, content, is_error}Endpoint: POST /v1/messagesParams: tools, messages[].content[] (tool_result)Status: DOCUMENTED · LIVE_VERIFIED13 compatible models |
Responses tools[type=function] {name, description, parameters} → function_call {call_id: call-…, name, arguments} → function_call_output; Chat tools[{type:function, function:{…}}] → message.tool_calls[] → {role: tool, tool_call_id}; ≤350 tools; parameters always strict; a client-side call ends the agentic request (fresh max_turns budget on the follow-up)Endpoint: POST /v1/responsesParams: tools[type=function], input[](function_call_output), parallel_tool_callsStatus: DOCUMENTED · LIVE_VERIFIEDall 7 Grok text models; also inside the voice session and /v1/messages |
tools[].functionDeclarations[] {name, description, parameters | parametersJsonSchema, response | responseJsonSchema, behavior} → parts[].functionCall {name, args, id} (+ mandatory thoughtSignature on Gemini 3) → you reply a user Content of functionResponse {name, id, response, parts[] (multimodal)}; max 512 declarationsEndpoint: POST /v1beta/models/{model}:generateContentParams: tools[].functionDeclarations, contents[].parts[].functionResponse, toolConfig.functionCallingConfigStatus: DOCUMENTED · LIVE_VERIFIEDInteractions: {type: function, name, description, parameters} + function_call/function_result steps; Live: toolCall/toolResponse messages |
yes | Same loop on all four. Arguments are a JSON string on OpenAI/xAI and an object on Anthropic (input) and Gemini (args); results are a top-level item (OpenAI/xAI), a leading tool_result block in a user turn (Anthropic) or functionResponse parts in a user Content (Gemini). Gemini 3 additionally requires echoing the thoughtSignature of the first call of each step (400 / MISSING_THOUGHT_SIGNATURE).Ref: docs/comparisons/tool-execution.md · docs/tools/xai/function-calling.md · docs/tools/gemini/function-calling.md |
| Free-form / grammar-constrained custom tools 1/4 |
tools[type=custom] {format: {type:text} | {type:grammar, syntax: lark|regex, definition}}Endpoint: POST /v1/responsesParams: tools[type=custom].formatStatus: DOCUMENTED · LIVE_VERIFIED |
— not offered | no custom tool type in the deserializer (422 'unknown variant'); no grammar modeEndpoint: POST /v1/responsesStatus: DOCUMENTED |
— not offered | no | OpenAI-only. Ref: tools.json · docs/xai/structured-outputs.md |
| Tool choice 4/4 |
tool_choice: auto|none|required string, {type:function,name}, hosted {type:web_search|mcp|shell|apply_patch|…}, {type:allowed_tools, mode, tools[]}Endpoint: POST /v1/responsesParams: tool_choice, tool_choice(allowed_tools)Status: DOCUMENTED · LIVE_VERIFIED |
tool_choice {type: auto|any|tool|none, name?, disable_parallel_tool_use?}Endpoint: POST /v1/messagesParams: tool_choice.type, tool_choice.disable_parallel_tool_useStatus: DOCUMENTED · LIVE_VERIFIEDany/tool → 400 on Fable 5.1 / Mythos 5.1 and with manual thinking |
tool_choice: auto|none|required or {type:function, name} (Responses) / {type:function, function:{name}} (Chat); /v1/messages: {type: auto|any|tool} (disable_parallel_tool_use → 400); no allowed_tools, no forcing of server toolsEndpoint: POST /v1/responsesParams: tool_choice, tool_choice.type, tool_choice.nameStatus: DOCUMENTED · LIVE_VERIFIED |
toolConfig.functionCallingConfig {mode: AUTO|ANY|NONE|VALIDATED, allowedFunctionNames[]} (ANY + names = forced subset); Interactions generation_config.tool_choice: auto|any|none|validated or {allowed_tools:{mode, tools[]}}; built-in tools cannot be forcedEndpoint: POST /v1beta/models/{model}:generateContentParams: toolConfig.functionCallingConfig.mode, toolConfig.functionCallingConfig.allowedFunctionNamesStatus: DOCUMENTED · LIVE_VERIFIEDLive setup.toolConfig → close 1007 |
yes | required ≈ any ≈ ANY; only OpenAI and Gemini (allowedFunctionNames / Interactions allowed_tools) can restrict to a subset; only OpenAI can force a hosted tool.Ref: parameters.json |
| Parallel tool calls 4/4 |
parallel_tool_calls (default true); hosted tools never batched with functionsEndpoint: POST /v1/responsesParams: parallel_tool_callsStatus: DOCUMENTED · LIVE_VERIFIED |
default on (Claude 4+); tool_choice.disable_parallel_tool_use:true to force ≤1; all results in one user messageEndpoint: POST /v1/messagesParams: tool_choice.disable_parallel_tool_useStatus: DOCUMENTED · LIVE_VERIFIED |
parallel_tool_calls (default true; false = at most one call) on Chat and ResponsesEndpoint: POST /v1/responsesParams: parallel_tool_callsStatus: DOCUMENTED · LIVE_VERIFIED/v1/messages disable_parallel_tool_use → 400 |
several functionCall parts in one model Content; answer all of them in one user Content (interleaving FC1,FR1,FC2 → 400); only the first call carries the thought signature; no on/off switchEndpoint: POST /v1beta/models/{model}:generateContentParams: contents[].parts[].functionCallStatus: DOCUMENTED · LIVE_VERIFIEDcompositional_function_calling chains calls across turns |
yes | Default-on everywhere; Gemini has no switch to disable it. Ref: docs/openai/tool-loop.md · docs/tools/anthropic/tool-use-loop.md · docs/tools/gemini/function-calling.md |
| Tool namespaces 1/4 |
tools[type=namespace] {name, description, tools[]}; calls carry namespaceEndpoint: POST /v1/responsesParams: tools[type=namespace]Status: DOCUMENTED · LIVE_VERIFIED |
no namespaces; toolsets (computer_toolset_20260801, mcp_toolset) group Anthropic-defined members onlyStatus: DOCUMENTED |
— not offered | — not offered | no | OpenAI-only. Ref: tools.json |
| Deferred tool loading / tool search 3/4 |
tools[type=tool_search] {execution: server|client} + defer_loading:true on function/custom/mcp; tool_search_call/tool_search_output items; additional_tools input itemEndpoint: POST /v1/responsesParams: tools[type=tool_search], tools[type=function].defer_loadingStatus: DOCUMENTED · LIVE_VERIFIEDgpt-5.4+ only; gpt-5.4-nano lacks it |
tool_search_tool_regex_20251119 / tool_search_tool_bm25_20251119 (server) + defer_loading:true (≤10,000 tools); tool_reference blocks expanded server-side; client-side search via tool_reference in tool_resultEndpoint: POST /v1/messagesParams: tools[].defer_loading, tools[].typeStatus: DOCUMENTED · LIVE_VERIFIEDGA no header; not cacheable together with cache_control |
tools[type=tool_search] {execution} + defer_loading:true on function/mcp tools → tool_search_call {arguments:{query, limit}} / tool_search_output {tools[]} — alpha: 403 'only available for alpha users'Endpoint: POST /v1/responsesParams: tools[type=tool_search], tools[].defer_loadingStatus: DOCUMENTED · ACCOUNT_RESTRICTEDlisted as compatible with all 7 text models |
— not offered | yes | OpenAI-shaped on xAI (gated), two algorithms on Anthropic; Gemini has no deferred loading (best practice: keep 10–20 active declarations). Ref: docs/tools/openai/tool-search.md · docs/tools/anthropic/tool-search.md · tools.json (xai tool_search) |
| Programmatic tool calling (model writes code that calls tools) 2/4 |
tools[type=programmatic_tool_calling] + allowed_callers:["programmatic"]; program/program_output items; nested calls caller.type: program; JavaScript in isolated V8Endpoint: POST /v1/responsesParams: tools[type=programmatic_tool_calling], tools[type=function].allowed_callersStatus: DOCUMENTED · UNVERIFIEDno compatible model list in tools.json |
code_execution_20260120+ + allowed_callers:["code_execution_20260120"]; Python await tool({...}) in the sandbox; caller {type, tool_id}; reply requires top-level containerEndpoint: POST /v1/messagesParams: tools[].allowed_callers, containerStatus: DOCUMENTED · LIVE_VERIFIEDnot Haiku 4.5 (400) |
— not offered | not a generateContent feature; nearest: compositional function calling (chained calls across turns) and code execution over functionResponse data; Antigravity agents script tools inside their sandboxEndpoint: POST /v1beta/models/{model}:generateContentStatus: DOCUMENTED |
yes | OpenAI and Anthropic only (JS vs Python). Ref: docs/tools/anthropic/programmatic-tool-calling.md · docs/openai/tool-loop.md §6 · docs/tools/gemini/function-calling.md |
| Async tools (model continues while a tool runs) 2/4 |
tools[type=function].async:true + wait tool / task handles (GPT-6 Astra+)Endpoint: POST /v1/responsesParams: tools[type=function].async, input[](function_call).asyncStatus: DOCUMENTED |
— not offered | — not offered | Live API only: functionDeclarations[].behavior: NON_BLOCKING (default on gemini-3.8-live) + toolResponse.functionResponses[].scheduling: INTERRUPT|WHEN_IDLE|SILENT, willContinue; generateContent → 400Endpoint: WSS BidiGenerateContentParams: setup.tools[].functionDeclarations[].behavior, toolResponse.functionResponses[].schedulingStatus: DOCUMENTED · LIVE_VERIFIEDgemini-3.8-live-extended-thinking: async only |
yes | Different scopes: OpenAI in text Responses, Gemini in the voice Live session. Ref: parameters.json · docs/gemini/live-api.md |
| Shell execution 3/4 |
tools[type=shell] hosted (environment.type: container_auto|container_reference) or local (type: local — you run commands); legacy local_shell → 400Endpoint: POST /v1/responsesParams: tools[type=shell].environmentStatus: DOCUMENTED · LIVE_VERIFIEDgpt-5.2+, codex, gpt-6-astra |
bash_20250124 client tool (you run a persistent bash) or server bash_code_execution sub-tool of code_execution_20250825+Endpoint: POST /v1/messagesParams: tools[].typeStatus: DOCUMENTED · LIVE_VERIFIEDall current models |
tools[type=shell] {environment:{type: local, skills:[{name, description, path}]}} → shell_call → you reply shell_call_output {stdout, stderr, outcome}; local only (no hosted container variant); hosted Python via code_interpreterEndpoint: POST /v1/responsesParams: tools[type=shell].environment, input[](shell_call_output)Status: DOCUMENTED · LIVE_VERIFIEDaccepted live, not invoked; Grok Build CLI runs its own sandbox |
no shell tool on generateContent; the Antigravity managed agent runs bash/python/node code_execution and file tools inside its Linux sandbox (Interactions API only, PREVIEW)Endpoint: POST /v1beta/interactionsParams: agent_config(antigravity)Status: DOCUMENTED · PREVIEWgemini-3.1-pro-preview-customtools is tuned for bash-style custom tools you define yourself |
yes | OpenAI has hosted + local, Anthropic hosted (code execution) + local (bash), xAI local only, Gemini only inside its managed agent. Ref: docs/tools/openai/shell.md · docs/tools/anthropic/bash.md · docs/xai/skills-api.md · docs/gemini/interactions-api.md |
| File editing tool 2/4 |
tools[type=apply_patch] → apply_patch_call {operation: create_file|update_file|delete_file, diff}; you reply apply_patch_call_outputEndpoint: POST /v1/responsesParams: tools[type=apply_patch]Status: DOCUMENTED · LIVE_VERIFIEDundocumented SSE events response.apply_patch_call_operation_diff.* (LIVE_DISCOVERED) |
text_editor_20250728 (name: str_replace_based_edit_tool, max_characters) → commands view/str_replace/create/insert; server variant text_editor_code_executionEndpoint: POST /v1/messagesParams: tools[].max_charactersStatus: DOCUMENTED · LIVE_VERIFIEDolder 20250429/20250124 → 400 on current models |
— not offered | no editor tool on generateContent; Antigravity 09-2026 built-ins write_to_file, replace_file_content, view_file, list_dir, find_by_name, grep_search (agent sandbox only)Endpoint: POST /v1beta/interactionsStatus: DOCUMENTED · PREVIEW05-2026 tool names deprecated → 2026-10-05 |
yes | Diff-based (OpenAI) vs command-based (Anthropic); Gemini only inside Antigravity; none on xAI. Ref: docs/tools/openai/apply-patch.md · docs/tools/anthropic/text-editor.md · docs/gemini/interactions-api.md |
| Memory tool (client-side persistent memory) 1/4 |
none in Responses; Agents API sessions persist items but no memory tool Status: DOCUMENTED |
memory_20250818 client tool (you store files under /memories)Endpoint: POST /v1/messagesParams: tools[].typeStatus: DOCUMENTED · LIVE_VERIFIEDManaged Agents add server-side memory stores |
— not offered | — not offered | no | Anthropic-only. Ref: docs/tools/anthropic/memory.md |
| Computer use 3/4 |
tools[type=computer] (current, gpt-5.4+/gpt-6) or computer_use_preview (+ computer-use-preview model, RETIRED); computer_call {action|actions[], pending_safety_checks} → computer_call_output {computer_screenshot, acknowledged_safety_checks}Endpoint: POST /v1/responsesParams: tools[type=computer], tools[type=computer_use_preview]Status: DOCUMENTED · PREVIEW · ACCOUNT_RESTRICTEDcomputer UNVERIFIED live |
computer_toolset_20260801 (GA, no header, Fable/Mythos/Opus 5/Sonnet 5/Opus 4.8) or beta computer_20251124 (computer-use-2025-11-24) / computer_20250124; member tools screenshot/zoom/click…; batch actionsEndpoint: POST /v1/messagesParams: tools[].display_width_px, tools[].enable_zoom, tools[].configsStatus: DOCUMENTED · LIVE_VERIFIED · BETA~4,500-token toolset definition |
— not offered | tools[].computerUse {environment: ENVIRONMENT_BROWSER|MOBILE|DESKTOP, excludedPredefinedFunctions[], enablePromptInjectionDetection, disabledSafetyPolicies[]} → predefined functionCalls (click, type, scroll, navigate, open_app…; coordinates 0–999) → you reply functionResponse with a screenshot; safety_decision: require_confirmation → safety_acknowledgement; Interactions {type: computer_use, environment: browser}Endpoint: POST /v1beta/models/{model}:generateContentParams: tools[].computerUse, tools[].computerUse.environmentStatus: DOCUMENTED · PREVIEW · ACCOUNT_RESTRICTEDgemini-3.8-flash recommended; legacy gemini-2.5-computer-use-preview-10-2025 browser-only; no free tier |
yes | Three client-executed screenshot loops (OpenAI, Anthropic, Gemini); none on xAI. Gemini normalises coordinates to 0–999 and adds mobile/desktop environments. Ref: docs/tools/openai/computer-use.md · docs/tools/anthropic/computer-use.md · docs/tools/gemini/computer-use.md |
| Browser use toolset 1/4 |
— not offered | browser_toolset_20260801 (client toolset, browser_state result blocks)Endpoint: POST /v1/messagesParams: tools[].typeStatus: DOCUMENTED · LIVE_VERIFIED≈6,600-token definition; 7 models |
— not offered | browser is an environment of the computer-use tool, not a separate toolsetEndpoint: POST /v1beta/models/{model}:generateContentStatus: DOCUMENTED |
no | Anthropic-only as a dedicated toolset. Ref: tools.json |
Tools — server side / hosted
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Web search 4/4 |
tools[type=web_search] {search_context_size, user_location, filters.allowed_domains, external_web_access, return_token_budget, search_content_types}; web_search_call item + url_citation annotations; $10/1k callsEndpoint: POST /v1/responsesParams: tools[type=web_search]Status: DOCUMENTED · LIVE_VERIFIEDChat: search models + web_search_options; preview variants LEGACY |
web_search_20260318|20260209|20250305 {max_uses, allowed_domains XOR blocked_domains, user_location}; server_tool_use + web_search_tool_result blocks, web_search_result_location citations; $10/1k searches; dynamic filtering via code execution (20260209+)Endpoint: POST /v1/messagesParams: tools[].max_uses, tools[].allowed_domains, tools[].blocked_domains, tools[].user_locationStatus: DOCUMENTED · LIVE_VERIFIED13 models; pause_turn after 10 iterations |
tools[type=web_search] {allowed_domains ≤5 XOR excluded_domains ≤5, enable_image_understanding, enable_image_search} → web_search_call {action: search|open_page|find_in_page} + url_citation annotations + inline [[N]](url) (disable via include: no_inline_citations); $5/1k successful calls; search_context_size → 400Endpoint: POST /v1/responsesParams: tools[type=web_search], tools[].allowed_domains, tools[].excluded_domains, tools[].enable_image_searchStatus: DOCUMENTED · LIVE_VERIFIEDResponses only; Chat Live Search → 410; usage server_side_tool_usage_details.web_search_calls |
tools[{googleSearch: {timeRangeFilter?, searchTypes?{webSearch, imageSearch}}}] → groundingMetadata {webSearchQueries, searchEntryPoint (must be displayed — ToS), groundingChunks[].web, groundingSupports}; Gemini 3.x 5,000 free queries/month then $14/1k queries, 2.5: 1,500 RPD free then $35/1k grounded prompts; legacy googleSearchRetrieval (dynamic retrieval) DEPRECATEDEndpoint: POST /v1beta/models/{model}:generateContentParams: tools[].googleSearch, tools[].googleSearch.searchTypes, tools[].googleSearchRetrievalStatus: DOCUMENTED · ACCOUNT_RESTRICTED429 limit: 0 on this free-tier key; Interactions {type: google_search}; the only tool allowed with functions in the Live API |
yes | Same task on all four with four price models ($10 / $10 / $5 per 1k calls; Gemini per query with a free monthly quota). Domain filters exist on OpenAI, Anthropic and xAI; Gemini offers time-range and image-search filters instead. Ref: docs/tools/openai/web-search.md · docs/tools/anthropic/web-search.md · docs/tools/xai/web-search.md · docs/tools/gemini/google-search-grounding.md |
| Social / vertical search (X posts, Google Maps) 2/4 |
— not offered | — not offered | tools[type=x_search] {allowed_x_handles ≤10 XOR excluded_x_handles, from_date, to_date, enable_image_understanding, enable_video_understanding} → live item custom_tool_call named x_keyword_search/x_semantic_search (docs: x_search_call) + url_citation to x.com posts; $5/1k calls until 2026-09-21, then $5/1k posts + $10/1k profiles fetchedEndpoint: POST /v1/responsesParams: tools[type=x_search], tools[].allowed_x_handles, tools[].from_date, tools[].to_dateStatus: DOCUMENTED · LIVE_VERIFIED · LIVE_DISCOVEREDalso in the voice session |
tools[{googleMaps: {enableWidget?}}] + toolConfig.retrievalConfig {latLng, languageCode} → groundingChunks[].maps {uri, title, placeId, text, placeAnswerSources}; GA; text only; not in Live; Gemini 3.x $14/1k queries after 5,000/month (tools table: $25/1k grounded prompts, 1,500 RPD free)Endpoint: POST /v1beta/models/{model}:generateContentParams: tools[].googleMaps, toolConfig.retrievalConfig.latLngStatus: DOCUMENTED · GA · LIVE_VERIFIEDpricing inconsistency flagged ( DOCUMENTATION_INCOMPLETE) |
no | Provider-specific data sources; no equivalent on OpenAI/Anthropic. Ref: docs/tools/xai/x-search.md · docs/tools/gemini/google-maps-grounding.md |
| Web fetch / URL context (retrieve specific URLs) 2/4 |
no dedicated fetch tool; web search may open pages; input_file.file_url fetches PDFs onlyStatus: DOCUMENTED |
web_fetch_20260318|20260309|20260209|20250910 {max_uses, allowed_domains, citations, max_content_tokens, use_cache, url_sources}; only URLs already in context; free (tokens only)Endpoint: POST /v1/messagesParams: tools[].max_content_tokens, tools[].url_sources, tools[].use_cacheStatus: DOCUMENTED · LIVE_VERIFIEDurl_not_in_prior_context error |
no fetch tool; web_search_call.action.type: open_page|find_in_page shows the search tool opening pages; input_file.file_url attaches a remote file (attachment search)Endpoint: POST /v1/responsesStatus: DOCUMENTED |
tools[{urlContext: {}}] (no options): ≤20 public URLs per request, ≤34 MB each (HTML/JSON/text/CSV/RTF/PNG/JPEG/PDF; no YouTube/Workspace/paywalls) → urlContextMetadata.urlMetadata[] {retrievedUrl, urlRetrievalStatus}; free tool, content billed as input (toolUsePromptTokenCount)Endpoint: POST /v1beta/models/{model}:generateContentParams: tools[].urlContextStatus: DOCUMENTED · GA · LIVE_VERIFIEDInteractions {type: url_context}; combinable with search/code exec/functions |
yes | Anthropic and Gemini offer a fetch tool; Anthropic restricts to URLs already in context, Gemini to any public URL you name. Ref: docs/tools/anthropic/web-fetch.md · docs/tools/gemini/url-context.md |
| File search / managed RAG 3/4 |
Vector Stores API (create, files, file_batches, search) + tools[type=file_search] {vector_store_ids, max_num_results, filters, ranking_options}; $2.50/1k calls + $0.10/GB/dayEndpoint: POST /v1/vector_stores · POST /v1/responsesParams: tools[type=file_search]Status: DOCUMENTED · LIVE_VERIFIED16 endpoints |
no vector store; patterns: search_result blocks (your RAG, citable), document blocks from Files API, code execution over uploaded filesEndpoint: POST /v1/messagesParams: messages[].content[] {type:'search_result'}, messages[].content[] {type:'document'}Status: DOCUMENTED · LIVE_VERIFIED |
Collections API (/v1/collections, documents added from Files ids, index_configuration.model_name: grok-embedding-small, chunk_configuration, POST /v1/documents/search {query, retrieval_mode: hybrid|semantic|keyword}) + tools[type=file_search|collections_search] {vector_store_ids (= collection ids), max_num_results, filters, ranking_options} → file_search_call {queries, results[{file_id, filename, score, text}]}; $2.50/1k calls + $0.10/GiB/day; implicit attachment_search over input_file parts $10/1kEndpoint: POST /v1/responsesParams: tools[type=file_search], tools[].vector_store_idsStatus: DOCUMENTED · LIVE_VERIFIED · ACCOUNT_RESTRICTED13 collection endpoints; documented on management-api.x.ai but working on api.x.ai (LIVE_DISCOVERED); search 404 while indexing |
File Search stores (/v1beta/fileSearchStores, :uploadToFileSearchStore resumable ≤100 MB, :importFile, documents, chunkingConfig, customMetadata[]; embedding model fixed at creation) + tools[{fileSearch: {fileSearchStoreNames[], metadataFilter (AIP-160), topK}}] → groundingChunks[].retrievedContext {title, text, pageNumber, customMetadata}; indexing $0.15/1M tokens once, storage + queries free; store quota 1 GB (free) … 1 TB (Tier 3)Endpoint: POST /v1beta/models/{model}:generateContentParams: tools[].fileSearch, tools[].fileSearch.fileSearchStoreNames, tools[].fileSearch.metadataFilterStatus: DOCUMENTED · PREVIEW · LIVE_VERIFIED12 endpoints; not combinable with Search/URL context; not in Live |
yes | Three hosted stores (OpenAI vector stores, xAI collections, Gemini file-search stores) with different billing (per call / per call + storage / per indexed token); Anthropic supplies citation plumbing instead of storage. Ref: docs/openai/vector-stores.md · docs/anthropic/citations.md · docs/xai/collections.md · docs/tools/gemini/file-search.md |
| Code execution (sandboxed Python) 4/4 |
tools[type=code_interpreter] {container: auto|cntr_id, file_ids, memory_limit 1g–64g, network_policy}; code_interpreter_call {code, outputs}; $0.03–$1.92 per 20-min session (per-minute since 2026-06-02)Endpoint: POST /v1/responsesParams: tools[type=code_interpreter].containerStatus: DOCUMENTED · LIVE_VERIFIED36 models |
code_execution_20260521|20260120|20250825 (bash + text editor sub-tools, Python 3.11, 5 GiB RAM, no internet); container {id, skills} reuse (30-day state); 1,550 free container-hours/org/month then $0.05/h; free with web_search/fetch 20260209+Endpoint: POST /v1/messagesParams: tools[].type, containerStatus: DOCUMENTED · LIVE_VERIFIED13 models; usage counter code_execution_requests missing live |
tools[type=code_interpreter] (alias code_execution) → code_interpreter_call {code, outputs[{type: logs, logs: <JSON string stdout/stderr/exit_code>} | {type: image, url}]} (outputs only with include: code_interpreter_call.outputs); Python + NumPy/Pandas/Matplotlib/SciPy, no network; $5/1k calls + tokens; no container objectEndpoint: POST /v1/responsesParams: tools[type=code_interpreter], includeStatus: DOCUMENTED · LIVE_VERIFIEDall 7 text models; gRPC rejects the code_interpreter alias |
tools[{codeExecution: {}}] → parts executableCode {language: PYTHON, code} + codeExecutionResult {outcome, output} (+ inlineData PNG for matplotlib); Python ≥3.10, fixed library set, 30 s per run, ≤5 retries, no pip/network; no fee — code and results billed as output then input tokensEndpoint: POST /v1beta/models/{model}:generateContentParams: tools[].codeExecutionStatus: DOCUMENTED · LIVE_VERIFIED12 models; Interactions {type: code_execution}; not in Live |
yes | Four sandboxes, four price models: per container-session (OpenAI), per container-hour with a free tier (Anthropic), per call (xAI), tokens only (Gemini). Only OpenAI/Anthropic expose a reusable container object. Ref: docs/tools/openai/code-interpreter.md · docs/tools/anthropic/code-execution.md · docs/tools/xai/code-execution.md · docs/tools/gemini/code-execution.md |
| Container management API 2/4 |
CRUD containers and container files, download outputs Endpoint: GET/POST/DELETE /v1/containers[/{id}/files]Status: DOCUMENTED · LIVE_VERIFIED9 endpoints; auto containers expire 20 min after last activity |
no container endpoints; container id returned on the message (container.id, expires_at) and reused via the container param; outputs via Files APIEndpoint: POST /v1/messages · GET /v1/files/{id}/contentParams: containerStatus: DOCUMENTED · LIVE_VERIFIED |
no container object; container param accepted for OpenAI compatibility, not neededEndpoint: POST /v1/responsesStatus: DOCUMENTED |
Environments API for Interactions agents (POST/GET/DELETE /v1beta/environments, files GET …/files/{path}, PUT /upload/v1beta/environments/{env}/files/{path}; sources repository/GCS/inline; network allowlist; idle 15 min, deleted after 7 days) — not usable by codeExecutionEndpoint: POST /v1beta/environmentsParams: environment, sources, networkStatus: DOCUMENTED · BETA · PREVIEW · LIVE_VERIFIED7 endpoints; sandbox compute unbilled during preview |
yes | OpenAI containers and Gemini environments are both first-class resources but serve different tools (code interpreter vs managed agents). Ref: docs/openai/containers.md · docs/tools/anthropic/code-execution.md · docs/gemini/interactions-api.md |
| Image generation as a tool / modality inside a text call 3/4 |
tools[type=image_generation] {model, quality, size, background, input_fidelity, partial_images, moderation}; image_generation_call item; streaming partialsEndpoint: POST /v1/responsesParams: tools[type=image_generation]Status: DOCUMENTED · LIVE_VERIFIED31 models |
— not offered | tools[type=image_generation] → image_generation_call billed at Imagine per-image rates (no call fee); accepted live, not invokedEndpoint: POST /v1/responsesParams: tools[type=image_generation]Status: DOCUMENTED · LIVE_VERIFIEDall 7 text models listed |
not a tool: generationConfig.responseModalities: [TEXT, IMAGE] + imageConfig {aspectRatio, imageSize 512|1K|2K|4K} on image models (gemini-3.1-flash-image, -lite-image, gemini-3-pro-image); Interactions response_format {type: image}Endpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.responseModalities, generationConfig.imageConfigStatus: DOCUMENTED · ACCOUNT_RESTRICTEDno free tier ( limit: 0 here); text models ignore IMAGE modality |
yes | Tool item on OpenAI/xAI, output modality on Gemini; Claude outputs text only. Ref: tools.json · docs/gemini/image-generation.md · docs/xai/images.md |
| Remote MCP servers 4/4 |
tools[type=mcp] {server_label, server_url|connector_id|tunnel_id, authorization, headers, allowed_tools, require_approval (default always), defer_loading, allowed_callers}; mcp_list_tools, mcp_call, mcp_approval_request/response items; no beta headerEndpoint: POST /v1/responsesParams: tools[type=mcp]Status: DOCUMENTED · LIVE_VERIFIED42 models; also Realtime and Agents API |
mcp_servers[] {type:url, url, name, authorization_token} (≤20) + tools[type=mcp_toolset] {mcp_server_name, default_config, configs}; mcp_tool_use/mcp_tool_result blocks; beta header mcp-client-2025-11-20Endpoint: POST /v1/messagesParams: mcp_servers, tools[].mcp_server_name, tools[].default_config, tools[].configsStatus: DOCUMENTED · BETA · LIVE_VERIFIEDno approvals in Messages; the error text advertises mcp-client-2026-09-15 which is rejected (inconsistency) |
tools[type=mcp] {server_url, server_label (required), server_description, allowed_tools[], authorization, headers} → mcp_call {server_label, name, arguments, output, error}; tokens only; no approval round-trip (require_approval silently accepted); also inside the voice sessionEndpoint: POST /v1/responsesParams: tools[type=mcp], tools[].server_url, tools[].server_label, tools[].allowed_toolsStatus: DOCUMENTED · LIVE_VERIFIEDLIVE_VERIFIED with DeepWiki; connector_id unsupported |
server-side tools[].mcpServers[] {name, streamableHttpTransport {url, headers, timeout}} in the discovery schema (UNVERIFIED, no guide) and Interactions {type: mcp_server, name, url, headers, allowed_tools} (documented; not on Gemini 3 per the overview); SDK-side MCP (mcpToTool(), Python ClientSession in tools) runs the calls in your process (BETA)Endpoint: POST /v1beta/interactionsParams: tools[].mcpServers, tools[](mcp_server)Status: DOCUMENTED · DOCUMENTATION_INCOMPLETE · UNVERIFIEDStreamable HTTP only; server names must not contain '-' |
yes | OpenAI GA with approvals; xAI GA without approvals; Anthropic beta without approvals; Gemini server-side MCP is documented but unverified — its verified path is SDK-side execution. Ref: docs/tools/openai/mcp-and-connectors.md · docs/tools/anthropic/mcp-connector.md · docs/tools/xai/mcp.md · docs/tools/gemini/mcp.md |
| Built-in SaaS connectors 1/4 |
connector_id: dropbox, gmail, googlecalendar, googledrive, microsoftteams, outlookcalendar, outlookemail, sharepoint (+ OAuth authorization) — deprecated for models after 2025-09-01Endpoint: POST /v1/responsesParams: tools[type=mcp].connector_idStatus: DOCUMENTED · DEPRECATED |
— not offered | connector_id documented as unsupportedEndpoint: POST /v1/responsesStatus: DOCUMENTED |
— not offered | no | OpenAI-only (deprecated). Ref: docs/tools/openai/mcp-and-connectors.md |
| Private-network MCP (tunnels) / agent credentials 3/4 |
Secure MCP Tunnel: outbound tunnel-client, tools[type=mcp].tunnel_id (tunnel_[a-z0-9]{32})Endpoint: POST /v1/responsesParams: tools[type=mcp].tunnel_idStatus: DOCUMENTED |
MCP tunnels API (/v1/tunnels, certificates, tokens) + tunnel agent; header mcp-tunnels-2026-06-22; WIF bearer with workspace:manage_tunnelsEndpoint: GET/POST /v1/tunnelsStatus: DOCUMENTED · BETA · PREVIEW · ACCOUNT_RESTRICTED10 endpoints; older /v1/organizations/tunnels deprecated |
— not offered | no tunnels; Credentials API (/v1beta/credentials: bearer_token, oauth2, environment_variable with injection_location and trusted_domains; secrets write-only) referenced from environment network allowlistsEndpoint: POST /v1beta/credentialsStatus: DOCUMENTED · BETA · PREVIEW5 endpoints, not tested |
no | Tunnels on OpenAI/Anthropic; Gemini solves the secret-injection half with credentials; nothing on xAI. Ref: endpoints.json (managed-agents tunnels, gemini credentials) · docs/tools/openai/mcp-and-connectors.md |
| Skills (packaged instructions + files) 3/4 |
Skills API (POST /v1/skills, versions, content zip) referenced from tools[type=shell].environment.skills[], containers, Agents sessions; ≤500 files, ≤25 MBEndpoint: POST /v1/skillsParams: tools[type=shell].environment.skillsStatus: DOCUMENTED · LIVE_VERIFIED · FAILED_VERIFICATIONversion-by-number endpoints returned 404 live; lists returned empty |
Skills API (/v1/skills, versions, content; GA no header) + Anthropic skills (pptx, xlsx, docx, pdf); used via container.skills[] with code execution and on Managed AgentsEndpoint: POST /v1/skillsParams: container.skillsStatus: DOCUMENTED · LIVE_VERIFIED9 endpoints all LIVE_VERIFIED |
Skills API GET/POST /v1/skills, GET/DELETE /v1/skills/{id}, GET …/content (zip; SKILL.md frontmatter name, description, when-to-use, paths, allowed-tools) — all 404 for this team (ACCOUNT_RESTRICTED); reach inference only via shell.environment.skills[]; Grok Build CLI also loads skills locallyEndpoint: POST /v1/skillsParams: tools[type=shell].environment.skillsStatus: DOCUMENTED · ACCOUNT_RESTRICTED5 endpoints, OpenAPI only |
no Skills API; custom Interactions agents mount .agents/skills/<name>/SKILL.md from their environment sourcesEndpoint: POST /v1beta/interactionsStatus: DOCUMENTED · PREVIEW |
yes | Same SKILL.md bundle model on OpenAI, Anthropic and xAI (xAI gated); Gemini uses environment files.Ref: docs/openai/skills-api.md · docs/anthropic/skills-api.md · docs/xai/skills-api.md · docs/gemini/interactions-api.md |
| Advisor (consult a stronger model mid-response) 1/4 |
— not offered | advisor_20260301 {model, max_tokens, caching} server tool, beta advisor-tool-2026-03-01Endpoint: POST /v1/messagesParams: tools[].model, tools[].max_tokens, tools[].cachingStatus: DOCUMENTED · BETA · LIVE_VERIFIED11 models |
— not offered | — not offered | no | Anthropic-only. Ref: tools.json |
| Citations / grounding metadata 4/4 |
output_text.annotations[]: url_citation, file_citation, container_file_citation, file_path (produced by web/file search, code interpreter)Endpoint: POST /v1/responsesParams: includeStatus: DOCUMENTED · LIVE_VERIFIEDno citations for plain documents |
citations {enabled:true} on document/search_result blocks → char_location, page_location, content_block_location, search_result_location, web_search_result_location; citations_delta when streamingEndpoint: POST /v1/messagesParams: messages[].content[].citations, tools[].citationsStatus: DOCUMENTED · LIVE_VERIFIEDall-or-nothing across documents |
output_text.annotations[] {type: url_citation, url, start_index, end_index, title} (live indices 0/0, title = URL) + inline markdown [[N]](url); collections citations collections://<cid>/files/<fid>; include: web_search_call.action.sourcesEndpoint: POST /v1/responsesParams: includeStatus: DOCUMENTED · LIVE_VERIFIEDusage.num_sources_used |
candidates[].groundingMetadata {groundingChunks[] (web|maps|retrievedContext), groundingSupports[] {segment, groundingChunkIndices, confidenceScores}, webSearchQueries, searchEntryPoint}, urlContextMetadata, citationMetadata (recitation); Interactions text_annotation_deltaEndpoint: POST /v1beta/models/{model}:generateContentParams: candidates[].groundingMetadata, candidates[].urlContextMetadataStatus: DOCUMENTED · LIVE_VERIFIEDstreaming chunks carry only new grounding chunks — accumulate |
yes | Anthropic cites any document you pass; OpenAI, xAI and Gemini cite only their own tool results (Gemini with segment-level supports and confidence scores). Ref: docs/anthropic/citations.md · docs/tools/gemini/google-search-grounding.md |
Multimodal input & media APIs
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Image input 4/4 |
input_image {image_url|file_id, detail: low|high|auto|original}; patch/tile token formulas; ≤1,500 images, ≤512 MBEndpoint: POST /v1/responsesParams: input[](message).content[](input_image)Status: DOCUMENTED · LIVE_VERIFIED |
image {source: base64|url|file, transformations}; ≤100 (200k models) / 600 (1M models) images; 10 MB each; tokens = ⌈w/28⌉×⌈h/28⌉ (2576 px hi-res on 4.7+)Endpoint: POST /v1/messagesParams: messages[].content[] {type:'image'}Status: DOCUMENTED · LIVE_VERIFIED |
Responses input_image {image_url (https or data URL), detail} (≥512 px; file_id variant → 400); Chat image_url {url, detail}; image tokens priced as text input (image_input = input price); all Grok 4 text models accept imagesEndpoint: POST /v1/responsesParams: input[](message).content[](input_image), messages[].content[](image_url)Status: DOCUMENTED · LIVE_VERIFIEDusage.prompt_tokens_details.image_tokens |
parts[].inlineData {mimeType, data} or fileData {fileUri} (Files API); mediaResolution LOW/MEDIUM/HIGH globally or per part (Gemini 3 adds ULTRA_HIGH, 2,240 tokens); docs 280/560/1,120/2,240 tokens by level (observed 256/529/1,089/2,209)Endpoint: POST /v1beta/models/{model}:generateContentParams: contents[].parts[].inlineData, contents[].parts[].fileData, generationConfig.mediaResolutionStatus: DOCUMENTED · LIVE_VERIFIEDpromptTokensDetails[] {modality: IMAGE} |
yes | Universal. Token accounting differs on every provider (patches/tiles, 28-px grid, text-price tokens, resolution levels). Ref: docs/openai/multimodal-input.md · docs/anthropic/vision-and-documents.md · docs/xai/chat-completions.md · docs/gemini/multimodal-input.md |
| PDF / document input 4/4 |
input_file {file_id|file_url|file_data+filename, detail}; text + page images in context; <50 MB combined; office formats via file_idEndpoint: POST /v1/responsesParams: input[](message).content[](input_file)Status: DOCUMENTED |
document {source: base64 pdf|url|file|text|content, title, context, citations}; ≤600 pages/request, 32 MB bodyEndpoint: POST /v1/messagesParams: messages[].content[] {type:'document'}Status: DOCUMENTED · LIVE_VERIFIEDcitable |
Responses input_file {file_id|file_url|file_data} → routed through the implicit attachment search tool ($10/1k calls); Chat file parts → 400; /v1/messages document → 422Endpoint: POST /v1/responsesParams: input[](message).content[](input_file)Status: DOCUMENTED · LIVE_VERIFIEDFiles API 50 MB (spec) / 512 MB (guide) |
inlineData {mimeType: application/pdf} or fileData (Files API, 2 GB); PDF pages billed at the image token rate (DOCUMENT modality, 560 tokens/page in countTokens vs IMAGE 520 observed in generateContent); pdf_input true on all 3.x text modelsEndpoint: POST /v1beta/models/{model}:generateContentParams: contents[].parts[].inlineData, contents[].parts[].fileDataStatus: DOCUMENTED · LIVE_VERIFIEDno citations for documents; URL context handles remote PDFs |
yes | Universal; only Anthropic makes documents citable, only xAI bills a per-call fee for attachments. Ref: docs/anthropic/vision-and-documents.md · docs/xai/files.md · docs/gemini/multimodal-input.md |
| Audio / video input (understanding) 2/4 |
Chat Completions input_audio {data, format} on gpt-audio-1.5, gpt-4o-audio-preview; Responses input_audio part in schema but UNVERIFIED; no video inputEndpoint: POST /v1/chat/completionsParams: messages[](user).content[](input_audio)Status: DOCUMENTED · LIVE_DISCOVERED |
— not offered | no audio/video parts on text models (audio only in the voice session and STT; grok-imagine-video accepts video/audio as generation inputs; view_x_video sub-tool understands X videos)Status: DOCUMENTED |
native: inlineData/fileData audio (WAV/MP3/AIFF/AAC/OGG/FLAC; ≈25–32 tokens/s) and video (≤1 fps frames + audio; videoMetadata {startOffset, endOffset, fps}; YouTube URLs via fileData) on every 3.x text model; gemini-embedding-2 embeds audio/video tooEndpoint: POST /v1beta/models/{model}:generateContentParams: contents[].parts[].inlineData, contents[].parts[].videoMetadataStatus: DOCUMENTED · LIVE_VERIFIEDaudio_input price rows on 2.5/3.1-lite/3-flash; 3.5+ single price |
yes | Gemini is the only provider with native audio and video understanding in the text API; OpenAI has audio-in on dedicated chat models. Ref: docs/openai/multimodal-input.md §4 · docs/gemini/multimodal-input.md |
| Text-to-speech 3/4 |
POST /v1/audio/speech (gpt-4o-mini-tts, tts-1, tts-1-hd; SSE speech.audio.delta) and Chat modalities:[text,audio]Endpoint: POST /v1/audio/speechParams: model, input, voice, instructionsStatus: DOCUMENTED · LIVE_VERIFIEDcustom voices ACCOUNT_RESTRICTED |
— not offered | POST /v1/tts {text ≤60,000 chars, voice_id, language (required), output_format {codec mp3|wav|pcm|mulaw|alaw, sample_rate, bit_rate}, speed, with_timestamps} + wss://api.x.ai/v1/tts streaming (text.delta → audio.delta); 28 built-in voices (GET /v1/tts/voices) + custom voices (/v1/custom-voices, Enterprise); $15 / 1M charactersEndpoint: POST /v1/ttsParams: text, voice_id, language, output_formatStatus: DOCUMENTED · LIVE_VERIFIED17 voice endpoints; voice models absent from GET /v1/models |
generateContent on TTS models (gemini-3.1-flash-tts-preview, gemini-2.5-flash|pro-preview-tts) with responseModalities: [AUDIO] + speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName (30 voices) or multiSpeakerVoiceConfig (≤2 speakers); raw PCM 24 kHz out; $1 text in / $20 audio out per 1M (≈ $0.03/min); free tier (10 req/day observed)Endpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.responseModalities, generationConfig.speechConfigStatus: DOCUMENTED · PREVIEW · LIVE_VERIFIEDstreaming TTS on 3.1 only; multi-speaker may return finishReason: OTHER |
yes | Dedicated endpoint (OpenAI, xAI) vs a generation modality (Gemini); none on Anthropic. Ref: docs/openai/audio.md · docs/xai/voice.md · docs/gemini/speech-generation.md |
| Speech-to-text / transcription 3/4 |
POST /v1/audio/transcriptions (gpt-transcribe, gpt-4o-transcribe(-diarize), whisper-1; streaming transcript.text.delta), POST /v1/audio/translations (whisper-1)Endpoint: POST /v1/audio/transcriptionsParams: model, file, response_format, streamStatus: DOCUMENTED · LIVE_VERIFIEDwhisper-1 & gpt-4o-transcribe shutdown 2027-02-26 |
— not offered | POST /v1/stt (multipart file ≤500 MB or url; language, diarize, keyterm, vad_threshold, model: grok-voice-transcribe-2.0|1.0) $0.10/h; wss://api.x.ai/v1/stt streaming (binary frames, interim_results, endpointing, smart_turn → transcript.partial/done) $0.20/hEndpoint: POST /v1/sttParams: file, url, language, diarize, modelStatus: DOCUMENTED · LIVE_VERIFIED25 languages; default model contradicts between release notes (1.0) and model page (2.0) |
gemini-3.5-transcribe via generateContent + generationConfig.audioTranscriptionConfig {languageCodes, customVocabulary, mode VERBATIM|SMART, diarization, wordTimestamp} → parts[].audioTranscription {text, speakerLabel, words[]} (LIVE_DISCOVERED shape); ≤1 h audio; ≈ $0.005/min blended; gemini-3.5-transcribe-live over the Live WebSocket (10-min sessions); gemini-3.5-live-translate-preview speech-to-speech translationEndpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.audioTranscriptionConfigStatus: DOCUMENTED · LIVE_VERIFIEDfree tier available |
yes | OpenAI/xAI dedicated endpoints (per minute / per hour); Gemini a dedicated model behind the generic endpoint; none on Anthropic. Ref: docs/openai/audio.md · docs/xai/voice.md · docs/gemini/transcription.md |
| Realtime speech-to-speech (WebSocket / WebRTC / SIP) 3/4 |
Realtime API GA: WebSocket wss://api.openai.com/v1/realtime?model=, WebRTC POST /v1/realtime/calls, SIP, ephemeral POST /v1/realtime/client_secrets; transcription & translation sessions; 60-min sessionsEndpoint: WS /v1/realtimeParams: session.update, response.createStatus: DOCUMENTED · LIVE_VERIFIEDlegacy /v1/realtime/sessions → 404 |
— not offered | wss://api.x.ai/v1/realtime?model=grok-voice-latest (= grok-voice-think-fast-2.0) with the OpenAI Realtime event vocabulary (session.update, input_audio_buffer.*, conversation.item.create, response.create → response.output_audio.delta, response.done; 39 server / 9 client events); tools in-session (function, web_search, x_search, file_search, mcp); reasoning.effort high|none; ephemeral POST /v1/realtime/client_secrets (≤3600 s); SIP via /v2/phone-numbers, /v1/realtime/calls/{id}/refer|hangup, webhook realtime.call.incoming; $0.08/min + $0.004 per text item; 120-min sessions; concurrent sessions 10–200 by tierEndpoint: WS wss://api.x.ai/v1/realtimeParams: session.update, conversation.item.create, response.createStatus: DOCUMENTED · LIVE_VERIFIEDtext turn LIVE_VERIFIED; audio UNVERIFIED; ping undocumented |
Live API wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent (setup → setupComplete; clientContent/realtimeInput {audio|video|text|activityStart|activityEnd}/toolResponse → serverContent {modelTurn, turnComplete, interrupted, inputTranscription, outputTranscription}, toolCall, goAway, sessionResumptionUpdate); models gemini-3.8-live (native audio, interleaved thinking, GA 2026-09-15), -extended-thinking, gemini-3.1-flash-live-preview, 2.5 native audio; ephemeral tokens POST /v1beta/auth_tokens on the …Constrained method; 16 kHz PCM in / 24 kHz out; ≈10-min connections, 15-min audio sessions (unlimited with contextWindowCompression), resumption handles 2 h; audio $3 in / $12 out per 1M (≈ $0.005 / $0.018 per min); free tierEndpoint: WSS BidiGenerateContentParams: setup, realtimeInput, toolResponse, setup.sessionResumption, setup.contextWindowCompressionStatus: DOCUMENTED · PREVIEW · LIVE_VERIFIED32 message types recorded; responseModalities: [TEXT] → close 1007; only googleSearch + functions as tools |
yes | Three voice stacks: xAI copies OpenAI's Realtime protocol (so client code ports almost verbatim), Gemini's Live API is a distinct message family; Anthropic has none. Only OpenAI and xAI offer WebRTC/SIP. Ref: docs/openai/realtime.md · docs/xai/voice.md · docs/gemini/live-api.md · docs/gemini/live-events.md |
| Full-duplex voice with backend delegation (Live) / voice agents 1/4 |
Live API: POST /v1/live/sessions (WebRTC), WS /v1/live/sessions, fork/attach/SIP; gpt-live-1 $0.05/min; delegation to Responses or clientEndpoint: POST /v1/live/sessionsParams: session.delegation, session.storeStatus: DOCUMENTED · LIVE_DISCOVERED |
— not offered | no separate delegation product; the realtime session itself runs server tools (web/x/file search, MCP) with reasoning.effortEndpoint: WS wss://api.x.ai/v1/realtimeStatus: DOCUMENTED |
no delegation layer; gemini-3.8-live-extended-thinking runs background reasoning and speaks fillers (interactionStatus: IN_PROGRESS); proactivity.proactiveAudio always on for 3.8Endpoint: WSS BidiGenerateContentStatus: DOCUMENTED |
no | OpenAI-only as a product; xAI and Gemini fold agentic behaviour into the voice session itself. Ref: docs/openai/live.md · docs/xai/voice.md · docs/gemini/live-api.md |
| Image generation API 3/4 |
POST /v1/images/generations, /edits (gpt-image-2, gpt-image-2.5-flare/sunburst; DALL·E RETIRED, gpt-image-1.x DEPRECATED); streaming partials; arbitrary sizes ≤4KEndpoint: POST /v1/images/generationsParams: prompt, size, quality, background, output_format, streamStatus: DOCUMENTED · LIVE_VERIFIED/variations RETIRED (404) |
— not offered | POST /v1/images/generations {model, prompt, n 1–10, response_format url|b64_json, aspect_ratio (incl. 21:9, 5:2, auto), resolution 1k|1.5k|2k, quality low|medium|auto, storage_options} and POST /v1/images/edits (JSON only: image {url|file_id} or images[] 2–5); grok-imagine-image $0.02, grok-imagine-image-2.0 $0.04–$0.08, grok-imagine-image-quality $0.05 (DEPRECATED → 2026-11-02); sync ~5–10 s; always JPEGEndpoint: POST /v1/images/generationsParams: prompt, n, aspect_ratio, resolution, quality, response_formatStatus: DOCUMENTED · LIVE_VERIFIEDno size/mask/seed; grok-2-image RETIRED; not on us.api.x.ai |
generateContent on image models with responseModalities + imageConfig {aspectRatio (14 ratios), imageSize 512|1K|2K|4K}; editing = pass the image as input and iterate; SynthID watermark; gemini-3.1-flash-image $0.045–$0.151, -lite-image $0.034, gemini-3-pro-image $0.134/$0.24; gemini-2.5-flash-image DEPRECATED → 2026-10-02; Imagen (:predict) RETIRED 2026-08-17; OpenAI-compat POST /v1beta/openai/images/generations subsetEndpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.responseModalities, generationConfig.imageConfig.aspectRatio, generationConfig.imageConfig.imageSizeStatus: DOCUMENTED · ACCOUNT_RESTRICTEDno free tier ( limit: 0); batch 50 % |
yes | OpenAI and xAI have dedicated image endpoints (token-priced vs per-image); Gemini generates images through the text endpoint; Claude produces no images. Ref: docs/openai/images.md · docs/xai/images.md · docs/gemini/image-generation.md |
| Video generation API 3/4 |
/v1/videos (Sora 2) — DEPRECATED, shutdown 2026-09-24; $0.10–$0.70/sEndpoint: POST /v1/videosParams: prompt, seconds, sizeStatus: DOCUMENTED · DEPRECATED · LIVE_VERIFIED |
— not offered | POST /v1/videos/generations {model, prompt, duration 1–15, aspect_ratio, resolution 480p|720p|1080p, generate_audio, image, reference_images[], reference_audios[] (1.5), last_frame (1.5)} → {request_id}; GET /v1/videos/{request_id} 202 pending → 200 {status: done, video:{url}}; POST /v1/videos/edits, /extensions; grok-imagine-video $0.05/s, grok-imagine-video-1.5 $0.08/s (1080p, reference-to-video)Endpoint: POST /v1/videos/generationsParams: prompt, duration, aspect_ratio, resolution, imageStatus: DOCUMENTED · LIVE_VERIFIED4 endpoints; also via Batch (URLs expire 1 h) |
Veo 3.1: POST /v1beta/models/veo-3.1-*:predictLongRunning {instances[{prompt, image, lastFrame, referenceImages[], video}], parameters {aspectRatio, resolution 720p|1080p|4k, durationSeconds 4|6|8, personGeneration, negativePrompt, seed}} → Operation; poll GET /v1beta/{name}; download files/{id}:download?alt=media; native audio; $0.40/$0.60 (Veo 3.1), $0.10–$0.30 (Fast), $0.05/$0.08 (Lite) per second; gemini-omni-1.1-flash video-out model via Interactions (≈ $0.10/s); OpenAI-compat POST /v1beta/openai/videosEndpoint: POST /v1beta/models/{model}:predictLongRunningParams: instances[].prompt, parameters.resolution, parameters.durationSecondsStatus: DOCUMENTED · PREVIEWVeo 2.0/3.0 RETIRED 2026-06-30; no free tier; 400 on this key |
yes | Async polling everywhere; OpenAI's is shutting down, xAI's is the cheapest per second, Gemini's the only one with 4K. Ref: docs/openai/video.md · docs/xai/videos.md · docs/gemini/video-generation.md |
| Music generation 1/4 |
— not offered | — not offered | — not offered | lyria-3.5 (GA, $0.08/song), lyria-3-clip-preview ($0.04 / 30-s clip), lyria-3-pro-preview via generateContent (MP3) or Interactions response_format {type: audio} (WAV); Lyria RealTime WebSocket …BidiGenerateMusic (weightedPrompts, musicGenerationConfig {bpm, density, brightness, scale…}, playbackControl) → 48 kHz stereo PCM chunks (experimental, unpriced)Endpoint: POST /v1beta/models/{model}:generateContentParams: contents[].parts[].text, musicGenerationConfig, playbackControlStatus: DOCUMENTED · ACCOUNT_RESTRICTEDno free tier for Lyria 3.x; RealTime LIVE_VERIFIED |
no | Gemini-only. Ref: docs/gemini/music-generation.md |
| Embeddings 2/4 |
text-embedding-3-small ($0.02/M), -3-large ($0.13/M), ada-002; dimensions, encoding_format; 8,192 tokens/input, 2,048 inputsEndpoint: POST /v1/embeddingsParams: input, model, dimensions, encoding_formatStatus: DOCUMENTED · LIVE_VERIFIED |
— not offered | POST /v1/embeddings {model, input, encoding_format, dimensions} documented (OpenAPI) but grok-embedding-small → 404 'does not exist or your team does not have access', GET /v1/embedding-models → {models: []}; no published price; used internally to index CollectionsEndpoint: POST /v1/embeddingsParams: input, model, dimensionsStatus: DOCUMENTED · ACCOUNT_RESTRICTEDgRPC Embedder.Embed exists |
POST /v1beta/models/gemini-embedding-2:embedContent {content, outputDimensionality 128–3072, taskType (001 only), embedContentConfig} and :batchEmbedContents {requests[]}; multimodal (text, ≤6 images, ≤180 s audio, ≤120 s video, PDF ≤6 pages → one vector); $0.20 text / $0.45 image / $6.50 audio / $12 video per 1M, batch 50 %; free tier; gemini-embedding-001 DEPRECATED → 2028-05-14; OpenAI-compat /v1beta/openai/embeddingsEndpoint: POST /v1beta/models/{model}:embedContentParams: content, outputDimensionality, taskTypeStatus: DOCUMENTED · LIVE_VERIFIED:asyncBatchEmbedContent ACCOUNT_RESTRICTED (400 FAILED_PRECONDITION) |
yes | OpenAI (text) and Gemini (multimodal) ship embeddings; xAI's endpoint is documented but inaccessible; Anthropic has none. Ref: docs/openai/embeddings.md · docs/gemini/embeddings.md · docs/models/xai-models.md |
| Moderation endpoint 1/4 |
POST /v1/moderations (omni-moderation-latest, text+image, free) and inline moderation {model, policy} on Responses/ChatEndpoint: POST /v1/moderationsParams: input, model, moderationStatus: DOCUMENTED · LIVE_VERIFIEDtext-moderation-* RETIRED |
— not offered | — not offered | no moderation route among the 125 Gemini endpoints; safetySettings thresholds and safetyRatings insteadStatus: DOCUMENTED |
no | OpenAI-only. Ref: docs/openai/moderation.md · docs/gemini/safety.md |
| Content provenance 1/4 |
POST /v1/content_provenance_checks (C2PA verification)Endpoint: POST /v1/content_provenance_checksStatus: DOCUMENTED · LIVE_VERIFIED |
generated media from code execution carry C2PA credentials; no verification endpoint Status: DOCUMENTED |
no provenance endpoint; respect_moderation flag on image/video resultsStatus: DOCUMENTED |
SynthID watermark on all generated images/video/audio (Nano Banana 2 Lite also C2PA); no verification endpoint on the Developer API Status: DOCUMENTED |
no | Only OpenAI exposes a check endpoint. Ref: endpoints.json · docs/gemini/image-generation.md |
Files & batch
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Files API 4/4 |
POST /v1/files (purposes user_data, batch, evals, assistants, vision, fine-tune), list/retrieve/delete/content; 512 MB/file; expires_after 1 h–30 d; batch files auto-expire 30 dEndpoint: POST /v1/filesParams: file, purpose, expires_afterStatus: DOCUMENTED · LIVE_VERIFIEDuser_data content not downloadable (400) |
POST /v1/files (any MIME; 500 MB/file; workspace-scoped), list (ids[]), metadata, content (tool-generated files only), delete; expires_in_seconds 1 h–90 d; GA no headerEndpoint: POST /v1/filesParams: file, expires_in_secondsStatus: DOCUMENTED · LIVE_VERIFIEDnot on Bedrock/Vertex |
POST /v1/files (multipart; purpose ignored; expires_after 1 h–30 d), list (filter AIP-160, sort_by), get, delete, GET …/content?format=original, public URLs (POST …/public-url, …/revoke; ≤50 MiB, ≤1,000 active/team), chunked upload files:initialize/:uploadChunks (UNVERIFIED); storage $0.025/GiB/day, downloads $0.20/GiB; used by Responses input_file, Collections, Imagine, BatchEndpoint: POST /v1/filesParams: file, expires_afterStatus: DOCUMENTED · LIVE_VERIFIED10 endpoints; disabled under ZDR |
POST /upload/v1beta/files (resumable X-Goog-Upload-* protocol, multipart or media), GET /v1beta/files[/{id}], DELETE, POST /v1beta/files:register {uris: gs://…} (GCS, 30 days), GET /v1beta/generatedFiles, download files/{id}:download?alt=media (generated files); 2 GB/file, 20 GB/project, 48 h TTL, free; used via fileData {fileUri}Endpoint: POST /upload/v1beta/filesParams: file, displayName, mimeTypeStatus: DOCUMENTED · LIVE_VERIFIEDunknown file → 403 (never 404); not available on Vertex AI |
yes | All four are blob stores referenced by id; TTLs range from 48 h (Gemini, fixed) to 90 d (Anthropic); only xAI prices storage and only xAI mints public URLs. Ref: docs/openai/files-and-uploads.md · docs/anthropic/files-api.md · docs/xai/files.md · docs/gemini/files.md |
| Multipart / resumable upload 3/4 |
Uploads API: ≤8 GB in ≤64 MB parts, 1 h TTL Endpoint: POST /v1/uploadsParams: filename, purpose, bytes, mime_typeStatus: DOCUMENTED · LIVE_VERIFIED |
— not offered | POST /v1/files:initialize + POST /v1/files:uploadChunks (documented, UNVERIFIED)Endpoint: POST /v1/files:initializeStatus: DOCUMENTED · UNVERIFIED |
Google resumable upload protocol on /upload/v1beta/files and /upload/v1beta/fileSearchStores/{s}:uploadToFileSearchStore (start → x-goog-upload-url → upload, finalize; 8 MiB chunk granularity); environments PUT /upload/v1beta/environments/{env}/files/{path} (≤2 GiB)Endpoint: POST /upload/v1beta/filesParams: X-Goog-Upload-Protocol, X-Goog-Upload-CommandStatus: DOCUMENTED · LIVE_VERIFIED |
yes | Three resumable schemes, all provider-specific. Ref: docs/openai/files-and-uploads.md · docs/xai/files.md · docs/gemini/files.md |
| Batch processing (async, discounted) 4/4 |
JSONL file (custom_id, method, url, body) → POST /v1/batches {input_file_id, endpoint, completion_window:'24h'}; 50,000 requests / 200 MB; endpoints responses, chat, embeddings, completions, moderations, images, videos; 50 % off; output file 30 dEndpoint: POST /v1/batchesParams: input_file_id, endpoint, completion_window, output_expires_afterStatus: DOCUMENTED · LIVE_VERIFIEDcancel LIVE_DISCOVERED 409 on terminal |
inline requests[] {custom_id, params} → POST /v1/messages/batches; 100,000 requests / 256 MB; 24 h expiry; results JSONL at results_url 29 d; 50 % off all token dims (stacks with caching); all Messages features incl. server tools; output-300k-2026-03-24 betaEndpoint: POST /v1/messages/batchesParams: requests, requests[].custom_id, requests[].paramsStatus: DOCUMENTED · LIVE_VERIFIED6 endpoints all LIVE_VERIFIED |
POST /v1/batches {name} → POST /v1/batches/{id}/requests {batch_requests[{batch_request_id, batch_request: {chat_get_completion | responses | image_generation | image_edit | video_generation | video_extension}}]} (inline, ≤25 MB/request) or JSONL via Files (input_file_id, ≤50,000 lines / 200 MB); GET …/requests, GET …/results, POST …:cancel; 20 % off on grok-4.3 / 4.20 only (grok-4.6 / 4.5 / build unsupported; Imagine at standard rates); bypasses rate limits; not with priority; live turnaround 11 sEndpoint: POST /v1/batchesParams: name, input_file_id, batch_requestsStatus: DOCUMENTED · LIVE_VERIFIED8 endpoints; DELETE → 405; responses bodies come back as chat_get_completion |
POST /v1beta/models/{model}:batchGenerateContent {batch:{displayName, inputConfig:{requests:{requests[]} | fileName}, priority, webhookConfig}} (inline ≤20 MB or JSONL Files ≤2 GB) → Operation batches/{id}; :asyncBatchEmbedContent; GET /v1beta/batches[/{id}], :cancel, DELETE, PATCH :update*Batch; states PENDING→RUNNING→SUCCEEDED|FAILED|CANCELLED|EXPIRED (>48 h); 50 %; target 24 h; 100 concurrent jobs; per-model enqueued-token caps; results 6 weeks; OpenAI-compat /v1beta/openai/batchesEndpoint: POST /v1beta/models/{model}:batchGenerateContentParams: batch.inputConfig.requests, batch.inputConfig.fileName, batch.displayNameStatus: DOCUMENTED · ACCOUNT_RESTRICTED9 endpoints; free tier → 400 FAILED_PRECONDITION; list LIVE_VERIFIED |
yes | Same 24-h / discounted idea on all four, with 50 % (OpenAI, Anthropic, Gemini) vs 20 % (xAI, three models only). Inline requests: Anthropic, xAI, Gemini; file-based: OpenAI, xAI, Gemini. Ref: docs/openai/batch.md · docs/anthropic/message-batches.md · docs/xai/batches.md · docs/gemini/batch.md |
Prompt caching
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Prompt / context caching 4/4 |
automatic prefix caching (≥1,024 tokens; exact prefix; prompt_cache_key routing hint); GPT-5.6+: prompt_cache_options {mode: implicit|explicit, ttl:'30m', prewarm} and per-part prompt_cache_breakpoint {mode:explicit} (≤4); prompt_cache_retention: in_memory|24h (deprecated)Endpoint: POST /v1/responsesParams: prompt_cache_key, prompt_cache_options, prompt_cache_retention, prompt_cache_breakpointStatus: DOCUMENTED · LIVE_VERIFIEDreads 0.1× (5.6+) / model-specific; writes 1.25× only on GPT-5.6+ |
explicit cache_control {type: ephemeral, ttl: 5m|1h} on system/tools/content blocks (≤4 breakpoints) or top-level automatic; 20-block lookback; min 512 (Fable/Mythos/Opus 5) · 1,024 (Sonnet 5/4.6/4.5, Opus 4.8) · 2,048 (Opus 4.7) · 4,096 (Haiku 4.5, Opus 4.6/4.5)Endpoint: POST /v1/messagesParams: cache_control, system[].cache_control, messages[].content[].cache_control, tools[].cache_controlStatus: DOCUMENTED · LIVE_VERIFIEDwrites 1.25× (5m) / 2× (1h); reads 0.1× (0.025× Fable 5.1) |
automatic prefix cache, no markers (cache_control on /v1/messages ignored); sticky routing via prompt_cache_key (Responses, Chat) or header x-grok-conv-id (Chat/gRPC); include prior reasoning_content / encrypted reasoning or use previous_response_id to keep hits; no TTL or minimum documented, no write charge; cached reads 0.15×–0.25× (grok-4.6 $0.50, grok-4.3 $0.20); long-context tier doubles cached price tooEndpoint: POST /v1/responsesParams: prompt_cache_key, x-grok-conv-idStatus: DOCUMENTED · LIVE_VERIFIEDlive: a fresh minimal request already shows ~192 cached tokens (hidden system prefix); cached tokens count toward TPM |
implicit caching (automatic on 2.5+, min prompt 4,096 tokens on 3.x Flash/3.1 Pro, 2,048 on 2.5; hits in usageMetadata.cachedContentTokenCount) and explicit cachedContents (POST /v1beta/cachedContents {model, contents, systemInstruction, tools, ttl|expireTime} → use cachedContent: cachedContents/{id}; default TTL 1 h, PATCH TTL only; min 1,024 tokens live); cached reads 0.1× everywhere; explicit storage $0.50–$4.50 per 1M tokens per hourEndpoint: POST /v1beta/models/{model}:generateContentParams: cachedContent, generationConfig, ttl, expireTimeStatus: DOCUMENTED · BETA · ACCOUNT_RESTRICTEDexplicit caching v1beta-only, not via Interactions; free tier limit: 0; no write fee |
yes | Two philosophies: implicit (OpenAI default, xAI, Gemini implicit) vs explicit breakpoints (Anthropic, OpenAI 5.6+, Gemini cachedContents). Write fees exist only on Anthropic and GPT-5.6+; Gemini charges storage time instead; xAI charges nothing but publishes no TTL.Ref: docs/comparisons/caching-and-reasoning.md · docs/xai/prompt-caching.md · docs/gemini/context-caching.md |
| Cache diagnostics 2/4 |
prompt_cache_options.comparison_response_id → prompt_cache_diagnostics {type: cache_hit|cache_miss{reason}} (GPT-5.6+)Endpoint: POST /v1/responsesParams: prompt_cache_options.comparison_response_idStatus: DOCUMENTEDreasons: model_changed, tools_changed, text_format_changed, … |
diagnostics.previous_message_id → response diagnostics cache-miss reasons; beta cache-diagnosis-2026-04-07Endpoint: POST /v1/messagesParams: diagnostics.previous_message_idStatus: DOCUMENTED · BETA |
no diagnostics; only usage.*cached_tokens counters (input_tokens_details.cached_tokens, prompt_tokens_details.cached_tokens, cache_read_input_tokens on /v1/messages)Endpoint: POST /v1/responsesStatus: DOCUMENTED |
no diagnostics; usageMetadata.cachedContentTokenCount + cacheTokensDetails[] onlyEndpoint: POST /v1beta/models/{model}:generateContentStatus: DOCUMENTED |
yes | OpenAI and Anthropic explain misses; xAI and Gemini only count hits. Ref: parameters.json |
Reasoning / thinking
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Reasoning control 4/4 |
reasoning {effort: none|minimal|low|medium|high|xhigh|max, summary: auto|concise|detailed, context: auto|current_turn|all_turns, mode: standard|pro}; reasoning output items with encrypted_content; usage.output_tokens_details.reasoning_tokensEndpoint: POST /v1/responsesParams: reasoning.effort, reasoning.summary, reasoning.context, reasoning.modeStatus: DOCUMENTED · LIVE_VERIFIEDChat: reasoning_effort only; effort set per model (gpt-6-astra rejects none) |
thinking {type: adaptive|enabled|disabled, budget_tokens, display: summarized|omitted|updates} + output_config.effort: low|medium|high|xhigh|max (default high); thinking/redacted_thinking blocks with signature; usage.output_tokens_details.thinking_tokensEndpoint: POST /v1/messagesParams: thinking.type, thinking.budget_tokens, thinking.display, output_config.effortStatus: DOCUMENTED · LIVE_VERIFIEDenabled (manual budget) 400 on 4.7+; adaptive 400 on 4.5 models |
Responses reasoning {effort: low|medium|high|xhigh, summary: auto|concise|detailed} (alias reasoning_effort); Chat reasoning_effort + message.reasoning_content (delta.reasoning_content when streaming); usage.completion_tokens_details.reasoning_tokens; reasoning is always on for grok-4.6/4.5/4.20-reasoning/build (reasoning_effort → 400 on 4.20-reasoning and build); grok-4.3 accepts none (LIVE_DISCOVERED, 0 reasoning tokens); grok-4.20-0309-non-reasoning has none; on grok-4.20-multi-agent-0309 effort selects 4 or 16 agentsEndpoint: POST /v1/responsesParams: reasoning.effort, reasoning.summary, reasoning_effortStatus: DOCUMENTED · LIVE_VERIFIEDmax_output_tokens documented to include reasoning but not enforced live; 'Reply with OK.' costs 70–180 reasoning tokens |
generationConfig.thinkingConfig {thinkingLevel: MINIMAL|LOW|MEDIUM|HIGH (Gemini 3), thinkingBudget: -1|0|N (2.5-era, LEGACY on 3.x; exclusive with level → 400), includeThoughts}; defaults high (3.1 Pro, 3 Flash) / medium (3.5–3.8 Flash) / minimal (Flash-Lite); MINIMAL → 400 on 3.8/3.7 Flash and 3.1 Pro; thought parts {text, thought: true}; usageMetadata.thoughtsTokenCount; Interactions generation_config.thinking_level + thinking_summaries: auto|noneEndpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.thinkingConfig.thinkingLevel, generationConfig.thinkingConfig.thinkingBudget, generationConfig.thinkingConfig.includeThoughtsStatus: DOCUMENTED · LIVE_VERIFIEDthinking cannot be disabled on 3.8/3.7 Flash, 3.1 Pro, 2.5 Pro; thinkingBudget: 0 works on 3.6/3.5 Flash; OpenAI-compat maps reasoning_effort minimal/low/medium/high |
yes | Four effort ladders overlap on low/medium/high: OpenAI adds none/minimal/xhigh/max, Anthropic xhigh/max, xAI xhigh, Gemini minimal. Off-switches: OpenAI none (most models), Anthropic disabled (not Fable/Mythos), xAI none on grok-4.3 only, Gemini thinkingBudget: 0 on two Flash models only. Only Anthropic and Gemini (2.5) expose a token budget.Ref: docs/openai/reasoning.md · docs/anthropic/thinking.md · docs/xai/reasoning.md · docs/gemini/thinking.md |
| Reasoning visibility 4/4 |
summaries only (reasoning.summary); raw content[] empty for GPT models; encrypted replay itemEndpoint: POST /v1/responsesParams: reasoning.summary, includeStatus: DOCUMENTED · LIVE_VERIFIED |
summarized text by default (≤4.6) or omitted (4.7+); display: updates progress text (Fable 5.x beta); signature always presentEndpoint: POST /v1/messagesParams: thinking.displayStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
Chat returns the full reasoning_content text by default; Responses returns reasoning items with summary[] (reasoning.summary accepted 'for compatibility'; the item is sometimes omitted); /v1/messages returns thinking blocks with an empty signatureEndpoint: POST /v1/responsesParams: reasoning.summary, includeStatus: DOCUMENTED · LIVE_VERIFIED |
includeThoughts: true → thought summaries as thought: true parts (not guaranteed on trivial prompts); thoughtSignature (opaque) on function-call and final parts; Interactions thought steps {signature, summary[]}Endpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.thinkingConfig.includeThoughts, contents[].parts[].thoughtSignatureStatus: DOCUMENTED · LIVE_VERIFIEDbilling is on full thoughts, not the summary |
yes | Only xAI (Chat Completions) exposes the raw chain of thought; the other three return summaries plus an opaque signature/encrypted blob. Ref: docs/anthropic/thinking.md §3 · docs/xai/reasoning.md · docs/gemini/thinking.md |
| Reasoning replay across turns 4/4 |
replay reasoning items verbatim (auto with previous_response_id); encrypted_content decrypted in memory when store:false; reasoning.context: all_turns (GPT-5.6 default)Endpoint: POST /v1/responsesParams: input[](reasoning).encrypted_content, reasoning.contextStatus: DOCUMENTED · LIVE_VERIFIED |
pass thinking blocks back unmodified within tool-use turns (400 if modified); prior-turn thinking kept for all turns on Opus 4.5+/Sonnet 4.6+/Fable, last turn on Haiku/Sonnet 4.5; clear_thinking_20251015 edit; beta thinking.block_bindingEndpoint: POST /v1/messagesParams: messages[].content[].signature, thinking.block_bindingStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
include: ["reasoning.encrypted_content"] (xai-sdk use_encrypted_content=True) → replay the reasoning items with encrypted_content when store:false (ZDR) — also the 'top cause of cache misses' when omitted; previous_response_id rehydrates automatically; Chat: resend reasoning_contentEndpoint: POST /v1/responsesParams: include, input[](reasoning).encrypted_content, previous_response_idStatus: DOCUMENTED · LIVE_VERIFIED |
mandatory on Gemini 3: echo thoughtSignature of the first functionCall of each step (400 'missing a thought_signature' / finishReason: MISSING_THOUGHT_SIGNATURE, even at minimal); text-part signatures recommended; tampering → 400 'Corrupted thought signature'; Interactions carry it in thought steps; 'thought preservation' across turns since 3.5 FlashEndpoint: POST /v1beta/models/{model}:generateContentParams: contents[].parts[].thoughtSignatureStatus: DOCUMENTED · LIVE_VERIFIEDdummy values skip_thought_signature_validator bypass validation (verified) |
yes | All four carry hidden reasoning as an opaque blob; Anthropic and Gemini validate it (400 on edits / omissions), OpenAI and xAI encrypt it. Ref: docs/anthropic/thinking.md §4 · docs/xai/reasoning.md · docs/gemini/thinking.md |
| Pro / extended-compute mode 3/4 |
reasoning.mode: pro (GPT-5.6) billed at standard rates; *-pro models (gpt-5.4-pro, gpt-5.5-pro)Endpoint: POST /v1/responsesParams: reasoning.modeStatus: DOCUMENTED |
effort: max is the top of the ladder; no separate modeParams: output_config.effortStatus: DOCUMENTED |
grok-4.20-multi-agent-0309 (BETA): one Responses call fans out to 4 (low/medium) or 16 (high/xhigh) agents; all agents' tokens billed at grok-4.20 rates; Responses-only (Chat → 400)Endpoint: POST /v1/responsesParams: reasoning.effortStatus: DOCUMENTED · BETA · LIVE_VERIFIED'Reply with OK.' cost 1,645 output tokens ≈ $0.0047 |
gemini-3.8-live-extended-thinking (Live) and Deep Research agents (deep-research-preview-04-2026, -max-; 60-min background runs, $1–7 per task estimated) rather than a mode; gemini-3.1-pro-preview is the deep-reasoning text modelEndpoint: POST /v1beta/interactionsParams: agent, agent_config(deep-research)Status: DOCUMENTED · PREVIEW · LIVE_VERIFIED |
no | Different vehicles: a mode/model (OpenAI), a multi-agent model (xAI), an agent (Gemini), a ladder value (Anthropic). Ref: parameters.json · docs/models/xai-models.md · docs/gemini/interactions-api.md |
| Change effort mid-conversation without breaking the cache 2/4 |
configuration_update input item {reasoning:{effort}} (gpt-6-astra)Endpoint: POST /v1/responsesParams: input[](configuration_update)Status: DOCUMENTED |
messages[].output_config.effort on role:system entries; beta mid-conversation-output-config-2026-07-01 (Fable 5.1, Mythos 5.1, Opus 5)Endpoint: POST /v1/messagesParams: messages[].output_config.effortStatus: DOCUMENTED · BETA |
per-request reasoning.effort only (automatic cache; no documented invalidation rule)Endpoint: POST /v1/responsesStatus: DOCUMENTED |
per-request thinkingLevel only; an explicit cachedContents prefix is unaffected by generationConfig changesEndpoint: POST /v1beta/models/{model}:generateContentStatus: DOCUMENTED |
yes | OpenAI and Anthropic only. Ref: parameters.json |
| Task-wide token budget (advisory) 3/4 |
none in Responses (max_tool_calls caps hosted tool calls; Agents API has session_budget_exceeded)Endpoint: POST /v1/responsesParams: max_tool_callsStatus: DOCUMENTED |
output_config.task_budget {type, total ≥20000, remaining}; beta task-budgets-2026-03-13 (Fable/Mythos/Opus 5/4.8/4.7)Endpoint: POST /v1/messagesParams: output_config.task_budgetStatus: DOCUMENTED · BETA |
max_turns (agentic turns per Responses request; default = server cap) — a turn budget, not tokensEndpoint: POST /v1/responsesParams: max_turnsStatus: DOCUMENTED · LIVE_VERIFIED |
Antigravity agent_config.max_total_tokens (Interactions, PREVIEW) → status incomplete when exhaustedEndpoint: POST /v1beta/interactionsParams: agent_config(antigravity).max_total_tokensStatus: DOCUMENTED · PREVIEW |
no | Three different budget units (tokens, turns, agent tokens); none on the OpenAI Responses API. Ref: compatibility/anthropic-feature-model-matrix.json · docs/xai/responses.md · docs/gemini/interactions-api.md |
Context management
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Context window 4/4 |
1,050,000 (GPT-5.4/5.5/5.6/6 Astra; >272k input = long-context pricing 2×/1.5×), 400k (5.x mini/nano, GPT-5–5.3), 200k (o-series), 128k (gpt-4o) Status: DOCUMENTED |
1,000,000 default on Claude 4.6+ (no header, no premium); 200k on Opus 4.5 / Sonnet 4.5 / Haiku 4.5; context-1m-2025-08-07 header RETIREDStatus: DOCUMENTED · LIVE_VERIFIED |
1,000,000 (grok-4.3, all grok-4.20 ids), 500,000 (grok-4.6, grok-4.5), 256,000 (grok-build-0.1); prompts ≥200,000 tokens switch the whole request to the 2× long-context tier Status: DOCUMENTED · LIVE_VERIFIEDusage.context_details on Responses |
1,048,576 on every Gemini 3.x / 2.5 text model (inputTokenLimit); 262,144 Gemma 4; 131,072 Live / image / agent models; Pro models bill >200k prompts at 2× input / 1.5× output, Flash models flatStatus: DOCUMENTED · LIVE_VERIFIED |
yes | ~1M on every flagship line; long-context premiums on OpenAI (>272k), xAI (≥200k, all tokens) and Gemini Pro (>200k); none on Anthropic or Gemini Flash. Ref: models.json |
| Max output tokens 4/4 |
128,000 on GPT-5.x/6 (272,000 on gpt-5-pro); max_output_tokens ≥16Endpoint: POST /v1/responsesParams: max_output_tokensStatus: DOCUMENTED · LIVE_VERIFIED |
128,000 on Claude 4.6+ (64,000 on 4.5 models); max_tokens required; 300,000 in Batches with output-300k-2026-03-24Endpoint: POST /v1/messagesParams: max_tokensStatus: DOCUMENTED · LIVE_VERIFIED |
no documented limit (max_output: null on every Grok record; grok-4.6 page: 'No text output limit'); Responses max_output_tokens default 128,000 (docs) and not enforced on reasoning live; Chat max_completion_tokens caps visible output onlyEndpoint: POST /v1/responsesParams: max_output_tokens, max_completion_tokensStatus: DOCUMENTED · LIVE_VERIFIEDmax_tokens DEPRECATED alias on Chat |
65,536 on every 3.x / 2.5 text model (outputTokenLimit); 32,768 Gemma 4 / image models; maxOutputTokens includes thinking tokens (hit → finishReason: MAX_TOKENS, possibly empty text)Endpoint: POST /v1beta/models/{model}:generateContentParams: generationConfig.maxOutputTokensStatus: DOCUMENTED · LIVE_VERIFIEDInteractions generation_config.max_output_tokens |
yes | 128k (OpenAI, Anthropic), 64k (Gemini), unpublished (xAI). Only Anthropic makes the cap mandatory. Ref: models.json |
| Server-side compaction (in-flight) 3/4 |
context_management: [{type: compaction, compact_threshold ≥1000}] → compaction output item, SSE response.compaction.compactingEndpoint: POST /v1/responsesParams: context_management[].compact_thresholdStatus: DOCUMENTED |
context_management.edits: [{type: compact_20260112, trigger ≥50000, pause_after_compaction, instructions}] → compaction block, stop_reason: compaction; beta compact-2026-01-12; 4.6+Endpoint: POST /v1/messagesParams: context_management.edits[]Status: DOCUMENTED · BETA · LIVE_VERIFIEDusage.iterations[] for billing |
context_management[] accepted but 'parsed but not yet executed' (compat only); use the stand-alone compact endpointEndpoint: POST /v1/responsesParams: context_managementStatus: DOCUMENTED |
Live API only: setup.contextWindowCompression {triggerTokens, slidingWindow {targetTokens}} (unlimited session length); Antigravity agents compact around ~135k automatically; nothing on generateContentEndpoint: WSS BidiGenerateContentParams: setup.contextWindowCompressionStatus: DOCUMENTED |
yes | OpenAI and Anthropic compact text conversations in-flight; Gemini only compresses Live sessions; xAI parses the field and ignores it. Ref: docs/openai/responses.md §6 · docs/anthropic/context-management.md §3 · docs/gemini/live-api.md |
| Stand-alone compaction request 3/4 |
POST /v1/responses/compact {model, input|previous_response_id} → response.compaction with encrypted compaction itemEndpoint: POST /v1/responses/compactParams: model, input, previous_response_id, instructionsStatus: DOCUMENTED · LIVE_VERIFIED |
compaction: {type: summarize} on Messages → single signed compaction block; beta compact-2026-09-04Endpoint: POST /v1/messagesParams: compaction, compaction.type, compaction.instructionsStatus: DOCUMENTED · BETA · LIVE_VERIFIEDblock must be sent first |
POST /v1/responses/compact {model, input} → {object: response.compaction, id: cmp_…, output:[{type: compaction, encrypted_content}], usage {…, dropped_message_count}}; put the output first in the next input; do not edit the blob; the pre-compaction conversation must still fit the window (May 2026 'Context Compaction API')Endpoint: POST /v1/responses/compactParams: model, inputStatus: DOCUMENTED · LIVE_VERIFIED48 parameter rows |
— not offered | yes | OpenAI-shaped on xAI (same endpoint and item); Anthropic as a Messages parameter; none on Gemini. Ref: docs/anthropic/context-management.md §3 · docs/xai/responses.md |
| Server-side context editing (clear old tool results / thinking) 1/4 |
no editing strategies; legacy truncation: auto drops oldest itemsEndpoint: POST /v1/responsesParams: truncationStatus: DOCUMENTED · LEGACY · LIVE_VERIFIED |
context_management.edits[]: clear_tool_uses_20250919, clear_thinking_20251015; beta context-management-2025-06-27; response context_management.applied_edits[]Endpoint: POST /v1/messagesParams: context_management.edits[].type, context_management.edits[].trigger, context_management.edits[].keepStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
truncation accepted (disabled echoed) — 'not supported, compatibility only'Endpoint: POST /v1/responsesParams: truncationStatus: DOCUMENTED |
— not offered | no | Anthropic-only. Ref: docs/anthropic/context-management.md §2 |
| Mid-conversation tool add/remove (cache-preserving) 2/4 |
additional_tools developer item (adds tools mid-thread); tool_choice: allowed_tools to restrictEndpoint: POST /v1/responsesParams: input[](additional_tools)Status: DOCUMENTED |
tool_addition/tool_removal blocks in role:system messages; beta mid-conversation-tool-changes-2026-07-01 (Fable 5, Mythos 5, Opus 4.8, Opus 5)Endpoint: POST /v1/messagesParams: messages[].roleStatus: DOCUMENTED · BETA |
follow-ups via previous_response_id 'may change tools/model' — no dedicated item; cache effect undocumentedEndpoint: POST /v1/responsesParams: previous_response_idStatus: DOCUMENTED |
resend the full tools[] each call (Interactions chaining also requires re-sending tools)Endpoint: POST /v1beta/models/{model}:generateContentStatus: DOCUMENTED |
yes | OpenAI and Anthropic only. Ref: anthropic-beta-headers.json |
Service tiers, limits, safety
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Service tiers / processing modes 4/4 |
service_tier: auto|default|flex|scale|priority|fast|ultrafast; flex = batch price, slower; fast = 2× (renamed from priority 2026-07-30); echoed in responseEndpoint: POST /v1/responsesParams: service_tierStatus: DOCUMENTED · LIVE_VERIFIED |
service_tier: auto|standard_only (Priority Tier commitments, no longer sold) + speed: fast (beta fast-mode-2026-02-01, Opus 5 / 4.8 only, 2× price); usage.service_tier standard|priority|batchEndpoint: POST /v1/messagesParams: service_tier, speedStatus: DOCUMENTED · BETA · PREVIEW · ACCOUNT_RESTRICTEDfast 429 'rate limit of 0' for our key |
service_tier: default|priority (Chat + Responses): priority = 2× on every token type (after the cache discount), billed only when the response echoes service_tier: priority; not combinable with BatchEndpoint: POST /v1/responsesParams: service_tierStatus: DOCUMENTED · LIVE_VERIFIEDlive responses always echoed default |
serviceTier: standard|flex|priority (service_tier on Interactions / OpenAI-compat): flex = 0.5× (1–15 min target, sheddable → 429), priority = 1.8× (0.3× rate limit, graceful downgrade to standard); echoed in usageMetadata.serviceTier and header X-Gemini-Service-TierEndpoint: POST /v1beta/models/{model}:generateContentParams: serviceTierStatus: DOCUMENTED · LIVE_VERIFIEDflex accepted live; ten text models listed |
yes | Cheaper tier: OpenAI flex, Gemini flex (both 50 %). Faster tier: OpenAI fast 2×, Anthropic fast 2× (two models), xAI priority 2×, Gemini priority 1.8×. Ref: docs/anthropic/service-tiers.md · docs/xai/pricing.md · docs/gemini/pricing.md · pricing.json |
| End-user identifier for abuse detection 4/4 |
safety_identifier (≤64 chars; replaces user) + prompt_cache_key; header OpenAI-Safety-Identifier on RealtimeEndpoint: POST /v1/responsesParams: safety_identifier, userStatus: DOCUMENTED · LIVE_VERIFIED |
metadata.user_id (≤512 chars, no PII); beta anthropic-user-profile-id headerEndpoint: POST /v1/messagesParams: metadata.user_id, anthropic-user-profile-idStatus: DOCUMENTED · LIVE_VERIFIED |
user and safety_identifier (both accepted on Chat and Responses); /v1/messages metadata.user_idEndpoint: POST /v1/responsesParams: safety_identifier, userStatus: DOCUMENTED · LIVE_VERIFIED |
labels {safety_identifier: …} (documented key; Cloud-label rules; accepted, not echoed) on generateContent and InteractionsEndpoint: POST /v1beta/models/{model}:generateContentParams: labelsStatus: DOCUMENTED · LIVE_VERIFIED |
yes | Same idea everywhere; the field name changes four times. Ref: parameters.json |
| Rate-limit headers 3/4 |
x-ratelimit-{limit,remaining,reset}-{requests,tokens} (+ -project-tokens), Retry-After, retry-after-ms, x-should-retryStatus: DOCUMENTED · LIVE_VERIFIEDreset as Go durations ( 6m0s) |
anthropic-ratelimit-{requests,tokens,input-tokens,output-tokens}-{limit,remaining,reset} (RFC 3339), retry-after, x-should-retry, anthropic-priority-*, anthropic-fast-*Status: DOCUMENTED · LIVE_VERIFIEDnot on GET /models or count_tokens |
undocumented but observed: x-ratelimit-limit-requests (7,200 grok-4.6/4.5, 1,800 grok-4.3/4.20/build — per minute), x-ratelimit-remaining-requests, x-ratelimit-limit-tokens (50,000,000 / 10,000,000 = documented T0 TPM), x-ratelimit-remaining-tokens; no reset, no Retry-After; x-request-id, x-zero-data-retentionStatus: DOCUMENTED · LIVE_DISCOVEREDabsent on /v1/responses multi-agent calls and catalogue GETs |
none — no x-ratelimit-* or Retry-After on 200/404/429; retry delay only inside the 429 message text ('Please retry in 54.2s') and google.rpc.RetryInfo; undocumented X-Gemini-Service-Tier headerStatus: DOCUMENTED · LIVE_DISCOVEREDno request-id header either — use responseId in the body |
yes | Headers on three providers (xAI's undocumented); Gemini puts the information in the 429 body. Ref: headers.json · rate-limits.json |
| Rate-limit tiers 4/4 |
Free, Tier 1–5 by cumulative spend ($5 → $1,000) with monthly usage caps $100 → $200k; per-model RPM/TPM/batch-queue tables; long-context tables >272k Status: DOCUMENTEDScale/Reserved Tier, Ultrafast preview above Tier 5 |
Start ($500/mo cap) / Build ($1,000) / Scale ($200k) / Custom; per model-class RPM/ITPM/OTPM (cache reads excluded from ITPM); Batches & count_tokens separate Status: DOCUMENTED · LIVE_DISCOVEREDobserved Scale-tier headers for our key |
Tier 0–4 by cumulative spend since 2026-01-01 ($0 / $50 / $250 / $1,000 / $5,000; Enterprise on request); per-model RPS (= RPM/60) and TPM (prompt + completion + reasoning + cached): grok-4.6/4.5 150→500 RPS, 50M→100M TPM; grok-4.3/4.20/build 37→208 RPS, 10M→85M TPM; multi-agent 9→56 RPS; Imagine RPS-only (6→100 images, 10→158 videos); voice concurrent sessions 10→200; Batch bypasses limits; per-key qps/qpm/tpm caps via Management APIStatus: DOCUMENTED · LIVE_DISCOVEREDconsole shows personalised limits |
Free / Tier 1 (billing linked; $250 cap, $10 per rolling 10 min) / Tier 2 ($100 paid + 3 days; $2,000; $50) / Tier 3 ($1,000 + 30 days; $20k–100k+; $200); dimensions RPM / TPM / RPD (+ IPM, TPD) per project; per-model matrix published only in AI Studio; Pro and media models unavailable on Free (limit: 0); priority 0.3× limits; batch enqueued-token caps per modelStatus: DOCUMENTED · LIVE_DISCOVEREDfree-tier RPM observed 15/min on flash-lite |
yes | Spend-based tiers everywhere; xAI is the only provider publishing exact per-model RPS/TPM per tier in the docs, Gemini the only one with a genuinely free tier. Ref: generated/rate-limits.json |
| Overload / capacity error 4/4 |
HTTP 503 server_is_overloaded (retryable, Retry-After)Status: DOCUMENTED |
HTTP 529 overloaded_error (also as SSE error event after 200); acceleration limits now 429Status: DOCUMENTED |
HTTP 429 (RPS/TPM/credits; gRPC RESOURCE_EXHAUSTED) and 5xx internal — no dedicated overload code; status.x.aiStatus: DOCUMENTEDno 429 triggered in this run |
HTTP 503 UNAVAILABLE ('The model is overloaded. Please try again later.'), 429 RESOURCE_EXHAUSTED when Flex capacity is shed, 504 DEADLINE_EXCEEDED for long Flex/Deep Research requestsStatus: DOCUMENTED |
yes | 503 (OpenAI, Gemini), 529 (Anthropic), 429 (xAI) for the same condition. Ref: errors.json |
| Error envelope 4/4 |
{error: {message, type, param, code}}; types invalid_request_error, rate_limit_error, insufficient_quota, server_error…; codes e.g. model_not_found, context_length_exceeded, previous_response_not_foundStatus: DOCUMENTED · LIVE_VERIFIEDempty-body 404 from Cloudflare on unknown URLs |
{type: error, error: {type, message}, request_id}; types invalid_request_error, authentication_error, permission_error, not_found_error, request_too_large, rate_limit_error, api_error, overloaded_error, billing_error, timeout_errorStatus: DOCUMENTED · LIVE_VERIFIEDno code field except error.details.error_code on some 429/529 |
{code: <kebab-case>, error: <message>} (invalid-argument, not-found, unauthenticated:no-credentials…); 422 = bare JSON string (serde message); Management API = gRPC-style {code: 16, message, details[]}; Realtime WS {type: error, error:{type, code, message}}; an invalid API key returns 400, not 401Status: DOCUMENTED · LIVE_VERIFIED12 error records; gRPC↔HTTP mapping 3→400, 16→401, 7→403, 5→404, 8→429 |
google.rpc Status: {error: {code: <http>, message, status: INVALID_ARGUMENT|FAILED_PRECONDITION|UNAUTHENTICATED|PERMISSION_DENIED|NOT_FOUND|ALREADY_EXISTS|RESOURCE_EXHAUSTED|INTERNAL|UNIMPLEMENTED|UNAVAILABLE|DEADLINE_EXCEEDED…, details[] (BadRequest.fieldViolations, QuotaFailure, RetryInfo, Help)}}; Interactions {error: {code: <snake_case>, message}}; soft failures at HTTP 200 (promptFeedback.blockReason, finishReason); Live WS close codes 1007/1008Status: DOCUMENTED · LIVE_VERIFIED18 error records; unknown File/Operation → 403, not 404 |
yes | Four envelopes; only Anthropic and OpenAI carry a request id in the body/header pair; Gemini has no request-id header at all. Ref: docs/errors/openai.md · docs/errors/anthropic.md · docs/xai/authentication-headers-errors.md · docs/errors/gemini.md |
| Idempotency key 0/4 |
not documented for api.openai.com (only Workspace Agents on api.chatgpt.com); dedupe via metadata/custom_idStatus: UNVERIFIED |
not documented; SDKs retry on 409 Status: UNVERIFIED |
not documented; public-url creation is idempotent by design; batch batch_request_id dedupes within a batchStatus: UNVERIFIED |
not documented; batch key per JSONL line; seed for determinismStatus: UNVERIFIED |
no | No provider offers request idempotency keys on the model APIs. Ref: headers.json |
| Per-request dollar cost in the response 1/4 |
token counts only (usage); costs via the Admin GET /v1/organization/costs reportStatus: DOCUMENTED |
token counts only; costs via /v1/organizations/cost_reportStatus: DOCUMENTED |
usage.cost_in_usd_ticks on every inference response (Chat, Responses, images, videos; 1 USD = 10^10 ticks; April 2026) — the effective price after cache, long-context, priority and regional multipliers; batch cost_breakdown (SDK/gRPC)Endpoint: POST /v1/responsesParams: usage.cost_in_usd_ticksStatus: DOCUMENTED · LIVE_VERIFIEDcatalogue endpoints return no usage/cost |
token counts only (usageMetadata, Interactions usage); costs in Google Cloud BillingEndpoint: POST /v1beta/models/{model}:generateContentStatus: DOCUMENTED |
no | xAI-only. Ref: docs/xai/pricing.md · docs/xai/responses.md |
Auth, versioning, SDKs, platform
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Authentication 4/4 |
Authorization: Bearer <sk-proj-…|sk-…|service-account key|WIF access token>; optional OpenAI-Organization, OpenAI-Project; Admin keys sk-admin-…Status: DOCUMENTED · LIVE_VERIFIEDWIF token exchange at auth.openai.com / mTLS |
x-api-key: sk-ant-api03-… or Authorization: Bearer <key|sk-ant-oat01-… OAuth/WIF token>; anthropic-workspace-id; Admin keys sk-ant-admin01-…Status: DOCUMENTED · LIVE_VERIFIEDWIF via POST /v1/oauth/token |
Authorization: Bearer xai-… on REST, WebSocket and gRPC metadata; separate Management key for management-api.x.ai (inference key → 401 code 16); ephemeral xai-realtime… client secrets for browsers (POST /v1/realtime/client_secrets); per-key ACLs api-key:endpoint:*, api-key:model:* (new keys have no access by default); mTLS host mtls.api.x.ai (enterprise)Status: DOCUMENTED · LIVE_VERIFIEDGET /v1/api-key, GET /v1/me introspection |
x-goog-api-key: AIza… (recommended; ?key= discouraged); Authorization: Bearer <GEMINI_API_KEY> mandatory on /v1beta/openai/*; OAuth/ADC Bearer alternative (+ x-goog-user-project); ephemeral tokens POST /v1beta/auth_tokens for the Live API; standard keys rejected from September 2026 in favour of service-account-bound auth keysStatus: DOCUMENTED · LIVE_VERIFIEDlimits are per project, not per key |
yes | Bearer everywhere except Gemini's native header; only OpenAI/Anthropic/xAI split admin vs inference keys; Gemini is the only one retiring a key type. Ref: headers.json · docs/openai/authentication-and-keys.md · docs/anthropic/admin-api.md §1 · docs/xai/authentication-headers-errors.md · docs/gemini/authentication-headers-versions.md |
| API version header / path version 2/4 |
none (server answers openai-version: 2020-10-01)Status: DOCUMENTED · LIVE_VERIFIED |
anthropic-version: 2023-06-01 required on every request (400 otherwise)Params: anthropic-versionStatus: DOCUMENTED · LIVE_VERIFIED |
no version header, no beta headers; features selected by body fields or base URL; Interactions-style Api-Revision does not existStatus: DOCUMENTED · LIVE_VERIFIED |
version in the URL path: /v1beta (86 methods, default for SDKs) vs /v1 (47; stable subset — no caching, tuning, Live, Files, agents); Interactions optional header Api-Revision: 2026-05-20Params: Api-RevisionStatus: DOCUMENTED · LIVE_VERIFIEDdocs say every model is in both versions; live /v1/models lists 22 vs 58 |
no | Header (Anthropic), path (Gemini), nothing (OpenAI, xAI). Ref: headers.json · docs/gemini/authentication-headers-versions.md |
| Beta opt-in header 2/4 |
OpenAI-Beta: agents=v1, chatkit_beta=v1, workspace_agent_runs=v1, responses_multi_agent=v1, legacy assistants=v2, realtime=v1; plus ?beta=true surface with body openai-beta[]Params: OpenAI-BetaStatus: DOCUMENTED400 invalid_beta when missing |
anthropic-beta: <feature>-<YYYY-MM-DD>[,…] (50 catalogued values; SDK betas=[…], client.beta.*); unknown → 400Params: anthropic-beta, betasStatus: DOCUMENTED · LIVE_VERIFIEDsee docs/faq.md for the current list |
none; alpha features answer 403 ('only available for alpha users' — tool_search) or 404 (ACL)Status: DOCUMENTED · LIVE_VERIFIED |
none; gating by /v1beta path and -preview / -exp model ids (kind: preview 45 records)Status: DOCUMENTED · LIVE_VERIFIED |
no | Anthropic gates parameters, OpenAI gates surfaces, Gemini gates by path/model id, xAI by account. Ref: generated/fragments/headers/anthropic-beta-headers.json · headers.json · docs/faq.md Q6 |
| Request correlation 3/4 |
response x-request-id; request X-Client-Request-Id (logged, not echoed); openai-processing-msStatus: DOCUMENTED · LIVE_VERIFIED |
response request-id (also request_id in error body); anthropic-organization-idStatus: DOCUMENTED · LIVE_VERIFIED |
response x-request-id (= chat.completion.id), Server-Timing (Cloudflare cfEdge/cfOrigin), CF-RAYStatus: DOCUMENTED · LIVE_VERIFIEDabsent on catalogue GETs |
no request-id header; responseId in the JSON body; Server-Timing: gfet4t7; dur=…Status: DOCUMENTED · LIVE_VERIFIED |
yes | Header on three providers, body field on Gemini. Ref: headers.json |
| Official SDKs 4/4 |
Python openai 3.16.2, Node openai 7.18/7.19, .NET, Java 4.65 (beta label), Go v3 (beta), Ruby, CLI, Agents SDK (py/ts), Azure libraries; retries 2, timeout 600 sStatus: DOCUMENTED · LIVE_VERIFIED |
Python anthropic 1.7.0, TS @anthropic-ai/sdk 0.126, Go, Java 2.63, Ruby, C# ≥10, PHP (beta), ant CLI 1.33; retries 2, timeout 10 min; cloud clients Bedrock/Vertex/AWS/FoundryStatus: DOCUMENTED · LIVE_VERIFIEDPython v1 removed sampling kwargs and completions |
Python xai-sdk 1.19 (gRPC, client.chat.create().sample()/stream()/defer()/parse(), response.cost_usd, timeout 1620 s) — no official Node SDK: use openai with baseURL: https://api.x.ai/v1, @ai-sdk/xai, langchain-xai, or the Anthropic SDK against /v1/messages (deprecated); Grok Build CLI (grok, BETA); protos xai-org/xai-proto for buf curl; gRPC api.x.ai:443 (11 services / 38 RPCs) lacks Responses items, voice, skillsStatus: DOCUMENTED · LIVE_VERIFIED9 SDK records |
Python google-genai 2.24 (genai.Client(); same SDK targets Vertex with vertexai=True), JS @google/genai 2.23, Go google.golang.org/genai, Java com.google.genai, C# Google.GenAI; Firebase AI Logic / Genkit / Vercel AI SDK; OpenAI SDK against /v1beta/openai; legacy google-generativeai, @google/generative-ai, Go/Dart/Swift/Android libs DEPRECATED 2025-11-30Status: DOCUMENTED · LIVE_VERIFIED8 SDK records; default api_version v1beta |
yes | Every provider has Python and Node coverage — xAI's Node path is the OpenAI SDK; xAI's own SDK is the only gRPC-first one. Ref: sdks.json · docs/xai/sdks.md · docs/gemini/sdks.md |
| Webhooks (platform events) 4/4 |
/v1/webhook_endpoints CRUD + rotate_secret, test; /v1/webhook_event_types; events batch., response., fine_tuning., eval.run., video., realtime/live incoming calls, safety., agent.session.*; Standard Webhooks signature (webhook-id/-timestamp/-signature, whsec_)Endpoint: GET /v1/webhook_endpointsStatus: DOCUMENTED · LIVE_VERIFIED28 event types |
Managed Agents only: endpoints registered in Console (no API), 44 event types (agent., session., deployment*., environment., vault*., memory_store.); same Standard Webhooks headers; ≤3 attempts, 5-min freshness; beta managed-agents-2026-04-01Status: DOCUMENTED · BETAno webhooks for Messages/Batches |
only the SIP voice webhook realtime.call.incoming (Standard Webhooks headers, HMAC-SHA256) registered via POST /v2/phone-numbers {webhook:{url}}; no batch/response webhooksEndpoint: POST /v2/phone-numbersStatus: DOCUMENTED |
POST/GET /v1/webhooks, GET|PATCH|DELETE /v1/webhooks/{id}, POST …/rotate_secret (documented on /v1, BETA) + per-request webhook_config {uris[], user_metadata} / Batch webhookConfig; events interaction.completed|failed|cancelled|requires_action, batch.succeeded|failed; JWKS-signedEndpoint: POST /v1/webhooksParams: webhook_configStatus: DOCUMENTED · BETA7 endpoints, not tested; :ping UNVERIFIED |
yes | OpenAI covers the widest event set; Gemini and Anthropic cover their agent/interaction resources; xAI only incoming calls. Ref: generated/webhook-events.json · docs/openai/webhooks.md · docs/anthropic/managed-agents.md §17.2 · docs/gemini/interactions-api.md |
| Administration API 3/4 |
124 operations under /v1/organization/* and /v1/projects/*: admin keys, users, invites, projects, service accounts, project API keys, groups, roles, certificates (mTLS), data retention, spend limits/alerts, audit logs, usage & costs; Admin API key sk-admin-…Endpoint: GET /v1/organization/*Status: DOCUMENTED · ACCOUNT_RESTRICTEDall probes 403/401 with a project key |
100 operations under /v1/organizations/*: me, users, invites, workspaces, api_keys, service accounts, federation (WIF), RBAC, spend limits, external keys, tunnels certs, usage_report & cost_report, analytics, Claude Code analytics; Admin key sk-ant-admin01-…Endpoint: GET /v1/organizations/*Status: DOCUMENTED · ACCOUNT_RESTRICTED · LIVE_VERIFIEDGET /v1/organizations/me works with a regular key |
Management API https://management-api.x.ai (21 endpoints, separate Management key, camelCase): API keys (create with acls[], qps/qpm/tpm, expireTime; rotate; delete; propagation), team models/endpoints ACLs, billing (billing-info, invoices, payment methods, postpaid spending limits, prepaid balance/top-up, POST …/usage aggregated by key/model/IP/cluster), audit events; teams/members/ZDR toggle are console-onlyEndpoint: GET /auth/teams/{teamId}/api-keysStatus: DOCUMENTED · ACCOUNT_RESTRICTEDinference key → 401 code 16 |
no admin API on the Developer API: keys/projects/quota live in AI Studio and Google Cloud IAM/Console; GET /v1beta/models is the only account-scoped readStatus: DOCUMENTED |
yes | Three admin surfaces (OpenAI, Anthropic, xAI — all inaccessible to this atlas's keys); Gemini delegates to Google Cloud. Ref: docs/openai/admin-api.md · docs/anthropic/admin-api.md · docs/xai/management-api.md |
| Usage & cost reporting 4/4 |
GET /v1/organization/usage/{completions,embeddings,images,…} + GET /v1/organization/costs (buckets 1m/1h/1d, group_by)Endpoint: GET /v1/organization/costsParams: start_time, bucket_width, group_byStatus: DOCUMENTED · ACCOUNT_RESTRICTED |
GET /v1/organizations/usage_report/messages, /usage_report/claude_code, /cost_report; analytics /analytics/{usage_report,cost_report,user_*}Endpoint: GET /v1/organizations/cost_reportParams: starting_at, bucket_width, group_byStatus: DOCUMENTED · ACCOUNT_RESTRICTED |
per-request usage.cost_in_usd_ticks (1 USD = 10^10 ticks) on every inference response + POST /v1/billing/teams/{team_id}/usage (Management API, aggregated by API key / model / IP / cluster / token type); batch cost_breakdown (SDK/gRPC)Endpoint: POST /v1/billing/teams/{team_id}/usageStatus: DOCUMENTED · ACCOUNT_RESTRICTED · LIVE_VERIFIEDcost ticks LIVE_VERIFIED; billing endpoint restricted |
no reporting endpoint; per-response usageMetadata (modality breakdown, thoughtsTokenCount, toolUsePromptTokenCount, cachedContentTokenCount, serviceTier) and Interactions usage {total_*_tokens, *_by_modality, grounding_tool_count}; billing dashboards in Google CloudEndpoint: POST /v1beta/models/{model}:generateContentParams: usageMetadataStatus: DOCUMENTED · LIVE_VERIFIED |
yes | Org-level reports on OpenAI/Anthropic/xAI; xAI is the only one returning the dollar cost per call; Gemini only per-response token detail. Ref: endpoints.json (admin, management) · docs/xai/pricing.md · docs/gemini/generate-content.md |
| Audit / compliance data access 3/4 |
GET /v1/organization/audit_logs (scope api.audit_logs.read); safety alerts/cases APIEndpoint: GET /v1/organization/audit_logsStatus: DOCUMENTED · ACCOUNT_RESTRICTED |
Compliance API (36 endpoints, Claude Enterprise Compliance Access Key): activities feed, chats, projects, files, sessions, roles, groups; DELETE operations Endpoint: GET /v1/compliance/activitiesStatus: DOCUMENTED · ACCOUNT_RESTRICTED |
GET /audit/teams/{teamId}/events (Management API; administrative events only)Endpoint: GET /audit/teams/{teamId}/eventsStatus: DOCUMENTED · ACCOUNT_RESTRICTED |
none on the Developer API (Cloud Audit Logs on Vertex AI) Status: DOCUMENTED |
yes | Admin-action logs on OpenAI/Anthropic/xAI; Anthropic additionally exports end-user content (Claude Enterprise). Ref: docs/anthropic/compliance-and-iam.md · docs/xai/management-api.md |
| Spend limits 3/4 |
org/project spend_limit, spend_alerts Admin endpoints; 429 insufficient_quota codes when hitEndpoint: POST /v1/organization/spend_limitStatus: DOCUMENTED |
/v1/organizations/spend_limits, /spend_limits/effective, increase requests approve/deny; tier cap → 429 without retry-after; self-set limit → 400Endpoint: POST /v1/organizations/spend_limitsStatus: DOCUMENTED · ACCOUNT_RESTRICTED |
GET/POST /v1/billing/teams/{team_id}/postpaid/spending-limits, prepaid balance / top-up (Management API); per-key qps/qpm/tpm capsEndpoint: POST /v1/billing/teams/{team_id}/postpaid/spending-limitsStatus: DOCUMENTED · ACCOUNT_RESTRICTED |
tier spend caps ($250 / $2,000 / $20k+) and rolling 10-minute spend limits ($10 / $50 / $200) are enforced by Google, not configurable via API (Cloud Billing budgets instead) Status: DOCUMENTED |
yes | API-settable on three providers; platform-imposed on Gemini. Ref: endpoints.json (admin) · rate-limits.json (gemini) |
| Cloud availability 4/4 |
Azure OpenAI / Microsoft Foundry (AzureOpenAI clients, Azure libraries); regional hosts us.|eu.|au.|jp.|in.api.openai.comStatus: DOCUMENTEDAzure surface not catalogued in this atlas |
Amazon Bedrock (Mantle bedrock-mantle.{region}.api.aws/anthropic/v1/messages and legacy InvokeModel), Google Cloud Vertex AI (rawPredict), Microsoft Foundry (/anthropic/v1/*), Claude Platform on AWS; feature gaps per platform (no Batches/Files/Skills/server tools on Bedrock/Vertex)Status: DOCUMENTED · UNVERIFIED24 cloud endpoint records |
Google Cloud Vertex AI Model Garden (partner model) and Microsoft Foundry (Azure) — OpenAI-compatible chat + Responses, cloud billing; also OpenRouter, Vercel AI Gateway, Cloudflare, Cursor; regional first-party endpoints https://us.api.x.ai/v1 (US-pinned, grok-4.6, 1.1×) and https://eu-west-1.api.x.ai/v1 (undocumented, grok-4.3, LIVE_DISCOVERED); clusters us-east-1, us-west-2, us-central-1, eu-west-1, us-saltlake-2Status: DOCUMENTED · LIVE_DISCOVERED |
Gemini Developer API (generativelanguage.googleapis.com, global, API key) vs Vertex AI / Gemini Enterprise Agent Platform ({location}-aiplatform.googleapis.com, IAM only, ~40 regions, data residency, ZDR, CMEK, VPC-SC, provisioned throughput, supervised tuning, RAG/Agent Engine); one SDK, two backends; Developer-API-only: Interactions, Live, Files, File Search, agents, free tierStatus: DOCUMENTEDVertex surface not catalogued in this atlas |
yes | Claude and Grok are resold on other clouds; OpenAI and Gemini have first-party cloud twins (Azure/Foundry, Vertex). Ref: docs/anthropic/cloud-providers.md · docs/openai/data-residency-and-regions.md · docs/xai/sdks.md · docs/gemini/vertex-vs-gemini-api.md |
| Data residency 3/4 |
project region at creation; regional hosts us./eu./au./jp./in.api.openai.com; 10 % uplift for models ≥2026-03-05; EU lacks background:trueStatus: DOCUMENTED · ACCOUNT_RESTRICTED |
inference_geo: us|global per request (Claude 4.6+; 1.1× multiplier for us); workspace default_inference_geo; usage.inference_geo echo; Bedrock/Vertex regional endpoints +10 %Endpoint: POST /v1/messagesParams: inference_geoStatus: DOCUMENTED · FAILED_VERIFICATION400 on Haiku 4.5 live |
base-URL choice: https://us.api.x.ai/v1 keeps request handling, inference, moderation and retained data in the US at 1.1× token prices (grok-4.6 only; no image/video/voice); global api.x.ai gives no region guaranteeStatus: DOCUMENTED · LIVE_VERIFIEDeu-west-1.api.x.ai serves grok-4.3 at global prices (undocumented) |
no residency option on the Developer API (global endpoint); residency, CMEK and ~40 regions are Vertex AI features Status: DOCUMENTED |
yes | Regional pinning costs ~10 % on OpenAI, Anthropic and xAI alike; Gemini requires moving to Vertex. Ref: docs/anthropic/regions-and-ips.md · docs/openai/data-residency-and-regions.md · docs/xai/pricing.md · docs/gemini/vertex-vs-gemini-api.md |
| Zero data retention / data-use terms 4/4 |
ZDR/Modified retention by approval; store:false keeps reasoning encrypted; Live recordings and Agents sessions excludedParams: storeStatus: DOCUMENTED |
Messages stateless → ZDR-eligible on Claude API (per model flag zero_data_retention_eligible); not eligible: Files API, MCP connector, programmatic tool calling, Managed Agents, Fable 5.1 / Mythos 5.1Status: DOCUMENTED |
team-wide ZDR toggle (console): response header x-zero-data-retention: true|false, GET /v1/me → zdr_status: no_zdr|zdr; disables store, previous_response_id, Files, Collections, Batch, deferred completions, stored media; default retention 30 days, not used for training; encrypted reasoning replay keeps ZDR + cachingStatus: DOCUMENTED · LIVE_VERIFIEDour team: no_zdr |
Unpaid Services (free tier): prompts/responses may be used to improve Google products, with human review; Paid Services: not used for training, 55-day abuse logging; ZDR not achievable on the Developer API (grounding stores 30 days; Interactions state unless store:false) — Vertex AI offers ZDR; EEA/UK/CH end-user apps must use Paid ServicesParams: storeStatus: DOCUMENTED |
yes | ZDR is a program (OpenAI), a per-model eligibility (Anthropic), a team switch (xAI) and unavailable on Gemini's Developer API. Ref: models.json (anthropic capabilities.zero_data_retention_eligible) · docs/xai/management-api.md · docs/gemini/authentication-headers-versions.md |
| OpenAI-compatibility layer 3/4 |
the native surface (Responses + Chat Completions) Endpoint: POST /v1/chat/completionsStatus: DOCUMENTED · LIVE_VERIFIED |
none — Anthropic exposes only its own Messages format (partners such as xAI implement it) Status: DOCUMENTED |
native: the whole inference API is OpenAI-shaped (/v1/chat/completions, /v1/responses incl. items and SSE events, /v1/batches, /v1/files, /v1/images/*, /v1/realtime events); rejected/ignored OpenAI fields: background, metadata (Responses 400), logit_bias 400, stop/penalties on reasoning models 400, logprobs ignored, store ignored on Chat, non-function tools 422, images.edit() multipart unsupported; extras: reasoning_effort, reasoning_content, deferred, max_turns, top_k, min_p, cost_in_usd_ticksEndpoint: POST /v1/responsesStatus: DOCUMENTED · LIVE_VERIFIEDplus an Anthropic-compatible /v1/messages (deprecated) |
/v1beta/openai/* (BETA): chat/completions (LIVE_VERIFIED), embeddings, models[/{id}], images/generations (subset), videos (Sora-style Veo), batches; Bearer key required; reasoning_effort → thinkingLevel mapping; extra_body.google {thinking_config, cached_content, safety_settings, tools[{google_search}]}; message.extra_content.google.thought_signature; unknown params silently ignored; no Responses, Assistants, audio, files, fine-tuning, moderationsEndpoint: POST /v1beta/openai/chat/completionsParams: extra_body.google.thinking_config, extra_body.google.cached_content, reasoning_effortStatus: DOCUMENTED · BETA · LIVE_VERIFIED7 endpoints; google.rpc error envelope |
yes | xAI is OpenAI-shaped by design; Gemini offers a partial adapter; Anthropic none. Ref: docs/xai/chat-completions.md · docs/xai/responses.md · docs/gemini/openai-compatibility.md |
Managed agents platforms
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Managed agent harness 4/4 |
Agents API (beta OpenAI-Beta: agents=v1): POST /v1/agents (model, instructions, reasoning, text, tools, multi_agent), sessions, turns, items, events, environments, templates, vaults; Codex harness; 34 + 9 endpointsEndpoint: POST /v1/agentsParams: model, instructions, tools, multi_agentStatus: DOCUMENTED · BETA · LIVE_VERIFIEDgpt-5.4-nano rejected; gpt-5.6-luna accepted |
Claude Managed Agents (beta managed-agents-2026-04-01): POST /v1/agents (name, model, system, tools, mcp_servers, skills, multiagent) with versions, /v1/sessions + events/stream, environments, deployments (cron), vaults, memory stores, dreams, tunnels, user profiles; 96 endpointsEndpoint: POST /v1/agentsParams: name, model, system, tools, mcp_servers, skills, multiagentStatus: DOCUMENTED · BETA · LIVE_VERIFIEDClaude 4.5+ models |
no agent resource: the Responses API agentic loop (server-side tools iterate inside one request, bounded by max_turns; stored responses + previous_response_id = the session); grok-4.20-multi-agent-0309 (BETA) as a built-in multi-agent model; Grok Build (grok-build-0.1 PREVIEW model + grok CLI BETA: TUI/headless/ACP, MCP, hooks, skills, subagents, sandbox) is a client-side harnessEndpoint: POST /v1/responsesParams: tools, max_turns, store, previous_response_idStatus: DOCUMENTED · LIVE_VERIFIED · BETASkills API 404 for this team |
Interactions API agents (agent instead of model): Deep Research (deep-research-preview-04-2026, -max-, agent_config {collaborative_planning, thinking_summaries, visualization}, background only, ≤60 min), Antigravity coding agent (antigravity-preview-09-2026, Linux sandbox 4 vCPU/16 GB, agent_config {model, max_total_tokens}, hooks), custom managed agents POST /v1beta/agents {base_agent, agent_config, system_instruction, tools, base_environment} (≤1,000/project, no versioning); environments, credentials, triggers (cron), webhooksEndpoint: POST /v1beta/interactionsParams: agent, agent_config, environment, backgroundStatus: DOCUMENTED · BETA · PREVIEW · LIVE_VERIFIEDDeep Research LIVE_VERIFIED (create → in_progress → cancel); custom agents / triggers / credentials not tested; sandbox compute unbilled in preview |
yes | Four different shapes: agent → session → events (OpenAI, Anthropic), a stateful request loop (xAI), agent-as-model inside Interactions (Gemini). Ref: docs/comparisons/agents-platforms.md · docs/xai/responses.md · docs/xai/grok-build.md · docs/gemini/interactions-api.md |
| Create an agent session / run 4/4 |
POST /v1/agents/sessions {agent|agent_id, environment (required: none|openai_hosted|self_hosted), input, stream, vault_ids} → 201 agent.sessionEndpoint: POST /v1/agents/sessionsParams: agent_id, environment, input, streamStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
POST /v1/sessions {agent, environment_id (required), initial_events[], resources[], vault_ids[], budget, title} → 200 sessionEndpoint: POST /v1/sessionsParams: agent, environment_id, initial_events, budgetStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
POST /v1/responses {model, input, tools[], max_turns, store} → the stored response is the session; continue with previous_response_idEndpoint: POST /v1/responsesParams: model, input, tools, max_turnsStatus: DOCUMENTED · LIVE_VERIFIED |
POST /v1beta/interactions {agent|model, input, background, store, environment, tools, webhook_config} → interaction {id, status, steps[]}; chain with previous_interaction_idEndpoint: POST /v1beta/interactionsParams: agent, input, background, previous_interaction_idStatus: DOCUMENTED · BETA · LIVE_VERIFIED/v1/interactions GA UNVERIFIED |
yes | See FAQ Q5 for the body-level comparison. Ref: endpoints.json · docs/faq.md Q5 |
| Send input / events to a session 4/4 |
POST /v1/agents/sessions/{id}/events with agent.session.input.message|tool_result|cancel → 202Endpoint: POST /v1/agents/sessions/{session_id}/eventsStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
POST /v1/sessions/{id}/events with user.message|interrupt|tool_confirmation|custom_tool_result|tool_result|define_outcome, system.messageEndpoint: POST /v1/sessions/{session_id}/eventsStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
a new POST /v1/responses with previous_response_id (user message or function_call_output / shell_call_output items)Endpoint: POST /v1/responsesParams: previous_response_id, input[](function_call_output)Status: DOCUMENTED · LIVE_VERIFIED |
a new POST /v1beta/interactions with previous_interaction_id and input (user text or function_result steps); POST …/cancel to stopEndpoint: POST /v1beta/interactionsParams: previous_interaction_id, inputStatus: DOCUMENTED · BETA · LIVE_VERIFIEDinteraction.requires_action for function calls |
yes | Dedicated event endpoints (OpenAI, Anthropic) vs chained requests (xAI, Gemini). Ref: streaming-events.json (client→server) |
| Session event stream 4/4 |
SSE via POST /sessions {stream:true} or GET …/events?stream=true; 31 event types agent.session.*, agent.output.*Endpoint: GET /v1/agents/sessions/{session_id}/eventsStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
GET /v1/sessions/{id}/events/stream (+ per-thread /threads/{id}/stream); 37 event types agent.*, session.*, span.*, event_start/deltaEndpoint: GET /v1/sessions/{session_id}/events/streamStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
the ordinary Responses SSE stream (24 recorded events incl. response.code_interpreter_call.*, response.web_search_call.*, response.mcp_call.*) or wss://api.x.ai/v1/responsesEndpoint: POST /v1/responsesParams: streamStatus: DOCUMENTED · LIVE_VERIFIED |
Interactions SSE (stream:true): interaction.created → interaction.status_update → step.start → step.delta {text|thought_signature|arguments_delta|function_result|*_call…} → step.stop → interaction.completed → done [DONE]; resumable with GET …?stream=true&last_event_id=; 15 recorded events (5 legacy names RETIRED 2026-06-08)Endpoint: POST /v1beta/interactionsParams: stream, last_event_idStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
yes | Ref: streaming-events.json |
| Hosted sandbox 3/4 |
environment.type: openai_hosted {packages, setup_commands, network, env, skills, plugins, files, environment_template_id}; /workspace; artifacts from /workspace/outputs; ~1 h idle expiry; container ratesEndpoint: POST /v1/agents/sessionsParams: environmentStatus: DOCUMENTED · BETA |
POST /v1/environments {type: cloud, …} (Ubuntu 24.04, 8 GB RAM, 10 GB disk, /workspace, /mnt/session/{uploads,outputs}, /mnt/memory); state kept 30 days; $0.08/session-hourEndpoint: POST /v1/environmentsStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
no hosted sandbox for agents (the code_interpreter tool runs Python without a container object; Grok Build sandboxes locally)Endpoint: POST /v1/responsesStatus: DOCUMENTED |
POST /v1beta/environments {sources[] (repository ≤500 MB | gcs ≤2 GB | inline ≤1 MB/file), network (unrestricted|disabled|allowlist with credentials), from_environment}; Antigravity sandbox 4 vCPU / 16 GB (Python 3.12, Node 22); files GET …/files/{path}, PUT /upload/…/files/{path}; idle after 15 min, deleted after 7 days; compute unbilled during previewEndpoint: POST /v1beta/environmentsParams: sources, networkStatus: DOCUMENTED · BETA · PREVIEW · LIVE_VERIFIEDlive GET showed storage.tier: free, 1 GiB project limit |
yes | Three hosted sandboxes (OpenAI, Anthropic, Gemini); none on xAI. Ref: docs/openai/agents-environments-and-vaults.md · docs/anthropic/managed-agents.md §8 · docs/gemini/interactions-api.md |
| Self-hosted execution 3/4 |
environment.type: self_hosted {workspace_directory} + codex exec-server --remote … executor with environment key CODEX_API_KEY; outbound onlyEndpoint: POST /v1/agents/sessionsParams: environmentStatus: DOCUMENTED · BETA |
type: self_hosted environment = work queue (/environments/{id}/work, poll/ack/heartbeat/stop) + worker (ant beta:worker, SDK EnvironmentWorker) with environment key sk-ant-oat01-…Endpoint: GET /v1/environments/{environment_id}/work/pollStatus: DOCUMENTED · BETA |
client-side by construction: shell {environment: local} and function tools end the server loop and hand control back; Grok Build CLI runs everything locallyEndpoint: POST /v1/responsesParams: tools[type=shell].environmentStatus: DOCUMENTED · LIVE_VERIFIED |
no self-hosted worker protocol; function calls (interaction.requires_action) are the only client-side hookEndpoint: POST /v1beta/interactionsStatus: DOCUMENTED |
yes | Worker protocols on OpenAI/Anthropic; plain tool round-trips on xAI/Gemini. Ref: endpoints.json (managed-agents) |
| Credential vaults / secrets 3/4 |
/v1/vaults, /vaults/{id}/credentials (rotate/delete); referenced by vault_ids[] and mcp.credential_idEndpoint: POST /v1/vaultsStatus: DOCUMENTED · BETA · LIVE_VERIFIED9 endpoints |
/v1/vaults, /vaults/{id}/credentials (mcp_oauth with refresh, environment_variable), mcp_oauth_validate; vault_credential.refresh_failed webhookEndpoint: POST /v1/vaultsStatus: DOCUMENTED · BETA · LIVE_VERIFIED13 endpoints |
pass authorization / headers on the mcp tool per request; no vaultEndpoint: POST /v1/responsesParams: tools[type=mcp].authorizationStatus: DOCUMENTED |
/v1beta/credentials (bearer_token, oauth2, environment_variable with injection_location, trusted_domains); write-only secrets; referenced from environment network allowlistsEndpoint: POST /v1beta/credentialsStatus: DOCUMENTED · BETA · PREVIEW5 endpoints, not tested |
yes | Three secret stores; xAI inlines credentials. Ref: endpoints.json |
| Multi-agent / subagents 3/4 |
multi_agent {enabled, max_concurrent_subagents (6)}; /sessions/{id}/subagents[/{id}/items|turns]; items create_subagent_call, wait_for_subagents_call…; also Responses beta responses_multi_agent=v1Endpoint: GET /v1/agents/sessions/{session_id}/subagentsParams: multi_agentStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
multiagent {type: coordinator, agents[] (≤20, incl. self and one advisor)}; threads (/sessions/{id}/threads, ≤25 concurrent); events agent.thread_message_*Endpoint: GET /v1/sessions/{session_id}/threadsParams: multiagentStatus: DOCUMENTED · BETA |
grok-4.20-multi-agent-0309 (BETA): the model itself fans out to 4 or 16 agents per reasoning.effort; no subagent resources; Grok Build CLI has client-side subagentsEndpoint: POST /v1/responsesParams: reasoning.effortStatus: DOCUMENTED · BETA · LIVE_VERIFIEDlive output carried \confidence{80} markers |
no sub-agents or delegation for custom agents (docs); Deep Research orchestrates internally Endpoint: POST /v1beta/interactionsStatus: DOCUMENTED |
yes | Orchestration APIs on OpenAI/Anthropic; a multi-agent model on xAI; internal only on Gemini. Ref: docs/comparisons/agents-platforms.md · docs/models/xai-models.md |
| Session budgets 3/4 |
no documented budget parameter; turn error code session_budget_exceeded existsStatus: DOCUMENTED · UNVERIFIED |
budget {type: limit, max_list_cost {amount (cents), currency: USD}} enforced between model requests; session.budget_reached webhookEndpoint: POST /v1/sessionsParams: budgetStatus: DOCUMENTED · BETA |
max_turns per Responses request (turn budget)Endpoint: POST /v1/responsesParams: max_turnsStatus: DOCUMENTED · LIVE_VERIFIED |
Antigravity agent_config.max_total_tokens → status: incompleteEndpoint: POST /v1beta/interactionsParams: agent_config(antigravity).max_total_tokensStatus: DOCUMENTED · PREVIEW |
no | Money (Anthropic), turns (xAI), tokens (Gemini) — not interchangeable. Ref: docs/anthropic/managed-agents.md §13 · docs/xai/responses.md · docs/gemini/interactions-api.md |
| Outcome grading (define outcome + rubric) 1/4 |
— not offered | user.define_outcome event → grader loop (max_iterations ≤20), span.outcome_evaluation_* events, outcome_evaluations[] on the sessionEndpoint: POST /v1/sessions/{session_id}/eventsStatus: DOCUMENTED · BETA |
— not offered | — not offered | no | Anthropic-only. Ref: docs/anthropic/managed-agents.md §13 |
| Server-side memory stores 1/4 |
— not offered | /v1/memory_stores (+ memories, versions, redact); mounted at /mnt/memory; beta agent-memory-2026-07-22; Dreams (/v1/dreams, dreaming-2026-04-21) reorganise themEndpoint: POST /v1/memory_storesStatus: DOCUMENTED · BETA · LIVE_VERIFIED14 + 5 endpoints |
— not offered | — not offered | no | Anthropic-only. Ref: endpoints.json |
| Scheduled runs 2/4 |
— not offered | /v1/deployments (cron) → /v1/deployment_runs; pause/unpause/run nowEndpoint: POST /v1/deploymentsStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
— not offered | /v1beta/triggers {schedule (cron), time_zone, display_name, max_consecutive_failures, execution_timeout_seconds, interaction}; PATCH {status: paused|active}; POST|GET …/executionsEndpoint: POST /v1beta/triggersParams: schedule, time_zone, interactionStatus: DOCUMENTED · BETA · PREVIEW7 endpoints, not tested |
yes | Anthropic deployments ≈ Gemini triggers. Ref: endpoints.json |
| Artifacts / deliverables 3/4 |
/sessions/{id}/artifacts[/{id}/content] from /workspace/outputs (≤200 MiB each)Endpoint: GET /v1/agents/sessions/{session_id}/artifactsStatus: DOCUMENTED · BETA · LIVE_VERIFIED |
files written to /mnt/session/outputs → GET /v1/files?scope_id=<session> + /contentEndpoint: GET /v1/filesStatus: DOCUMENTED · BETA |
tool outputs (code_interpreter_call.outputs, image/video file_output) via include or the Files API; no artifact resourceEndpoint: POST /v1/responsesParams: includeStatus: DOCUMENTED |
environment files: GET /v1beta/environments/{env}/files/{path}?alt=media (file bytes or tar), legacy GET /v1beta/files/environment-{env}:download; Deep Research reports as model_output steps (text/image)Endpoint: GET /v1beta/environments/{environment}/files/{path}Status: DOCUMENTED · BETA · PREVIEW · LIVE_VERIFIEDpartial 200 live |
yes | Ref: docs/openai/agents-environments-and-vaults.md §2 · docs/gemini/interactions-api.md |
| Embeddable chat UI & workspace agents 1/4 |
ChatKit (OpenAI-Beta: chatkit_beta=v1, /v1/chatkit/sessions|threads) and Workspace Agents (api.chatgpt.com/v1/workspace_agents/{id}/trigger)Endpoint: POST /v1/chatkit/sessionsStatus: DOCUMENTED · BETA · LIVE_VERIFIEDAgent Builder shuts down 2026-11-30 |
— not offered | Grok Apps / Grok Bot integrations are consumer products, not an embeddable API (docs/xai/grok-apps-and-integrations.md) Status: DOCUMENTED |
AI Studio and Firebase AI Logic are builder/SDK surfaces, not an embeddable chat API Status: DOCUMENTED |
no | OpenAI-only. Ref: docs/openai/chatkit.md · docs/openai/workspace-agents.md |
| Client-side agent framework / CLI 4/4 |
Agents SDK (openai-agents, @openai/agents) — loop over Responses API; handoffs, guardrails, tracingStatus: DOCUMENTED |
Claude Agent SDK / ant CLI (ant beta:sessions, ant apply); SDK tool_runner helpersStatus: DOCUMENTEDClaude Agent SDK not catalogued in this atlas beyond the migration mapping |
Grok Build CLI (curl -fsSL https://x.ai/cli/install.sh | bash; grok, grok -p … --output-format json, grok agent stdio (ACP); ~/.grok/config.toml api_backend = chat_completions|responses|messages; MCP servers, hooks, skills, plugins, subagents, Landlock/Seatbelt sandbox; reads CLAUDE.md/AGENTS.md) — BETA; xai-sdk tool_runner-style helpers absentStatus: DOCUMENTED · BETA |
google-genai SDK automatic function calling (Python callables, maximum_remote_calls 10) and mcpToTool(); Genkit, Firebase AI Logic, Vercel AI SDK, LangGraph/CrewAI/LlamaIndex integrations; Antigravity is the hosted coding agentStatus: DOCUMENTED · BETA |
yes | Ref: docs/openai/agents-sdk.md · docs/anthropic/cli.md · docs/xai/grok-build.md · docs/gemini/sdks.md |
Model customisation & evaluation
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Fine-tuning / tuning 1/4 |
/v1/fine_tuning/jobs (SFT, DPO, RFT, vision) on gpt-4.1*, gpt-4o*, o4-mini — DEPRECATED: no new jobs after 2027-01-06; our key 403 training_not_availableEndpoint: POST /v1/fine_tuning/jobsParams: model, training_file, methodStatus: DOCUMENTED · DEPRECATED · ACCOUNT_RESTRICTEDno GPT-5.x/6 fine-tuning |
— not offered | none offered by xAI (capability matrix row 'fine-tuning: none') Status: DOCUMENTED |
tunedModels.* (12 endpoints) still in the v1beta discovery document but RETIRED on the Developer API since the Gemini 1.5 Flash-001 deprecation (May 2025): create/list/get → 501 UNIMPLEMENTED; no tunable 3.x/2.x model; use Vertex AI supervised tuningEndpoint: POST /v1beta/tunedModelsStatus: DOCUMENTED · DEPRECATED · RETIREDGemma 4 tuning: not available |
no | Only OpenAI still accepts jobs, and only until 2027-01-06. Ref: docs/openai/fine-tuning.md · docs/gemini/tuning.md · docs/models/xai-models.md |
| Evals 1/4 |
/v1/evals, runs, output items; graders — DEPRECATED: read-only 2026-10-31, shutdown 2026-11-30Endpoint: POST /v1/evalsStatus: DOCUMENTED · DEPRECATED · LIVE_VERIFIED12 endpoints |
— not offered | — not offered | no evals API on the Developer API (Vertex AI evaluation service) Status: DOCUMENTED |
no | OpenAI-only (deprecated). Ref: docs/openai/evals.md |
| Graders (standalone) 1/4 |
/v1/fine_tuning/alpha/graders/{run,validate}Endpoint: POST /v1/fine_tuning/alpha/graders/runStatus: DOCUMENTED · DEPRECATED · BETA · LIVE_VERIFIED |
— not offered | — not offered | — not offered | no | OpenAI-only. Ref: docs/openai/graders.md |
| Stored completions / distillation 1/4 |
Chat store:true + metadata → list/retrieve/update/delete stored completions; distillation via SFTEndpoint: GET /v1/chat/completionsParams: store, metadataStatus: DOCUMENTED · LIVE_VERIFIED |
— not offered | Chat store/metadata silently accepted; no stored-completion retrieval endpoints (Responses store instead)Endpoint: POST /v1/chat/completionsParams: storeStatus: DOCUMENTED · LIVE_DISCOVERED |
— not offered | no | OpenAI-only. Ref: docs/openai/chat-completions.md · parameters.json (xai store) |
| Reusable prompt templates 1/4 |
prompt {id, version, variables} on Responses — DEPRECATED, shutdown 2026-11-30Endpoint: POST /v1/responsesParams: promptStatus: DOCUMENTED · DEPRECATED |
— not offered | — not offered | no prompt registry on the API (AI Studio saves prompts client-side); cachedContents reuse a prefixStatus: DOCUMENTED |
no | OpenAI-only (deprecated). Ref: docs/openai/deprecations.md |
Legacy / retired surfaces
| Feature | OpenAI (how / endpoint / params / status) | Anthropic | xAI | Gemini | Portable | Notes on differences |
|---|---|---|---|---|---|---|
| Retired agent / thread APIs 0/4 |
Assistants API RETIRED 2026-08-26 — 404 empty body; 23 operations mapped to Conversations/Responses/Agents Endpoint: GET /v1/assistantsStatus: RETIRED · DOCUMENTEDOpenAI-Beta: assistants=v2 historical |
— not offered | — not offered | Interactions v1beta legacy schema (outputs → steps, old SSE names interaction.start, content.*) removed 2026-06-08; total_reasoning_tokens → total_thought_tokensEndpoint: POST /v1beta/interactionsStatus: RETIRED · DOCUMENTED |
no | Ref: docs/openai/assistants-retired.md · docs/gemini/interactions-api.md |
| Retired media models / endpoints 0/4 |
DALL·E 2/3 RETIRED 2026-05-12; /v1/images/variations 404; Sora 2 + Videos API shut down 2026-09-24Endpoint: POST /v1/images/variationsStatus: RETIRED · DEPRECATED |
— not offered | grok-2-image(-1212) RETIRED (404); grok-imagine-image-pro → -quality (DEPRECATED, retires 2026-11-02 → grok-imagine-image-2.0 low); Live Search 410Endpoint: POST /v1/images/generationsStatus: RETIRED · DEPRECATED |
Imagen 3/4 (:predict) RETIRED 2026-08-17; Veo 2.0/3.0 RETIRED 2026-06-30; gemini-2.5-flash-image DEPRECATED → 2026-10-02; gemini-omni-flash-preview → 2026-09-30; half-cascade Live models RETIRED 2025-12-09Endpoint: POST /v1beta/models/{model}:predictStatus: RETIRED · DEPRECATED |
no | Ref: docs/openai/images.md · docs/xai/deprecations-and-release-notes.md · docs/gemini/deprecations-and-changelog.md |
| Retired text models 0/4 |
gpt-3.5-turbo-instruct/babbage/davinci shut down 2026-09-28; o1/o3-mini/o4-mini/gpt-4/gpt-4-turbo 2026-10-23; gpt-5 2025 snapshots, o3 2026-12-11; codex ≤5.2, deep-research, computer-use-preview retired Status: DEPRECATED · RETIRED |
RETIRED on the Claude API: Opus 4.1 (2026-08-05), Sonnet 4 / Opus 4 (2026-06-15), Claude 3.x, 2.x, 1.x, Instant; some still served on Bedrock/Vertex Endpoint: POST /v1/messagesStatus: RETIRED404 not_found_error |
RETIRED 2026-05-15 with redirects: grok-3, grok-4-0709, grok-4-fast-*, grok-4-1-fast-* → grok-4.3 (billed at 4.3 rates; GET /v1/models/grok-3 returns the grok-4.3 object), grok-code-fast-1 → grok-build-0.1; LEGACY/UNVERIFIED: grok-2-*, grok-3-mini, grok-beta, grok-vision-beta, grok-4-latestEndpoint: GET /v1/models/{model_id}Status: RETIRED · LEGACYno consolidated deprecations page; ~60-day notice observed |
Gemini 2.0 Flash/-Lite RETIRED 2026-06-01; 2.5 previews 2025-11/2026-03; gemini-3-pro-preview 2026-03-09 (id repointed to 3.1 Pro), gemini-3.1-flash-lite-preview 2026-05-25; Gemini 2.5 Pro/Flash/Flash-Lite 'no longer available to new users' (404, undocumented); gemini-3.1-flash-lite DEPRECATED → 2027-05-07; gemini-embedding-001 → 2028-05-14; text-embedding-004 2026-01-14Endpoint: POST /v1beta/models/{model}:generateContentStatus: RETIRED · DEPRECATED · ACCOUNT_RESTRICTEDshut-down ids still appear in GET /v1beta/models |
no | Retirement means 404 on Anthropic/Gemini, a dated shutdown on OpenAI, and a silent redirect on xAI. Ref: deprecations.json · models.json |
| Retired / deprecated beta headers, parameters and SDKs 4/4 |
OpenAI-Beta: realtime=v1 (legacy /v1/realtime/sessions → 404), assistants=v2; prompt_cache_retention, truncation: auto, user (→ safety_identifier) legacyStatus: LEGACY · FAILED_VERIFICATION |
RETIRED: context-1m-2025-08-07, computer-use-2024-10-22, max-tokens-3-5-sonnet-2024-07-15; DEPRECATED: mcp-client-2025-04-04; LEGACY (GA, header optional): prompt-caching, message-batches, pdfs, token-counting, files-api, skills, structured-outputs, extended-cache-ttl, code-execution-2025-05-22, interleaved/fine-grained streaming, effort-2025-11-24…Status: RETIRED · DEPRECATED · LEGACY |
no headers to retire; DEPRECATED params/features: max_tokens (→ max_completion_tokens), logprobs/top_logprobs (ignored ≥4.20), Anthropic-compatible /v1/messages, x_search per-call billing (→ 2026-09-21), zdr_status: pii_scrubbing, Management teamId (→ scope/scopeId); LEGACY: /v1/completions, /v1/complete, Live SearchStatus: DEPRECATED · LEGACY · RETIRED |
DEPRECATED: temperature/topP/topK guidance (2026-07-21), thinkingBudget (LEGACY on 3.x), HARM_CATEGORY_CIVIC_INTEGRITY (→ enableEnhancedCivicAnswers), googleSearchRetrieval, standard API keys (rejected from Sept 2026), legacy SDKs google-generativeai / @google/generative-ai (2025-11-30); LEGACY: PaLM methods, responseSchema/responseJsonSchema (→ responseFormat), mediaChunks, speechStateStatus: DEPRECATED · LEGACY |
no | Ref: generated/fragments/headers/anthropic-beta-headers.json · deprecations.json (xai, gemini api_features) |
| Retired model ids keep resolving (redirect aliases) 1/4 |
retired ids fail with model_not_found; shutdown_date exposed on GET /v1/modelsEndpoint: GET /v1/models/{model}Status: DOCUMENTED · LIVE_VERIFIED |
retired ids → 404 not_found_error (some still served on Bedrock/Vertex)Endpoint: POST /v1/messagesStatus: DOCUMENTED |
retired slugs redirect to their replacement and are billed at the replacement's price: grok-3, grok-4-0709, grok-4-fast-*, grok-4-1-fast-* → grok-4.3 (verified: GET /v1/models/grok-3 returns the grok-4.3 object), grok-code-fast-1 → grok-build-0.1, grok-imagine-image-pro → -quality → -2.0 low; response.model reveals the targetEndpoint: GET /v1/models/{model_id}Status: DOCUMENTED · RETIRED · LIVE_VERIFIED6 retired_redirect records in models.json |
shut-down ids stay listed in GET /v1beta/models but generation fails (404); one documented repoint: gemini-3-pro-preview → gemini-3.1-pro-preview (2026-03-09); -latest aliases hot-swap targetsEndpoint: GET /v1beta/modelsStatus: DOCUMENTED · LIVE_DISCOVERED |
no | xAI-only behaviour (silent redirect); Gemini has one documented id repoint. Ref: deprecations.json (xai) · docs/xai/deprecations-and-release-notes.md |
Reading the matrix programmatically
# all features unique to xAI
jq '[.records[] | select(.providers_supporting == ["xai"]) | .feature]' generated/compatibility/cross-provider-feature-matrix.json
# features on all four providers that are marked portable
jq '[.records[] | select(.provider_count == 4 and .portable) | .feature]' generated/compatibility/cross-provider-feature-matrix.json
# every Gemini cell that is ACCOUNT_RESTRICTED (paid-tier feature probed with a free key)
jq '[.records[] | select(.gemini.status|index("ACCOUNT_RESTRICTED")) | {feature, endpoint: .gemini.endpoint}]' generated/compatibility/cross-provider-feature-matrix.json
# every beta-gated Anthropic cell
jq '[.records[] | select(.anthropic.status|index("BETA")) | {feature, endpoint: .anthropic.endpoint}]' generated/compatibility/cross-provider-feature-matrix.jsonRelated pages: index · models · state management · tool execution · streaming · agents platforms · pricing · caching and reasoning · realtime and media · endpoint catalogue · FAQ.