FAQ — answering the owner's questions from the atlas data (4 providers)
Status: every table on this page is computed from generated/*.json on 2026-09-18 (the jq/python used is shown so the answer can be re-derived after the next build); generator scripts/generators/synth/faq.py. Statuses are the record statuses. Where a question needs interpretation, the interpretation is stated. Providers: openai, anthropic, xai, gemini.
Sources: generated/models.json, generated/tools.json, generated/parameters.json, generated/endpoints.json, generated/streaming-events.json, generated/webhook-events.json, generated/pricing.json, generated/fragments/headers/anthropic-beta-headers.json, generated/examples-manifest.json, plus the docs pages linked per answer.
Last verified: 2026-09-18
Questions: 1 tool X + structured outputs + caching (Anthropic) · 2 image + reasoning + web search + streaming (OpenAI) · 2b the same combination on all four providers · 3 exact JSON for feature Y · 4 SSE / WebSocket events per call type · 5 endpoint that creates an agent session · 6 features needing a beta header / beta path · 7 which model supports combination Z · 8 which provider offers X and at what price · 9 where things live in the atlas
Q1. Which Anthropic model accepts tool X with structured outputs and prompt caching?
Interpretation: model records where capabilities.structured_outputs_json, capabilities.strict_tool_use and capabilities.prompt_caching are all true, crossed with the versioned tool types listed in each model's tools[]. Retired ids excluded. All 13 active Claude models satisfy the SO + caching precondition, so the answer is really the tool column.
jq -r '.[] | select(.provider=="anthropic" and (.status|index("RETIRED")|not) and .capabilities.structured_outputs_json==true and .capabilities.prompt_caching==true) | [.id, ([.tools[].type]|join(","))] | @tsv' generated/models.json| Model | Status | SO | strict | cache min tok | 1h TTL | web_search | web_fetch | code_exec | tool_search | mcp_toolset | computer toolset GA | computer beta | browser | bash | text_editor | memory | advisor |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
claude-fable-5-1 |
DOCUMENTED · LIVE_VERIFIED |
✔ | ✔ | 512 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
claude-mythos-5-1 |
DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW |
✔ | ✔ | 512 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
claude-fable-5 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
✔ | ✔ | 512 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
claude-mythos-5 |
DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW |
✔ | ✔ | 512 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
claude-mythos-preview |
DOCUMENTED · ACCOUNT_RESTRICTED · DEPRECATED · PREVIEW |
✔ | ✔ | 2048 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | ✘ | ✘ | ✔ | ✔ | ✔ | ✘ |
claude-opus-5 |
DOCUMENTED · LIVE_VERIFIED |
✔ | ✔ | 512 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
claude-opus-4-8 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
✔ | ✔ | 1024 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
claude-opus-4-7 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
✔ | ✔ | 2048 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | ✔ | ✘ | ✔ | ✔ | ✔ | ✔ |
claude-opus-4-6 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
✔ | ✔ | 4096 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | ✔ | ✘ | ✔ | ✔ | ✔ | ✔ |
claude-opus-4-5-20251101 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
✔ | ✔ | 4096 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | ✔ | ✘ | ✔ | ✔ | ✔ | ✔ |
claude-sonnet-5 |
DOCUMENTED · LIVE_VERIFIED |
✔ | ✔ | 1024 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
claude-sonnet-4-6 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
✔ | ✔ | 1024 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | ✔ | ✘ | ✔ | ✔ | ✔ | ✔ |
claude-sonnet-4-5-20250929 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
✔ | ✔ | 1024 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | ✔ | ✘ | ✔ | ✔ | ✔ | ✔ |
claude-haiku-4-5-20251001 |
DOCUMENTED · LIVE_VERIFIED |
✔ | ✔ | 4096 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✘ | ✔ | ✘ | ✔ | ✔ | ✔ | ✔ |
Caveats from the records: strict is rejected on toolsets (mcp_toolset, computer_toolset_20260801, browser_toolset_20260801) and on programmatic callers; changing output_config.format invalidates the prompt cache (docs/anthropic/prompt-caching.md); tool_choice any/tool is 400 on Fable 5.1 / Mythos 5.1; mcp_toolset and advisor_20260301 still need beta headers (mcp-client-2025-11-20, advisor-tool-2026-03-01); Mythos ids are ACCOUNT_RESTRICTED (invite only). Tool-level compatibility lists: generated/compatibility/anthropic-tool-model-matrix.json, generated/compatibility/model-tool-matrix.json.
Q2. Which OpenAI models accept image input + reasoning + web search + streaming?
Interpretation: capabilities.image_in, capabilities.reasoning, capabilities.tool_web_search and capabilities.streaming all true in generated/models.json. Web search here means the hosted web_search tool on /v1/responses (Chat Completions only has search models). Retired ids are listed separately.
jq -r '.[] | select(.provider=="openai" and .capabilities.image_in==true and .capabilities.reasoning==true and .capabilities.tool_web_search==true and .capabilities.streaming==true) | [.id, (.status|join(",")), (.context_window|tostring), (.capabilities.reasoning_effort_values//[]|join("/"))] | @tsv' generated/models.json| Model | Status | Context | Max out | Effort values | Structured outputs | Prompt caching | Code interp. | Computer | MCP | Std price in / out |
|---|---|---|---|---|---|---|---|---|---|---|
gpt-5 |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | minimal/low/medium/high | ✔ | ✔ | ✔ | ✘ | ✔ | 1.25 / 10.0 |
gpt-5-2025-08-07 |
DOCUMENTED · LIVE_VERIFIED · DEPRECATED |
400000 | 128000 | minimal/low/medium/high | ✔ | ✔ | ✔ | ✘ | ✔ | 1.25 / 10.0 |
gpt-5-mini |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | (model default only) | ✔ | ✘ | ✔ | ✘ | ✔ | 0.25 / 2.0 |
gpt-5-mini-2025-08-07 |
DOCUMENTED · LIVE_VERIFIED · DEPRECATED |
400000 | 128000 | (model default only) | ✔ | ✘ | ✔ | ✘ | ✔ | 0.25 / 2.0 |
gpt-5-nano |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | (model default only) | ✔ | ✔ | ✔ | ✘ | ✔ | 0.05 / 0.4 |
gpt-5-nano-2025-08-07 |
DOCUMENTED · LIVE_VERIFIED · DEPRECATED |
400000 | 128000 | (model default only) | ✔ | ✔ | ✔ | ✘ | ✔ | 0.05 / 0.4 |
gpt-5-pro |
DOCUMENTED · LIVE_VERIFIED |
400000 | 272000 | high | ✔ | ✘ | ✘ | ✘ | ✔ | 15.0 / 120.0 |
gpt-5-pro-2025-10-06 |
DOCUMENTED · LIVE_VERIFIED · DEPRECATED |
400000 | 272000 | high | ✔ | ✘ | ✘ | ✘ | ✔ | 15.0 / 120.0 |
gpt-5.1 |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | none/low/medium/high | ✔ | ✔ | ✔ | ✘ | ✔ | 1.25 / 10.0 |
gpt-5.1-2025-11-13 |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | none/low/medium/high | ✔ | ✔ | ✔ | ✘ | ✔ | 1.25 / 10.0 |
gpt-5.2 |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | none/low/medium/high/xhigh | ✔ | ✔ | ✔ | ✘ | ✔ | 1.75 / 14.0 |
gpt-5.2-2025-12-11 |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | none/low/medium/high/xhigh | ✔ | ✔ | ✔ | ✘ | ✔ | 1.75 / 14.0 |
gpt-5.2-pro |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | (model default only) | ✘ | ✘ | ✘ | ✘ | ✔ | 21.0 / 168.0 |
gpt-5.2-pro-2025-12-11 |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | (model default only) | ✘ | ✘ | ✘ | ✘ | ✔ | 21.0 / 168.0 |
gpt-5.3-codex |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | low/medium/high/xhigh | ✔ | ✔ | ✘ | ✘ | ✘ | 1.75 / 14.0 |
gpt-5.4 |
DOCUMENTED · LIVE_VERIFIED |
1050000 | 128000 | none/low/medium/high/xhigh | ✔ | ✔ | ✔ | ✔ | ✔ | 2.5 / 15.0 |
gpt-5.4-2026-03-05 |
DOCUMENTED · LIVE_VERIFIED |
1050000 | 128000 | none/low/medium/high/xhigh | ✔ | ✔ | ✔ | ✔ | ✔ | 2.5 / 15.0 |
gpt-5.4-mini |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | none/low/medium/high/xhigh | ✔ | ✔ | ✔ | ✔ | ✔ | 0.75 / 4.5 |
gpt-5.4-mini-2026-03-17 |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | none/low/medium/high/xhigh | ✔ | ✔ | ✔ | ✔ | ✔ | 0.75 / 4.5 |
gpt-5.4-nano |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | none/low/medium/high/xhigh | ✔ | ✔ | ✔ | ✘ | ✔ | 0.2 / 1.25 |
gpt-5.4-nano-2026-03-17 |
DOCUMENTED · LIVE_VERIFIED |
400000 | 128000 | none/low/medium/high/xhigh | ✔ | ✔ | ✔ | ✘ | ✔ | 0.2 / 1.25 |
gpt-5.4-pro |
DOCUMENTED · LIVE_VERIFIED |
1050000 | 128000 | medium/high/xhigh | ✘ | ✘ | ✘ | ✔ | ✔ | 30.0 / 180.0 |
gpt-5.4-pro-2026-03-05 |
DOCUMENTED · LIVE_VERIFIED |
1050000 | 128000 | medium/high/xhigh | ✘ | ✘ | ✘ | ✔ | ✔ | 30.0 / 180.0 |
gpt-5.5 |
DOCUMENTED · LIVE_VERIFIED |
1050000 | 128000 | none/low/medium/high/xhigh | ✔ | ✔ | ✔ | ✔ | ✔ | 5.0 / 30.0 |
gpt-5.5-2026-04-23 |
DOCUMENTED · LIVE_VERIFIED |
1050000 | 128000 | none/low/medium/high/xhigh | ✔ | ✔ | ✔ | ✔ | ✔ | 5.0 / 30.0 |
gpt-5.6 |
DOCUMENTED · LIVE_VERIFIED |
1050000 | 128000 | none/low/medium/high/xhigh/max | ✔ | ✔ | ✔ | ✔ | ✔ | 4.0 / 20.0 |
gpt-5.6-cyber |
DOCUMENTED · ACCOUNT_RESTRICTED |
400000 | 128000 | (model default only) | ✔ | ✔ | ✔ | ✔ | ✔ | 12.5 / 75.0 |
gpt-5.6-luna |
DOCUMENTED · LIVE_VERIFIED |
1050000 | 128000 | none/low/medium/high/xhigh/max | ✔ | ✔ | ✔ | ✔ | ✔ | 0.2 / 1.2 |
gpt-5.6-sol |
DOCUMENTED · LIVE_VERIFIED |
1050000 | 128000 | none/low/medium/high/xhigh/max | ✔ | ✔ | ✔ | ✔ | ✔ | 4.0 / 20.0 |
gpt-5.6-terra |
DOCUMENTED · LIVE_VERIFIED |
1050000 | 128000 | none/low/medium/high/xhigh/max | ✔ | ✔ | ✔ | ✔ | ✔ | 2.0 / 12.0 |
gpt-6-astra |
DOCUMENTED · LIVE_VERIFIED |
1050000 | 128000 | low/medium/high/xhigh/max | ✔ | ✔ | ✔ | ✔ | ✔ | 10.0 / 50.0 |
gpt-daybreak-blue-latest |
DOCUMENTED · ACCOUNT_RESTRICTED |
1050000 | 128000 | (model default only) | ✔ | ✔ | ✔ | ✔ | ✔ | — / — |
gpt-daybreak-red-latest |
DOCUMENTED |
400000 | 128000 | (model default only) | ✔ | ✔ | ✔ | ✔ | ✔ | — / — |
o3 |
DOCUMENTED · LIVE_VERIFIED |
200000 | 100000 | (model default only) | ✔ | ✔ | ✔ | ✘ | ✔ | 2.0 / 8.0 |
o3-2025-04-16 |
DOCUMENTED · LIVE_VERIFIED · DEPRECATED |
200000 | 100000 | (model default only) | ✔ | ✔ | ✔ | ✘ | ✔ | 2.0 / 8.0 |
o4-mini |
DOCUMENTED · LIVE_VERIFIED · DEPRECATED |
200000 | 100000 | (model default only) | ✔ | ✔ | ✔ | ✘ | ✔ | 1.1 / 4.4 |
o4-mini-2025-04-16 |
DOCUMENTED · LIVE_VERIFIED · DEPRECATED |
200000 | 100000 | (model default only) | ✔ | ✔ | ✔ | ✘ | ✔ | 1.1 / 4.4 |
Also matching but RETIRED (excluded): gpt-5-codex, gpt-5.1-codex, gpt-5.1-codex-max, gpt-5.1-codex-mini, gpt-5.2-codex, o3-deep-research, o3-deep-research-2025-06-26, o4-mini-deep-research, o4-mini-deep-research-2025-06-26. Caveats: o4-mini, o3-2025-04-16, gpt-5-*-2025-08-07 snapshots are DEPRECATED with shutdown dates (see docs/openai/deprecations.md); web search is not supported with gpt-5 reasoning.effort: minimal and may degrade with none (docs/tools/openai/web-search.md); gpt-5.6-cyber, gpt-daybreak-* are gated (ACCOUNT_RESTRICTED/DOCUMENTED).
Q2b. The same combination on all four providers — image input + reasoning + vendor web search + structured outputs + streaming
Interpretation per provider (capability keys differ): openai image_in, reasoning, tool_web_search, structured_outputs, streaming; anthropic vision, (thinking_adaptive or thinking_extended_manual_budget), web_search, structured_outputs_json (streaming is universal); xai image_input, reasoning, web_search, structured_outputs, streaming; gemini image_input, thinking (true or a 'Supported…' string), google_search_grounding, structured_output (streaming is universal). Retired ids, aliases and id-only records excluded; xAI retired_redirect/legacy kinds and Gemini alias/agent kinds excluded.
jq -r '.[] | select(.provider=="xai" and .kind=="model" and .capabilities.image_input==true and .capabilities.reasoning==true and .capabilities.web_search==true and .capabilities.structured_outputs==true) | .id' generated/models.json
jq -r '.[] | select(.provider=="gemini" and (.kind=="stable" or .kind=="preview") and (.status|index("RETIRED")|not) and .capabilities.image_input==true and (.capabilities.thinking==true or (.capabilities.thinking|type=="string" and startswith("Supported"))) and .capabilities.google_search_grounding==true and .capabilities.structured_output==true) | .id' generated/models.json| Provider | Matching models (status other than DOCUMENTED/LIVE_*) | Reasoning control | Web search tool | Structured output parameter |
|---|---|---|---|---|
| openai | gpt-5, gpt-5-2025-08-07 (DEPRECATED), gpt-5-mini, gpt-5-mini-2025-08-07 (DEPRECATED), gpt-5-nano, gpt-5-nano-2025-08-07 (DEPRECATED), gpt-5-pro, gpt-5-pro-2025-10-06 (DEPRECATED), gpt-5.1, gpt-5.1-2025-11-13, gpt-5.2, gpt-5.2-2025-12-11, gpt-5.3-codex, gpt-5.4, gpt-5.4-2026-03-05, gpt-5.4-mini, gpt-5.4-mini-2026-03-17, gpt-5.4-nano, gpt-5.4-nano-2026-03-17, gpt-5.5, gpt-5.5-2026-04-23, gpt-5.6, gpt-5.6-cyber (ACCOUNT_RESTRICTED), gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, gpt-daybreak-blue-latest (ACCOUNT_RESTRICTED), gpt-daybreak-red-latest, o3, o3-2025-04-16 (DEPRECATED), o4-mini (DEPRECATED), o4-mini-2025-04-16 (DEPRECATED) |
reasoning.effort (Responses) / reasoning_effort (Chat) |
tools[type=web_search] |
text.format {type: json_schema} |
| anthropic | claude-fable-5-1, claude-mythos-5-1 (ACCOUNT_RESTRICTED/PREVIEW), claude-fable-5 (LEGACY), claude-mythos-5 (ACCOUNT_RESTRICTED/PREVIEW), claude-mythos-preview (ACCOUNT_RESTRICTED/DEPRECATED/PREVIEW), claude-opus-5, claude-opus-4-8 (LEGACY), claude-opus-4-7 (LEGACY), claude-opus-4-6 (LEGACY), claude-opus-4-5-20251101 (LEGACY), claude-sonnet-5, claude-sonnet-4-6 (LEGACY), claude-sonnet-4-5-20250929 (LEGACY), claude-haiku-4-5-20251001 |
thinking {type: adaptive | enabled} + output_config.effort |
tools[type=web_search_20260318] |
output_config.format {type: json_schema} |
| xai | grok-4.6 |
reasoning.effort / reasoning_effort (low…xhigh; none LIVE_DISCOVERED on grok-4.3; 400 on 4.20-reasoning / build) |
tools[type=web_search] ($5/1k) + x_search |
text.format {type: json_schema} / response_format.json_schema |
| gemini | gemini-2.5-flash, gemini-2.5-pro, gemini-2.5-flash-lite (ACCOUNT_RESTRICTED), gemini-3-flash-preview (PREVIEW), gemini-3.1-pro-preview (PREVIEW/ACCOUNT_RESTRICTED), gemini-3.1-pro-preview-customtools (PREVIEW), gemini-3.1-flash-lite (DEPRECATED), gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash, gemini-3.7-flash, gemini-3.8-flash, gemini-robotics-er-2-preview (PREVIEW), gemini-robotics-er-2-streaming-preview (PREVIEW), gemini-2.5-flash-preview-09-2025 (PREVIEW/UNVERIFIED) |
generationConfig.thinkingConfig {thinkingLevel | thinkingBudget, includeThoughts} |
tools[{googleSearch:{}}] (ACCOUNT_RESTRICTED on this key; 5,000 free/month on 3.x then $14/1k) |
generationConfig.responseMimeType: application/json + responseJsonSchema / responseSchema |
Reading notes: OpenAI snapshots (gpt-5.4-2026-03-05…) appear as separate callable ids; Anthropic aliases (claude-haiku-4-5) are excluded (capabilities live on the snapshot); Gemini image-output models (gemini-3.1-flash-image, gemini-3-pro-image) support Google Search grounding and thinking but not structured output, so they drop out; gemini-2.5-* match on paper but are ACCOUNT_RESTRICTED ('no longer available to new users'); xAI grok-4.20-0309-non-reasoning drops out on reasoning: false.
Q3. What is the exact JSON to call feature Y?
Three lookups, in order of precision:
- Parameter rows —
generated/parameters.json(7,862 rows) is keyed byendpoint(METHOD /path) and dottedparameterpath (Gemini uses the REST camelCase names, e.g.generationConfig.thinkingConfig.thinkingLevel), withtype,required,default,enum,beta_header,status,description.
# every parameter of the Anthropic structured-output block
jq '[.[] | select(.provider=="anthropic" and .endpoint=="POST /v1/messages" and (.parameter|startswith("output_config.format")))]' generated/parameters.json
# the OpenAI equivalent
jq '[.[] | select(.provider=="openai" and .endpoint=="POST /v1/responses" and (.parameter|startswith("text.format")))]' generated/parameters.json
# xAI Responses reasoning + caching knobs
jq '[.[] | select(.provider=="xai" and .endpoint=="POST /v1/responses" and (.parameter|test("^(reasoning|prompt_cache_key|store|previous_response_id|max_turns)")))]' generated/parameters.json
# Gemini thinking and structured-output knobs
jq '[.[] | select(.provider=="gemini" and .endpoint=="POST /v1beta/models/{model}:generateContent" and (.parameter|test("thinkingConfig|responseJsonSchema|responseMimeType|cachedContent")))]' generated/parameters.json
# all tool entry shapes accepted by OpenAI Responses
jq -r '.[] | select(.provider=="openai" and .endpoint=="POST /v1/responses" and (.parameter|test("^tools\\[\\]\\("))) | .parameter' generated/parameters.json- Tool records —
generated/tools.jsonhasparameters_schema,result_shape,examples{curl,python,typescript}per exact tooltype(69 records: OpenAI 19, Anthropic 27, xAI 12, Gemini 11).
jq '.[] | select(.type=="web_search_20260318") | {parameters_schema, result_shape, beta_header, examples}' generated/tools.json
jq '.[] | select(.provider=="xai" and .type=="x_search") | {parameters_schema, result_shape, billing}' generated/tools.json
jq '.[] | select(.provider=="gemini" and .type=="googleSearch") | {parameters_schema, result_shape, billing, limitations}' generated/tools.json- Runnable examples —
generated/examples-manifest.jsonmaps features to files underexamples/<provider>/<area>/with their verification status. Selection:
| Feature | OpenAI | Anthropic | xAI | Gemini |
|---|---|---|---|---|
| Minimal call | examples/openai/responses/ |
examples/anthropic/messages/ |
examples/xai/responses/minimal.py (LIVE_VERIFIED) |
examples/gemini/generate-content/minimal.py (LIVE_VERIFIED) |
| Structured outputs | examples/openai/responses/structured_output.py (LIVE_VERIFIED) |
examples/anthropic/structured-output/ |
examples/xai/structured-output/json_schema.sh (LIVE_VERIFIED) |
examples/gemini/structured-output/json_schema.py (LIVE_VERIFIED) |
| Streaming | examples/openai/responses/streaming.py (LIVE_VERIFIED) |
examples/anthropic/streaming/sdk_stream.py (LIVE_VERIFIED) |
examples/xai/streaming/responses_stream.py (LIVE_VERIFIED) |
examples/gemini/streaming/stream.py (LIVE_VERIFIED) |
| Multi-turn state | examples/openai/responses/previous_response_id.py (LIVE_VERIFIED) |
examples/shared/tool-loop/anthropic_tool_loop.py (LIVE_VERIFIED) |
examples/xai/responses/compact.py (LIVE_VERIFIED) |
examples/gemini/generate-content/system_multiturn.py (LIVE_VERIFIED) |
| Tool loop | examples/shared/tool-loop/openai_tool_loop.py (not in manifest) |
examples/shared/tool-loop/anthropic_tool_loop.py (LIVE_VERIFIED) |
examples/xai/responses/tool_loop.py (LIVE_VERIFIED) |
examples/shared/tool-loop/gemini_tool_loop.py (LIVE_VERIFIED) |
| Web search | examples/openai/tools/web-search/ |
examples/anthropic/web-search/web_search.py (LIVE_VERIFIED) |
examples/xai/tools/web-search/web_search.py (LIVE_VERIFIED) |
examples/gemini/tools/google-search/grounding.py (FAILED_VERIFICATION) |
| Code execution | examples/openai/tools/code-interpreter/ |
examples/anthropic/code-execution/code_execution.py (LIVE_VERIFIED) |
examples/xai/tools/code-execution/code_interpreter.py (LIVE_VERIFIED) |
examples/gemini/tools/code-execution/code_execution.py (LIVE_VERIFIED) |
| MCP | examples/openai/tools/mcp-and-connectors/ |
examples/anthropic/mcp/mcp_connector.py (LIVE_VERIFIED) |
examples/xai/tools/mcp/deepwiki.py (LIVE_VERIFIED) |
docs/tools/gemini/mcp.md |
| File search / RAG | docs/openai/vector-stores.md |
docs/anthropic/citations.md |
examples/xai/tools/collections-search/file_search.py (LIVE_VERIFIED) |
examples/gemini/file-search/file_search_lifecycle.py (LIVE_VERIFIED) |
| Batch | examples/openai/batch/batch-lifecycle.py (LIVE_VERIFIED) |
examples/anthropic/batch/batch_lifecycle.py (LIVE_VERIFIED) |
examples/xai/batches/lifecycle.py (LIVE_VERIFIED) |
examples/gemini/batch/batch_inline.py (UNVERIFIED) |
| Agents / sessions | examples/openai/agents/ |
examples/anthropic/agents/create_agent_session_stream.py (LIVE_VERIFIED) |
examples/xai/responses/tool_loop.py (LIVE_VERIFIED) |
examples/gemini/interactions/interactions_basic.py (LIVE_VERIFIED) |
| Prompt / context caching | docs/openai/prompt-caching.md §5 (live pair) |
examples/anthropic/prompt-caching/ |
docs/xai/prompt-caching.md |
examples/gemini/context-caching/cache_lifecycle.py (LIVE_VERIFIED) |
| Reasoning / thinking | examples/openai/responses/reasoning_summary.py (LIVE_VERIFIED) |
examples/anthropic/thinking/ |
docs/xai/reasoning.md |
examples/gemini/thinking/thinking.py (LIVE_VERIFIED) |
| Image input | examples/openai/responses/image_input.py (LIVE_VERIFIED) |
examples/anthropic/vision/ |
examples/xai/responses/image_input.sh (LIVE_VERIFIED) |
examples/gemini/multimodal/inline_image_pdf.py (LIVE_VERIFIED) |
| Realtime / Live voice | docs/openai/realtime.md |
— | docs/xai/voice.md |
examples/gemini/live/live_ws_audio.py (LIVE_VERIFIED) |
| Image generation | docs/openai/images.md |
— | docs/xai/images.md |
examples/gemini/image-generation/generate.py (UNVERIFIED) |
| Embeddings | docs/openai/embeddings.md |
— | docs/xai/index.md (ACCOUNT_RESTRICTED) |
examples/gemini/embeddings/embed.py (LIVE_VERIFIED) |
| Skills | docs/openai/skills-api.md |
examples/anthropic/skills/skills_lifecycle.py (LIVE_VERIFIED) |
docs/xai/skills-api.md (ACCOUNT_RESTRICTED) |
— |
| Webhook verification | examples/openai/webhooks/offline_test.py (LIVE_VERIFIED) |
— | docs/xai/voice.md (SIP realtime.call.incoming) |
docs/gemini/interactions-api.md (webhook_config) |
Minimal verified bodies (from the live probes recorded in the docs):
// OpenAI — POST /v1/responses (structured output, LIVE_VERIFIED gpt-5.4-nano)
{"model":"gpt-5.4-nano","input":"Reply with OK.","max_output_tokens":32,
"text":{"format":{"type":"json_schema","name":"ok_reply","strict":true,
"schema":{"type":"object","properties":{"answer":{"type":"string"}},"required":["answer"],"additionalProperties":false}}}}// Anthropic — POST /v1/messages (structured output, LIVE_VERIFIED claude-haiku-4-5-20251001)
{"model":"claude-haiku-4-5-20251001","max_tokens":100,
"messages":[{"role":"user","content":"Reply with OK."}],
"output_config":{"format":{"type":"json_schema",
"schema":{"type":"object","properties":{"answer":{"type":"string"}},"required":["answer"],"additionalProperties":false}}}}// xAI — POST /v1/responses (structured output, LIVE_VERIFIED grok-4.3; reasoning tokens are billed on every call)
{"model":"grok-4.3","input":"Reply with OK.","max_output_tokens":64,"reasoning":{"effort":"low"},
"text":{"format":{"type":"json_schema","name":"ok_reply","strict":true,
"schema":{"type":"object","properties":{"answer":{"type":"string"}},"required":["answer"],"additionalProperties":false}}}}// Gemini — POST /v1beta/models/gemini-3.5-flash-lite:generateContent (structured output, LIVE_VERIFIED; header x-goog-api-key)
{"contents":[{"role":"user","parts":[{"text":"Reply with OK."}]}],
"generationConfig":{"maxOutputTokens":64,"thinkingConfig":{"thinkingLevel":"minimal"},
"responseMimeType":"application/json",
"responseJsonSchema":{"type":"object","properties":{"answer":{"type":"string"}},"required":["answer"]}}}Per-topic parameter mappings with full JSON quads: state-management, tool-execution, caching-and-reasoning, streaming, agents-platforms, realtime-and-media.
Q4. Which SSE / WebSocket events can arrive during this call?
Computed from generated/streaming-events.json (506 records: openai 282, anthropic 68, xai 94, gemini 62; api field = call type, direction = server→client unless noted). Event names are listed exactly as recorded; ✔ marks events observed live in this run (LIVE_VERIFIED).
jq -r '.[] | select(.provider=="xai" and .api=="responses") | .event' generated/streaming-events.json
jq -r '.[] | select(.provider=="gemini" and .api=="live" and .direction=="server→client") | .event' generated/streaming-events.json| Provider | Call type (api) |
Direction | # | Events (✔ = LIVE_VERIFIED) |
|---|---|---|---|---|
| openai | responses |
server→client | 63 | response.audio.delta, response.audio.done, response.audio.transcript.delta, response.audio.transcript.done, response.code_interpreter_call_code.delta, response.code_interpreter_call_code.done, response.code_interpreter_call.completed, response.code_interpreter_call.in_progress, response.code_interpreter_call.interpreting, response.compaction.compacting, ✔response.completed, ✔response.content_part.added, ✔response.content_part.done, ✔response.created, error, ✔response.file_search_call.completed, ✔response.file_search_call.in_progress, ✔response.file_search_call.searching, ✔response.function_call_arguments.delta, ✔response.function_call_arguments.done, ✔response.shell_call_command.added, ✔response.shell_call_command.delta, ✔response.shell_call_command.done, response.shell_call_output_content.delta, response.shell_call_output_content.done, ✔response.in_progress, response.failed, response.incomplete, ✔response.output_item.added, ✔response.output_item.done, response.reasoning_summary_part.added, response.reasoning_summary_part.done, response.reasoning_summary_text.delta, response.reasoning_summary_text.done, response.reasoning_text.delta, response.reasoning_text.done, response.refusal.delta, response.refusal.done, ✔response.output_text.delta, ✔response.output_text.done, ✔response.web_search_call.completed, ✔response.web_search_call.in_progress, ✔response.web_search_call.searching, response.image_generation_call.completed, response.image_generation_call.generating, response.image_generation_call.in_progress, response.image_generation_call.partial_image, ✔response.mcp_call_arguments.delta, ✔response.mcp_call_arguments.done, ✔response.mcp_call.completed, response.mcp_call.failed, ✔response.mcp_call.in_progress, ✔response.mcp_list_tools.completed, response.mcp_list_tools.failed, ✔response.mcp_list_tools.in_progress, response.output_text.annotation.added, ✔response.queued, ✔response.custom_tool_call_input.delta, ✔response.custom_tool_call_input.done, ✔_sequence_observed, ✔response.apply_patch_call_operation_diff.delta, ✔response.apply_patch_call_operation_diff.done, ✔response.output_item.added / response.output_item.done (tool items) |
| openai | responses-websocket |
server→client | 69 | response.audio.delta, response.audio.done, response.audio.transcript.delta, response.audio.transcript.done, response.code_interpreter_call_code.delta, response.code_interpreter_call_code.done, response.code_interpreter_call.completed, response.code_interpreter_call.in_progress, response.code_interpreter_call.interpreting, response.compaction.compacting, response.completed, response.content_part.added, response.content_part.done, response.created, response.file_search_call.completed, response.file_search_call.in_progress, response.file_search_call.searching, response.function_call_arguments.delta, response.function_call_arguments.done, response.shell_call_command.added, response.shell_call_command.delta, response.shell_call_command.done, response.shell_call_output_content.delta, response.shell_call_output_content.done, response.in_progress, response.failed, response.incomplete, response.output_item.added, response.output_item.done, response.reasoning_summary_part.added, response.reasoning_summary_part.done, response.reasoning_summary_text.delta, response.reasoning_summary_text.done, response.reasoning_text.delta, response.reasoning_text.done, response.refusal.delta, response.refusal.done, response.output_text.delta, response.output_text.done, response.web_search_call.completed, response.web_search_call.in_progress, response.web_search_call.searching, response.image_generation_call.completed, response.image_generation_call.generating, response.image_generation_call.in_progress, response.image_generation_call.partial_image, response.mcp_call_arguments.delta, response.mcp_call_arguments.done, response.mcp_call.completed, response.mcp_call.failed, response.mcp_call.in_progress, response.mcp_list_tools.completed, response.mcp_list_tools.failed, response.mcp_list_tools.in_progress, response.output_text.annotation.added, response.queued, response.custom_tool_call_input.delta, response.custom_tool_call_input.done, error, response.steer.accepted, response.steer.pending, response.steer.failed, response.inject.created, response.inject.failed, error (code=previous_response_not_found), error (code=invalid_stream_id), error (code=websocket_stream_limit_reached), error (code=websocket_connection_limit_reached), _connection_limits |
| openai | responses-websocket |
client→server | 3 | response.create, response.steer, response.inject |
| openai | chat_completions |
server→client | 2 | ✔chat.completion.chunk, ✔[DONE] |
| openai | completions |
server→client | 1 | ✔text_completion (streamed) |
| openai | realtime |
server→client | 45 | ✔session.created, ✔session.updated, ✔conversation.item.added, ✔conversation.item.done, conversation.item.retrieved, conversation.item.input_audio_transcription.completed, conversation.item.input_audio_transcription.delta, conversation.item.input_audio_transcription.segment, conversation.item.input_audio_transcription.failed, conversation.item.truncated, conversation.item.deleted, input_audio_buffer.committed, input_audio_buffer.dtmf_event_received, input_audio_buffer.cleared, input_audio_buffer.speech_started, input_audio_buffer.speech_stopped, input_audio_buffer.timeout_triggered, output_audio_buffer.started, output_audio_buffer.stopped, output_audio_buffer.cleared, ✔response.created, ✔response.done, ✔response.output_item.added, ✔response.output_item.done, ✔response.content_part.added, ✔response.content_part.done, ✔response.output_text.delta, ✔response.output_text.done, response.output_audio_transcript.delta, response.output_audio_transcript.done, response.output_audio.delta, response.output_audio.done, response.function_call_arguments.delta, response.function_call_arguments.done, response.mcp_call_arguments.delta, response.mcp_call_arguments.done, response.mcp_call.in_progress, response.mcp_call.completed, response.mcp_call.failed, mcp_list_tools.in_progress, mcp_list_tools.completed, mcp_list_tools.failed, rate_limits.updated, conversation.created, conversation.item.created |
| openai | realtime |
client→server | 11 | ✔session.update, input_audio_buffer.append, input_audio_buffer.commit, input_audio_buffer.clear, ✔conversation.item.create, conversation.item.retrieve, conversation.item.truncate, conversation.item.delete, ✔response.create, response.cancel, output_audio_buffer.clear |
| openai | realtime-translation |
server→client | 6 | session.created, session.updated, session.closed, session.input_transcript.delta, session.output_transcript.delta, session.output_audio.delta |
| openai | realtime-translation |
client→server | 3 | session.update, session.input_audio_buffer.append, session.close |
| openai | live |
server→client | 20 | session.started, session.updated, session.input_audio.muted, session.input_audio.unmuted, session.instructions.appended, session.thinking.appended, session.commentary.appended, session.output_audio.delta, session.input_transcript.delta, session.output_transcript.delta, session.delegation.created, response.event, session.usage.updated, session.closed, session.input_audio.append, transport.dtmf.received, transport.dtmf.send, transport.ringing, transport.answered, transport.failed |
| openai | live |
client→server | 11 | session.start, session.update, session.input_audio.append, session.input_audio.mute, session.input_audio.unmute, session.instructions.append, session.thinking.append, session.commentary.append, response.item.create, response.create, session.close |
| openai | audio-transcriptions |
server→client | 3 | ✔transcript.text.segment, ✔transcript.text.delta, ✔transcript.text.done |
| openai | audio-speech |
server→client | 2 | ✔speech.audio.delta, ✔speech.audio.done |
| anthropic | messages |
server→client | 23 | ✔content_block_start (thinking), ✔content_block_delta: thinking_delta, ✔content_block_delta: signature_delta, content_block_start (redacted_thinking), ✔content_block_delta: citations_delta, content_block_start/delta/stop (compaction, threshold), content_block_start (compaction, on-demand), ✔message_delta (usage.output_tokens_details / context_management), ✔message_start (usage cache fields / input_transformations), ✔content_block_delta: text_delta (structured outputs), ✔message_start, ✔content_block_start, ✔content_block_delta / text_delta, ✔content_block_delta / input_json_delta, ✔content_block_delta / thinking_delta, ✔content_block_delta / signature_delta, content_block_delta / citations_delta, ✔content_block_stop, ✔message_delta, ✔message_stop, ✔ping, error, (unknown event types) |
| anthropic | managed-agents |
server→client | 37 | user.message, user.interrupt, user.tool_confirmation, user.custom_tool_result, agent.custom_tool_use, agent.message, agent.thinking, agent.mcp_tool_use, agent.mcp_tool_result, agent.tool_use, agent.tool_result, agent.thread_message_received, agent.thread_message_sent, agent.thread_context_compacted, session.error, session.status_rescheduled, session.status_running, session.status_idle, session.status_terminated, session.thread_created, span.outcome_evaluation_start, span.outcome_evaluation_end, span.model_request_start, span.model_request_end, span.outcome_evaluation_ongoing, user.define_outcome, session.deleted, session.thread_status_running, session.thread_status_idle, session.thread_status_terminated, user.tool_result, session.thread_status_rescheduled, session.updated, event_start, event_delta, system.message, session.usage |
| anthropic | managed-agents |
client→server | 7 | user.message, user.interrupt, user.tool_confirmation, user.custom_tool_result, user.define_outcome, user.tool_result, system.message |
| xai | responses |
server→client | 24 | ✔response.created, ✔response.in_progress, ✔response.output_item.added, ✔response.reasoning_summary_part.added, ✔response.reasoning_summary_text.delta, ✔response.reasoning_summary_text.done, ✔response.reasoning_summary_part.done, ✔response.output_item.done, ✔response.content_part.added, ✔response.output_text.delta, ✔response.output_text.done, ✔response.content_part.done, ✔response.function_call_arguments.delta, ✔response.function_call_arguments.done, ✔response.code_interpreter_call.in_progress, ✔response.code_interpreter_call_code.delta, ✔response.code_interpreter_call_code.done, ✔response.code_interpreter_call.interpreting, ✔response.code_interpreter_call.completed, ✔response.completed, response.reasoning_text.delta, response.web_search_call.in_progress / .searching / .completed, response.mcp_call.* / response.file_search_call.*, response.incomplete / response.failed / error |
| xai | chat_completions |
server→client | 3 | ✔chat.completion.chunk, ✔chat.completion.chunk (usage chunk), ✔[DONE] |
| xai | messages |
server→client | 6 | ✔message_start, ✔content_block_start, ✔content_block_delta, ✔content_block_stop, ✔message_delta, ✔message_stop |
| xai | realtime |
server→client | 39 | ✔session.created, ✔conversation.created, ✔session.updated, input_audio_buffer.speech_started, input_audio_buffer.speech_stopped, input_audio_buffer.committed, input_audio_buffer.timeout_triggered, input_audio_buffer.cleared, conversation.item.deleted, ✔conversation.item.added, conversation.item.truncated, conversation.item.input_audio_transcription.completed, conversation.item.input_audio_transcription.updated, input_audio_buffer.dtmf_event_received, ✔response.created, ✔response.output_item.added, ✔response.output_item.done, ✔response.content_part.added, ✔response.content_part.done, ✔response.output_audio_transcript.delta, ✔response.output_audio_transcript.done, ✔response.output_audio.delta, ✔response.output_audio.done, response.text.delta, response.output_text.delta, response.function_call_arguments.delta, response.function_call_arguments.done, mcp_list_tools.in_progress, mcp_list_tools.completed, mcp_list_tools.failed, response.mcp_call_arguments.delta, response.mcp_call_arguments.done, response.mcp_call.in_progress, response.mcp_call.completed, response.mcp_call.failed, ✔response.done, error, ping, response.audio.delta |
| xai | realtime |
client→server | 9 | ✔session.update, input_audio_buffer.append, input_audio_buffer.commit, ✔conversation.item.create, input_audio_buffer.clear, conversation.item.delete, conversation.item.truncate, ✔response.create, response.cancel |
| xai | realtime |
server→client (webhook) | 1 | realtime.call.incoming |
| xai | tts |
client→server | 2 | text.delta, text.done |
| xai | tts |
server→client | 3 | audio.delta, audio.done, error |
| xai | stt |
client→server | 3 | Binary frame (audio), finalize, audio.done |
| xai | stt |
server→client | 4 | transcript.created, transcript.partial, transcript.done, error |
| gemini | generate-content |
server→client | 8 | ✔sse_framing, ✔json_array_framing, ✔chunk, ✔thought_chunk, ✔final_chunk, ✔structured_output_chunks, prompt_blocked, ✔single_event_alt_sse_on_generateContent |
| gemini | interactions |
server→client | 15 | ✔interaction.created, ✔interaction.status_update, interaction.in_progress, interaction.requires_action, ✔step.start, step.delta, ✔step.stop, ✔interaction.completed, error, ✔done, interaction.start, content.start, content.delta, content.stop, interaction.complete |
| gemini | live |
server→client | 21 | ✔setupComplete, ✔serverContent, ✔serverContent.modelTurn, ✔serverContent.generationComplete, ✔serverContent.turnComplete, serverContent.interactionStatus, serverContent.interrupted, serverContent.groundingMetadata, serverContent.inputTranscription, serverContent.interimInputTranscription, ✔serverContent.outputTranscription, serverContent.urlContextMetadata, serverContent.waitingForInput, serverContent.speechState, ✔toolCall, toolCallCancellation, ✔usageMetadata, goAway, ✔sessionResumptionUpdate, voiceActivity, voiceActivityDetectionSignal |
| gemini | live |
client→server | 11 | ✔setup, ✔clientContent, realtimeInput, realtimeInput.audio, realtimeInput.video, realtimeInput.text, realtimeInput.activityStart, realtimeInput.activityEnd, realtimeInput.audioStreamEnd, realtimeInput.mediaChunks, toolResponse |
| gemini | lyria-realtime |
server→client | 3 | setupComplete, serverContent.audioChunks[], filteredPrompt |
| gemini | lyria-realtime |
client→server | 4 | setup, clientContent, musicGenerationConfig, playbackControl |
| anthropic | complete-legacy |
server→client | 1 | completion (legacy /v1/complete) |
| openai | agents |
server→client | 31 | error, agent.session.environment.ready, agent.session.environment.reset, agent.output.command_execution_output.delta, ✔agent.session.created, ✔agent.session.turn.created, ✔agent.session.turn.in_progress, ✔agent.session.turn.completed, agent.session.turn.failed, agent.session.turn.cancelled, ✔agent.session.turn.item.added, ✔agent.session.idle, ✔agent.session.in_progress, agent.session.requires_action, agent.session.failed, agent.session.environment.pending, agent.session.environment.connected, agent.session.environment.disconnected, agent.session.environment.failed, agent.session.subagent.created, agent.session.subagent.active, agent.session.subagent.closed, ✔agent.session.turn.item.done, ✔agent.session.turn.content_part.added, ✔agent.session.turn.content_part.done, ✔agent.session.turn.output_text.delta, ✔agent.session.turn.output_text.done, agent.session.turn.reasoning_summary_part.added, agent.session.turn.reasoning_summary_part.done, agent.session.turn.reasoning_summary_text.delta, agent.session.turn.reasoning_summary_text.done |
| openai | agents |
client→server | 3 | ✔agent.session.input.message, ✔agent.session.input.cancel, agent.session.input.tool_result |
| openai | agents-webhooks |
server→client (webhook) | 5 | agent.session.created, agent.session.action_required, agent.session.in_progress, agent.session.idle, agent.session.failed |
| openai | images |
server→client | 4 | image_generation.partial_image, image_generation.completed, image_edit.partial_image, image_edit.completed |
How to read the four core streams: OpenAI Responses is event: <type> + data: with sequence_number, terminal response.completed|failed|incomplete|error, no [DONE]; Anthropic Messages is event: <type> + data: (data.type == event), fixed skeleton message_start → content_block_start/delta/stop × N → message_delta → message_stop, ping anywhere, no [DONE]; xAI reuses both shapes — /v1/responses emits the OpenAI Responses event names (with sequence_number, no [DONE]), /v1/chat/completions emits data-only chat.completion.chunk + data: [DONE], and /v1/messages emits the Anthropic skeleton (no ping); Gemini streamGenerateContent has no event names at all — a JSON array of GenerateContentResponse chunks by default or data: lines with ?alt=sse (thought parts flagged thought: true, end detected by finishReason), while the Interactions API uses named events (interaction.start … content.delta … interaction.complete). Details: streaming comparison.
The Anthropic messages group is the advanced fragment (thinking/citations/compaction variants of the same core events); the POST /v1/messages (stream=true) group is the core catalogue — both describe one stream. Voice WebSockets (OpenAI Realtime/Live, xAI realtime/tts/stt, Gemini Live/Lyria RealTime) are bidirectional message families, not SSE.
Q5. Which endpoint creates an agent session on each provider?
jq '.[] | select((.provider=="openai" and .path=="/v1/agents/sessions") or (.provider=="anthropic" and .path=="/v1/sessions") or (.provider=="gemini" and .path=="/v1beta/interactions") or (.provider=="xai" and .path=="/v1/responses")) | select(.method=="POST") | {provider, method, path, status, beta_header, verification}' generated/endpoints.json| Provider | Endpoint | Beta gate | Status | Live result | Required body | Notes |
|---|---|---|---|---|---|---|
| openai | POST /v1/agents/sessions |
OpenAI-Beta: agents=v1 |
DOCUMENTED · BETA · LIVE_VERIFIED |
success HTTP 201 | environment (required: none | openai_hosted | self_hosted), agent or agent_id, input (required when environment is none), stream, vault_ids, metadata |
201 agent.session; first turn starts immediately when input is given; send later input via POST …/sessions/{id}/events (agent.session.input.message); SSE with stream:true or GET …/events?stream=true |
| anthropic | POST /v1/sessions |
managed-agents-2026-04-01 |
DOCUMENTED · BETA · LIVE_VERIFIED |
success HTTP 200 | agent (id or {type:agent,id,version} or agent_with_overrides), environment_id (required), optional initial_events[] (≤50), resources[], vault_ids[], budget, title, metadata |
200 session (sesn_…); starts idle unless initial_events; send user.message via POST /v1/sessions/{id}/events; stream GET /v1/sessions/{id}/events/stream; agents and environments must exist first (POST /v1/agents, POST /v1/environments) |
| xai | POST /v1/responses |
none (no beta headers on xAI) | DOCUMENTED · LIVE_VERIFIED |
success HTTP 200 | model, input; agentic behaviour comes from tools[] (server-side web_search, x_search, code_interpreter, file_search, mcp, image_generation + your function/shell tools), max_turns, store (default true), previous_response_id |
xAI has no separate agent/session resource: the Responses call itself runs the server-side agentic loop (bounded by max_turns) and the stored response (store:true, 30-day retention) is the session — continue with previous_response_id, inspect with GET /v1/responses/{id} / …/input_items, shrink with POST /v1/responses/compact. grok-4.20-multi-agent-0309 (BETA) fans out to 4 or 16 agents inside one call. Grok Build is a CLI product, not an API |
| gemini | POST /v1beta/interactions |
/v1beta path version |
DOCUMENTED · BETA · LIVE_VERIFIED |
success HTTP 200 | model or agent (deep-research-preview-04-2026, deep-research-max-preview-04-2026, antigravity-preview-09-2026, or a custom agent id from POST /v1beta/agents), input, optional previous_interaction_id, store (default true), stream, background, tools[], generation_config, agent_config, environment, webhook_config |
200 interaction (id, status, steps[]/outputs); stateful chaining via previous_interaction_id; poll/resume with GET /v1beta/interactions/{id}, POST …/cancel, DELETE; long runs with background:true (Deep Research); sandboxed agents need POST /v1beta/environments (LIVE_VERIFIED) and can be scheduled with /v1beta/triggers (PREVIEW). /v1/interactions is documented GA but UNVERIFIED here. The Live API WebSocket (BidiGenerateContent, setup message) is the other session-shaped surface (voice) |
Prerequisite creates: OpenAI POST /v1/agents (LIVE_VERIFIED 201) or an inline agent; Anthropic POST /v1/agents (LIVE_VERIFIED 200) and POST /v1/environments (LIVE_VERIFIED 200); Gemini POST /v1beta/agents (custom managed agents, PREVIEW, not tested) and POST /v1beta/environments (LIVE_VERIFIED). Not to be confused with: OpenAI Conversations (POST /v1/conversations — a state container for Responses, not an agent), Realtime client secrets (POST /v1/realtime/client_secrets), ChatKit sessions, the retired Assistants threads/runs; xAI POST /v1/realtime/client_secrets (voice session token); Gemini POST /v1beta/auth_tokens (ephemeral Live token) and cachedContents (a prefix cache, not a session). Full comparison: agents-platforms.
Q6. Which feature needs a beta header (or a beta path / preview model)?
The four providers gate pre-GA features differently: Anthropic by anthropic-beta header values (per parameter/tool/endpoint), OpenAI by OpenAI-Beta header values (per surface), Gemini by the URL version (/v1beta vs /v1) and -preview / -exp model ids (no header exists), xAI by nothing visible — no version or beta headers; gated features answer 403 ('alpha users') or 404 (ACL) and are recorded ACCOUNT_RESTRICTED.
Anthropic — anthropic-beta values that are still gating (status BETA / PREVIEW / DOCUMENTED-only), from generated/fragments/headers/anthropic-beta-headers.json, cross-referenced with the beta_header field of generated/parameters.json, tools.json and endpoints.json
| Header value | Feature | Status | Models / scope | Gated parameters (sample) | Gated endpoints | Gated tool types |
|---|---|---|---|---|---|---|
computer-use-2025-01-24 |
computer_20250124 tool version | BETA · LEGACY |
claude-sonnet-4-5-20250929, claude-haiku-4-5-20251001, claude-opus-4-1-20250805 (retired), claude-sonnet-4-20250514 (retired) … | — | — | computer_20250124 |
computer-use-2025-11-24 |
computer_20251124 tool version (zoom, batch actions precursor) | BETA |
claude-fable-5-1, claude-mythos-5-1, claude-fable-5, claude-mythos-5 … | POST /v1/messages → tools[].enable_zoom |
— | computer_20251124 |
mcp-client-2025-11-20 |
MCP connector (mcp_toolset tool type) | BETA |
— | POST /v1/messages → mcp_servers[].authorization_token; POST /v1/messages → mcp_servers[].name; POST /v1/messages → mcp_servers[].tool_configuration … (+5) |
— | mcp_toolset |
dev-full-thinking-2025-05-14 |
Full (unsummarized) thinking - developer/dev-only | LIVE_DISCOVERED |
— | — | — | — |
interleaved-thinking-2025-05-14 |
Interleaved thinking between tool calls with manual extended thinking | BETA |
models using thinking type enabled (4.5 family, 4.6 manual mode); automatic (no header) with adaptive thinking; not supported on Haiku 4.5 | — | — | — |
context-management-2025-06-27 |
Context editing (clear_tool_uses_20250919, clear_thinking_20251015 strategies) | BETA |
all supported Claude models (context_management capability true on all 11 live ids) | POST /v1/messages → context_management.edits[].clear_at_least; POST /v1/messages → context_management.edits[].clear_tool_inputs; POST /v1/messages → context_management.edits[].exclude_tools … (+5) |
— | — |
model-context-window-exceeded-2025-08-26 |
Stop reason model_context_window_exceeded / request max possible tokens | DOCUMENTED |
— | POST /v1/messages → anthropic-beta: model-context-window-exceeded-2025-08-26 |
— | — |
fast-mode-2026-02-01 |
Fast mode (speed: "fast") | PREVIEW |
claude-opus-5, claude-opus-4-8 | POST /v1/messages → speed |
— | — |
output-300k-2026-03-24 |
300k max_tokens on the Message Batches API | BETA |
claude-opus-5, claude-opus-4-8, claude-opus-4-7, claude-opus-4-6 … | — | — | — |
user-profiles-2026-03-24 |
User profiles API (/v1/user_profiles) | DOCUMENTED · BETA |
— | — | — | — |
user-profiles-2026-08-18 |
User profiles API revision | DOCUMENTED · BETA |
— | GET /v1/user_profiles → limit; GET /v1/user_profiles → order_by; GET /v1/user_profiles → order … (+30) |
GET /v1/user_profiles; GET /v1/user_profiles/{user_profile_id}; POST /v1/user_profiles … (+2) |
— |
user-profiles-2026-09-04 |
User profiles API revision | DOCUMENTED · BETA |
— | POST /v1/messages → anthropic-user-profile-id; POST /v1/messages/count_tokens → anthropic-user-profile-id |
— | — |
advisor-tool-2026-03-01 |
Advisor tool (advisor_20260301 server tool) | BETA |
— | POST /v1/messages → tools[].caching; POST /v1/messages → tools[].max_tokens; POST /v1/messages → tools[].model |
— | advisor_20260301 |
managed-agents-2026-04-01 |
Claude Managed Agents (agents, sessions, environments, deployments, vaults, skills in sessions, multiagent, outcomes, memory public beta 2026-04-23) | BETA |
— | DELETE /v1/environments/{environment_id} → environment_id; DELETE /v1/sessions/{session_id} → session_id; DELETE /v1/sessions/{session_id}/resources/{resource_id} → resource_id … (+586) |
DELETE /v1/environments/{environment_id}; DELETE /v1/sessions/{session_id}; DELETE /v1/sessions/{session_id}/resources/{resource_id} … (+59) |
— |
cache-diagnosis-2026-04-07 |
Cache diagnostics (diagnostics.previous_message_id -> cache_miss_reason) | BETA |
— | POST /v1/messages → diagnostics.previous_message_id; POST /v1/messages → diagnostics |
— | — |
dreaming-2026-04-21 |
Dreams (memory store reorganization) for Managed Agents | PREVIEW |
Fable 5, Sonnet 5 (2026-07-10), Opus 5 (2026-08-01) | GET /v1/dreams → created_at[gt]; GET /v1/dreams → created_at[lt]; GET /v1/dreams → include_archived … (+17) |
GET /v1/dreams; GET /v1/dreams/{dream_id}; POST /v1/dreams … (+2) |
— |
server-side-fallback-2026-06-01 |
fallbacks parameter (re-run refused requests on another model) | BETA |
— | — | — | — |
server-side-fallback-2026-07-01 |
fallbacks with "default" mode (Anthropic-recommended fallback by refusal category) | BETA |
— | POST /v1/messages → fallbacks |
— | — |
fallback-credit-2026-06-01 |
Fallback credit (billing credit when a fallback runs) | BETA |
— | — | — | — |
fallback-credit-2026-07-01 |
Fallback credit revision | BETA |
— | POST /v1/messages → fallback_credit_token |
— | — |
agent-memory-2026-07-22 |
Memory stores list semantics (stable server order, depth 0/1, path_prefix segment match) | BETA |
— | DELETE /v1/memory_stores/{memory_store_id} → memory_store_id; DELETE /v1/memory_stores/{memory_store_id}/memories/{memory_id} → expected_content_sha256; DELETE /v1/memory_stores/{memory_store_id}/memories/{memory_id} → memory_id … (+52) |
DELETE /v1/memory_stores/{memory_store_id}; DELETE /v1/memory_stores/{memory_store_id}/memories/{memory_id}; GET /v1/memory_stores … (+11) |
— |
mid-conversation-tool-changes-2026-07-01 |
Add/remove tools mid-conversation (tool_addition / tool_removal blocks in role:system messages) preserving cache | BETA |
claude-fable-5, claude-mythos-5, claude-opus-4-8, claude-opus-5 | — | — | — |
compact-2026-01-12 |
Server-side compaction (context_management compact_20260112 strategy) | BETA |
claude-fable-5-1, claude-mythos-5-1, claude-fable-5, claude-mythos-5 … | POST /v1/messages → context_management.edits[].instructions; POST /v1/messages → context_management.edits[].pause_after_compaction |
— | — |
compact-2026-09-04 |
Compact on demand (top-level compaction parameter -> signed compaction block) | BETA |
— | POST /v1/messages → compaction.instructions; POST /v1/messages → compaction.type; POST /v1/messages → compaction |
— | — |
mcp-tunnels-2026-06-22 |
MCP tunnels API (/v1/tunnels) for private-network MCP servers | PREVIEW |
— | GET /v1/tunnels → include_archived; GET /v1/tunnels → limit; GET /v1/tunnels → page … (+16) |
GET /v1/tunnels; GET /v1/tunnels/{tunnel_id}; GET /v1/tunnels/{tunnel_id}/certificates … (+7) |
— |
task-budgets-2026-03-13 |
Task budgets (advisory token budget for a whole agentic loop) | BETA |
claude-fable-5-1, claude-mythos-5-1, claude-fable-5, claude-mythos-5 … | POST /v1/messages → output_config.task_budget.remaining; POST /v1/messages → output_config.task_budget.total; POST /v1/messages → output_config.task_budget.type … (+1) |
— | — |
thinking-display-updates-2026-08-18 |
thinking.display: "updates" (progress updates between tool calls as text) | BETA |
claude-fable-5-1, claude-mythos-5-1, claude-fable-5 | — | — | — |
mid-conversation-output-config-2026-07-01 |
Per-message effort (output_config.effort on role:system messages) | BETA |
claude-fable-5-1, claude-mythos-5-1, claude-opus-5 | POST /v1/messages → messages[].output_config.effort |
— | — |
thinking-binding-controls-2026-08-01 |
Preserved-thinking binding: input_transformations response field + thinking.block_binding.prefix_mismatch_behavior | BETA |
Fable 5.1 / Mythos 5.1 (prefix check enforced for accounts created on/after 2026-08-31) | POST /v1/messages → thinking.block_binding.prefix_mismatch_behavior; POST /v1/messages → thinking.block_binding |
— | — |
mid-conversation-system-clear-at-2026-08-21 |
Turn-scoped system messages (clear_at: "next_user_message") | BETA |
— | — | — | — |
Graduated (header now optional — LEGACY): message-batches-2024-09-24, prompt-caching-2024-07-31, pdfs-2024-09-25, token-counting-2024-11-01, computer-use-2025-01-24, token-efficient-tools-2025-02-19, output-128k-2025-02-19, files-api-2025-04-14, code-execution-2025-05-22, extended-cache-ttl-2025-04-11, skills-2025-10-02, thinking-token-count-2026-05-13, structured-outputs-2025-11-13, ce-user-management-2026-07-13, fine-grained-tool-streaming-2025-05-14, search-results-2025-06-09, effort-2025-11-24. Retired / deprecated: computer-use-2024-10-22, mcp-client-2025-04-04, context-1m-2025-08-07, max-tokens-3-5-sonnet-2024-07-15.
Live findings worth knowing: mcp_servers without a header → 400 that names mcp-client-2026-09-15, which is itself rejected (use mcp-client-2025-11-20); memory-store calls reject the pair agent-memory-2026-07-22 + managed-agents-2026-04-01; speed and context_management are rejected with 400 'Extra inputs are not permitted' without their headers; output_format (legacy) → 400 even with structured-outputs-2025-11-13.
OpenAI — OpenAI-Beta values (from generated/endpoints.json / headers.json)
| Header value | Surface | Endpoints | Status |
|---|---|---|---|
OpenAI-Beta: agents=v1 |
agents-platform/agents | 43 (e.g. GET /v1/agents/environments/{environment_id}) |
BETA · DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED |
OpenAI-Beta: assistants=v2 (historical) |
agents-platform/assistants | 23 (e.g. GET /v1/assistants) |
DOCUMENTED · RETIRED |
OpenAI-Beta: chatkit_beta=v1 |
agents-platform/chatkit | 6 (e.g. POST /v1/chatkit/sessions/{session_id}/cancel) |
BETA · DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED |
OpenAI-Beta: realtime=v1 |
realtime | 2 (e.g. POST /v1/realtime/sessions) |
BETA · DOCUMENTED · FAILED_VERIFICATION · LEGACY |
OpenAI-Beta: workspace_agent_runs=v1 (required to get agent_trigger_run_id / use runs polling) |
agents-platform/workspace-agents | 2 (e.g. POST /v1/workspace_agents/{id}/trigger) |
BETA · DOCUMENTED · UNVERIFIED |
openai-beta (optional, array of strings: "Optional beta features to enable for this request") |
responses | 1 (e.g. POST /v1/responses?beta=true) |
BETA · DOCUMENTED · LIVE_VERIFIED |
Everything else on OpenAI (Responses, hosted tools incl. MCP, structured outputs, prompt caching, Batch, Files, Webhooks, Realtime GA) needs no beta header; POST /v1/responses?beta=true exposes a beta schema (multi_agent) with an optional body-level openai-beta[]. OpenAI-Beta: realtime=v1 is legacy (the beta endpoints return 404).
Gemini — /v1beta vs /v1, -preview models, BETA/PREVIEW endpoints (from generated/endpoints.json, models.json)
# Gemini surfaces that exist only under /v1beta
jq -r '[.[] | select(.provider=="gemini" and (.path|startswith("/v1beta/")))] | group_by(.api_family) | .[] | "\(.[0].api_family)\t\(length)"' generated/endpoints.json
# preview / experimental Gemini model ids still served
jq -r '.[] | select(.provider=="gemini" and (.kind=="preview" or .kind=="experimental") and (.status|index("RETIRED")|not)) | .id' generated/models.json- Stable
/v1surface (9 paths recorded):/v1/interactions,/v1/interactions/{id},/v1/interactions/{id}/cancel,/v1/models,/v1/models/{model}:generateContent,/v1/webhooks,/v1/webhooks/{id},/v1/webhooks/{id}/rotate_secret,/v1/webhooks/{id}:ping. Everything else — including explicit context caching (cachedContents), Files, File Search stores, Batch, embeddings,countTokens, Live API (v1betaWebSocket), ephemeralauth_tokens, VeopredictLongRunning, the OpenAI-compatibility layer and the managed-agents preview (agents,environments,credentials,triggers) — is v1beta-only (families:agents,auth-tokens,batch,caching,credentials,dynamic,embeddings,environments,file-search,files,image-generation,legacy-palm,openai-compat,operations,tokens,transcription,triggers,tuning,video-generation)./v1/cachedContentsreturns 404. - Endpoints recorded BETA/PREVIEW (64):
agents(4),auth-tokens(1),caching(5),credentials(5),dynamic(2),environments(7),file-search(12),interactions(4),live(2),music-generation(1),openai-compat(7),triggers(7),video-generation(1),webhooks(6).POST /v1beta/interactionsis BETA + LIVE_VERIFIED; its/v1/interactionstwin is documented GA but UNVERIFIED. - Preview / experimental model ids still served (25):
gemini-2.5-flash-preview-tts(PREVIEW),gemini-2.5-pro-preview-tts(PREVIEW),gemini-3-flash-preview(PREVIEW),gemini-3.1-pro-preview(PREVIEW/ACCOUNT_RESTRICTED),gemini-3.1-pro-preview-customtools(PREVIEW),gemini-omni-flash-preview(PREVIEW/DEPRECATED),lyria-3-clip-preview(PREVIEW),lyria-3-pro-preview(PREVIEW),gemini-3.1-flash-tts-preview(PREVIEW),gemini-robotics-er-2-preview(PREVIEW),gemini-2.5-computer-use-preview-10-2025(PREVIEW),gemini-embedding-2-preview(PREVIEW),veo-3.1-generate-preview(PREVIEW),veo-3.1-fast-generate-preview(PREVIEW),veo-3.1-lite-generate-preview(PREVIEW),gemini-2.5-flash-native-audio-preview-09-2025(PREVIEW),gemini-2.5-flash-native-audio-preview-12-2025(PREVIEW),gemini-3.1-flash-live-preview(PREVIEW),gemini-robotics-er-2-streaming-preview(PREVIEW),gemini-3.5-live-translate-preview(PREVIEW),lyria-realtime-exp(BETA),gemini-2.0-flash-exp(BETA/UNVERIFIED),gemini-2.5-flash-preview-09-2025(PREVIEW/UNVERIFIED),lyria-3.5-clip-preview(PREVIEW/UNVERIFIED),lyria-3.5-pro-preview(PREVIEW/UNVERIFIED). Preview models 'may be used in production' but carry more restrictive limits and can be deprecated with two weeks' notice;-latestaliases (gemini-flash-latest,gemini-pro-latest,gemini-flash-lite-latest) are hot-swapped. - Gated by account tier rather than header: Pro models and paid-only media models return 429
limit: 0on the free tier (gemini-3.1-pro-preview,gemini-pro-latest,cachedContentscreate — allACCOUNT_RESTRICTEDhere); Gemini 2.5 models return 404 'no longer available to new users'.
xAI — no beta headers; gating is by ACL / alpha program (from generated/endpoints.json, tools.json, headers.json)
headers.jsonrecords no versioning / beta headers for xAI: features are selected by body fields (service_tier,reasoning_effort,store,prompt_cache_key,include) or by base URL (us.api.x.ai,management-api.x.ai).- ACCOUNT_RESTRICTED endpoints (41):
collections(11),embeddings(1),management(21),models(2),skills(5),voice(1) — the Management API needs a separate Management key (401 code 16 with an inference key),/v1/embeddingsandgrok-embedding-smallreturn 404 for this team,/v1/skillsreturns 404, custom-voice creation is Enterprise-only, and Collections are documented onmanagement-api.x.aialthough the same paths answered onapi.x.ai(LIVE_DISCOVERED). - Tools:
tool_search→DOCUMENTED·ACCOUNT_RESTRICTED;live_search→DOCUMENTED·RETIRED.tool_search(+defer_loading) answers 403 'only available for alpha users';live_search/search_parameters/web_search_optionsanswer 410. - Models flagged BETA/PREVIEW:
grok-4.20-multi-agent-0309(BETA),grok-build-0.1(PREVIEW) (multi-agent is Responses-only; Grok Build is 'early access'). Regional hosteu-west-1.api.x.aiis undocumented (LIVE_DISCOVERED).
Q7. Which model supports combination Z?
Recipe: filter generated/models.json on capabilities.* booleans. Key families — OpenAI: image_in, reasoning, structured_outputs, function_calling, prompt_caching, tool_web_search, tool_code_interpreter, tool_mcp, tool_computer_use, tool_hosted_shell, tool_apply_patch, tool_skills, tool_tool_search, fine_tuning, batch, audio_in/out, realtime; Anthropic: vision, pdf_input, structured_outputs_json, strict_tool_use, prompt_caching, web_search, web_fetch, code_execution, programmatic_tool_calling, computer_use_toolset_ga, browser_use, mcp_connector, agent_skills, tool_search, effort_parameter, thinking_adaptive, context_1m_default_no_beta, fast_mode_speed_fast, batch_api, zero_data_retention_eligible; xAI: reasoning, reasoning_can_be_disabled, reasoning_effort_levels[], function_calling, structured_outputs, prompt_caching_automatic, batch_api, priority_processing, context_compaction, websocket_responses, web_search, x_search, code_execution, collections_search, remote_mcp, image_generation_tool, files_attachments, encrypted_reasoning_content, anthropic_messages_compat; Gemini: thinking, thinking_level_param, thinking_levels[], thought_signatures, structured_output, function_calling, parallel_function_calling, compositional_function_calling, google_search_grounding, google_maps_grounding, url_context, code_execution, computer_use, file_search, context_caching_explicit, context_caching_implicit, batch_api, flex_inference, priority_inference, live_api, interactions_api, openai_compatible_chat, audio_input, video_input, pdf_input, image_output, audio_output, embeddings, tuning. Flat matrices: generated/compatibility/model-capability-matrix.json (+ .csv), gemini-feature-model-matrix.json, gemini-tool-model-matrix.json.
# generic pattern
jq -r '.[] | select(.provider=="anthropic" and .capabilities.context_1m_default_no_beta==true and .capabilities.computer_use_toolset_ga==true and .capabilities.web_search==true and (.status|index("RETIRED")|not)) | .id' generated/models.json
jq -r '.[] | select(.provider=="gemini" and .capabilities.live_api==true and .capabilities.function_calling==true and .capabilities.google_search_grounding==true and (.status|index("RETIRED")|not)) | .id' generated/models.json| Combination | Provider | Matching models (status) |
|---|---|---|
| Z1 (OpenAI): ≥1M context + structured outputs + prompt caching + web search + code interpreter + computer use | openai | gpt-5.4, gpt-5.4-2026-03-05, gpt-5.5, gpt-5.5-2026-04-23, gpt-5.6, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, gpt-daybreak-blue-latest (ACCOUNT_RESTRICTED) |
| Z1 (Anthropic): 1M context (no header) + structured outputs + web search + code execution + computer-use toolset GA | anthropic | claude-fable-5-1, claude-mythos-5-1 (ACCOUNT_RESTRICTED/PREVIEW), claude-fable-5 (LEGACY), claude-mythos-5 (ACCOUNT_RESTRICTED/PREVIEW), claude-opus-5, claude-opus-4-8 (LEGACY), claude-sonnet-5 |
| Z1 (xAI): ≥1M context + structured outputs + automatic prompt caching + web_search + code_execution + batch | xai | — none |
| Z1 (Gemini): ≥1M context + structured output + explicit context caching + Google Search grounding + code execution + computer use | gemini | gemini-3-flash-preview (PREVIEW), gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash, gemini-3.7-flash, gemini-3.8-flash |
Z2 (OpenAI): reasoning can be switched off (effort: none accepted) + function calling + image input |
openai | gpt-5.1, gpt-5.1-2025-11-13, gpt-5.2, gpt-5.2-2025-12-11, gpt-5.4, gpt-5.4-2026-03-05, gpt-5.4-mini, gpt-5.4-mini-2026-03-17, gpt-5.4-nano, gpt-5.4-nano-2026-03-17, gpt-5.5, gpt-5.5-2026-04-23, gpt-5.6, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra |
| Z2 (Anthropic): thinking can be disabled + effort parameter + tool use | anthropic | claude-mythos-preview (ACCOUNT_RESTRICTED/DEPRECATED/PREVIEW), claude-opus-5, claude-opus-4-8 (LEGACY), claude-opus-4-7 (LEGACY), claude-opus-4-6 (LEGACY), claude-opus-4-5-20251101 (LEGACY), claude-sonnet-5, claude-sonnet-4-6 (LEGACY) |
Z2 (xAI): reasoning can be disabled (reasoning_effort: none) or is absent + function calling + image input |
xai | grok-4.3, grok-4.20-0309-non-reasoning |
Z2 (Gemini): thinkingLevel parameter with minimal accepted + function calling + image input |
gemini | gemini-3-flash-preview (PREVIEW), gemini-3.1-flash-lite (DEPRECATED), gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash |
| Z3 (OpenAI): fine-tuning supported (self-serve winding down, 2027-01-06) | openai | babbage-002 (DEPRECATED/LEGACY), davinci-002 (DEPRECATED/LEGACY), gpt-3.5-turbo (DEPRECATED/LEGACY), gpt-3.5-turbo-0125 (DEPRECATED/LEGACY), gpt-3.5-turbo-1106 (DEPRECATED/LEGACY), gpt-3.5-turbo-instruct (DEPRECATED/LEGACY), gpt-4 (DEPRECATED/LEGACY), gpt-4-0613 (DEPRECATED/LEGACY), gpt-4.1, gpt-4.1-2025-04-14, gpt-4.1-mini, gpt-4.1-mini-2025-04-14, gpt-4.1-nano (DEPRECATED), gpt-4.1-nano-2025-04-14 (DEPRECATED), gpt-4o, gpt-4o-2024-05-13 (DEPRECATED/LEGACY), gpt-4o-2024-08-06, gpt-4o-2024-11-20, gpt-4o-mini, gpt-4o-mini-2024-07-18, o4-mini (DEPRECATED), o4-mini-2025-04-16 (DEPRECATED) |
| Z3 (Gemini): tuning supported on the Developer API | gemini | — none |
| Z4 (OpenAI): audio in + audio out (chat or realtime) | openai | gpt-4o-audio-preview-2024-12-17 (PREVIEW), gpt-4o-audio-preview-2025-06-03 (PREVIEW), gpt-4o-mini-audio-preview-2024-12-17 (PREVIEW), gpt-4o-mini-realtime-preview-2024-12-17 (PREVIEW), gpt-audio (DEPRECATED), gpt-audio-1.5, gpt-audio-2025-08-28, gpt-audio-mini (DEPRECATED), gpt-audio-mini-2025-12-15, gpt-live-1, gpt-realtime (DEPRECATED), gpt-realtime-1.5, gpt-realtime-2, gpt-realtime-2.1, gpt-realtime-2.1-mini, gpt-realtime-2025-08-28, gpt-realtime-mini (DEPRECATED), gpt-realtime-mini-2025-12-15, gpt-realtime-translate |
Z4 (Anthropic): speed: fast supported |
anthropic | claude-opus-5, claude-opus-4-8 (LEGACY) |
| Z4 (xAI): speech-to-speech + function calling + web/x search in the voice session | xai | grok-voice-think-fast-2.0 |
| Z4 (Gemini): Live API (audio in + audio out) + function calling + Google Search | gemini | gemini-2.5-flash-native-audio-preview-12-2025 (PREVIEW), gemini-3.1-flash-live-preview (PREVIEW), gemini-3.8-live |
| Z5 (OpenAI): flex tier + batch + 24h extended cache retention documented | openai | gpt-5, gpt-5-2025-08-07 (DEPRECATED), gpt-5.1, gpt-5.1-2025-11-13, gpt-5.2, gpt-5.2-2025-12-11, gpt-5.4, gpt-5.4-2026-03-05, gpt-5.5, gpt-5.5-2026-04-23, gpt-5.5-pro, gpt-5.5-pro-2026-04-23, gpt-5.6, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra |
| Z5 (Anthropic): zero-data-retention eligible + programmatic tool calling + MCP connector | anthropic | claude-mythos-preview (ACCOUNT_RESTRICTED/DEPRECATED/PREVIEW), claude-opus-5, claude-opus-4-8 (LEGACY), claude-opus-4-7 (LEGACY), claude-opus-4-6 (LEGACY), claude-opus-4-5-20251101 (LEGACY), claude-sonnet-5, claude-sonnet-4-6 (LEGACY), claude-sonnet-4-5-20250929 (LEGACY) |
| Z5 (xAI): priority processing + WebSocket Responses + context compaction + remote MCP | xai | — none |
| Z5 (Gemini): flex + priority + batch tiers + implicit caching | gemini | gemini-2.5-flash, gemini-2.5-pro, gemini-2.5-flash-lite (ACCOUNT_RESTRICTED), gemini-3-flash-preview (PREVIEW), gemini-3.1-pro-preview (PREVIEW/ACCOUNT_RESTRICTED), gemini-3.1-pro-preview-customtools (PREVIEW), gemini-3.1-flash-lite (DEPRECATED), gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash, gemini-3.7-flash, gemini-3.8-flash |
| Z6 (OpenAI): skills + hosted shell + apply_patch + tool search (full coding toolset) | openai | gpt-5.4, gpt-5.4-2026-03-05, gpt-5.4-mini, gpt-5.4-mini-2026-03-17, gpt-5.5, gpt-5.5-2026-04-23, gpt-5.6, gpt-5.6-cyber (ACCOUNT_RESTRICTED), gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, gpt-daybreak-blue-latest (ACCOUNT_RESTRICTED), gpt-daybreak-red-latest |
| Z6 (Anthropic): per-message effort (beta) + task budgets (beta) + adaptive thinking | anthropic | claude-fable-5-1, claude-mythos-5-1 (ACCOUNT_RESTRICTED/PREVIEW), claude-opus-5 |
| Z6 (xAI): x_search + collections_search + image_generation tool + files attachments (full Grok agent toolset) | xai | — none |
| Z6 (Gemini): Google Maps grounding + URL context + File Search + thought signatures (full Gemini 3 toolset) | gemini | gemini-3-flash-preview (PREVIEW), gemini-3.1-pro-preview (PREVIEW/ACCOUNT_RESTRICTED), gemini-3.1-pro-preview-customtools (PREVIEW), gemini-3.1-flash-lite (DEPRECATED), gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash, gemini-3.7-flash, gemini-3.8-flash |
| Z7 (Gemini): free tier + Google Search grounding + structured output (zero-cost prototyping) | gemini | gemini-2.5-flash, gemini-2.5-pro, gemini-2.5-flash-lite (ACCOUNT_RESTRICTED), gemini-3-flash-preview (PREVIEW), gemini-3.1-flash-lite (DEPRECATED), gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash, gemini-3.7-flash, gemini-3.8-flash, gemini-robotics-er-2-preview (PREVIEW) |
Z7 (xAI): OpenAI-compatible chat + Anthropic-compatible /v1/messages + deferred completions |
xai | grok-4.6, grok-4.5, grok-4.3, grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning, grok-build-0.1 (PREVIEW) |
Snapshot ids (e.g. gpt-5.4-2026-03-05) share their alias's capabilities and appear in the raw query output; the table above keeps them because they are distinct callable ids. Capability flags marked "unknown" in the source never match a == true filter — check the model page before excluding a model on that basis (many Gemini media/agent records and Gemma 4 carry "unknown" for tool flags). Gemini computer_use / file_search may be the string "Supported (Preview)" / "Supported (AI Studio only)" — treated as true above (✔*).
Q8. Which provider offers X, and at what price?
Computed from generated/pricing.json (tool/service rows) and the pricing blocks of generated/models.json. Units are quoted as recorded; '—' = no documented surface. Statuses are those of the underlying tool/model record.
# every per-call tool price across providers
jq -r '.[] | select(.unit|test("call|search|request|prompt";"i")) | [.provider, .model_or_service, .dimension, (.price|tostring), .unit, (.tier//"")] | @tsv' generated/pricing.json| Capability | OpenAI | Anthropic | xAI | Gemini |
|---|---|---|---|---|
| Web search (per 1k) | $10 per 1k calls (web_search; preview $25/1k on non-reasoning models) |
$10 per 1,000 searches (web_search_2026*; failed searches free) |
$5 per 1K calls (web_search, image search included) |
Gemini 3.x: $14 per 1K requests after 5,000 free/month (billed per query); Gemini 2.5: $35 per 1K grounded prompts after 1,500 RPD free; tool googleSearch ACCOUNT_RESTRICTED on this key |
| Social / maps search | — | — | x_search $5 per 1K calls until 2026-09-21, then $5 per 1K posts + $10 per 1K profiles |
googleMaps $14 per 1K search queries (Gemini 3.x, 5,000 free/month); $25 / 1k grounded prompts on 2.5 |
| URL fetch / context | — (web search opens pages) | web_fetch $0 per fetch (content tokens) |
— (open_page action of web search) |
urlContext $0 per call + content billed as input tokens |
| Code execution | $0.03 per 20-minute session per container (by size) (1 GB; 4 GB $0.12, 16 GB $0.48, 64 GB $1.92; per-minute billing since 2026-06-02) | $0.05 per hour per container after 1,550 free container-hours/org/month; free alongside web_search/web_fetch 20260209+ | $5 per 1K calls (code_interpreter / code_execution) |
$0 per call — codeExecution has no fee; code + results billed as tokens |
| Managed RAG / file search | $2.5 per 1k calls + storage $0.1 per GB per day (1 GB free) | — (you retrieve; search_result blocks) |
file_search / collections_search $2.5 per 1K calls + storage $0.1 per GiB per day; implicit attachment_search $10 per 1K calls |
fileSearch indexing $0.15 per 1M tokens once; storage and queries free; retrieved chunks billed as input |
| Image generation (per image) | gpt-image-2: $0.006 per image (low 1024²) … $0.211 per image (high 1024²); token rates $5 text in / $8 image in / $30 image out per 1M | — | grok-imagine-image $0.02; grok-imagine-image-2.0 $0.04–$0.08 (quality × resolution); grok-imagine-image-quality $0.05 (DEPRECATED) | gemini-3.1-flash-lite-image $0.0336 (1K); gemini-3.1-flash-image $0.045–$0.151 (0.5K–4K); gemini-3-pro-image $0.134 (1K/2K) / $0.24 (4K); batch 50 % |
| Realtime voice (audio minutes) | gpt-realtime-2.1 audio $32 per 1M tokens in / $64 per 1M tokens out (≈ $0.06 in + $0.24 out per minute at ~10 tok/s... token-based); gpt-live-1 $0.05 per minute (billed per second) | — | grok-voice-think-fast-2.0 $0.08 / min ($4.8 / h) audio sent or received + $0.004 per text item | gemini-3.8-live audio in $3.0 / 1M (≈ $0.005 / min), audio out $12.0 / 1M (≈ $0.018 / min); free tier available |
| Text-to-speech | gpt-4o-mini-tts $0.6 per 1M tokens text in + $12 per 1M tokens audio out; tts-1 $15 / 1M chars | — | $15 per 1M characters (POST /v1/tts, 28 voices, language required) |
gemini-3.1-flash-tts-preview $1.0 text in + $20.0 audio out per 1M (25 tok/s ≈ $0.030 / min); free tier |
| Speech-to-text | gpt-transcribe $0.0045 per minute; whisper-1 $0.006 / min (→ 2027-02-26) | — | $0.1 per hour REST, $0.2 per hour streaming (grok-voice-transcribe-2.0 / 1.0) | gemini-3.5-transcribe audio in $2.0 / 1M (≈ $0.003 / min) + text out $12.0 / 1M (blended blended ~$0.005/min); free tier |
| Embeddings (per 1M tokens) | text-embedding-3-small $0.02 per 1M tokens, -large $0.13 per 1M tokens | — | grok-embedding-small: no published price (ACCOUNT_RESTRICTED, 404) |
gemini-embedding-2 text $0.2 (batch $0.1); image $0.45, audio $6.5, video $12.0; free tier |
| Video generation (per second) | sora-2 $0.10 (720p), sora-2-pro $0.30–$0.70 — shutdown 2026-09-24 | — | grok-imagine-video $0.05; grok-imagine-video-1.5 $0.08 | Veo 3.1 $0.4 (720p/1080p) / $0.6 (4K); Veo 3.1 Fast $0.1 / $0.12 / $0.3; Veo 3.1 Lite $0.05 / $0.08 |
| Music generation | — | — | — | lyria-3.5 $0.08 per song (full length); lyria-3-clip-preview $0.04 per song (30 s clip) |
| Token counting | free (POST /v1/responses/input_tokens) |
free (POST /v1/messages/count_tokens) |
free (POST /v1/tokenize-text) |
free (:countTokens) |
Provider-specific billing quirks that change the arithmetic: xAI bills reasoning tokens on every Grok call and applies its ≥200k-token long-context rate to the whole request; Gemini's 3.6–3.8 Flash prices double on 2027-01-01 and its free tier covers most Flash/Live/TTS/embedding models; Anthropic and OpenAI charge cache writes (1.25×) on their newest models while xAI and Gemini implicit caching have none (Gemini explicit caches pay storage per token-hour). Full tables and cost models: pricing comparison.
Q9. Where do I look for … ?
| Question | File(s) |
|---|---|
| Every endpoint, with auth/beta/verification | generated/endpoints.json (870 records) · docs/endpoints/index.md · by-status |
| Every parameter of an endpoint | generated/parameters.json (7,862 rows) filtered by endpoint |
| Model facts (context, output, cutoff, prices, capabilities, cloud ids) | generated/models.json (396 records) · docs/comparisons/models.md · docs/models/{openai,anthropic,xai,gemini}-models.md |
| Tool definitions & result shapes | generated/tools.json (69 records) · docs/tools/{openai,anthropic,xai,gemini}/*.md |
| Prices | generated/pricing.json (1,522 rows) · docs/comparisons/pricing.md · docs/{openai,anthropic,xai,gemini}/pricing.md |
| Errors and retry rules | generated/errors.json · docs/errors/{openai,anthropic,gemini}.md · docs/xai/authentication-headers-errors.md |
| Headers (request/response/webhook) | generated/headers.json · generated/fragments/headers/{anthropic-beta-headers,anthropic-headers,openai-headers,xai-headers,gemini-headers}.json |
| Streaming & webhook events | generated/streaming-events.json (506) · generated/webhook-events.json (72; OpenAI + Anthropic only — xAI has the single SIP realtime.call.incoming, Gemini Interactions use webhook_config / /v1/webhooks) |
| Deprecations / retirements | generated/deprecations.json · docs/openai/deprecations.md · docs/anthropic/deprecations.md · docs/xai/deprecations-and-release-notes.md · docs/gemini/deprecations-and-changelog.md |
| Feature-by-feature comparison (4 providers) | docs/comparisons/features.md · generated/compatibility/cross-provider-feature-matrix.json |
| Rate limits | generated/rate-limits.json · docs/openai/rate-limits.md · docs/anthropic/rate-limits.md · docs/xai/rate-limits.md · docs/gemini/rate-limits.md |
| SDKs | generated/sdks.json · docs/openai/sdks.md · docs/anthropic/sdks.md · docs/xai/sdks.md · docs/gemini/sdks.md |
| Cloud platform availability | generated/compatibility/anthropic-feature-platform-matrix.json · docs/anthropic/cloud-providers.md · docs/openai/data-residency-and-regions.md · docs/gemini/vertex-vs-gemini-api.md · docs/xai/sdks.md (Vertex Model Garden / Foundry) |
| OpenAI-compatibility layers | docs/xai/chat-completions.md · docs/xai/messages-compat.md · docs/gemini/openai-compatibility.md |
| Runnable examples and their status | generated/examples-manifest.json · examples/ |
| Live calls made by this atlas (cost, status) | reports/live-requests.jsonl |