API Atlas — Final report (OpenAI + Anthropic + xAI + Gemini, runs of 2026-09-18/19 UTC)
Snapshot of the public API surface of OpenAI, Anthropic, xAI (Grok) and Google Gemini, built from the
current official documentation (downloaded, hashed, kept under sources/), the official machine-readable specs
(OpenAI OpenAPI v2.3.0, xAI OpenAPI 3.1, Google discovery documents v1beta rev. 20260918 and v1), the official SDK
surfaces, and 5,811 real API calls made with the project's own keys (every one logged, sanitized, in
reports/live-requests.jsonl). Machine-readable outputs are in generated/; human docs in docs/; runnable examples in
examples/; smoke tests in tests/. Numbers below come from reports/coverage.json and generated/build-report.json.
1. Coverage
| Metric | OpenAI | Anthropic | xAI | Gemini | Total |
|---|---|---|---|---|---|
| Official documentation pages (Markdown twins / split export, HTTP 200) | 551 | 696 | 182 | 291 | 1,720 |
| Machine-readable spec | OpenAPI (352 ops) + SDK api.md |
SDK api.md (201 endpoints) |
OpenAPI (38 paths, 44 ops) + gRPC stubs (11 services) | discovery v1beta (86 methods) + v1 (47) + SDK types | — |
| Model records (canonical + snapshots + aliases + retired/id-only) | 230 | 36 | 33 | 97 | 396 |
| Model ids returned live by the models endpoint | 136 | 11 | 12 (+7/3/2 typed) | 58 (v1beta) / 22 (v1) | 217 |
Model records LIVE_VERIFIED |
137 | 21 | 12 | 10 | 180 |
| Endpoint records | 359 | 282 | 104 | 125 | 870 |
Tool records (exact type strings) |
19 | 27 | 12 | 11 | 69 |
| Parameter records (dotted paths) | 4,215 | 2,040 | 891 | 716 | 7,862 |
| Streaming / WebSocket event records | 282 | 68 | 94 | 62 | 506 |
| Error records | 50 | 46 | 12 | 18 | 126 |
| Header records | 33 | 22 | 15 | 16 | 86 |
| Price records | 717 | 158 | 120 | 527 | 1,522 |
| Deprecation records | 165 | 1 (nested: 14 models + 35 items) | 1 (nested) | 1 (nested) | 168 |
| Object/schema records | 189 | 224 | 80 | 213 | 706 |
| SDK records | 9 | 9 | 9 | 8 | 35 |
| Webhook / audit-log event types | 28 / 149 | 44 / — | — | — | 72 / 149 |
Examples in manifests (LIVE_VERIFIED where run) |
95 | 95 | 39 (+17 media) | 82 | 311 |
Example files (.sh / .py / .ts / .html) |
— | — | — | — | 425 |
| Test files | — | — | — | — | 63 |
Test results (.venv/bin/python -m pytest tests) |
— | — | — | — | 448 passed, 67 skipped (gated), 2 flaky tests fixed and re-run green (full run 8 min 18 s on 2026-09-19) |
| Live API calls | 2,515 | 1,735 | 588 | 973 | 5,811 |
| Estimated spend (upper bound; several agents logged conservative estimates) | ≈ $0.30 | ≈ $0.33 | ≈ $0.80 (actual ≈ $0.45) | ≈ $0.26 (actual ≈ $0.02, free tier) | ≈ $1.68 logged / ≈ $1.1 actual |
| Human documentation pages | — | — | — | — | 228 Markdown files |
| Cross-provider feature matrix | — | — | — | — | 125 features: on all 4 = 45, on 3 = 34, on 2 = 16, unique = 26, legacy = 4 |
Endpoint statuses (an endpoint can carry several): OpenAI 123 LIVE_VERIFIED, 18 ACCOUNT_RESTRICTED, 56 BETA, 24 RETIRED;
Anthropic 47 LIVE_VERIFIED, 59 ACCOUNT_RESTRICTED, 111 BETA; xAI and Gemini per-status buckets are in
docs/endpoints/by-status.md.
2. What was actually exercised live (highlights)
OpenAI — Responses (create/stream/retrieve/delete/cancel/input_items/input_tokens/compact/background + resume), Chat Completions (incl. store → retrieve → messages → delete), legacy Completions, Conversations, function calling (forced, strict, streaming, round-trip), custom grammar tools, web_search, file_search (vector store round trip), code_interpreter + Containers API, MCP (approval flow), shell, apply_patch, tool_search, Skills API, Agents API (agent → session → streamed turn → items/turns/subagents/artifacts → delete; environment templates; vault list), ChatKit list, Realtime (client_secrets, one WebSocket session with full event sequence), Audio (TTS, transcription variants, translation), Images (one generation + validation errors), provenance checks (C2PA detected), Embeddings, Moderation, Files, Uploads, Vector Stores, Batch (5 batches end to end), Graders, Evals, webhook event types, ~100 deliberate error probes, all Admin GET endpoints (403/401 = ACCOUNT_RESTRICTED).
Anthropic — Messages (system/multi-turn/prefill/stop sequences/metadata/service_tier/streaming with exact SSE order), count_tokens, custom tools (all tool_choice modes, parallel, fine-grained streaming), web_search, web_fetch, code_execution (+ container reuse, skills), tool_search, programmatic tool calling, computer use round trip, text editor / bash / memory shapes, MCP connector, advisor, prompt caching (5 m / 1 h TTL numbers), extended/adaptive/interleaved thinking, effort, fast mode (429 quota 0), 1M context header, context editing + compaction, structured outputs, citations (all location types), vision, PDF, Files API, Message Batches (2 batches end to end), Skills API, Managed Agents (environment → agent → session → streamed events round trip), Admin GET /v1/organizations/me (200) and 21 other Admin endpoints (401), Models API pagination, ~50 deliberate error probes.
xAI — Models catalogue endpoints (models, language/image/video/embedding models, api-key, me, tokenize-text; redirect aliases grok-3/grok-4-fast/grok-4-0709 → grok-4.3 observed), regional hosts us.api.x.ai and undocumented eu-west-1.api.x.ai, Chat Completions (reasoning_content, reasoning_effort per model, structured output, function calling, image input, deferred completions, streaming), Responses API (minimal/stream/previous_response_id/store/compact/input_items; server tools web_search, x_search, code_interpreter, mcp, file_search shape; 422 enum list of accepted tool types), Anthropic-compatible /v1/messages (text/image/tool_use/thinking blocks, streaming), legacy completions (400 for all live models), Files (upload/list/get/content/public-url/revoke/delete), Collections on api.x.ai (create/add/search/delete), Batches (create/list/get/add/cancel), image generation + edit, one 1-second video job (202 pending → 200), Voice realtime WebSocket text turn (18 events), TTS/STT/voices/client_secrets, grok-build-0.1, Management API (401 = ACCOUNT_RESTRICTED), Skills API (404 = ACCOUNT_RESTRICTED).
Gemini — Models list (v1beta paginated, v1), model get incl. the "no longer available to new users" 404 for 2.5 models, generateContent on 7 models (thoughtSignature, serviceTier), streaming (?alt=sse and JSON-array framing), system instructions and every generationConfig knob (with 400s recorded for those "not enabled for this model"), structured output (responseFormat / responseSchema / responseJsonSchema / enum), thinking (budget, levels, disabled), safety settings, countTokens (text/system/image/tools/PDF/YouTube), embeddings (001 and 2, dimensions), Files (resumable/multipart/raw uploads, register, use, delete), function calling (ANY/NONE/VALIDATED, parallel, JSON-schema params, thoughtSignature echo rule with the 400 on omission), urlContext, codeExecution, googleMaps grounding, File Search (store → upload → operation → documents → grounded query → delete), Interactions API (create/stream/chain/get/delete/cancel; Deep Research background create + cancel), Environments CRUD, ephemeral auth tokens (+ constrained Live method), Live API audio WebSocket turn sequence (text modality rejected), TTS (two models, streaming), transcription model (undocumented audioTranscription part), Lyria RealTime WebSocket (2 audio chunks), OpenAI-compatibility layer (chat + models), legacy PaLM methods (404/501), tunedModels (501). Paid-tier features refused with 429 quota 0 / 400 FAILED_PRECONDITION: Pro models, image generation, Google Search grounding, context caching, Batch, Lyria 3, Veo (validation-only probes).
3. Missing coverage (honest list)
| Area | Why | Status recorded |
|---|---|---|
| OpenAI Admin API (124 endpoints), usage/costs, audit logs, certificates, external storage, WIF | no Admin key (Missing scopes: api.management.read …) |
ACCOUNT_RESTRICTED; shapes from the OpenAPI spec |
| OpenAI computer-use-preview, codex-mini-latest, gpt-5.6-cyber, Daybreak, gpt-oss hosted; Anthropic Mythos ids | 404 model_not_found — gated or retired |
ACCOUNT_RESTRICTED / RETIRED per docs |
| OpenAI video create/remix/edit, image edits/streaming, deep research runs, fine-tuning job creation (403), Realtime WebRTC/SIP, Live sessions, voice consents, Responses WebSocket, Agents hosted sandboxes | cost / account / needs peer | DOCUMENTED / UNVERIFIED / FAILED_VERIFICATION |
Anthropic fast mode, per-message effort, Managed Agents tunnels/dreams/user profiles, Admin API beyond /me, cloud platforms |
account gating / no cloud credentials | ACCOUNT_RESTRICTED / DOCUMENTED |
xAI Management API (billing, audit, keys, ACLs), Skills API, embedding models, tool_search (alpha), grok-4.6/4.5 server-tool matrix, video edits/extensions, custom voices, SIP, Collections indexed-search happy path |
401/404/403 gating or cost; indexing did not finish in 75 s | ACCOUNT_RESTRICTED / DOCUMENTED / UNVERIFIED |
Gemini paid-tier features on this key: Pro models, native image generation, Google Search grounding, explicit context caching, Batch, Lyria 3, Veo generation, computer-use legacy model; Vertex AI; Interactions credentials/agents/triggers/webhooks; dynamic/* methods |
the Gemini key is free tier (429 limit: 0 / 400 FAILED_PRECONDITION); no Vertex credentials; not exercised |
ACCOUNT_RESTRICTED / DOCUMENTED |
| Anthropic MCP connector and Gemini/xAI MCP against our own local MCP server | needs a public URL | DOCUMENTED + local server test |
4. Major findings
Present on all four: JSON-Schema function calling with a client-side loop; streaming; structured JSON output; reasoning with some control (effort / thinking / reasoning_effort / thinkingConfig); prompt/context caching in some form; files APIs; image input; server-side web search on the three that expose tools inside generation (OpenAI, Anthropic, xAI) plus Gemini grounding (paid); code execution sandboxes; MCP in some form; batch processing (OpenAI, Anthropic, xAI −20 %, Gemini paid); agents platforms (Agents API, Managed Agents, xAI Responses agentic loop, Gemini Interactions API); OpenAI-style wire formats reused by xAI natively and by Gemini through its compatibility layer.
Unique to OpenAI: Chat Completions + Conversations as first-class stateful surfaces, background mode with resumable streams, WebSocket Responses, Realtime/Live speech-to-speech over WebRTC/SIP, moderation endpoint, provenance checks, vector stores, fine-tuning/evals/graders (winding down), ChatKit/Workspace Agents, 149 audit-log event types, project/service-account/role model.
Unique to Anthropic: web_fetch, native citations, server-side context editing/compaction, adaptive/interleaved thinking with signed blocks, count_tokens endpoint, memory/advisor/browser tools, programmatic tool calling, Managed Agents memory stores and scheduled deployments, required anthropic-version and 50 catalogued beta headers.
Unique to xAI: per-request cost_in_usd_ticks, X (Twitter) search as a server tool, an Anthropic-compatible Messages endpoint next to the OpenAI-compatible ones, retired-model redirect aliases, regional endpoints with a 1.1× US price uplift, a 500k-context model priced in two tiers around 200k prompt tokens, multi-agent model variant on Responses.
Unique to Gemini: a documented free tier, native audio/video/PDF understanding on text models with per-modality token accounting, music generation (Lyria 3 + RealTime), Google Search/Maps grounding, URL context, thoughtSignature mandatory echo, Interactions API with Deep Research agents and environments, storage-priced (per token-hour) context caching, alt=sse vs JSON-array streaming.
Legacy and retirements observed live: OpenAI Assistants gone, DALL·E gone, Sora/Videos shutdown 2026-09-24, local_shell unsupported, Evals shutdown 2026-11-30, self-serve fine-tuning ends 2027-01-06; Anthropic /v1/complete 400 "deprecated", Claude 3.x ids 404; xAI Live Search 410, /v1/completions and /v1/complete 400 for every live model, grok-2-image 404, grok-3/4-fast silently redirected; Gemini 2.5 models 404 for new users, Imagen predict 404, tunedModels 501, PaLM methods 404/501.
Undocumented-but-live (LIVE_DISCOVERED): OpenAI phase/billing/tool_usage/prompt_cache_retention/personality/shutdown_date; Anthropic stop_details/container/usage.cache_creation.*/usage.inference_geo/tool_use.caller/file_id on citations; xAI eu-west-1.api.x.ai, x-ratelimit-* headers, usage.context_details, video polling HTTP 202, realtime ping and billable_audio_seconds, reasoning_effort: none on grok-4.3, store:false responses still retrievable, x_search emitting custom_tool_call; Gemini Model.thinking/maxTemperature, X-Gemini-Service-Tier header, audioTranscription part, generateContent?alt=sse, files:register, gemini-flash-latest → 3.8-flash.
Documentation vs live discrepancies worth knowing: o3 priced differently on two OpenAI pages; xAI catalogue prices are ticks (1e-10 USD) while docs say cents; xAI reasoning_effort documented for models that return 400; Gemini explicit-cache minimum 1,024 live vs 4,096 in docs; Gemini v1 exposes only 22 GA ids although docs say all models are on both versions; Gemini Maps/File Search prices differ between pages; Anthropic mcp-client-2026-09-15 advertised but rejected. Full lists: docs/comparisons/models.md §9 and each domain page.
5. Confidence by section
| Section | Confidence | Basis |
|---|---|---|
| OpenAI (all areas) | HIGH, except Admin (MEDIUM: spec-only shapes, all 403) and WebRTC/SIP/Live/video generation (MEDIUM) | 2,515 live calls, OpenAPI spec |
| Anthropic (all areas) | HIGH, except fast mode / Admin beyond /me / clouds (MEDIUM) |
1,735 live calls |
| xAI models, pricing, chat, Responses, tools, files, collections, batches, media, voice | HIGH for exercised paths; MEDIUM for gRPC (proto stubs), Management API, Skills, video edits, custom voices | 588 live calls, OpenAPI + WebSocket JSON specs |
| Gemini core generation, streaming, thinking, structured output, files, embeddings, function calling, File Search, Interactions, Live, TTS, transcription | HIGH (free-tier models) | 973 live calls, discovery doc |
| Gemini paid-tier features (Pro models, image/video/music generation, Search grounding, caching, Batch, tuning) | MEDIUM (docs + discovery schemas + validation-error probes only) | key is free tier |
| Cross-provider comparisons, FAQ, endpoint catalogue | HIGH (generated from the merged data) with the listed inconsistencies | derived |
| Security, resilience, multi-provider abstraction (4 adapters), feature detection, capability graph, streaming parsers | HIGH (143 offline tests + 22 live adapter calls) | code + tests |
6. Known data-quality issues to fix in a follow-up pass
Detailed in docs/comparisons/models.md §9. Main ones: capability flag naming differs between providers; xAI max_output and most knowledge_cutoff values are null; a few endpoint records marked success without a logged call; duplicate Gemini price records with conflicting values (Search $14 vs $35 per 1k, Maps $14 vs $25) kept as-is because both appear on official pages; Lyria RealTime events with null status; composite http_status values in some error records. None affect the exact JSON shapes or the tool type strings.
7. How to keep it current
python3 scripts/update_atlas.py snapshots generated/, refreshes the four llms.txt indexes, re-crawls all pages (SHA-256 diff; xAI pages are re-split from llms-full.txt by scripts/split_xai_llms_full.py if the site keeps refusing direct .md fetches), re-runs the four model-listing endpoints (scripts/discover_*_models.py), rebuilds generated/ and writes reports/changes.md with NEW/REMOVED MODELS, ENDPOINTS, TOOLS, PRICING and CONTEXT WINDOW CHANGES, NEW BETA FEATURES, DEPRECATIONS, live drift and the official pages whose content changed. Curated fragments are never rewritten automatically; every domain generator is preserved under scripts/generators/ and generated/fragments/_builders/.