SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
16.1 KB

# API Atlas — Final report (OpenAI + Anthropic + xAI + Gemini, runs of 2026-09-18/19 UTC)

Snapshot of the public API surface of OpenAI, Anthropic, xAI (Grok) and Google Gemini, built from the current official documentation (downloaded, hashed, kept under sources/), the official machine-readable specs (OpenAI OpenAPI v2.3.0, xAI OpenAPI 3.1, Google discovery documents v1beta rev. 20260918 and v1), the official SDK surfaces, and 5,811 real API calls made with the project's own keys (every one logged, sanitized, in reports/live-requests.jsonl). Machine-readable outputs are in generated/; human docs in docs/; runnable examples in examples/; smoke tests in tests/. Numbers below come from reports/coverage.json and generated/build-report.json.

# 1. Coverage

Metric OpenAI Anthropic xAI Gemini Total
Official documentation pages (Markdown twins / split export, HTTP 200) 551 696 182 291 1,720
Machine-readable spec OpenAPI (352 ops) + SDK api.md SDK api.md (201 endpoints) OpenAPI (38 paths, 44 ops) + gRPC stubs (11 services) discovery v1beta (86 methods) + v1 (47) + SDK types —
Model records (canonical + snapshots + aliases + retired/id-only) 230 36 33 97 396
Model ids returned live by the models endpoint 136 11 12 (+7/3/2 typed) 58 (v1beta) / 22 (v1) 217
Model records LIVE_VERIFIED 137 21 12 10 180
Endpoint records 359 282 104 125 870
Tool records (exact type strings) 19 27 12 11 69
Parameter records (dotted paths) 4,215 2,040 891 716 7,862
Streaming / WebSocket event records 282 68 94 62 506
Error records 50 46 12 18 126
Header records 33 22 15 16 86
Price records 717 158 120 527 1,522
Deprecation records 165 1 (nested: 14 models + 35 items) 1 (nested) 1 (nested) 168
Object/schema records 189 224 80 213 706
SDK records 9 9 9 8 35
Webhook / audit-log event types 28 / 149 44 / — — — 72 / 149
Examples in manifests (LIVE_VERIFIED where run) 95 95 39 (+17 media) 82 311
Example files (.sh / .py / .ts / .html) — — — — 425
Test files — — — — 63
Test results (.venv/bin/python -m pytest tests) — — — — 448 passed, 67 skipped (gated), 2 flaky tests fixed and re-run green (full run 8 min 18 s on 2026-09-19)
Live API calls 2,515 1,735 588 973 5,811
Estimated spend (upper bound; several agents logged conservative estimates) ≈ $0.30 ≈ $0.33 ≈ $0.80 (actual ≈ $0.45) ≈ $0.26 (actual ≈ $0.02, free tier) ≈ $1.68 logged / ≈ $1.1 actual
Human documentation pages — — — — 228 Markdown files
Cross-provider feature matrix — — — — 125 features: on all 4 = 45, on 3 = 34, on 2 = 16, unique = 26, legacy = 4

Endpoint statuses (an endpoint can carry several): OpenAI 123 LIVE_VERIFIED, 18 ACCOUNT_RESTRICTED, 56 BETA, 24 RETIRED; Anthropic 47 LIVE_VERIFIED, 59 ACCOUNT_RESTRICTED, 111 BETA; xAI and Gemini per-status buckets are in docs/endpoints/by-status.md.

# 2. What was actually exercised live (highlights)

OpenAI — Responses (create/stream/retrieve/delete/cancel/input_items/input_tokens/compact/background + resume), Chat Completions (incl. store → retrieve → messages → delete), legacy Completions, Conversations, function calling (forced, strict, streaming, round-trip), custom grammar tools, web_search, file_search (vector store round trip), code_interpreter + Containers API, MCP (approval flow), shell, apply_patch, tool_search, Skills API, Agents API (agent → session → streamed turn → items/turns/subagents/artifacts → delete; environment templates; vault list), ChatKit list, Realtime (client_secrets, one WebSocket session with full event sequence), Audio (TTS, transcription variants, translation), Images (one generation + validation errors), provenance checks (C2PA detected), Embeddings, Moderation, Files, Uploads, Vector Stores, Batch (5 batches end to end), Graders, Evals, webhook event types, ~100 deliberate error probes, all Admin GET endpoints (403/401 = ACCOUNT_RESTRICTED).

Anthropic — Messages (system/multi-turn/prefill/stop sequences/metadata/service_tier/streaming with exact SSE order), count_tokens, custom tools (all tool_choice modes, parallel, fine-grained streaming), web_search, web_fetch, code_execution (+ container reuse, skills), tool_search, programmatic tool calling, computer use round trip, text editor / bash / memory shapes, MCP connector, advisor, prompt caching (5 m / 1 h TTL numbers), extended/adaptive/interleaved thinking, effort, fast mode (429 quota 0), 1M context header, context editing + compaction, structured outputs, citations (all location types), vision, PDF, Files API, Message Batches (2 batches end to end), Skills API, Managed Agents (environment → agent → session → streamed events round trip), Admin GET /v1/organizations/me (200) and 21 other Admin endpoints (401), Models API pagination, ~50 deliberate error probes.

xAI — Models catalogue endpoints (models, language/image/video/embedding models, api-key, me, tokenize-text; redirect aliases grok-3/grok-4-fast/grok-4-0709 → grok-4.3 observed), regional hosts us.api.x.ai and undocumented eu-west-1.api.x.ai, Chat Completions (reasoning_content, reasoning_effort per model, structured output, function calling, image input, deferred completions, streaming), Responses API (minimal/stream/previous_response_id/store/compact/input_items; server tools web_search, x_search, code_interpreter, mcp, file_search shape; 422 enum list of accepted tool types), Anthropic-compatible /v1/messages (text/image/tool_use/thinking blocks, streaming), legacy completions (400 for all live models), Files (upload/list/get/content/public-url/revoke/delete), Collections on api.x.ai (create/add/search/delete), Batches (create/list/get/add/cancel), image generation + edit, one 1-second video job (202 pending → 200), Voice realtime WebSocket text turn (18 events), TTS/STT/voices/client_secrets, grok-build-0.1, Management API (401 = ACCOUNT_RESTRICTED), Skills API (404 = ACCOUNT_RESTRICTED).

Gemini — Models list (v1beta paginated, v1), model get incl. the "no longer available to new users" 404 for 2.5 models, generateContent on 7 models (thoughtSignature, serviceTier), streaming (?alt=sse and JSON-array framing), system instructions and every generationConfig knob (with 400s recorded for those "not enabled for this model"), structured output (responseFormat / responseSchema / responseJsonSchema / enum), thinking (budget, levels, disabled), safety settings, countTokens (text/system/image/tools/PDF/YouTube), embeddings (001 and 2, dimensions), Files (resumable/multipart/raw uploads, register, use, delete), function calling (ANY/NONE/VALIDATED, parallel, JSON-schema params, thoughtSignature echo rule with the 400 on omission), urlContext, codeExecution, googleMaps grounding, File Search (store → upload → operation → documents → grounded query → delete), Interactions API (create/stream/chain/get/delete/cancel; Deep Research background create + cancel), Environments CRUD, ephemeral auth tokens (+ constrained Live method), Live API audio WebSocket turn sequence (text modality rejected), TTS (two models, streaming), transcription model (undocumented audioTranscription part), Lyria RealTime WebSocket (2 audio chunks), OpenAI-compatibility layer (chat + models), legacy PaLM methods (404/501), tunedModels (501). Paid-tier features refused with 429 quota 0 / 400 FAILED_PRECONDITION: Pro models, image generation, Google Search grounding, context caching, Batch, Lyria 3, Veo (validation-only probes).

# 3. Missing coverage (honest list)

Area Why Status recorded
OpenAI Admin API (124 endpoints), usage/costs, audit logs, certificates, external storage, WIF no Admin key (Missing scopes: api.management.read …) ACCOUNT_RESTRICTED; shapes from the OpenAPI spec
OpenAI computer-use-preview, codex-mini-latest, gpt-5.6-cyber, Daybreak, gpt-oss hosted; Anthropic Mythos ids 404 model_not_found — gated or retired ACCOUNT_RESTRICTED / RETIRED per docs
OpenAI video create/remix/edit, image edits/streaming, deep research runs, fine-tuning job creation (403), Realtime WebRTC/SIP, Live sessions, voice consents, Responses WebSocket, Agents hosted sandboxes cost / account / needs peer DOCUMENTED / UNVERIFIED / FAILED_VERIFICATION
Anthropic fast mode, per-message effort, Managed Agents tunnels/dreams/user profiles, Admin API beyond /me, cloud platforms account gating / no cloud credentials ACCOUNT_RESTRICTED / DOCUMENTED
xAI Management API (billing, audit, keys, ACLs), Skills API, embedding models, tool_search (alpha), grok-4.6/4.5 server-tool matrix, video edits/extensions, custom voices, SIP, Collections indexed-search happy path 401/404/403 gating or cost; indexing did not finish in 75 s ACCOUNT_RESTRICTED / DOCUMENTED / UNVERIFIED
Gemini paid-tier features on this key: Pro models, native image generation, Google Search grounding, explicit context caching, Batch, Lyria 3, Veo generation, computer-use legacy model; Vertex AI; Interactions credentials/agents/triggers/webhooks; dynamic/* methods the Gemini key is free tier (429 limit: 0 / 400 FAILED_PRECONDITION); no Vertex credentials; not exercised ACCOUNT_RESTRICTED / DOCUMENTED
Anthropic MCP connector and Gemini/xAI MCP against our own local MCP server needs a public URL DOCUMENTED + local server test

# 4. Major findings

Present on all four: JSON-Schema function calling with a client-side loop; streaming; structured JSON output; reasoning with some control (effort / thinking / reasoning_effort / thinkingConfig); prompt/context caching in some form; files APIs; image input; server-side web search on the three that expose tools inside generation (OpenAI, Anthropic, xAI) plus Gemini grounding (paid); code execution sandboxes; MCP in some form; batch processing (OpenAI, Anthropic, xAI −20 %, Gemini paid); agents platforms (Agents API, Managed Agents, xAI Responses agentic loop, Gemini Interactions API); OpenAI-style wire formats reused by xAI natively and by Gemini through its compatibility layer.

Unique to OpenAI: Chat Completions + Conversations as first-class stateful surfaces, background mode with resumable streams, WebSocket Responses, Realtime/Live speech-to-speech over WebRTC/SIP, moderation endpoint, provenance checks, vector stores, fine-tuning/evals/graders (winding down), ChatKit/Workspace Agents, 149 audit-log event types, project/service-account/role model. Unique to Anthropic: web_fetch, native citations, server-side context editing/compaction, adaptive/interleaved thinking with signed blocks, count_tokens endpoint, memory/advisor/browser tools, programmatic tool calling, Managed Agents memory stores and scheduled deployments, required anthropic-version and 50 catalogued beta headers. Unique to xAI: per-request cost_in_usd_ticks, X (Twitter) search as a server tool, an Anthropic-compatible Messages endpoint next to the OpenAI-compatible ones, retired-model redirect aliases, regional endpoints with a 1.1× US price uplift, a 500k-context model priced in two tiers around 200k prompt tokens, multi-agent model variant on Responses. Unique to Gemini: a documented free tier, native audio/video/PDF understanding on text models with per-modality token accounting, music generation (Lyria 3 + RealTime), Google Search/Maps grounding, URL context, thoughtSignature mandatory echo, Interactions API with Deep Research agents and environments, storage-priced (per token-hour) context caching, alt=sse vs JSON-array streaming.

Legacy and retirements observed live: OpenAI Assistants gone, DALL·E gone, Sora/Videos shutdown 2026-09-24, local_shell unsupported, Evals shutdown 2026-11-30, self-serve fine-tuning ends 2027-01-06; Anthropic /v1/complete 400 "deprecated", Claude 3.x ids 404; xAI Live Search 410, /v1/completions and /v1/complete 400 for every live model, grok-2-image 404, grok-3/4-fast silently redirected; Gemini 2.5 models 404 for new users, Imagen predict 404, tunedModels 501, PaLM methods 404/501.

Undocumented-but-live (LIVE_DISCOVERED): OpenAI phase/billing/tool_usage/prompt_cache_retention/personality/shutdown_date; Anthropic stop_details/container/usage.cache_creation.*/usage.inference_geo/tool_use.caller/file_id on citations; xAI eu-west-1.api.x.ai, x-ratelimit-* headers, usage.context_details, video polling HTTP 202, realtime ping and billable_audio_seconds, reasoning_effort: none on grok-4.3, store:false responses still retrievable, x_search emitting custom_tool_call; Gemini Model.thinking/maxTemperature, X-Gemini-Service-Tier header, audioTranscription part, generateContent?alt=sse, files:register, gemini-flash-latest → 3.8-flash.

Documentation vs live discrepancies worth knowing: o3 priced differently on two OpenAI pages; xAI catalogue prices are ticks (1e-10 USD) while docs say cents; xAI reasoning_effort documented for models that return 400; Gemini explicit-cache minimum 1,024 live vs 4,096 in docs; Gemini v1 exposes only 22 GA ids although docs say all models are on both versions; Gemini Maps/File Search prices differ between pages; Anthropic mcp-client-2026-09-15 advertised but rejected. Full lists: docs/comparisons/models.md §9 and each domain page.

# 5. Confidence by section

Section Confidence Basis
OpenAI (all areas) HIGH, except Admin (MEDIUM: spec-only shapes, all 403) and WebRTC/SIP/Live/video generation (MEDIUM) 2,515 live calls, OpenAPI spec
Anthropic (all areas) HIGH, except fast mode / Admin beyond /me / clouds (MEDIUM) 1,735 live calls
xAI models, pricing, chat, Responses, tools, files, collections, batches, media, voice HIGH for exercised paths; MEDIUM for gRPC (proto stubs), Management API, Skills, video edits, custom voices 588 live calls, OpenAPI + WebSocket JSON specs
Gemini core generation, streaming, thinking, structured output, files, embeddings, function calling, File Search, Interactions, Live, TTS, transcription HIGH (free-tier models) 973 live calls, discovery doc
Gemini paid-tier features (Pro models, image/video/music generation, Search grounding, caching, Batch, tuning) MEDIUM (docs + discovery schemas + validation-error probes only) key is free tier
Cross-provider comparisons, FAQ, endpoint catalogue HIGH (generated from the merged data) with the listed inconsistencies derived
Security, resilience, multi-provider abstraction (4 adapters), feature detection, capability graph, streaming parsers HIGH (143 offline tests + 22 live adapter calls) code + tests

# 6. Known data-quality issues to fix in a follow-up pass

Detailed in docs/comparisons/models.md §9. Main ones: capability flag naming differs between providers; xAI max_output and most knowledge_cutoff values are null; a few endpoint records marked success without a logged call; duplicate Gemini price records with conflicting values (Search $14 vs $35 per 1k, Maps $14 vs $25) kept as-is because both appear on official pages; Lyria RealTime events with null status; composite http_status values in some error records. None affect the exact JSON shapes or the tool type strings.

# 7. How to keep it current

python3 scripts/update_atlas.py snapshots generated/, refreshes the four llms.txt indexes, re-crawls all pages (SHA-256 diff; xAI pages are re-split from llms-full.txt by scripts/split_xai_llms_full.py if the site keeps refusing direct .md fetches), re-runs the four model-listing endpoints (scripts/discover_*_models.py), rebuilds generated/ and writes reports/changes.md with NEW/REMOVED MODELS, ENDPOINTS, TOOLS, PRICING and CONTEXT WINDOW CHANGES, NEW BETA FEATURES, DEPRECATIONS, live drift and the official pages whose content changed. Curated fragments are never rewritten automatically; every domain generator is preserved under scripts/generators/ and generated/fragments/_builders/.