Repository and data model of the API Atlas (4 providers)
Status: DOCUMENTED (describes conventions in CLAUDE.md and the intended behaviour of scripts/update_atlas.py, which the orchestrator writes)
Sources: CLAUDE.md; generated/fragments/** (as produced by the domain agents on 2026-09-18/19 for openai, anthropic, xai, gemini); sources/*/pages-manifest.json; scripts/live.py, scripts/lib.sh, scripts/build_generated.py (generated/build-report.json: 396 models, 870 endpoints, 69 tools, 506 streaming events after the 2026-09-19 rebuild)
Last verified: 2026-09-19
1. Pipeline
sources/ immutable inputs per provider (see §1.1): docs as Markdown twins, OpenAPI spec or Google discovery
│ document, SDK surfaces, sanitized live model lists
│ grep / parse / minimal live calls (scripts/live.py {openai,anthropic,xai,gemini}_request, scripts/lib.sh oai|ant|xai|gemini
│ → reports/live-requests.jsonl)
▼
generated/fragments/<domain>/<provider>-<topic>.json ← the ONLY place agents write machine-readable records
│ scripts/build_generated.py (validate against schemas/*.schema.json, merge, dedupe, sort)
▼
generated/{models,endpoints,parameters,tools,streaming-events,errors,headers,pricing,examples,…}.json (+ CSV twins)
│ derived views
├─ generated/compatibility/* matrices (model × endpoint, model × tool, model × capability)
├─ scripts/export_capability_graph.py → generated/capability-graph.{json,mmd,dot} (4 provider nodes)
└─ examples/shared/feature-detection → runtime registry (`supports()`), consumed by apps/agents
docs/ human pages; every fact in docs/ must also exist in generated/
reports/ live-requests.jsonl (every live call), changes.md (diff between runs), final-report.md, coverage1.1 Sources layout — what each provider gives us
| OpenAI | Anthropic | xAI | Gemini | |
|---|---|---|---|---|
| Docs pages | sources/openai/pages/api/docs/**, pages/api/reference/** (Markdown twins of developers.openai.com) |
sources/anthropic/pages/** (api, build-with-claude, agents-and-tools, …) |
sources/xai/pages/** — 182 pages split from the official llms-full.txt (developers/** incl. rest-api-reference/, grpc-api-reference/, tools/, models/; build/, grok/, grok-bot/, console/); the split is ours, the text is verbatim |
sources/gemini/pages/gemini-api/docs/** (guides) + sources/gemini/pages/api/** (REST reference); files are .md.txt twins of ai.google.dev pages (Google serves <page>.md.txt) |
| Manifest | pages-manifest.json |
pages-manifest.json |
pages-manifest.json (maps page → offset in llms-full.txt) |
pages-manifest.json |
| Combined export | llms-full.txt (7.5 MB) |
llms-full.txt (35 MB, frontmatter per page) |
official llms-full.txt (the only upstream form; pages are derived) |
none (per-page .md.txt) |
| Machine spec | OpenAPI openapi/openapi-master.yaml (352 operations, openapi-master-ops.json) |
none official (SDK api.md surfaces) |
OpenAPI openapi/openapi.json (38 paths, openapi-ops.json) — REST inference API only; Management API + gRPC documented in pages |
Google API discovery documents discovery-v1beta.json (86 methods, full request/response schemas, rev. 20260918) + discovery-v1.json (47 methods); method list openapi/discovery-v1beta-methods.json. Not OpenAPI: schemas are Google-JSON-schema-like (camelCase, enum/enumDescriptions, $ref by name), mediaUpload describes the resumable protocol |
| SDK surfaces | openapi/{python,node}-sdk-api.md |
openapi/{python,node}-sdk-api.md |
openapi/python-sdk-readme.md (xai-sdk, gRPC; no official Node SDK) |
openapi/python-genai-types.py (authoritative field names incl. Vertex-only ones), READMEs |
| Live discovery | models-api-raw.json (136 ids) |
models-api-raw.json (11 ids) |
models-api-raw.json (12 ids), language-models-raw.json, image-generation-models-raw.json (typed catalogues with prices in ticks) |
models-api-raw.json (58 models, supportedGenerationMethods, token limits) |
| Canonical URL to cite | https://developers.openai.com/api/docs/... / .../api/reference/... |
https://platform.claude.com/docs/en/... |
https://docs.x.ai/<path> |
https://ai.google.dev/gemini-api/docs/<slug> or https://ai.google.dev/api/<slug> (drop .md.txt) |
| Auth in helpers | Authorization: Bearer |
x-api-key + anthropic-version |
Authorization: Bearer xai-… (+ optional XAI_MANAGEMENT_KEY for management-api.x.ai, which we do not have) |
x-goog-api-key header — never ?key= (the helper masks key= in anything saved) |
| Log path normalisation | as sent | as sent | as sent | models/<id> → models/{model} so reports/live-requests.jsonl groups by method, not by model |
Provider-specific record conventions that the shared layer understands: xAI kind: model | alias | retired_redirect | legacy | service with redirect targets in verification.request_note; Gemini kind: stable | preview | alias | experimental | agent, -latest records with dict aliases {alias, resolves_to_live, history}, dict-shaped tools[] {type, category, support}; xAI usage in cost_in_usd_ticks; Gemini usageMetadata (see feature-detection.md, capability-graph.md, multi-provider-abstraction.md).
2. Record types (summary — authoritative list in CLAUDE.md, JSON Schemas in schemas/)
| type | key | notable fields |
|---|---|---|
model |
provider + id |
aliases[], snapshots[], canonical_model, record_kind (model | snapshot | id_only), capabilities{} (tri-state), endpoints[], tools[], pricing{}, beta_headers[], availability{}, verification{} |
endpoint |
provider + method + path |
api_family, auth, beta_header, request{}, response{}, streaming{supported, events_ref}, pagination, idempotency, destructive, sdk{python,node} |
parameter |
provider + endpoint + parameter (dotted path) |
location, type, required, default, enum[], compatible_models[], beta_header |
tool |
provider + type (exact versioned JSON type string) |
category (client | server | hosted | programmatic | mcp), compatible_models[], compatible_endpoints[], parameters_schema, streaming_events[], security, billing |
streaming_event |
provider + api + event |
direction, schema, example |
error |
provider + http_status + type + code |
retryable, recommended_action |
header |
provider + name + direction |
|
price |
provider + model_or_service + dimension + tier |
price, unit, retrieved_at |
example |
file |
language, status, verified_at |
Every record carries status[], sources[{url, retrieved_at}] and (where applicable) verification{method, verified_at, result, http_status, request_note}.
Status vocabulary (exact strings)
DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED · BETA · PREVIEW · LEGACY · DEPRECATED · RETIRED · ACCOUNT_RESTRICTED · UNVERIFIED · FAILED_VERIFICATION. A record may carry several. A 403/404 with our key never means "does not exist" (→ ACCOUNT_RESTRICTED, not RETIRED).
Provenance
sources[].urlis the canonical public URL (https://developers.openai.com/api/docs/...,https://platform.claude.com/docs/en/...,https://docs.x.ai/...,https://ai.google.dev/...), not the.md/.md.txttwin path.- Live evidence is a line in
reports/live-requests.jsonl(ts, provider ∈ openai|anthropic|xai|gemini, method, path, status, est_cost_usd, note) plus a sanitized excerpt intmp-live/(gitignored). Secrets are masked byscripts/live.py::mask(patternssk-…,sk-ant-…,xai-…,AIza…,AQ.…,?key=…,Bearer …,x-api-key,x-goog-api-key) before anything is written. retrieved_aton every source andlast_verifiedon every record make staleness queryable.ACCOUNT_RESTRICTEDis a status, not an absence: xAI Management API (needs a Management key), xAI embeddings, Gemini Pro models / explicit caching / Search grounding / Batch on the free-tier key.
3. Fragments → merged files
Agents write generated/fragments/<domain>/<provider>-<topic>.json (a list of records, optionally with a sibling *.meta.json; e.g. models/xai-models.json, errors/gemini-errors.json, streaming-events/gemini-interactions.json). build_generated.py (orchestrator-owned):
- loads every fragment, validates each record against
schemas/<type>.schema.json, rejects on error (report inreports/coverage); - merges by key; on conflict prefers the record with the stronger status (
LIVE_VERIFIED>LIVE_DISCOVERED>DOCUMENTED) and unionssources[],status[]; - writes
generated/<type>.json(+ CSV for tabular types) — never edit merged files by hand; - rebuilds
generated/compatibility/*matrices.
Consumers in this domain are shape-tolerant (list, {records: []}, or id-keyed object) so they work before the merger is finalised.
4. scripts/update_atlas.py — intended behaviour (written by the orchestrator)
update_atlas.py [--offline] [--providers openai,anthropic,xai,gemini] [--domains models,endpoints,...] [--budget-usd 2.00]- Refresh sources —
crawl_docs.pyre-fetches pages listed insources/*/pages-manifest.json(ETag / content hash): OpenAI/Anthropic Markdown twins + OpenAPI spec +llms-full.txt; xAI re-downloads the officialllms-full.txtand re-splits it intopages/**(+openapi/openapi.json); Gemini re-fetches the.md.txttwins and both discovery documents ($discovery/rest?version=v1beta|v1, comparerevision).discover_models.pycalls the four model-list endpoints (GET /v1/models×3,GET /v1beta/models+ xAI typed catalogues) →sources/*/models-api-raw.json. Diffs go toreports/changes.md(new/removed models, changed pages, moved-latestaliases, new redirects). - Re-extract fragments — re-run the domain extractors (or the agents) only for changed sources; keep untouched fragments.
- Verify — re-run cheap verification (
pytest -q, cheap smoke tests; expensive ones behindRUN_*_TESTS) within--budget-usd; updateverification{}andlast_verifiedon the records touched; append toreports/live-requests.jsonl. - Build —
build_generated.py,export_capability_graph.py, matrices,verify_links.py(everysources[].urlresolves),reports/coverage. - Report —
reports/final-report.md(counts per status, drift since last run, failures verbatim).
--offline skips 1 and 3 (docs-only rebuild). Idempotent: running twice without upstream changes produces byte-identical generated/ (deterministic ordering).
5. Consuming the atlas
| Use | What to load | Notes |
|---|---|---|
| RAG / assistant grounding | docs/**/*.md chunked by heading + the matching generated/*.json records as metadata (provider, status, last_verified) |
Filter out RETIRED/DEPRECATED unless asked; always surface status and sources[].url in answers. |
| Comparator (OpenAI vs Anthropic vs xAI vs Gemini) | generated/models.json, tools.json, errors.json, pricing.json + Registry.compare() |
Map capability synonyms explicitly (CAPABILITY_SYNONYMS in feature-detection.md); compare tool types after normalisation (googleSearch ≡ google_search). |
| Code generator | generated/endpoints.json + parameters.json (+ examples/** headers for verified snippets) |
Emit only parameters whose compatible_models[] include the target; add beta_header automatically (Anthropic); put the Gemini model in the path; never emit ?key=. |
| Provider-choosing agent | Registry.models_with(...), generated/pricing.json, capability-graph.json |
Pick by required capabilities → price → LIVE_VERIFIED status → not ACCOUNT_RESTRICTED for your tier; fall back per resilience.md. |
| Runtime guards | examples/shared/feature-detection/supports.* |
Ship the two JSON files with your app; reload on atlas updates (Gemini -latest aliases and xAI redirects move). |
6. Security invariants (apply to every consumer)
Keys only in .env (mode 600); scripts/live.py/lib.sh inject and mask them; no key ever in docs/, examples/, tests/, sources/, generated/, reports/. Gemini keys go in the x-goog-api-key header, never in a URL; xAI Management keys (if ever obtained) are a separate secret. Admin/Management endpoints are GET/LIST only. Minimal payloads and cheapest models for verification (gpt-5.4-nano, claude-haiku-4-5-20251001, grok-4.3, gemini-3.5-flash-lite); every live call logged. See docs/security/.