Python 88.3%
TypeScript 7.6%
Shell 4.1%
1# Multi-provider abstraction — OpenAI Responses ↔ Anthropic Messages ↔ xAI (Responses / Chat Completions) ↔ Gemini generateContent23**Status:** LIVE_VERIFIED for `generate` + `stream` in all four adapters and both languages ("Reply with OK.", ≤ 16 output tokens, all HTTP 200 — 2026-09-18 for OpenAI/Anthropic, 2026-09-19 for xAI `grok-4.3` (Responses in py+ts, Chat Completions in py) and Gemini `gemini-3.5-flash-lite`; see `reports/live-requests.jsonl` notes `llm_provider.*` / `llmProvider.*`). `count_tokens`, `use_tools`, `structured_output`, `upload_file`, `web_search`: DOCUMENTED (spec/discovery-derived, offline-tested with mock transports in `tests/shared/test_llm_provider.py`, 34 tests). Gemini `web_search` (`googleSearch`) is ACCOUNT_RESTRICTED on this key (free tier → 429 `limit: 0`): the adapter returns a typed `stop_reason == "restricted"` result.4**Sources:** OpenAI OpenAPI spec (`CreateResponse`, `TextResponseFormatConfiguration`), https://developers.openai.com/api/reference/responses/create · https://platform.claude.com/docs/en/api/messages, …/build-with-claude/structured-outputs, …/agents-and-tools/tool-use/web-search-tool · xAI: https://docs.x.ai/developers/rest-api-reference/inference/responses, …/inference/chat-completions, …/other (tokenize-text), …/files/upload, https://docs.x.ai/developers/tools/web-search, `sources/xai/openapi/openapi.json` (`ModelRequest`, `TokenizeRequest`, `WebSearchFilters`) · Gemini: https://ai.google.dev/api/generate-content, https://ai.google.dev/api/tokens, https://ai.google.dev/api/files, https://ai.google.dev/gemini-api/docs/function-calling, …/thought-signatures, …/structured-output, …/google-search, https://ai.google.dev/api/interactions, `sources/gemini/discovery-v1beta.json` (`GenerateContentRequest`, `FunctionCallingConfig`, `CountTokensRequest`) · atlas pages `docs/xai/{responses,streaming,structured-outputs,files,tool-loop}.md`, `docs/gemini/{generate-content,streaming,structured-outputs,files,token-counting,tool-loop,interactions-api}.md`5**Last verified:** 2026-09-1967Implementation: `examples/shared/provider-abstraction/llm_provider.py` and `llmProvider.ts` (built on the resilient client; no SDK needed). `get_provider("openai"|"anthropic"|"xai"|"gemini")`; `XAIProvider(api="responses"|"chat")`; `GeminiProvider(api_version="v1beta"|"v1")`.89## 1. Interface1011```12LLMProvider13 generate(req) -> GenerateResult {text, reasoning, tool_calls[], stop_reason, usage, response_id, assistant_content, error, raw}14 stream(req) -> StreamEvent* {text_delta | reasoning_delta | tool_call_start | tool_call_delta | tool_call_done | usage | done | error | raw}15 count_tokens(req) -> int | None OpenAI POST /v1/responses/input_tokens · Anthropic POST /v1/messages/count_tokens ·16 xAI POST /v1/tokenize-text (text only) · Gemini POST models/{m}:countTokens {generateContentRequest}17 use_tools(req, impls) -> GenerateResult portable call/result loop (max_rounds, errors contained, never leaks stack traces;18 replays the provider-native assistant turn verbatim when the provider needs it — Gemini)19 structured_output(req, schema) -> dict OpenAI text.format json_schema(strict) · Anthropic output_config.format · xAI text.format /20 response_format json_schema · Gemini generationConfig.responseMimeType+responseJsonSchema21 upload_file(path) -> FileRef OpenAI POST /v1/files (purpose) · Anthropic POST /v1/files (beta header) · xAI POST /v1/files22 (purpose ignored) · Gemini resumable upload (start → upload,finalize) → File.uri23 web_search(req) -> GenerateResult OpenAI {type:"web_search"} · Anthropic {type:"web_search_20250305"} · xAI {type:"web_search",24 allowed_domains|excluded_domains} · Gemini {googleSearch:{}} (→ "restricted" on this key)25```2627`GenerateRequest.extensions` (merged verbatim into the body) and `extra_headers` are the **only** way to use provider-specific features. The adapters never translate an extension into "the closest thing" on another provider — that is the caller's decision.2829## 2. Portable common denominator — parameter ↔ parameter3031| Concept | Portable field | OpenAI Responses | Anthropic Messages | xAI Responses (default) / Chat Completions (`api="chat"`) | Gemini generateContent |32|---|---|---|---|---|---|33| Endpoint | — | `POST /v1/responses` | `POST /v1/messages` | `POST /v1/responses` / `POST /v1/chat/completions` | `POST /v1beta/models/{model}:generateContent` (model in the **path**, not the body) |34| Auth | client | `Authorization: Bearer` | `x-api-key` + `anthropic-version` | `Authorization: Bearer xai-…` (no version/beta headers) | `x-goog-api-key` (never `?key=`) |35| Model | `model` | `model` | `model` | `model` | path segment `models/{id}`; response `modelVersion` |36| Conversation | `messages[] {user/assistant/tool}` | `input[]` items: `{role, content}`, `function_call`, `function_call_output` | `messages[]` (user/assistant; tool results = `tool_result` blocks in a **user** turn) | Responses: as OpenAI · Chat: `messages[]` with `tool_calls[]` / `{role:"tool", tool_call_id}` | `contents[] {role: user\|model, parts[]}`; tool results = `functionResponse` parts in a **user** Content (all calls of a turn in ONE Content) |37| System prompt | `system` | `instructions` | `system` | `instructions` (400 if combined with `previous_response_id`) / `messages[0].role="system"` | `systemInstruction {parts:[{text}]}` |38| Output cap | `max_tokens` | `max_output_tokens` | `max_tokens` (**required**) | `max_output_tokens` (includes reasoning; live not enforced on reasoning) / `max_tokens` | `generationConfig.maxOutputTokens` (includes thinking tokens → a small cap can yield empty text + `MAX_TOKENS`) |39| Sampling | `temperature` | `temperature` | `temperature` | `temperature` (+ xAI-only `top_k`, `min_p`) | `generationConfig.temperature` [0, 2] (Gemini 3: keep 1.0) |40| Tool definitions | `ToolDef{name, description, parameters, strict}` | `tools[]: {type:"function", name, …, strict}` | `tools[]: {name, description, input_schema, strict?}` | as OpenAI (`strict` accepted, **implicitly always true**) / `tools[]: {type:"function", function:{…}}` | `tools:[{functionDeclarations:[{name, description, parametersJsonSchema}]}]` (or `parameters` = OpenAPI subset, UPPERCASE types) |41| Tool choice | `"auto" \| "none" \| "required" \| {name}` | `"auto"`/`"none"`/`"required"`/`{type:"function", name}` | `{type:"auto"}`/`{type:"none"}`/`{type:"any"}`/`{type:"tool", name}` | as OpenAI / `{type:"function", function:{name}}` | `toolConfig.functionCallingConfig.mode: AUTO \| NONE \| ANY (+ allowedFunctionNames)` (also `VALIDATED`) |42| Tool call (model → you) | `ToolCall{id, name, arguments, thought_signature?}` | item `function_call {call_id, name, arguments (string)}` | block `tool_use {id, name, input (object)}` | as OpenAI (`call-…` ids) / `message.tool_calls[] {id, function:{name, arguments}}` | part `functionCall {name, args (object), id}` + sibling **`thoughtSignature`** on the part |43| Tool result (you → model) | `Message(role="tool", tool_call_id, tool_name, content, is_error)` | `function_call_output {call_id, output}` — no error flag → `[TOOL ERROR]` prefix | `tool_result {tool_use_id, content, is_error}` | as OpenAI / `{role:"tool", tool_call_id, content}` | `functionResponse {name, id?, response: OBJECT}` — scalars wrapped as `{result}`, errors as `{error}` |44| JSON-schema output | `json_schema{name, schema, strict}` | `text.format = {type:"json_schema", name, schema, strict}` | `output_config.format = {type:"json_schema", schema}` | `text.format` (same) / `response_format = {type:"json_schema", json_schema:{name, schema, strict}}` | `generationConfig.responseMimeType: "application/json"` + `responseJsonSchema` (new form: `responseFormat.text {mimeType: APPLICATION_JSON, schema}`) |45| Streaming text | `text_delta` | `response.output_text.delta` | `content_block_delta/text_delta` | same as OpenAI / `choices[].delta.content` | chunk `candidates[0].content.parts[].text` (non-thought) |46| Streaming reasoning | `reasoning_delta` | `response.reasoning_summary_text.delta` | `thinking_delta` | `response.reasoning_summary_text.delta` / `delta.reasoning_content` | parts with `thought: true` (only with `thinkingConfig.includeThoughts`) |47| Streaming tool args | `tool_call_delta` | `function_call_arguments.delta` → `.done` | `input_json_delta` → `content_block_stop` | one delta with the **whole** JSON / one chunk with the whole `tool_calls[]` | never partial: `functionCall.args` arrives whole in one part |48| Usage | `Usage{input, output, cached_input, cache_write, total}` | `input_tokens`, `output_tokens`, `input_tokens_details.cached_tokens` | `input_tokens` (excl. cache) + `cache_read/creation_input_tokens` — normalised to "input includes cached" | Responses: as OpenAI, `output_tokens` **includes** reasoning · Chat: `prompt_tokens`, `completion_tokens` (**excludes** reasoning) + `completion_tokens_details.reasoning_tokens` — normalised to billed output | `usageMetadata.promptTokenCount` (incl. cached), `candidatesTokenCount` + `thoughtsTokenCount` (both billed → output), `cachedContentTokenCount`, `totalTokenCount` |49| Stop reason | `end \| max_tokens \| tool_use \| stop_sequence \| refusal \| incomplete \| restricted \| other` | `status` + `incomplete_details.reason`, refusal part | `stop_reason` enum | as OpenAI / `finish_reason: stop\|length\|tool_calls\|content_filter` | `finishReason`: `STOP`, `MAX_TOKENS`, `SAFETY`/`RECITATION`/`BLOCKLIST`/`PROHIBITED_CONTENT`/`SPII`/`IMAGE_*`/`LANGUAGE` → refusal, `MALFORMED_FUNCTION_CALL`/`MISSING_THOUGHT_SIGNATURE`/`UNEXPECTED_TOOL_CALL`/`TOO_MANY_TOOL_CALLS`/`MALFORMED_RESPONSE` → other; `promptFeedback.blockReason` (no candidates, HTTP 200) → refusal |50| Stop sequences | `stop[]` | **none** → dropped | `stop_sequences` | Responses: **none** → dropped / Chat: `stop` (≤ 4) | `generationConfig.stopSequences` (≤ 5, else 400); hit → `finishReason: STOP` (indistinguishable from a natural end) |51| Files in prompt | `Message.file_ids` | `{type:"input_file", file_id}` | `{type:"document", source:{type:"file", file_id}}` | Responses: `{type:"input_file", file_id}` (turns the request into an "attachment search", $10/1k) / Chat: **400** | `{fileData:{fileUri: File.uri}}` (`FileRef.id` = the URI) |52| Request metadata | `extensions.metadata` | `metadata` (16 pairs) | `metadata.user_id` | Responses **400 "Argument not supported: metadata"** (response `metadata` = `{system_fingerprint}`) / Chat: ignored | `labels` (Cloud label rules; documented key `safety_identifier`), accepted, not echoed |5354## 3. State management — four different answers5556| | OpenAI Responses | Anthropic Messages | xAI Responses | Gemini generateContent | Gemini Interactions API (not wrapped here) |57|---|---|---|---|---|---|58| Server-side state | `previous_response_id` (needs `store: true`, default) or `conversation` id | **none** — stateless, replay `messages[]` | `previous_response_id` (30-day retention; `store:false` responses are still retrievable — LIVE_DISCOVERED); alt. replay `output[]` incl. `reasoning.encrypted_content`; `POST /v1/responses/compact` | **none** — stateless, replay `contents[]` | `previous_interaction_id` (default `store: true`; `store:false` = stateless `input[]` of Steps; 400 if the previous interaction is still `in_progress`); `background: true` + polling/resumable SSE |59| Reasoning carry-over | reasoning items / `reasoning.encrypted_content` | `thinking` blocks with `signature` (echo verbatim) | reasoning items / `encrypted_content` (replay verbatim) | **`thoughtSignature`** on parts — echo the model turn VERBATIM (`GenerateResult.assistant_content` → `Message.raw_content`); dropping it → 400 "Function call is missing a thought_signature"; only the first parallel call carries one | `thought` step `signature` (`step.delta` `thought_signature`) |60| Restrictions | `store:false` disables chaining | — | `instructions` + `previous_response_id` together → 400; `background` unsupported (400) | last turn must be `user`; roles are `user`/`model` (not `assistant`/`system`) | `store:false` incompatible with `background` and later chaining; snake_case only |61| ZDR | Zero Data Retention org setting | ZDR arrangement | team-wide ZDR disables `store`/`previous_response_id`, Files, Collections, Batch | not fully achievable (Search/Maps grounding stores 30 days) | `store:false`; not a ZDR guarantee |6263## 4. Provider-specific extensions (pass through `extensions` / `extra_headers`)6465**OpenAI Responses**: `previous_response_id`, `conversation`, `reasoning {effort, summary}`, `include[]`, `background`, `store`, `service_tier`, `truncation`, `max_tool_calls`, `parallel_tool_calls`, `prompt`, `prompt_cache_key`, `safety_identifier`, `top_logprobs`, `context_management`, built-in tools (`web_search`, `file_search`, `code_interpreter`, `image_generation`, `mcp`, `computer_use_preview`, `shell`, `apply_patch`), `text.verbosity`.6667**Anthropic Messages**: `thinking {type, budget_tokens}`, `output_config {effort, format}`, `cache_control`, `service_tier`, `speed`, `top_k`, `container`, `inference_geo`, `metadata.user_id`, `tool_choice.disable_parallel_tool_use`, server tools (`web_search_*`, `web_fetch_*`, `code_execution_*`, `computer_toolset_*`, `text_editor_*`, `bash_*`, `memory_*`, `tool_search_tool_*`, `mcp_toolset`), `mcp_servers` (+ `anthropic-beta: mcp-client-2025-11-20`), `context_management` (+ beta header).6869**xAI Responses** (`ModelRequest`, docs/xai/responses.md): `previous_response_id`, `store`, `include[]` (`reasoning.encrypted_content`, `web_search_call.action.sources`, `code_interpreter_call.outputs`, `file_search_call.results`, `no_inline_citations`), `max_turns` (agentic turns per request), `reasoning {effort: low|medium|high|xhigh, summary}` / `reasoning_effort`, `top_k`, `min_p`, `service_tier: default|priority`, `prompt_cache_key` (= `x-grok-conv-id`), `safety_identifier`, `parallel_tool_calls`, server tools (`web_search`, `x_search`, `code_interpreter`/`code_execution`, `file_search`/`collections_search`, `mcp` (no approval round-trip), `image_generation`, `shell`, `tool_search` (403 alpha)). Rejected: `metadata`, `background`, `search_parameters` (retired). **xAI Chat Completions**: `reasoning_effort`, `stream_options.include_usage`, `x-grok-conv-id` header, `deferred: true` (poll `GET /v1/chat/deferred-completion/{id}`), `response_format.json_object`; **no** server-side tools (422/410), **no** files (400). Anthropic-compatible `POST /v1/messages` exists but is DEPRECATED (not wrapped).7071**Gemini generateContent** (docs/gemini/generate-content.md): `generationConfig.{thinkingConfig{thinkingLevel|thinkingBudget, includeThoughts}, topP, topK, seed, candidateCount(=1), responseModalities, speechConfig, imageConfig, mediaResolution, responseSchema (OpenAPI subset), responseFormat (new)}`, `safetySettings[]`, `cachedContent` (`cachedContents/{id}`, explicit caching — ACCOUNT_RESTRICTED on free tier), `serviceTier: standard|flex|priority`, `labels`, `store`, `toolConfig.{includeServerSideToolInvocations, retrievalConfig}`, server tools (`googleSearch{searchTypes, timeRangeFilter}`, `googleMaps`, `urlContext`, `codeExecution`, `fileSearch{fileSearchStoreNames}`, `computerUse{environment, disabledSafetyPolicies}`, `mcpServers[]` (schema only, UNVERIFIED on generateContent)), per-part `videoMetadata`, `mediaResolution`, `mediaProcessing`. The adapter merges `extensions["generationConfig"]` into its own `generationConfig` and appends `extensions["tools"]` to the function declarations; everything else is top-level.7273## 5. Lossy conversions (be explicit with users)74751. `stop[]` → OpenAI Responses and xAI Responses: **dropped** (no parameter). xAI Chat keeps ≤ 4; Gemini keeps ≤ 5 but reports the hit as a plain `STOP` (no `stop_sequence` stop reason).762. Tool result `is_error` → OpenAI/xAI: no field, `[TOOL ERROR]` prefix; Gemini: no field, the adapter sends `{"error": …}` as the `functionResponse.response` object (the documented pattern).773. `Message.tool_name` is required by Gemini (`functionResponse.name`) — the portable loop fills it from the `ToolCall`; callers building messages by hand must set it.784. Gemini `functionResponse.response` must be an **object**: string/number results are wrapped as `{"result": …}`; JSON-object strings are parsed back.795. OpenAI `previous_response_id`/`conversation`, xAI `previous_response_id`, Gemini Interactions `previous_interaction_id` → Anthropic / Gemini generateContent: replay the full history yourself. Reverse direction: the stateless histories contain provider-native reasoning artefacts (`thoughtSignature`, `thinking` signatures, `encrypted_content`) that do **not** transfer between providers — start a fresh conversation when switching.806. Reasoning controls: OpenAI `reasoning.effort` ↔ Anthropic `output_config.effort`/`thinking.budget_tokens` ↔ xAI `reasoning.effort` (`low|medium|high|xhigh`; grok-4.3 defaults to `low` and **always bills reasoning tokens**) ↔ Gemini `thinkingConfig.thinkingLevel` XOR `thinkingBudget` (both → 400). Same intent, four scales — **not** auto-mapped.817. Usage: Anthropic `input_tokens` excludes cache reads; xAI Chat `completion_tokens` excludes reasoning while xAI Responses `output_tokens` includes it; Gemini `candidatesTokenCount` excludes `thoughtsTokenCount`. The adapter normalises to "input includes cached, output = everything billed as output" — read `raw` for the provider view.828. Structured output: OpenAI needs `strict:true` + `additionalProperties:false` + all `required`; Anthropic has its own keyword limits; xAI defaults `additionalProperties` to **false** and rejects `items` arrays / `minContains`; Gemini ignores unsupported keywords silently (`minLength`), rejects unknown keys in the legacy `responseSchema`, supports `$defs/$ref/anyOf/null`, and truncates JSON on `MAX_TOKENS` (the adapter raises). One schema may need four variants.839. Web search: Anthropic `max_uses`/`allowed_domains|blocked_domains`/`user_location`; OpenAI different filters; xAI `allowed_domains|excluded_domains` (≤ 5, mutually exclusive, `search_context_size` → 400) and no `max_uses` (bound with `max_turns`); Gemini `googleSearch` has no domain filters at all (only `timeRangeFilter`/`searchTypes`) and is unavailable on free-tier keys (429 `limit: 0` → `restricted`).8410. Files: OpenAI `purpose` required; Anthropic beta header, `size_bytes`; xAI `purpose` ignored (echoed `""`), 50/512 MB, team-scoped; Gemini 2-step resumable protocol, `purpose` has no equivalent, files auto-delete after **48 h**, referenced by `uri` (not id), 2 GB.8511. Streaming stop reasons arrive at different moments: OpenAI/xAI in the terminal `response.completed`; Anthropic in `message_delta`; xAI Chat in the last `choices[].finish_reason` before `[DONE]`; Gemini only on the **last chunk** (`finishReason`) — a connection closed without it is `incomplete`.8612. Anthropic requires `max_tokens`; the adapter defaults to 1024 everywhere. On Gemini and xAI Responses that cap also covers reasoning/thinking tokens: 16 tokens was enough for "OK." live, but thinking models can spend the whole budget on thoughts (`finishReason: MAX_TOKENS`, empty text).8713. Token counting is not comparable: OpenAI/Anthropic count the full request; xAI `/v1/tokenize-text` counts a **text string only** (the adapter concatenates system + message text; tools/images excluded); Gemini's wrapper form counts system + tools + media (free, no billing).8814. Gemini roles are `user`/`model`; an `assistant` message with `raw_content` is replayed verbatim (the only lossless path); without it the adapter rebuilds `functionCall` parts and re-attaches `ToolCall.thought_signature` (text-part signatures are lost).8990## 6. Verification record9192| Adapter | Language | generate | stream | Where logged |93|---|---|---|---|---|94| OpenAI (gpt-5.4-nano) | Python / TS | HTTP 200, "OK." (2026-09-18) | HTTP 200, `response.completed` | note `llm_provider.* / llmProvider.*` |95| Anthropic (claude-haiku-4-5-20251001) | Python / TS | HTTP 200, "OK." (2026-09-18) | HTTP 200, `message_stop` | idem |96| xAI Responses (grok-4.3) | Python | HTTP 200, "OK", usage 196/178 (178 = 2 visible + reasoning) | HTTP 200, "OK.", 11 events, `response.completed`, usage 196/189 | note `llm_provider.* grok-4.3 … verify xai` (2026-09-19) |97| xAI Responses (grok-4.3) | TypeScript | HTTP 200, "OK", usage 196/142 | HTTP 200, "OK.", 11 events, usage 196/100 | note `llmProvider.* … verify ts xai` |98| xAI Chat Completions (grok-4.3) | Python | HTTP 200, "OK.", usage 196/2 (+ reasoning in details) | HTTP 200, "OK.", 4 events incl. usage chunk + `[DONE]` | note `… verify xai-chat` |99| Gemini (gemini-3.5-flash-lite) | Python / TS | HTTP 200, "OK.", usage 5/2 | HTTP 200 `?alt=sse`, "OK.", 4 events, `finishReason: STOP` | note `… verify gemini` / `… verify ts gemini` |10010110 xAI/Gemini verification calls (6 py + 4 ts), estimated ≤ $0.0053 with the client's deliberately high default price table (actual at list prices: grok-4.3 $1.25/$2.50 per 1M, gemini-3.5-flash-lite $0.30/$2.50 per 1M → ≈ $0.001).102