Multi-provider abstraction — OpenAI Responses ↔ Anthropic Messages ↔ xAI (Responses / Chat Completions) ↔ Gemini generateContent
Status: LIVE_VERIFIED for generate + stream in all four adapters and both languages ("Reply with OK.", ≤ 16 output tokens, all HTTP 200 — 2026-09-18 for OpenAI/Anthropic, 2026-09-19 for xAI grok-4.3 (Responses in py+ts, Chat Completions in py) and Gemini gemini-3.5-flash-lite; see reports/live-requests.jsonl notes llm_provider.* / llmProvider.*). count_tokens, use_tools, structured_output, upload_file, web_search: DOCUMENTED (spec/discovery-derived, offline-tested with mock transports in tests/shared/test_llm_provider.py, 34 tests). Gemini web_search (googleSearch) is ACCOUNT_RESTRICTED on this key (free tier → 429 limit: 0): the adapter returns a typed stop_reason == "restricted" result.
Sources: OpenAI OpenAPI spec (CreateResponse, TextResponseFormatConfiguration), https://developers.openai.com/api/reference/responses/create · https://platform.claude.com/docs/en/api/messages, …/build-with-claude/structured-outputs, …/agents-and-tools/tool-use/web-search-tool · xAI: https://docs.x.ai/developers/rest-api-reference/inference/responses, …/inference/chat-completions, …/other (tokenize-text), …/files/upload, https://docs.x.ai/developers/tools/web-search, sources/xai/openapi/openapi.json (ModelRequest, TokenizeRequest, WebSearchFilters) · Gemini: https://ai.google.dev/api/generate-content, https://ai.google.dev/api/tokens, https://ai.google.dev/api/files, https://ai.google.dev/gemini-api/docs/function-calling, …/thought-signatures, …/structured-output, …/google-search, https://ai.google.dev/api/interactions, sources/gemini/discovery-v1beta.json (GenerateContentRequest, FunctionCallingConfig, CountTokensRequest) · atlas pages docs/xai/{responses,streaming,structured-outputs,files,tool-loop}.md, docs/gemini/{generate-content,streaming,structured-outputs,files,token-counting,tool-loop,interactions-api}.md
Last verified: 2026-09-19
Implementation: examples/shared/provider-abstraction/llm_provider.py and llmProvider.ts (built on the resilient client; no SDK needed). get_provider("openai"|"anthropic"|"xai"|"gemini"); XAIProvider(api="responses"|"chat"); GeminiProvider(api_version="v1beta"|"v1").
1. Interface
LLMProvider
generate(req) -> GenerateResult {text, reasoning, tool_calls[], stop_reason, usage, response_id, assistant_content, error, raw}
stream(req) -> StreamEvent* {text_delta | reasoning_delta | tool_call_start | tool_call_delta | tool_call_done | usage | done | error | raw}
count_tokens(req) -> int | None OpenAI POST /v1/responses/input_tokens · Anthropic POST /v1/messages/count_tokens ·
xAI POST /v1/tokenize-text (text only) · Gemini POST models/{m}:countTokens {generateContentRequest}
use_tools(req, impls) -> GenerateResult portable call/result loop (max_rounds, errors contained, never leaks stack traces;
replays the provider-native assistant turn verbatim when the provider needs it — Gemini)
structured_output(req, schema) -> dict OpenAI text.format json_schema(strict) · Anthropic output_config.format · xAI text.format /
response_format json_schema · Gemini generationConfig.responseMimeType+responseJsonSchema
upload_file(path) -> FileRef OpenAI POST /v1/files (purpose) · Anthropic POST /v1/files (beta header) · xAI POST /v1/files
(purpose ignored) · Gemini resumable upload (start → upload,finalize) → File.uri
web_search(req) -> GenerateResult OpenAI {type:"web_search"} · Anthropic {type:"web_search_20250305"} · xAI {type:"web_search",
allowed_domains|excluded_domains} · Gemini {googleSearch:{}} (→ "restricted" on this key)GenerateRequest.extensions (merged verbatim into the body) and extra_headers are the only way to use provider-specific features. The adapters never translate an extension into "the closest thing" on another provider — that is the caller's decision.
2. Portable common denominator — parameter ↔ parameter
| Concept | Portable field | OpenAI Responses | Anthropic Messages | xAI Responses (default) / Chat Completions (api="chat") |
Gemini generateContent |
|---|---|---|---|---|---|
| Endpoint | — | POST /v1/responses |
POST /v1/messages |
POST /v1/responses / POST /v1/chat/completions |
POST /v1beta/models/{model}:generateContent (model in the path, not the body) |
| Auth | client | Authorization: Bearer |
x-api-key + anthropic-version |
Authorization: Bearer xai-… (no version/beta headers) |
x-goog-api-key (never ?key=) |
| Model | model |
model |
model |
model |
path segment models/{id}; response modelVersion |
| Conversation | messages[] {user/assistant/tool} |
input[] items: {role, content}, function_call, function_call_output |
messages[] (user/assistant; tool results = tool_result blocks in a user turn) |
Responses: as OpenAI · Chat: messages[] with tool_calls[] / {role:"tool", tool_call_id} |
contents[] {role: user|model, parts[]}; tool results = functionResponse parts in a user Content (all calls of a turn in ONE Content) |
| System prompt | system |
instructions |
system |
instructions (400 if combined with previous_response_id) / messages[0].role="system" |
systemInstruction {parts:[{text}]} |
| Output cap | max_tokens |
max_output_tokens |
max_tokens (required) |
max_output_tokens (includes reasoning; live not enforced on reasoning) / max_tokens |
generationConfig.maxOutputTokens (includes thinking tokens → a small cap can yield empty text + MAX_TOKENS) |
| Sampling | temperature |
temperature |
temperature |
temperature (+ xAI-only top_k, min_p) |
generationConfig.temperature [0, 2] (Gemini 3: keep 1.0) |
| Tool definitions | ToolDef{name, description, parameters, strict} |
tools[]: {type:"function", name, …, strict} |
tools[]: {name, description, input_schema, strict?} |
as OpenAI (strict accepted, implicitly always true) / tools[]: {type:"function", function:{…}} |
tools:[{functionDeclarations:[{name, description, parametersJsonSchema}]}] (or parameters = OpenAPI subset, UPPERCASE types) |
| Tool choice | "auto" | "none" | "required" | {name} |
"auto"/"none"/"required"/{type:"function", name} |
{type:"auto"}/{type:"none"}/{type:"any"}/{type:"tool", name} |
as OpenAI / {type:"function", function:{name}} |
toolConfig.functionCallingConfig.mode: AUTO | NONE | ANY (+ allowedFunctionNames) (also VALIDATED) |
| Tool call (model → you) | ToolCall{id, name, arguments, thought_signature?} |
item function_call {call_id, name, arguments (string)} |
block tool_use {id, name, input (object)} |
as OpenAI (call-… ids) / message.tool_calls[] {id, function:{name, arguments}} |
part functionCall {name, args (object), id} + sibling thoughtSignature on the part |
| Tool result (you → model) | Message(role="tool", tool_call_id, tool_name, content, is_error) |
function_call_output {call_id, output} — no error flag → [TOOL ERROR] prefix |
tool_result {tool_use_id, content, is_error} |
as OpenAI / {role:"tool", tool_call_id, content} |
functionResponse {name, id?, response: OBJECT} — scalars wrapped as {result}, errors as {error} |
| JSON-schema output | json_schema{name, schema, strict} |
text.format = {type:"json_schema", name, schema, strict} |
output_config.format = {type:"json_schema", schema} |
text.format (same) / response_format = {type:"json_schema", json_schema:{name, schema, strict}} |
generationConfig.responseMimeType: "application/json" + responseJsonSchema (new form: responseFormat.text {mimeType: APPLICATION_JSON, schema}) |
| Streaming text | text_delta |
response.output_text.delta |
content_block_delta/text_delta |
same as OpenAI / choices[].delta.content |
chunk candidates[0].content.parts[].text (non-thought) |
| Streaming reasoning | reasoning_delta |
response.reasoning_summary_text.delta |
thinking_delta |
response.reasoning_summary_text.delta / delta.reasoning_content |
parts with thought: true (only with thinkingConfig.includeThoughts) |
| Streaming tool args | tool_call_delta |
function_call_arguments.delta → .done |
input_json_delta → content_block_stop |
one delta with the whole JSON / one chunk with the whole tool_calls[] |
never partial: functionCall.args arrives whole in one part |
| Usage | Usage{input, output, cached_input, cache_write, total} |
input_tokens, output_tokens, input_tokens_details.cached_tokens |
input_tokens (excl. cache) + cache_read/creation_input_tokens — normalised to "input includes cached" |
Responses: as OpenAI, output_tokens includes reasoning · Chat: prompt_tokens, completion_tokens (excludes reasoning) + completion_tokens_details.reasoning_tokens — normalised to billed output |
usageMetadata.promptTokenCount (incl. cached), candidatesTokenCount + thoughtsTokenCount (both billed → output), cachedContentTokenCount, totalTokenCount |
| Stop reason | end | max_tokens | tool_use | stop_sequence | refusal | incomplete | restricted | other |
status + incomplete_details.reason, refusal part |
stop_reason enum |
as OpenAI / finish_reason: stop|length|tool_calls|content_filter |
finishReason: STOP, MAX_TOKENS, SAFETY/RECITATION/BLOCKLIST/PROHIBITED_CONTENT/SPII/IMAGE_*/LANGUAGE → refusal, MALFORMED_FUNCTION_CALL/MISSING_THOUGHT_SIGNATURE/UNEXPECTED_TOOL_CALL/TOO_MANY_TOOL_CALLS/MALFORMED_RESPONSE → other; promptFeedback.blockReason (no candidates, HTTP 200) → refusal |
| Stop sequences | stop[] |
none → dropped | stop_sequences |
Responses: none → dropped / Chat: stop (≤ 4) |
generationConfig.stopSequences (≤ 5, else 400); hit → finishReason: STOP (indistinguishable from a natural end) |
| Files in prompt | Message.file_ids |
{type:"input_file", file_id} |
{type:"document", source:{type:"file", file_id}} |
Responses: {type:"input_file", file_id} (turns the request into an "attachment search", $10/1k) / Chat: 400 |
{fileData:{fileUri: File.uri}} (FileRef.id = the URI) |
| Request metadata | extensions.metadata |
metadata (16 pairs) |
metadata.user_id |
Responses 400 "Argument not supported: metadata" (response metadata = {system_fingerprint}) / Chat: ignored |
labels (Cloud label rules; documented key safety_identifier), accepted, not echoed |
3. State management — four different answers
| OpenAI Responses | Anthropic Messages | xAI Responses | Gemini generateContent | Gemini Interactions API (not wrapped here) | |
|---|---|---|---|---|---|
| Server-side state | previous_response_id (needs store: true, default) or conversation id |
none — stateless, replay messages[] |
previous_response_id (30-day retention; store:false responses are still retrievable — LIVE_DISCOVERED); alt. replay output[] incl. reasoning.encrypted_content; POST /v1/responses/compact |
none — stateless, replay contents[] |
previous_interaction_id (default store: true; store:false = stateless input[] of Steps; 400 if the previous interaction is still in_progress); background: true + polling/resumable SSE |
| Reasoning carry-over | reasoning items / reasoning.encrypted_content |
thinking blocks with signature (echo verbatim) |
reasoning items / encrypted_content (replay verbatim) |
thoughtSignature on parts — echo the model turn VERBATIM (GenerateResult.assistant_content → Message.raw_content); dropping it → 400 "Function call is missing a thought_signature"; only the first parallel call carries one |
thought step signature (step.delta thought_signature) |
| Restrictions | store:false disables chaining |
— | instructions + previous_response_id together → 400; background unsupported (400) |
last turn must be user; roles are user/model (not assistant/system) |
store:false incompatible with background and later chaining; snake_case only |
| ZDR | Zero Data Retention org setting | ZDR arrangement | team-wide ZDR disables store/previous_response_id, Files, Collections, Batch |
not fully achievable (Search/Maps grounding stores 30 days) | store:false; not a ZDR guarantee |
4. Provider-specific extensions (pass through extensions / extra_headers)
OpenAI Responses: previous_response_id, conversation, reasoning {effort, summary}, include[], background, store, service_tier, truncation, max_tool_calls, parallel_tool_calls, prompt, prompt_cache_key, safety_identifier, top_logprobs, context_management, built-in tools (web_search, file_search, code_interpreter, image_generation, mcp, computer_use_preview, shell, apply_patch), text.verbosity.
Anthropic Messages: thinking {type, budget_tokens}, output_config {effort, format}, cache_control, service_tier, speed, top_k, container, inference_geo, metadata.user_id, tool_choice.disable_parallel_tool_use, server tools (web_search_*, web_fetch_*, code_execution_*, computer_toolset_*, text_editor_*, bash_*, memory_*, tool_search_tool_*, mcp_toolset), mcp_servers (+ anthropic-beta: mcp-client-2025-11-20), context_management (+ beta header).
xAI Responses (ModelRequest, docs/xai/responses.md): previous_response_id, store, include[] (reasoning.encrypted_content, web_search_call.action.sources, code_interpreter_call.outputs, file_search_call.results, no_inline_citations), max_turns (agentic turns per request), reasoning {effort: low|medium|high|xhigh, summary} / reasoning_effort, top_k, min_p, service_tier: default|priority, prompt_cache_key (= x-grok-conv-id), safety_identifier, parallel_tool_calls, server tools (web_search, x_search, code_interpreter/code_execution, file_search/collections_search, mcp (no approval round-trip), image_generation, shell, tool_search (403 alpha)). Rejected: metadata, background, search_parameters (retired). xAI Chat Completions: reasoning_effort, stream_options.include_usage, x-grok-conv-id header, deferred: true (poll GET /v1/chat/deferred-completion/{id}), response_format.json_object; no server-side tools (422/410), no files (400). Anthropic-compatible POST /v1/messages exists but is DEPRECATED (not wrapped).
Gemini generateContent (docs/gemini/generate-content.md): generationConfig.{thinkingConfig{thinkingLevel|thinkingBudget, includeThoughts}, topP, topK, seed, candidateCount(=1), responseModalities, speechConfig, imageConfig, mediaResolution, responseSchema (OpenAPI subset), responseFormat (new)}, safetySettings[], cachedContent (cachedContents/{id}, explicit caching — ACCOUNT_RESTRICTED on free tier), serviceTier: standard|flex|priority, labels, store, toolConfig.{includeServerSideToolInvocations, retrievalConfig}, server tools (googleSearch{searchTypes, timeRangeFilter}, googleMaps, urlContext, codeExecution, fileSearch{fileSearchStoreNames}, computerUse{environment, disabledSafetyPolicies}, mcpServers[] (schema only, UNVERIFIED on generateContent)), per-part videoMetadata, mediaResolution, mediaProcessing. The adapter merges extensions["generationConfig"] into its own generationConfig and appends extensions["tools"] to the function declarations; everything else is top-level.
5. Lossy conversions (be explicit with users)
stop[]→ OpenAI Responses and xAI Responses: dropped (no parameter). xAI Chat keeps ≤ 4; Gemini keeps ≤ 5 but reports the hit as a plainSTOP(nostop_sequencestop reason).- Tool result
is_error→ OpenAI/xAI: no field,[TOOL ERROR]prefix; Gemini: no field, the adapter sends{"error": …}as thefunctionResponse.responseobject (the documented pattern). Message.tool_nameis required by Gemini (functionResponse.name) — the portable loop fills it from theToolCall; callers building messages by hand must set it.- Gemini
functionResponse.responsemust be an object: string/number results are wrapped as{"result": …}; JSON-object strings are parsed back. - OpenAI
previous_response_id/conversation, xAIprevious_response_id, Gemini Interactionsprevious_interaction_id→ Anthropic / Gemini generateContent: replay the full history yourself. Reverse direction: the stateless histories contain provider-native reasoning artefacts (thoughtSignature,thinkingsignatures,encrypted_content) that do not transfer between providers — start a fresh conversation when switching. - Reasoning controls: OpenAI
reasoning.effort↔ Anthropicoutput_config.effort/thinking.budget_tokens↔ xAIreasoning.effort(low|medium|high|xhigh; grok-4.3 defaults tolowand always bills reasoning tokens) ↔ GeminithinkingConfig.thinkingLevelXORthinkingBudget(both → 400). Same intent, four scales — not auto-mapped. - Usage: Anthropic
input_tokensexcludes cache reads; xAI Chatcompletion_tokensexcludes reasoning while xAI Responsesoutput_tokensincludes it; GeminicandidatesTokenCountexcludesthoughtsTokenCount. The adapter normalises to "input includes cached, output = everything billed as output" — readrawfor the provider view. - Structured output: OpenAI needs
strict:true+additionalProperties:false+ allrequired; Anthropic has its own keyword limits; xAI defaultsadditionalPropertiesto false and rejectsitemsarrays /minContains; Gemini ignores unsupported keywords silently (minLength), rejects unknown keys in the legacyresponseSchema, supports$defs/$ref/anyOf/null, and truncates JSON onMAX_TOKENS(the adapter raises). One schema may need four variants. - Web search: Anthropic
max_uses/allowed_domains|blocked_domains/user_location; OpenAI different filters; xAIallowed_domains|excluded_domains(≤ 5, mutually exclusive,search_context_size→ 400) and nomax_uses(bound withmax_turns); GeminigoogleSearchhas no domain filters at all (onlytimeRangeFilter/searchTypes) and is unavailable on free-tier keys (429limit: 0→restricted). - Files: OpenAI
purposerequired; Anthropic beta header,size_bytes; xAIpurposeignored (echoed""), 50/512 MB, team-scoped; Gemini 2-step resumable protocol,purposehas no equivalent, files auto-delete after 48 h, referenced byuri(not id), 2 GB. - Streaming stop reasons arrive at different moments: OpenAI/xAI in the terminal
response.completed; Anthropic inmessage_delta; xAI Chat in the lastchoices[].finish_reasonbefore[DONE]; Gemini only on the last chunk (finishReason) — a connection closed without it isincomplete. - Anthropic requires
max_tokens; the adapter defaults to 1024 everywhere. On Gemini and xAI Responses that cap also covers reasoning/thinking tokens: 16 tokens was enough for "OK." live, but thinking models can spend the whole budget on thoughts (finishReason: MAX_TOKENS, empty text). - Token counting is not comparable: OpenAI/Anthropic count the full request; xAI
/v1/tokenize-textcounts a text string only (the adapter concatenates system + message text; tools/images excluded); Gemini's wrapper form counts system + tools + media (free, no billing). - Gemini roles are
user/model; anassistantmessage withraw_contentis replayed verbatim (the only lossless path); without it the adapter rebuildsfunctionCallparts and re-attachesToolCall.thought_signature(text-part signatures are lost).
6. Verification record
| Adapter | Language | generate | stream | Where logged |
|---|---|---|---|---|
| OpenAI (gpt-5.4-nano) | Python / TS | HTTP 200, "OK." (2026-09-18) | HTTP 200, response.completed |
note llm_provider.* / llmProvider.* |
| Anthropic (claude-haiku-4-5-20251001) | Python / TS | HTTP 200, "OK." (2026-09-18) | HTTP 200, message_stop |
idem |
| xAI Responses (grok-4.3) | Python | HTTP 200, "OK", usage 196/178 (178 = 2 visible + reasoning) | HTTP 200, "OK.", 11 events, response.completed, usage 196/189 |
note llm_provider.* grok-4.3 … verify xai (2026-09-19) |
| xAI Responses (grok-4.3) | TypeScript | HTTP 200, "OK", usage 196/142 | HTTP 200, "OK.", 11 events, usage 196/100 | note llmProvider.* … verify ts xai |
| xAI Chat Completions (grok-4.3) | Python | HTTP 200, "OK.", usage 196/2 (+ reasoning in details) | HTTP 200, "OK.", 4 events incl. usage chunk + [DONE] |
note … verify xai-chat |
| Gemini (gemini-3.5-flash-lite) | Python / TS | HTTP 200, "OK.", usage 5/2 | HTTP 200 ?alt=sse, "OK.", 4 events, finishReason: STOP |
note … verify gemini / … verify ts gemini |
10 xAI/Gemini verification calls (6 py + 4 ts), estimated ≤ $0.0053 with the client's deliberately high default price table (actual at list prices: grok-4.3 $1.25/$2.50 per 1M, gemini-3.5-flash-lite $0.30/$2.50 per 1M → ≈ $0.001).