SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
21.0 KB

# Multi-provider abstraction — OpenAI Responses ↔ Anthropic Messages ↔ xAI (Responses / Chat Completions) ↔ Gemini generateContent

Status: LIVE_VERIFIED for generate + stream in all four adapters and both languages ("Reply with OK.", ≤ 16 output tokens, all HTTP 200 — 2026-09-18 for OpenAI/Anthropic, 2026-09-19 for xAI grok-4.3 (Responses in py+ts, Chat Completions in py) and Gemini gemini-3.5-flash-lite; see reports/live-requests.jsonl notes llm_provider.* / llmProvider.*). count_tokens, use_tools, structured_output, upload_file, web_search: DOCUMENTED (spec/discovery-derived, offline-tested with mock transports in tests/shared/test_llm_provider.py, 34 tests). Gemini web_search (googleSearch) is ACCOUNT_RESTRICTED on this key (free tier → 429 limit: 0): the adapter returns a typed stop_reason == "restricted" result. Sources: OpenAI OpenAPI spec (CreateResponse, TextResponseFormatConfiguration), https://developers.openai.com/api/reference/responses/create · https://platform.claude.com/docs/en/api/messages, …/build-with-claude/structured-outputs, …/agents-and-tools/tool-use/web-search-tool · xAI: https://docs.x.ai/developers/rest-api-reference/inference/responses, …/inference/chat-completions, …/other (tokenize-text), …/files/upload, https://docs.x.ai/developers/tools/web-search, sources/xai/openapi/openapi.json (ModelRequest, TokenizeRequest, WebSearchFilters) · Gemini: https://ai.google.dev/api/generate-content, https://ai.google.dev/api/tokens, https://ai.google.dev/api/files, https://ai.google.dev/gemini-api/docs/function-calling, …/thought-signatures, …/structured-output, …/google-search, https://ai.google.dev/api/interactions, sources/gemini/discovery-v1beta.json (GenerateContentRequest, FunctionCallingConfig, CountTokensRequest) · atlas pages docs/xai/{responses,streaming,structured-outputs,files,tool-loop}.md, docs/gemini/{generate-content,streaming,structured-outputs,files,token-counting,tool-loop,interactions-api}.md Last verified: 2026-09-19

Implementation: examples/shared/provider-abstraction/llm_provider.py and llmProvider.ts (built on the resilient client; no SDK needed). get_provider("openai"|"anthropic"|"xai"|"gemini"); XAIProvider(api="responses"|"chat"); GeminiProvider(api_version="v1beta"|"v1").

# 1. Interface

text
LLMProvider
  generate(req)            -> GenerateResult {text, reasoning, tool_calls[], stop_reason, usage, response_id, assistant_content, error, raw}
  stream(req)              -> StreamEvent* {text_delta | reasoning_delta | tool_call_start | tool_call_delta | tool_call_done | usage | done | error | raw}
  count_tokens(req)        -> int | None        OpenAI POST /v1/responses/input_tokens · Anthropic POST /v1/messages/count_tokens ·
                                                xAI POST /v1/tokenize-text (text only) · Gemini POST models/{m}:countTokens {generateContentRequest}
  use_tools(req, impls)    -> GenerateResult    portable call/result loop (max_rounds, errors contained, never leaks stack traces;
                                                replays the provider-native assistant turn verbatim when the provider needs it — Gemini)
  structured_output(req, schema) -> dict        OpenAI text.format json_schema(strict) · Anthropic output_config.format · xAI text.format /
                                                response_format json_schema · Gemini generationConfig.responseMimeType+responseJsonSchema
  upload_file(path)        -> FileRef           OpenAI POST /v1/files (purpose) · Anthropic POST /v1/files (beta header) · xAI POST /v1/files
                                                (purpose ignored) · Gemini resumable upload (start → upload,finalize) → File.uri
  web_search(req)          -> GenerateResult    OpenAI {type:"web_search"} · Anthropic {type:"web_search_20250305"} · xAI {type:"web_search",
                                                allowed_domains|excluded_domains} · Gemini {googleSearch:{}} (→ "restricted" on this key)

GenerateRequest.extensions (merged verbatim into the body) and extra_headers are the only way to use provider-specific features. The adapters never translate an extension into "the closest thing" on another provider — that is the caller's decision.

# 2. Portable common denominator — parameter ↔ parameter

Concept Portable field OpenAI Responses Anthropic Messages xAI Responses (default) / Chat Completions (api="chat") Gemini generateContent
Endpoint — POST /v1/responses POST /v1/messages POST /v1/responses / POST /v1/chat/completions POST /v1beta/models/{model}:generateContent (model in the path, not the body)
Auth client Authorization: Bearer x-api-key + anthropic-version Authorization: Bearer xai-… (no version/beta headers) x-goog-api-key (never ?key=)
Model model model model model path segment models/{id}; response modelVersion
Conversation messages[] {user/assistant/tool} input[] items: {role, content}, function_call, function_call_output messages[] (user/assistant; tool results = tool_result blocks in a user turn) Responses: as OpenAI · Chat: messages[] with tool_calls[] / {role:"tool", tool_call_id} contents[] {role: user|model, parts[]}; tool results = functionResponse parts in a user Content (all calls of a turn in ONE Content)
System prompt system instructions system instructions (400 if combined with previous_response_id) / messages[0].role="system" systemInstruction {parts:[{text}]}
Output cap max_tokens max_output_tokens max_tokens (required) max_output_tokens (includes reasoning; live not enforced on reasoning) / max_tokens generationConfig.maxOutputTokens (includes thinking tokens → a small cap can yield empty text + MAX_TOKENS)
Sampling temperature temperature temperature temperature (+ xAI-only top_k, min_p) generationConfig.temperature [0, 2] (Gemini 3: keep 1.0)
Tool definitions ToolDef{name, description, parameters, strict} tools[]: {type:"function", name, …, strict} tools[]: {name, description, input_schema, strict?} as OpenAI (strict accepted, implicitly always true) / tools[]: {type:"function", function:{…}} tools:[{functionDeclarations:[{name, description, parametersJsonSchema}]}] (or parameters = OpenAPI subset, UPPERCASE types)
Tool choice "auto" | "none" | "required" | {name} "auto"/"none"/"required"/{type:"function", name} {type:"auto"}/{type:"none"}/{type:"any"}/{type:"tool", name} as OpenAI / {type:"function", function:{name}} toolConfig.functionCallingConfig.mode: AUTO | NONE | ANY (+ allowedFunctionNames) (also VALIDATED)
Tool call (model → you) ToolCall{id, name, arguments, thought_signature?} item function_call {call_id, name, arguments (string)} block tool_use {id, name, input (object)} as OpenAI (call-… ids) / message.tool_calls[] {id, function:{name, arguments}} part functionCall {name, args (object), id} + sibling thoughtSignature on the part
Tool result (you → model) Message(role="tool", tool_call_id, tool_name, content, is_error) function_call_output {call_id, output} — no error flag → [TOOL ERROR] prefix tool_result {tool_use_id, content, is_error} as OpenAI / {role:"tool", tool_call_id, content} functionResponse {name, id?, response: OBJECT} — scalars wrapped as {result}, errors as {error}
JSON-schema output json_schema{name, schema, strict} text.format = {type:"json_schema", name, schema, strict} output_config.format = {type:"json_schema", schema} text.format (same) / response_format = {type:"json_schema", json_schema:{name, schema, strict}} generationConfig.responseMimeType: "application/json" + responseJsonSchema (new form: responseFormat.text {mimeType: APPLICATION_JSON, schema})
Streaming text text_delta response.output_text.delta content_block_delta/text_delta same as OpenAI / choices[].delta.content chunk candidates[0].content.parts[].text (non-thought)
Streaming reasoning reasoning_delta response.reasoning_summary_text.delta thinking_delta response.reasoning_summary_text.delta / delta.reasoning_content parts with thought: true (only with thinkingConfig.includeThoughts)
Streaming tool args tool_call_delta function_call_arguments.delta → .done input_json_delta → content_block_stop one delta with the whole JSON / one chunk with the whole tool_calls[] never partial: functionCall.args arrives whole in one part
Usage Usage{input, output, cached_input, cache_write, total} input_tokens, output_tokens, input_tokens_details.cached_tokens input_tokens (excl. cache) + cache_read/creation_input_tokens — normalised to "input includes cached" Responses: as OpenAI, output_tokens includes reasoning · Chat: prompt_tokens, completion_tokens (excludes reasoning) + completion_tokens_details.reasoning_tokens — normalised to billed output usageMetadata.promptTokenCount (incl. cached), candidatesTokenCount + thoughtsTokenCount (both billed → output), cachedContentTokenCount, totalTokenCount
Stop reason end | max_tokens | tool_use | stop_sequence | refusal | incomplete | restricted | other status + incomplete_details.reason, refusal part stop_reason enum as OpenAI / finish_reason: stop|length|tool_calls|content_filter finishReason: STOP, MAX_TOKENS, SAFETY/RECITATION/BLOCKLIST/PROHIBITED_CONTENT/SPII/IMAGE_*/LANGUAGE → refusal, MALFORMED_FUNCTION_CALL/MISSING_THOUGHT_SIGNATURE/UNEXPECTED_TOOL_CALL/TOO_MANY_TOOL_CALLS/MALFORMED_RESPONSE → other; promptFeedback.blockReason (no candidates, HTTP 200) → refusal
Stop sequences stop[] none → dropped stop_sequences Responses: none → dropped / Chat: stop (≤ 4) generationConfig.stopSequences (≤ 5, else 400); hit → finishReason: STOP (indistinguishable from a natural end)
Files in prompt Message.file_ids {type:"input_file", file_id} {type:"document", source:{type:"file", file_id}} Responses: {type:"input_file", file_id} (turns the request into an "attachment search", $10/1k) / Chat: 400 {fileData:{fileUri: File.uri}} (FileRef.id = the URI)
Request metadata extensions.metadata metadata (16 pairs) metadata.user_id Responses 400 "Argument not supported: metadata" (response metadata = {system_fingerprint}) / Chat: ignored labels (Cloud label rules; documented key safety_identifier), accepted, not echoed

# 3. State management — four different answers

OpenAI Responses Anthropic Messages xAI Responses Gemini generateContent Gemini Interactions API (not wrapped here)
Server-side state previous_response_id (needs store: true, default) or conversation id none — stateless, replay messages[] previous_response_id (30-day retention; store:false responses are still retrievable — LIVE_DISCOVERED); alt. replay output[] incl. reasoning.encrypted_content; POST /v1/responses/compact none — stateless, replay contents[] previous_interaction_id (default store: true; store:false = stateless input[] of Steps; 400 if the previous interaction is still in_progress); background: true + polling/resumable SSE
Reasoning carry-over reasoning items / reasoning.encrypted_content thinking blocks with signature (echo verbatim) reasoning items / encrypted_content (replay verbatim) thoughtSignature on parts — echo the model turn VERBATIM (GenerateResult.assistant_content → Message.raw_content); dropping it → 400 "Function call is missing a thought_signature"; only the first parallel call carries one thought step signature (step.delta thought_signature)
Restrictions store:false disables chaining — instructions + previous_response_id together → 400; background unsupported (400) last turn must be user; roles are user/model (not assistant/system) store:false incompatible with background and later chaining; snake_case only
ZDR Zero Data Retention org setting ZDR arrangement team-wide ZDR disables store/previous_response_id, Files, Collections, Batch not fully achievable (Search/Maps grounding stores 30 days) store:false; not a ZDR guarantee

# 4. Provider-specific extensions (pass through extensions / extra_headers)

OpenAI Responses: previous_response_id, conversation, reasoning {effort, summary}, include[], background, store, service_tier, truncation, max_tool_calls, parallel_tool_calls, prompt, prompt_cache_key, safety_identifier, top_logprobs, context_management, built-in tools (web_search, file_search, code_interpreter, image_generation, mcp, computer_use_preview, shell, apply_patch), text.verbosity.

Anthropic Messages: thinking {type, budget_tokens}, output_config {effort, format}, cache_control, service_tier, speed, top_k, container, inference_geo, metadata.user_id, tool_choice.disable_parallel_tool_use, server tools (web_search_*, web_fetch_*, code_execution_*, computer_toolset_*, text_editor_*, bash_*, memory_*, tool_search_tool_*, mcp_toolset), mcp_servers (+ anthropic-beta: mcp-client-2025-11-20), context_management (+ beta header).

xAI Responses (ModelRequest, docs/xai/responses.md): previous_response_id, store, include[] (reasoning.encrypted_content, web_search_call.action.sources, code_interpreter_call.outputs, file_search_call.results, no_inline_citations), max_turns (agentic turns per request), reasoning {effort: low|medium|high|xhigh, summary} / reasoning_effort, top_k, min_p, service_tier: default|priority, prompt_cache_key (= x-grok-conv-id), safety_identifier, parallel_tool_calls, server tools (web_search, x_search, code_interpreter/code_execution, file_search/collections_search, mcp (no approval round-trip), image_generation, shell, tool_search (403 alpha)). Rejected: metadata, background, search_parameters (retired). xAI Chat Completions: reasoning_effort, stream_options.include_usage, x-grok-conv-id header, deferred: true (poll GET /v1/chat/deferred-completion/{id}), response_format.json_object; no server-side tools (422/410), no files (400). Anthropic-compatible POST /v1/messages exists but is DEPRECATED (not wrapped).

Gemini generateContent (docs/gemini/generate-content.md): generationConfig.{thinkingConfig{thinkingLevel|thinkingBudget, includeThoughts}, topP, topK, seed, candidateCount(=1), responseModalities, speechConfig, imageConfig, mediaResolution, responseSchema (OpenAPI subset), responseFormat (new)}, safetySettings[], cachedContent (cachedContents/{id}, explicit caching — ACCOUNT_RESTRICTED on free tier), serviceTier: standard|flex|priority, labels, store, toolConfig.{includeServerSideToolInvocations, retrievalConfig}, server tools (googleSearch{searchTypes, timeRangeFilter}, googleMaps, urlContext, codeExecution, fileSearch{fileSearchStoreNames}, computerUse{environment, disabledSafetyPolicies}, mcpServers[] (schema only, UNVERIFIED on generateContent)), per-part videoMetadata, mediaResolution, mediaProcessing. The adapter merges extensions["generationConfig"] into its own generationConfig and appends extensions["tools"] to the function declarations; everything else is top-level.

# 5. Lossy conversions (be explicit with users)

  1. stop[] → OpenAI Responses and xAI Responses: dropped (no parameter). xAI Chat keeps ≤ 4; Gemini keeps ≤ 5 but reports the hit as a plain STOP (no stop_sequence stop reason).
  2. Tool result is_error → OpenAI/xAI: no field, [TOOL ERROR] prefix; Gemini: no field, the adapter sends {"error": …} as the functionResponse.response object (the documented pattern).
  3. Message.tool_name is required by Gemini (functionResponse.name) — the portable loop fills it from the ToolCall; callers building messages by hand must set it.
  4. Gemini functionResponse.response must be an object: string/number results are wrapped as {"result": …}; JSON-object strings are parsed back.
  5. OpenAI previous_response_id/conversation, xAI previous_response_id, Gemini Interactions previous_interaction_id → Anthropic / Gemini generateContent: replay the full history yourself. Reverse direction: the stateless histories contain provider-native reasoning artefacts (thoughtSignature, thinking signatures, encrypted_content) that do not transfer between providers — start a fresh conversation when switching.
  6. Reasoning controls: OpenAI reasoning.effort ↔ Anthropic output_config.effort/thinking.budget_tokens ↔ xAI reasoning.effort (low|medium|high|xhigh; grok-4.3 defaults to low and always bills reasoning tokens) ↔ Gemini thinkingConfig.thinkingLevel XOR thinkingBudget (both → 400). Same intent, four scales — not auto-mapped.
  7. Usage: Anthropic input_tokens excludes cache reads; xAI Chat completion_tokens excludes reasoning while xAI Responses output_tokens includes it; Gemini candidatesTokenCount excludes thoughtsTokenCount. The adapter normalises to "input includes cached, output = everything billed as output" — read raw for the provider view.
  8. Structured output: OpenAI needs strict:true + additionalProperties:false + all required; Anthropic has its own keyword limits; xAI defaults additionalProperties to false and rejects items arrays / minContains; Gemini ignores unsupported keywords silently (minLength), rejects unknown keys in the legacy responseSchema, supports $defs/$ref/anyOf/null, and truncates JSON on MAX_TOKENS (the adapter raises). One schema may need four variants.
  9. Web search: Anthropic max_uses/allowed_domains|blocked_domains/user_location; OpenAI different filters; xAI allowed_domains|excluded_domains (≤ 5, mutually exclusive, search_context_size → 400) and no max_uses (bound with max_turns); Gemini googleSearch has no domain filters at all (only timeRangeFilter/searchTypes) and is unavailable on free-tier keys (429 limit: 0 → restricted).
  10. Files: OpenAI purpose required; Anthropic beta header, size_bytes; xAI purpose ignored (echoed ""), 50/512 MB, team-scoped; Gemini 2-step resumable protocol, purpose has no equivalent, files auto-delete after 48 h, referenced by uri (not id), 2 GB.
  11. Streaming stop reasons arrive at different moments: OpenAI/xAI in the terminal response.completed; Anthropic in message_delta; xAI Chat in the last choices[].finish_reason before [DONE]; Gemini only on the last chunk (finishReason) — a connection closed without it is incomplete.
  12. Anthropic requires max_tokens; the adapter defaults to 1024 everywhere. On Gemini and xAI Responses that cap also covers reasoning/thinking tokens: 16 tokens was enough for "OK." live, but thinking models can spend the whole budget on thoughts (finishReason: MAX_TOKENS, empty text).
  13. Token counting is not comparable: OpenAI/Anthropic count the full request; xAI /v1/tokenize-text counts a text string only (the adapter concatenates system + message text; tools/images excluded); Gemini's wrapper form counts system + tools + media (free, no billing).
  14. Gemini roles are user/model; an assistant message with raw_content is replayed verbatim (the only lossless path); without it the adapter rebuilds functionCall parts and re-attaches ToolCall.thought_signature (text-part signatures are lost).

# 6. Verification record

Adapter Language generate stream Where logged
OpenAI (gpt-5.4-nano) Python / TS HTTP 200, "OK." (2026-09-18) HTTP 200, response.completed note llm_provider.* / llmProvider.*
Anthropic (claude-haiku-4-5-20251001) Python / TS HTTP 200, "OK." (2026-09-18) HTTP 200, message_stop idem
xAI Responses (grok-4.3) Python HTTP 200, "OK", usage 196/178 (178 = 2 visible + reasoning) HTTP 200, "OK.", 11 events, response.completed, usage 196/189 note llm_provider.* grok-4.3 … verify xai (2026-09-19)
xAI Responses (grok-4.3) TypeScript HTTP 200, "OK", usage 196/142 HTTP 200, "OK.", 11 events, usage 196/100 note llmProvider.* … verify ts xai
xAI Chat Completions (grok-4.3) Python HTTP 200, "OK.", usage 196/2 (+ reasoning in details) HTTP 200, "OK.", 4 events incl. usage chunk + [DONE] note … verify xai-chat
Gemini (gemini-3.5-flash-lite) Python / TS HTTP 200, "OK.", usage 5/2 HTTP 200 ?alt=sse, "OK.", 4 events, finishReason: STOP note … verify gemini / … verify ts gemini

10 xAI/Gemini verification calls (6 py + 4 ts), estimated ≤ $0.0053 with the client's deliberately high default price table (actual at list prices: grok-4.3 $1.25/$2.50 per 1M, gemini-3.5-flash-lite $0.30/$2.50 per 1M → ≈ $0.001).