SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
23.9 KB

# Prompt caching and reasoning — parameter-level mapping across four providers

Status: parameters from generated/parameters.json (prompt_cache_*, reasoning.*, cache_control*, thinking.*, output_config.*, xAI prompt_cache_key / reasoning_effort / include, Gemini cachedContent, generationConfig.thinkingConfig.*), per-model facts from generated/models.json and generated/compatibility/{anthropic,gemini}-feature-model-matrix.json, prices from generated/pricing.json; live observations quoted from docs/openai/prompt-caching.md, docs/openai/reasoning.md, docs/anthropic/prompt-caching.md, docs/anthropic/thinking.md, docs/anthropic/effort.md, docs/xai/prompt-caching.md, docs/xai/reasoning.md, docs/gemini/context-caching.md, docs/gemini/thinking.md (2026-09-18/19). Gemini explicit caching is ACCOUNT_RESTRICTED on this free-tier key (429 … limit=0); size validation and CRUD shapes were LIVE_VERIFIED. Sources: https://developers.openai.com/api/docs/guides/prompt-caching · https://developers.openai.com/api/docs/guides/reasoning · https://platform.claude.com/docs/en/build-with-claude/prompt-caching · https://platform.claude.com/docs/en/build-with-claude/extended-thinking · https://platform.claude.com/docs/en/build-with-claude/effort · https://docs.x.ai/developers/model-capabilities/text/prompt-caching · https://docs.x.ai/developers/model-capabilities/text/reasoning · https://ai.google.dev/gemini-api/docs/caching · https://ai.google.dev/gemini-api/docs/thinking Last verified: 2026-09-18

# Part A — Prompt / context caching

# A1. Mechanics

Topic OpenAI Anthropic xAI Gemini
Activation automatic on every request ≥ minimum (implicit); GPT-5.6+ optional explicit mode explicit: cache_control: {type: ephemeral} on ≤4 blocks or top-level automatic cache_control automatic prefix cache, no markers; cache_control on /v1/messages is ignored implicit (automatic, 2.5+) and explicit cachedContents resources referenced by cachedContent
What is cached rendered prefix (hidden system content, instructions, tools, text.format, history); exact prefix match prefix tools → system → messages up to the marked block; byte-identical; 20-block lookback prompt prefix incl. the hidden system prefix (a fresh minimal call already shows ~192 cached tokens), instructions, tools, history incl. prior reasoning_content / encrypted reasoning implicit: the leading prefix of contents (put shared content first); explicit: contents, systemInstruction, tools, toolConfig stored under a cachedContents/{id} name for one model
Minimum 1,024 tokens on GPT-5.6+ (model-dependent before); hits in 128-token increments per model: 512 (Fable/Mythos/Opus 5) · 1,024 (Sonnet 5/4.6/4.5, Opus 4.8) · 2,048 (Opus 4.7) · 4,096 (Haiku 4.5, Opus 4.6/4.5) not documented; live hits at 192 and 896 tokens implicit 4,096 (3.8/3.7/3.6/3.5 Flash, 3.1 Pro), 2,048 (2.5 Flash/Pro); explicit 1,024 live (min_total_token_count=1024; docs quote 4,096)
Lifetime in-memory 5–10 min (≤1 h) or 24h retention; GPT-5.6+ prompt_cache_options.ttl: "30m" ttl: "5m" (default) or "1h"; hits refresh for free not documented; caches do not carry across hosts implicit: undocumented; explicit: ttl (default 1 h) or expireTime, mutable via PATCH; no min/max documented
Routing prompt_cache_key (stable per prefix) none needed prompt_cache_key (Responses, Chat — "plumbed to x-grok-conv-id") or header x-grok-conv-id (Chat/gRPC): sticky routing to the server holding the cache; top cause of misses = omitting prior reasoning or not using previous_response_id none needed (cachedContent names the cache explicitly)
Explicit breakpoints GPT-5.6+: prompt_cache_breakpoint: {mode: "explicit"} on parts (≤4) with prompt_cache_options.mode: "explicit" ≤4 cache_control blocks in system/tools/messages none one explicit cache per request (cachedContent); the cache must use the same model; do not repeat the cached systemInstruction/tools in the request
Pre-warm prompt_cache_options.prewarm: true max_tokens: 0 request with a breakpoint none documented POST /v1beta/cachedContents creates (and bills storage for) the prefix before first use
Diagnostics prompt_cache_options.comparison_response_id → prompt_cache_diagnostics {cache_hit | cache_miss {reason}} (5.6+) beta cache-diagnosis-2026-04-07: diagnostics.previous_message_id → miss reasons none none
Usage fields usage.input_tokens_details.cached_tokens, .cache_write_tokens; Chat prompt_tokens_details.cached_tokens usage.cache_read_input_tokens, cache_creation_input_tokens, cache_creation.ephemeral_{5m,1h}_input_tokens; input_tokens excludes cached Responses usage.input_tokens_details.cached_tokens; Chat usage.prompt_tokens_details.cached_tokens; /v1/messages usage.cache_read_input_tokens (cache_creation_input_tokens always 0); xai-sdk cached_prompt_text_tokens usageMetadata.cachedContentTokenCount (+ cacheTokensDetails[] {modality, tokenCount}); promptTokenCount includes cached tokens; countTokens with generateContentRequest.cachedContent reports it too
Rate-limit accounting cached tokens count toward TPM cache_read_input_tokens excluded from ITPM cached tokens count toward TPM and toward the 200k long-context threshold cached tokens count toward the context window and standard rate limits
Invalidation any change before the breakpoint (model, key, tier, tools, text.format, effort, verbosity, compaction) any byte before the breakpoint; tool_choice, images, output_config.format, thinking/effort changes, speed any prefix change; different host (routing key); store:false without replayed encrypted reasoning implicit: prefix change; explicit: different model → error; changing generationConfig does not touch the cache
Where unavailable pro models, some minis top-level automatic caching not on legacy Bedrock disabled under ZDR? (no — encrypted reasoning replay keeps caching); Batch −20 % applies to cached tokens too Interactions API (explicit caching not supported); free tier (limit: 0); /v1/cachedContents (404 — v1beta only)

# A2. Prices

OpenAI Anthropic xAI Gemini
read cached_input (GPT-5.6+/6: 0.1×; o3 0.25×; gpt-4o 0.5×) 0.1× input (0.025× Fable 5.1 / Mythos 5.1) grok-4.6 $0.50 (0.25×), grok-4.5 $0.30 (0.15×), grok-4.3 / 4.20 $0.20 (0.16×), grok-build $0.20 (0.2×); ≥200k prompts: cached price doubles too; priority ×2 and US ×1.1 applied after the cache discount 0.1× input on every priced model (gemini-3.8-flash $0.075, gemini-3.5-flash $0.15, gemini-3.5-flash-lite $0.03, gemini-3.1-pro-preview $0.20 / $0.40 >200k, gemini-2.5-pro $0.125 / $0.25)
write free before GPT-5.6; 1.25× on GPT-5.6+ (cache_write: sol $5, terra $2.5, luna $0.25, astra $12.5) 5 m 1.25×, 1 h 2× (Opus 5 $6.25 / $10; Sonnet 5 $2.5 / $4; Haiku 4.5 $1.25 / $2) none implicit: none; explicit: storage per 1M tokens per hour — $0.50 (3.6–3.8 Flash intro; $1.00 from 2027), $1.00 (3.5 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite, 3 Flash, 2.5 Flash), $1.80 (Flash priority), $4.50 (3.1 Pro, 2.5 Pro), $8.10 (Pro priority)
batch own cached_input (half of standard) 50 % on all cache dimensions −20 % on cached tokens (grok-4.3 / 4.20 only) batch/flex cached_input halves too (3.8 Flash $0.0375; 3.5 Flash flex $0.08 — data inconsistency)
worked example pricing.md §7 idem; live Haiku: 5,809-token write $0.0073, read $0.00058 idem (long-context rates for a 900k prefix); live: 920-token prompt twice → 192 → 896 cached tokens idem; explicit storage of 0.9M tokens for one hour on gemini-3.8-flash = $0.45

# A3. The same cached call on all four providers

json
// OpenAI (GPT-5.6+ explicit mode; on older models omit prompt_cache_options — caching is implicit)
{"model": "gpt-5.6-terra", "prompt_cache_key": "tenant-42",
 "prompt_cache_options": {"mode": "explicit", "ttl": "30m"},
 "input": [{"role": "developer", "content": [{"type": "input_text", "text": "<~2k tokens of stable policy>", "prompt_cache_breakpoint": {"mode": "explicit"}}]},
           {"role": "user", "content": "Question of the day"}],
 "max_output_tokens": 200}
json
// Anthropic (explicit breakpoint on the system block, 1-hour TTL)
{"model": "claude-sonnet-5", "max_tokens": 200,
 "system": [{"type": "text", "text": "<~2k tokens of stable policy>", "cache_control": {"type": "ephemeral", "ttl": "1h"}}],
 "messages": [{"role": "user", "content": "Question of the day"}]}
json
// xAI (automatic; the routing key keeps every request of this tenant on the same server)
{"model": "grok-4.3", "prompt_cache_key": "tenant-42",
 "instructions": "<~2k tokens of stable policy>",
 "input": [{"role": "user", "content": "Question of the day"}],
 "reasoning": {"effort": "low"}, "max_output_tokens": 200}
json
// Gemini — explicit: POST /v1beta/cachedContents (once) …
{"model": "models/gemini-3.8-flash", "displayName": "tenant-42-policy", "ttl": "3600s",
 "systemInstruction": {"parts": [{"text": "<~4k tokens of stable policy>"}]}}
// … then POST /v1beta/models/gemini-3.8-flash:generateContent (do NOT repeat the cached systemInstruction)
{"cachedContent": "cachedContents/abc123",
 "contents": [{"role": "user", "parts": [{"text": "Question of the day"}]}],
 "generationConfig": {"maxOutputTokens": 200}}
// implicit alternative: send the same ≥4,096-token prefix first in `contents` on every call and watch usageMetadata.cachedContentTokenCount

Anthropic shortcut equivalent to the implicit modes: top-level "cache_control": {"type": "ephemeral"}.

# A4. Portability notes

  • Ordering discipline is identical everywhere: stable content first (instructions, tools, schema), volatile content last.
  • Keys: OpenAI and xAI key the cache per prompt_cache_key (xAI: sticky routing); Anthropic per exact bytes; Gemini explicit caches per resource name and implicit per prefix.
  • Token accounting: Anthropic's input_tokens excludes cached tokens; OpenAI's, xAI's and Gemini's promptTokenCount include them (cached_tokens / cachedContentTokenCount is a subset).
  • Reasoning and caching interact on every provider: OpenAI/xAI replay encrypted_content, Anthropic keeps thinking blocks, Gemini keeps thoughtSignature — dropping them breaks the prefix (xAI: "top cause of cache misses").
  • Compaction resets the prefix on OpenAI/Anthropic/xAI (compaction item/block); keep a breakpoint or a cachedContents resource on the system prompt so only the summary is rewritten.
  • Cost of a miss: only Anthropic (1.25×/2×) and GPT-5.6+ (1.25×) charge writes; Gemini charges storage time for explicit caches; xAI never charges for filling the cache.

# Part B — Reasoning / thinking

# B1. Controls

Control OpenAI (POST /v1/responses) Anthropic (POST /v1/messages) xAI (POST /v1/responses; Chat) Gemini (generateContent; Interactions)
on/off reasoning.effort: "none" (GPT-5.1–5.6 minis/flagships; rejected by gpt-6-astra); non-reasoning models reject reasoning (400) thinking: {type: "disabled"} (Opus 5 except xhigh/max, Sonnet 5, 4.x); impossible on Fable / Mythos always on for grok-4.6 / 4.5 / 4.20-reasoning / build; reasoning_effort: "none" works on grok-4.3 only (LIVE_DISCOVERED → 0 reasoning tokens); grok-4.20-0309-non-reasoning has none; reasoning_effort on 4.20-reasoning / build → 400 thinkingConfig.thinkingBudget: 0 on 3.6 / 3.5 Flash (and 2.5 Flash / Flash-Lite); cannot disable on 3.8 / 3.7 Flash, 3.1 Pro, 2.5 Pro; Flash-Lite budget: 0 → 400 (use minimal)
depth reasoning.effort: none | minimal | low | medium | high | xhigh | max (per model) output_config.effort: low | medium | high (default) | xhigh | max (400 on Sonnet 4.5 / Haiku 4.5) reasoning.effort / reasoning_effort: low | medium | high | xhigh (grok-4.6 default high; grok-4.3 default low; grok-4.5 xhigh = high); on grok-4.20-multi-agent-0309 it selects 4 or 16 agents thinkingConfig.thinkingLevel: MINIMAL | LOW | MEDIUM | HIGH (case-insensitive; Gemini 3 only) — defaults high (3.1 Pro, 3 Flash), medium (3.5–3.8 Flash), minimal (Flash-Lite, 3.1 image models); MINIMAL → 400 on 3.8 / 3.7 Flash and 3.1 Pro; Interactions generation_config.thinking_level
explicit token budget none (max_output_tokens includes reasoning) thinking: {type: "enabled", budget_tokens ≥ 1024} — 4.5 models only (400 on 4.7+) none (max_output_tokens documented to include reasoning, not enforced live; Chat max_completion_tokens caps visible output only) thinkingConfig.thinkingBudget int32 [-1, 65535] (2.5-era; -1 dynamic; LEGACY on 3.x, accepted except 3.1 Pro; exclusive with thinkingLevel → 400); 2.5 Pro 128–32,768, 2.5 Flash 0–24,576, 2.5 Flash-Lite 512–24,576
adaptive implicit (effort is the only knob) thinking: {type: "adaptive"} (4.6+; 400 on 4.5) implicit (model decides depth within the effort) implicit at each level; thinkingBudget: -1 = dynamic on 2.5
visibility reasoning.summary: auto | concise | detailed; raw reasoning never returned thinking.display: summarized | omitted | updates Responses reasoning.summary (accepted 'for compatibility') → summary[]; Chat returns full reasoning_content; /v1/messages thinking blocks (empty signature) thinkingConfig.includeThoughts: true → summary parts {text, thought: true} (not guaranteed); Interactions thinking_summaries: auto | none → thought.summary[]
replay context reasoning.context: auto | current_turn | all_turns model-specific preservation; clear_thinking_20251015; thinking.block_binding (beta) replay reasoning items (encrypted when store:false) or previous_response_id echo thoughtSignature on the first functionCall of each step (mandatory, even at minimal) and on text parts (recommended); "thought preservation" across turns since 3.5 Flash
extra compute mode reasoning.mode: standard | pro (GPT-5.6) — (effort: max) grok-4.20-multi-agent-0309 (4/16 agents) Deep Research agents; gemini-3.8-live-extended-thinking
mid-conversation change configuration_update item (gpt-6-astra) messages[].output_config.effort on role: system (beta) per request per request
task-level budget max_tool_calls (hosted tools) output_config.task_budget (beta) max_turns Antigravity agent_config.max_total_tokens
Chat / compat layers reasoning_effort only on Chat Completions n/a Chat reasoning_effort (same values); /v1/messages thinking param ignored OpenAI-compat reasoning_effort → minimal (3.1 Pro: low; 2.5: budget 1,024) / low / medium (8,192) / high (24,576); "none" only on 2.5 non-Pro; extra_body.google.thinking_config

# B2. Response shapes

OpenAI Anthropic xAI Gemini
item/block/part {type: "reasoning", id: "rs_…", summary: [{type: "summary_text", text}], content: [], encrypted_content, status} before the message {type: "thinking", thinking: "<summary>", signature} / {type: "redacted_thinking", data} Responses {type: "reasoning", id: "rs_…", summary: [{type: "summary_text", text}], encrypted_content?} (sometimes omitted); Chat message.reasoning_content (full text); /v1/messages {type: "thinking", thinking, signature: ""} parts[] {text: "<summary>", thought: true} (only with includeThoughts); parts[].thoughtSignature (opaque base64) on function-call and final parts; Interactions {type: "thought", signature, summary[]} steps
tokens usage.output_tokens_details.reasoning_tokens usage.output_tokens_details.thinking_tokens Responses usage.output_tokens_details.reasoning_tokens (output_tokens = visible + reasoning); Chat usage.completion_tokens_details.reasoning_tokens (completion_tokens excludes, total_tokens includes reasoning); cost_in_usd_ticks usageMetadata.thoughtsTokenCount (billed as output; candidatesTokenCount may be absent when only thoughts were produced)
truncation status: incomplete, incomplete_details.reason: max_output_tokens stop_reason: max_tokens incomplete_details.reason: max_output_tokens (live: cap 32 → still completed with 138 reasoning tokens — not enforced on reasoning) finishReason: MAX_TOKENS with possibly empty text (thinking consumed the budget; live maxOutputTokens: 8)
streaming response.reasoning_summary_part.added/done, response.reasoning_summary_text.delta/done, response.reasoning_text.delta/done content_block_start(thinking) → thinking_delta* → signature_delta → content_block_stop Responses: same event names as OpenAI (summary events LIVE_VERIFIED); Chat: delta.reasoning_content before delta.content {text, thought: true} chunks first; final chunk carries an empty-text part with thoughtSignature; Interactions step.start(thought), step.delta {thought_summary | thought_signature}
replay rules send items back verbatim (auto with previous_response_id); encrypted_content when store: false send thinking blocks back unmodified in tool-use turns (400 if edited) send reasoning items back (encrypted with include: ["reasoning.encrypted_content"]); Chat: resend reasoning_content send thoughtSignature back unmodified (400 'Corrupted thought signature' if tampered; dummy skip_thought_signature_validator bypasses — verified); interleaving FC/FR pairs → 400
incompatibilities temperature/top_p rejected by reasoning models manual thinking ⟂ forced tool_choice, prefill, temperature≠1, top_k; Fable 5.1 / Mythos 5.1 reject forced tool choice stop, presence_penalty, frequency_penalty → 400 on reasoning models; logit_bias 400; sampling params accepted thinkingLevel + thinkingBudget → 400; MINIMAL on 3.8/3.7 Flash, 3.1 Pro → 400; Live 3.8: omit thinkingConfig; candidateCount > 1 → 400

# B3. Per-model reasoning modes

OpenAI model effort values (default) Anthropic model thinking (default) · effort xAI model reasoning · effort Gemini model thinkingLevel (default) · budget
gpt-6-astra low, medium, high, xhigh, max (no none) claude-fable-5-1 / claude-mythos-5-1 adaptive always on · low–max · per-message effort (beta) grok-4.6 always on · low, medium, high, xhigh gemini-3.1-pro-preview low, medium, high (no minimal) · budget not accepted
gpt-5.6-sol/terra/luna none, low, medium, high, xhigh, max · mode: pro claude-fable-5 / claude-mythos-5 adaptive always on · low–max grok-4.5 always on · low, medium, high (xhigh = high) gemini-3.8-flash, gemini-3.7-flash low, medium, high (minimal → 400) · cannot disable
gpt-5.5 none, low, medium, high, xhigh claude-opus-5 adaptive on (disable except xhigh/max) · low–max grok-4.3 optional · low, medium, high, xhigh, none (LIVE_DISCOVERED) gemini-3.6-flash, gemini-3.5-flash minimal, low, medium, high · thinkingBudget: 0 disables
gpt-5.5-pro / gpt-5.4-pro medium, high, xhigh claude-sonnet-5 adaptive on · low–max grok-4.20-0309-reasoning always on · reasoning_effort → 400 gemini-3.5-flash-lite, gemini-3.1-flash-lite minimal, low, medium, high · budget 0 → 400
gpt-5.4, -mini, -nano none, low, medium, high, xhigh claude-opus-4-8 / -4-7 adaptive off by default · low–max grok-4.20-0309-non-reasoning no reasoning (0 tokens; reasoning_effort → 400) gemini-3-flash-preview minimal, low, medium, high
gpt-5.3-codex low, medium, high, xhigh claude-opus-4-6 / claude-sonnet-4-6 adaptive · low, medium, high, max grok-4.20-multi-agent-0309 multi-agent · effort = 4 (low/medium) or 16 (high/xhigh) agents, default medium gemini-3.1-flash-image, -lite-image MINIMAL / HIGH only (always on, billed)
gpt-5.2 / gpt-5.1 none, low, medium, high(, xhigh) claude-opus-4-5-20251101 manual enabled {budget_tokens} · low, medium, high grok-build-0.1 always on · reasoning_effort → 400 gemini-2.5-pro thinkingBudget 128–32,768 (cannot disable)
gpt-5 / -mini / -nano minimal, low, medium, high claude-sonnet-4-5-20250929, claude-haiku-4-5-20251001 manual enabled only · no effort (400) — — gemini-2.5-flash / -flash-lite budget 0–24,576 / 512–24,576 (Lite off by default)
o3 / o3-pro / o4-mini model default (reasoning_effort on Chat) — — — — gemini-3.8-live / -extended-thinking interleaved (omit config) / LOW–HIGH + includeThoughts
gpt-4.1*, gpt-4o*, chat-latest non-reasoning (reject reasoning) none current — — — gemma-4-* live thinking: true (thoughtsTokenCount: 5), levels unknown

# B4. The same "think hard, show a summary" request

json
// OpenAI
{"model": "gpt-5.6-terra", "input": "How many r's in strawberry? Explain briefly.",
 "reasoning": {"effort": "high", "summary": "auto"}, "max_output_tokens": 600}
json
// Anthropic (adaptive model)
{"model": "claude-sonnet-5", "max_tokens": 600,
 "thinking": {"type": "adaptive", "display": "summarized"}, "output_config": {"effort": "high"},
 "messages": [{"role": "user", "content": "How many r's in strawberry? Explain briefly."}]}
json
// xAI — Responses (summary) …
{"model": "grok-4.6", "input": "How many r's in strawberry? Explain briefly.",
 "reasoning": {"effort": "high", "summary": "auto"}, "max_output_tokens": 600}
// … or Chat Completions (full reasoning_content text)
{"model": "grok-4.6", "reasoning_effort": "high",
 "messages": [{"role": "user", "content": "How many r's in strawberry? Explain briefly."}]}
json
// Gemini
{"contents": [{"role": "user", "parts": [{"text": "How many r's in strawberry? Explain briefly."}]}],
 "generationConfig": {"thinkingConfig": {"thinkingLevel": "high", "includeThoughts": true}, "maxOutputTokens": 600}}

Read the summary from output[].type == "reasoning" → summary[].text (OpenAI, xAI Responses), choices[0].message.reasoning_content (xAI Chat), content[].type == "thinking" → thinking (Anthropic), or parts[] with thought: true (Gemini); billed tokens from reasoning_tokens / thinking_tokens / reasoning_tokens / thoughtsTokenCount.

# B5. Caching × reasoning interactions

OpenAI Anthropic xAI Gemini
changing effort invalidates the cache (reasoning_effort_changed); configuration_update on gpt-6-astra model-specific; per-message effort (beta) preserves the prefix undocumented (automatic cache) generationConfig is not part of an explicit cachedContents; implicit prefix unaffected
reasoning items in the prefix replayed items are cacheable input prior thinking blocks are billed input and cache-friendly on models that keep all turns replayed reasoning/reasoning_content is required for cache hits and is billed as input echoed thoughtSignature counts as input; signatures are part of the replayed history
minimum-prefix effects effort/verbosity affect the hidden prefix on pre-5.6 models thinking config participates in the cache key hidden system prefix (~192 tokens) is cached across users implicit minimum 4,096 tokens on 3.x Flash / 3.1 Pro

Related: pricing · state-management · features · docs/openai/prompt-caching.md · docs/anthropic/prompt-caching.md · docs/xai/prompt-caching.md · docs/gemini/context-caching.md · docs/openai/reasoning.md · docs/anthropic/thinking.md · docs/xai/reasoning.md · docs/gemini/thinking.md.