# Prompt caching and reasoning — parameter-level mapping across four providers
Status: parameters from generated/parameters.json (prompt_cache_*, reasoning.*, cache_control*, thinking.*, output_config.*, xAI prompt_cache_key / reasoning_effort / include, Gemini cachedContent, generationConfig.thinkingConfig.*), per-model facts from generated/models.json and generated/compatibility/{anthropic,gemini}-feature-model-matrix.json, prices from generated/pricing.json; live observations quoted from docs/openai/prompt-caching.md, docs/openai/reasoning.md, docs/anthropic/prompt-caching.md, docs/anthropic/thinking.md, docs/anthropic/effort.md, docs/xai/prompt-caching.md, docs/xai/reasoning.md, docs/gemini/context-caching.md, docs/gemini/thinking.md (2026-09-18/19). Gemini explicit caching is ACCOUNT_RESTRICTED on this free-tier key (429 … limit=0); size validation and CRUD shapes were LIVE_VERIFIED.
Sources: https://developers.openai.com/api/docs/guides/prompt-caching · https://developers.openai.com/api/docs/guides/reasoning · https://platform.claude.com/docs/en/build-with-claude/prompt-caching · https://platform.claude.com/docs/en/build-with-claude/extended-thinking · https://platform.claude.com/docs/en/build-with-claude/effort · https://docs.x.ai/developers/model-capabilities/text/prompt-caching · https://docs.x.ai/developers/model-capabilities/text/reasoning · https://ai.google.dev/gemini-api/docs/caching · https://ai.google.dev/gemini-api/docs/thinking
Last verified: 2026-09-18
# Part A — Prompt / context caching
# A1. Mechanics
| Topic |
OpenAI |
Anthropic |
xAI |
Gemini |
| Activation |
automatic on every request ≥ minimum (implicit); GPT-5.6+ optional explicit mode |
explicit: cache_control: {type: ephemeral} on ≤4 blocks or top-level automatic cache_control |
automatic prefix cache, no markers; cache_control on /v1/messages is ignored |
implicit (automatic, 2.5+) and explicit cachedContents resources referenced by cachedContent |
| What is cached |
rendered prefix (hidden system content, instructions, tools, text.format, history); exact prefix match |
prefix tools → system → messages up to the marked block; byte-identical; 20-block lookback |
prompt prefix incl. the hidden system prefix (a fresh minimal call already shows ~192 cached tokens), instructions, tools, history incl. prior reasoning_content / encrypted reasoning |
implicit: the leading prefix of contents (put shared content first); explicit: contents, systemInstruction, tools, toolConfig stored under a cachedContents/{id} name for one model |
| Minimum |
1,024 tokens on GPT-5.6+ (model-dependent before); hits in 128-token increments |
per model: 512 (Fable/Mythos/Opus 5) · 1,024 (Sonnet 5/4.6/4.5, Opus 4.8) · 2,048 (Opus 4.7) · 4,096 (Haiku 4.5, Opus 4.6/4.5) |
not documented; live hits at 192 and 896 tokens |
implicit 4,096 (3.8/3.7/3.6/3.5 Flash, 3.1 Pro), 2,048 (2.5 Flash/Pro); explicit 1,024 live (min_total_token_count=1024; docs quote 4,096) |
| Lifetime |
in-memory 5–10 min (≤1 h) or 24h retention; GPT-5.6+ prompt_cache_options.ttl: "30m" |
ttl: "5m" (default) or "1h"; hits refresh for free |
not documented; caches do not carry across hosts |
implicit: undocumented; explicit: ttl (default 1 h) or expireTime, mutable via PATCH; no min/max documented |
| Routing |
prompt_cache_key (stable per prefix) |
none needed |
prompt_cache_key (Responses, Chat — "plumbed to x-grok-conv-id") or header x-grok-conv-id (Chat/gRPC): sticky routing to the server holding the cache; top cause of misses = omitting prior reasoning or not using previous_response_id |
none needed (cachedContent names the cache explicitly) |
| Explicit breakpoints |
GPT-5.6+: prompt_cache_breakpoint: {mode: "explicit"} on parts (≤4) with prompt_cache_options.mode: "explicit" |
≤4 cache_control blocks in system/tools/messages |
none |
one explicit cache per request (cachedContent); the cache must use the same model; do not repeat the cached systemInstruction/tools in the request |
| Pre-warm |
prompt_cache_options.prewarm: true |
max_tokens: 0 request with a breakpoint |
none documented |
POST /v1beta/cachedContents creates (and bills storage for) the prefix before first use |
| Diagnostics |
prompt_cache_options.comparison_response_id → prompt_cache_diagnostics {cache_hit | cache_miss {reason}} (5.6+) |
beta cache-diagnosis-2026-04-07: diagnostics.previous_message_id → miss reasons |
none |
none |
| Usage fields |
usage.input_tokens_details.cached_tokens, .cache_write_tokens; Chat prompt_tokens_details.cached_tokens |
usage.cache_read_input_tokens, cache_creation_input_tokens, cache_creation.ephemeral_{5m,1h}_input_tokens; input_tokens excludes cached |
Responses usage.input_tokens_details.cached_tokens; Chat usage.prompt_tokens_details.cached_tokens; /v1/messages usage.cache_read_input_tokens (cache_creation_input_tokens always 0); xai-sdk cached_prompt_text_tokens |
usageMetadata.cachedContentTokenCount (+ cacheTokensDetails[] {modality, tokenCount}); promptTokenCount includes cached tokens; countTokens with generateContentRequest.cachedContent reports it too |
| Rate-limit accounting |
cached tokens count toward TPM |
cache_read_input_tokens excluded from ITPM |
cached tokens count toward TPM and toward the 200k long-context threshold |
cached tokens count toward the context window and standard rate limits |
| Invalidation |
any change before the breakpoint (model, key, tier, tools, text.format, effort, verbosity, compaction) |
any byte before the breakpoint; tool_choice, images, output_config.format, thinking/effort changes, speed |
any prefix change; different host (routing key); store:false without replayed encrypted reasoning |
implicit: prefix change; explicit: different model → error; changing generationConfig does not touch the cache |
| Where unavailable |
pro models, some minis |
top-level automatic caching not on legacy Bedrock |
disabled under ZDR? (no — encrypted reasoning replay keeps caching); Batch −20 % applies to cached tokens too |
Interactions API (explicit caching not supported); free tier (limit: 0); /v1/cachedContents (404 — v1beta only) |
# A2. Prices
|
OpenAI |
Anthropic |
xAI |
Gemini |
| read |
cached_input (GPT-5.6+/6: 0.1×; o3 0.25×; gpt-4o 0.5×) |
0.1× input (0.025× Fable 5.1 / Mythos 5.1) |
grok-4.6 $0.50 (0.25×), grok-4.5 $0.30 (0.15×), grok-4.3 / 4.20 $0.20 (0.16×), grok-build $0.20 (0.2×); ≥200k prompts: cached price doubles too; priority ×2 and US ×1.1 applied after the cache discount |
0.1× input on every priced model (gemini-3.8-flash $0.075, gemini-3.5-flash $0.15, gemini-3.5-flash-lite $0.03, gemini-3.1-pro-preview $0.20 / $0.40 >200k, gemini-2.5-pro $0.125 / $0.25) |
| write |
free before GPT-5.6; 1.25× on GPT-5.6+ (cache_write: sol $5, terra $2.5, luna $0.25, astra $12.5) |
5 m 1.25×, 1 h 2× (Opus 5 $6.25 / $10; Sonnet 5 $2.5 / $4; Haiku 4.5 $1.25 / $2) |
none |
implicit: none; explicit: storage per 1M tokens per hour — $0.50 (3.6–3.8 Flash intro; $1.00 from 2027), $1.00 (3.5 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite, 3 Flash, 2.5 Flash), $1.80 (Flash priority), $4.50 (3.1 Pro, 2.5 Pro), $8.10 (Pro priority) |
| batch |
own cached_input (half of standard) |
50 % on all cache dimensions |
−20 % on cached tokens (grok-4.3 / 4.20 only) |
batch/flex cached_input halves too (3.8 Flash $0.0375; 3.5 Flash flex $0.08 — data inconsistency) |
| worked example |
pricing.md §7 |
idem; live Haiku: 5,809-token write $0.0073, read $0.00058 |
idem (long-context rates for a 900k prefix); live: 920-token prompt twice → 192 → 896 cached tokens |
idem; explicit storage of 0.9M tokens for one hour on gemini-3.8-flash = $0.45 |
# A3. The same cached call on all four providers
// OpenAI (GPT-5.6+ explicit mode; on older models omit prompt_cache_options — caching is implicit)
{"model": "gpt-5.6-terra", "prompt_cache_key": "tenant-42",
"prompt_cache_options": {"mode": "explicit", "ttl": "30m"},
"input": [{"role": "developer", "content": [{"type": "input_text", "text": "<~2k tokens of stable policy>", "prompt_cache_breakpoint": {"mode": "explicit"}}]},
{"role": "user", "content": "Question of the day"}],
"max_output_tokens": 200}
// Anthropic (explicit breakpoint on the system block, 1-hour TTL)
{"model": "claude-sonnet-5", "max_tokens": 200,
"system": [{"type": "text", "text": "<~2k tokens of stable policy>", "cache_control": {"type": "ephemeral", "ttl": "1h"}}],
"messages": [{"role": "user", "content": "Question of the day"}]}
// xAI (automatic; the routing key keeps every request of this tenant on the same server)
{"model": "grok-4.3", "prompt_cache_key": "tenant-42",
"instructions": "<~2k tokens of stable policy>",
"input": [{"role": "user", "content": "Question of the day"}],
"reasoning": {"effort": "low"}, "max_output_tokens": 200}
// Gemini — explicit: POST /v1beta/cachedContents (once) …
{"model": "models/gemini-3.8-flash", "displayName": "tenant-42-policy", "ttl": "3600s",
"systemInstruction": {"parts": [{"text": "<~4k tokens of stable policy>"}]}}
// … then POST /v1beta/models/gemini-3.8-flash:generateContent (do NOT repeat the cached systemInstruction)
{"cachedContent": "cachedContents/abc123",
"contents": [{"role": "user", "parts": [{"text": "Question of the day"}]}],
"generationConfig": {"maxOutputTokens": 200}}
// implicit alternative: send the same ≥4,096-token prefix first in `contents` on every call and watch usageMetadata.cachedContentTokenCount
Anthropic shortcut equivalent to the implicit modes: top-level "cache_control": {"type": "ephemeral"}.
# A4. Portability notes
- Ordering discipline is identical everywhere: stable content first (instructions, tools, schema), volatile content last.
- Keys: OpenAI and xAI key the cache per
prompt_cache_key (xAI: sticky routing); Anthropic per exact bytes; Gemini explicit caches per resource name and implicit per prefix.
- Token accounting: Anthropic's
input_tokens excludes cached tokens; OpenAI's, xAI's and Gemini's promptTokenCount include them (cached_tokens / cachedContentTokenCount is a subset).
- Reasoning and caching interact on every provider: OpenAI/xAI replay
encrypted_content, Anthropic keeps thinking blocks, Gemini keeps thoughtSignature — dropping them breaks the prefix (xAI: "top cause of cache misses").
- Compaction resets the prefix on OpenAI/Anthropic/xAI (
compaction item/block); keep a breakpoint or a cachedContents resource on the system prompt so only the summary is rewritten.
- Cost of a miss: only Anthropic (1.25×/2×) and GPT-5.6+ (1.25×) charge writes; Gemini charges storage time for explicit caches; xAI never charges for filling the cache.
# Part B — Reasoning / thinking
# B1. Controls
| Control |
OpenAI (POST /v1/responses) |
Anthropic (POST /v1/messages) |
xAI (POST /v1/responses; Chat) |
Gemini (generateContent; Interactions) |
| on/off |
reasoning.effort: "none" (GPT-5.1–5.6 minis/flagships; rejected by gpt-6-astra); non-reasoning models reject reasoning (400) |
thinking: {type: "disabled"} (Opus 5 except xhigh/max, Sonnet 5, 4.x); impossible on Fable / Mythos |
always on for grok-4.6 / 4.5 / 4.20-reasoning / build; reasoning_effort: "none" works on grok-4.3 only (LIVE_DISCOVERED → 0 reasoning tokens); grok-4.20-0309-non-reasoning has none; reasoning_effort on 4.20-reasoning / build → 400 |
thinkingConfig.thinkingBudget: 0 on 3.6 / 3.5 Flash (and 2.5 Flash / Flash-Lite); cannot disable on 3.8 / 3.7 Flash, 3.1 Pro, 2.5 Pro; Flash-Lite budget: 0 → 400 (use minimal) |
| depth |
reasoning.effort: none | minimal | low | medium | high | xhigh | max (per model) |
output_config.effort: low | medium | high (default) | xhigh | max (400 on Sonnet 4.5 / Haiku 4.5) |
reasoning.effort / reasoning_effort: low | medium | high | xhigh (grok-4.6 default high; grok-4.3 default low; grok-4.5 xhigh = high); on grok-4.20-multi-agent-0309 it selects 4 or 16 agents |
thinkingConfig.thinkingLevel: MINIMAL | LOW | MEDIUM | HIGH (case-insensitive; Gemini 3 only) — defaults high (3.1 Pro, 3 Flash), medium (3.5–3.8 Flash), minimal (Flash-Lite, 3.1 image models); MINIMAL → 400 on 3.8 / 3.7 Flash and 3.1 Pro; Interactions generation_config.thinking_level |
| explicit token budget |
none (max_output_tokens includes reasoning) |
thinking: {type: "enabled", budget_tokens ≥ 1024} — 4.5 models only (400 on 4.7+) |
none (max_output_tokens documented to include reasoning, not enforced live; Chat max_completion_tokens caps visible output only) |
thinkingConfig.thinkingBudget int32 [-1, 65535] (2.5-era; -1 dynamic; LEGACY on 3.x, accepted except 3.1 Pro; exclusive with thinkingLevel → 400); 2.5 Pro 128–32,768, 2.5 Flash 0–24,576, 2.5 Flash-Lite 512–24,576 |
| adaptive |
implicit (effort is the only knob) |
thinking: {type: "adaptive"} (4.6+; 400 on 4.5) |
implicit (model decides depth within the effort) |
implicit at each level; thinkingBudget: -1 = dynamic on 2.5 |
| visibility |
reasoning.summary: auto | concise | detailed; raw reasoning never returned |
thinking.display: summarized | omitted | updates |
Responses reasoning.summary (accepted 'for compatibility') → summary[]; Chat returns full reasoning_content; /v1/messages thinking blocks (empty signature) |
thinkingConfig.includeThoughts: true → summary parts {text, thought: true} (not guaranteed); Interactions thinking_summaries: auto | none → thought.summary[] |
| replay context |
reasoning.context: auto | current_turn | all_turns |
model-specific preservation; clear_thinking_20251015; thinking.block_binding (beta) |
replay reasoning items (encrypted when store:false) or previous_response_id |
echo thoughtSignature on the first functionCall of each step (mandatory, even at minimal) and on text parts (recommended); "thought preservation" across turns since 3.5 Flash |
| extra compute mode |
reasoning.mode: standard | pro (GPT-5.6) |
— (effort: max) |
grok-4.20-multi-agent-0309 (4/16 agents) |
Deep Research agents; gemini-3.8-live-extended-thinking |
| mid-conversation change |
configuration_update item (gpt-6-astra) |
messages[].output_config.effort on role: system (beta) |
per request |
per request |
| task-level budget |
max_tool_calls (hosted tools) |
output_config.task_budget (beta) |
max_turns |
Antigravity agent_config.max_total_tokens |
| Chat / compat layers |
reasoning_effort only on Chat Completions |
n/a |
Chat reasoning_effort (same values); /v1/messages thinking param ignored |
OpenAI-compat reasoning_effort → minimal (3.1 Pro: low; 2.5: budget 1,024) / low / medium (8,192) / high (24,576); "none" only on 2.5 non-Pro; extra_body.google.thinking_config |
# B2. Response shapes
|
OpenAI |
Anthropic |
xAI |
Gemini |
| item/block/part |
{type: "reasoning", id: "rs_…", summary: [{type: "summary_text", text}], content: [], encrypted_content, status} before the message |
{type: "thinking", thinking: "<summary>", signature} / {type: "redacted_thinking", data} |
Responses {type: "reasoning", id: "rs_…", summary: [{type: "summary_text", text}], encrypted_content?} (sometimes omitted); Chat message.reasoning_content (full text); /v1/messages {type: "thinking", thinking, signature: ""} |
parts[] {text: "<summary>", thought: true} (only with includeThoughts); parts[].thoughtSignature (opaque base64) on function-call and final parts; Interactions {type: "thought", signature, summary[]} steps |
| tokens |
usage.output_tokens_details.reasoning_tokens |
usage.output_tokens_details.thinking_tokens |
Responses usage.output_tokens_details.reasoning_tokens (output_tokens = visible + reasoning); Chat usage.completion_tokens_details.reasoning_tokens (completion_tokens excludes, total_tokens includes reasoning); cost_in_usd_ticks |
usageMetadata.thoughtsTokenCount (billed as output; candidatesTokenCount may be absent when only thoughts were produced) |
| truncation |
status: incomplete, incomplete_details.reason: max_output_tokens |
stop_reason: max_tokens |
incomplete_details.reason: max_output_tokens (live: cap 32 → still completed with 138 reasoning tokens — not enforced on reasoning) |
finishReason: MAX_TOKENS with possibly empty text (thinking consumed the budget; live maxOutputTokens: 8) |
| streaming |
response.reasoning_summary_part.added/done, response.reasoning_summary_text.delta/done, response.reasoning_text.delta/done |
content_block_start(thinking) → thinking_delta* → signature_delta → content_block_stop |
Responses: same event names as OpenAI (summary events LIVE_VERIFIED); Chat: delta.reasoning_content before delta.content |
{text, thought: true} chunks first; final chunk carries an empty-text part with thoughtSignature; Interactions step.start(thought), step.delta {thought_summary | thought_signature} |
| replay rules |
send items back verbatim (auto with previous_response_id); encrypted_content when store: false |
send thinking blocks back unmodified in tool-use turns (400 if edited) |
send reasoning items back (encrypted with include: ["reasoning.encrypted_content"]); Chat: resend reasoning_content |
send thoughtSignature back unmodified (400 'Corrupted thought signature' if tampered; dummy skip_thought_signature_validator bypasses — verified); interleaving FC/FR pairs → 400 |
| incompatibilities |
temperature/top_p rejected by reasoning models |
manual thinking ⟂ forced tool_choice, prefill, temperature≠1, top_k; Fable 5.1 / Mythos 5.1 reject forced tool choice |
stop, presence_penalty, frequency_penalty → 400 on reasoning models; logit_bias 400; sampling params accepted |
thinkingLevel + thinkingBudget → 400; MINIMAL on 3.8/3.7 Flash, 3.1 Pro → 400; Live 3.8: omit thinkingConfig; candidateCount > 1 → 400 |
# B3. Per-model reasoning modes
| OpenAI model |
effort values (default) |
Anthropic model |
thinking (default) · effort |
xAI model |
reasoning · effort |
Gemini model |
thinkingLevel (default) · budget |
gpt-6-astra |
low, medium, high, xhigh, max (no none) |
claude-fable-5-1 / claude-mythos-5-1 |
adaptive always on · low–max · per-message effort (beta) |
grok-4.6 |
always on · low, medium, high, xhigh |
gemini-3.1-pro-preview |
low, medium, high (no minimal) · budget not accepted |
gpt-5.6-sol/terra/luna |
none, low, medium, high, xhigh, max · mode: pro |
claude-fable-5 / claude-mythos-5 |
adaptive always on · low–max |
grok-4.5 |
always on · low, medium, high (xhigh = high) |
gemini-3.8-flash, gemini-3.7-flash |
low, medium, high (minimal → 400) · cannot disable |
gpt-5.5 |
none, low, medium, high, xhigh |
claude-opus-5 |
adaptive on (disable except xhigh/max) · low–max |
grok-4.3 |
optional · low, medium, high, xhigh, none (LIVE_DISCOVERED) |
gemini-3.6-flash, gemini-3.5-flash |
minimal, low, medium, high · thinkingBudget: 0 disables |
gpt-5.5-pro / gpt-5.4-pro |
medium, high, xhigh |
claude-sonnet-5 |
adaptive on · low–max |
grok-4.20-0309-reasoning |
always on · reasoning_effort → 400 |
gemini-3.5-flash-lite, gemini-3.1-flash-lite |
minimal, low, medium, high · budget 0 → 400 |
gpt-5.4, -mini, -nano |
none, low, medium, high, xhigh |
claude-opus-4-8 / -4-7 |
adaptive off by default · low–max |
grok-4.20-0309-non-reasoning |
no reasoning (0 tokens; reasoning_effort → 400) |
gemini-3-flash-preview |
minimal, low, medium, high |
gpt-5.3-codex |
low, medium, high, xhigh |
claude-opus-4-6 / claude-sonnet-4-6 |
adaptive · low, medium, high, max |
grok-4.20-multi-agent-0309 |
multi-agent · effort = 4 (low/medium) or 16 (high/xhigh) agents, default medium |
gemini-3.1-flash-image, -lite-image |
MINIMAL / HIGH only (always on, billed) |
gpt-5.2 / gpt-5.1 |
none, low, medium, high(, xhigh) |
claude-opus-4-5-20251101 |
manual enabled {budget_tokens} · low, medium, high |
grok-build-0.1 |
always on · reasoning_effort → 400 |
gemini-2.5-pro |
thinkingBudget 128–32,768 (cannot disable) |
gpt-5 / -mini / -nano |
minimal, low, medium, high |
claude-sonnet-4-5-20250929, claude-haiku-4-5-20251001 |
manual enabled only · no effort (400) |
— |
— |
gemini-2.5-flash / -flash-lite |
budget 0–24,576 / 512–24,576 (Lite off by default) |
o3 / o3-pro / o4-mini |
model default (reasoning_effort on Chat) |
— |
— |
— |
— |
gemini-3.8-live / -extended-thinking |
interleaved (omit config) / LOW–HIGH + includeThoughts |
gpt-4.1*, gpt-4o*, chat-latest |
non-reasoning (reject reasoning) |
none current |
— |
— |
— |
gemma-4-* |
live thinking: true (thoughtsTokenCount: 5), levels unknown |
# B4. The same "think hard, show a summary" request
// OpenAI
{"model": "gpt-5.6-terra", "input": "How many r's in strawberry? Explain briefly.",
"reasoning": {"effort": "high", "summary": "auto"}, "max_output_tokens": 600}
// Anthropic (adaptive model)
{"model": "claude-sonnet-5", "max_tokens": 600,
"thinking": {"type": "adaptive", "display": "summarized"}, "output_config": {"effort": "high"},
"messages": [{"role": "user", "content": "How many r's in strawberry? Explain briefly."}]}
// xAI — Responses (summary) …
{"model": "grok-4.6", "input": "How many r's in strawberry? Explain briefly.",
"reasoning": {"effort": "high", "summary": "auto"}, "max_output_tokens": 600}
// … or Chat Completions (full reasoning_content text)
{"model": "grok-4.6", "reasoning_effort": "high",
"messages": [{"role": "user", "content": "How many r's in strawberry? Explain briefly."}]}
// Gemini
{"contents": [{"role": "user", "parts": [{"text": "How many r's in strawberry? Explain briefly."}]}],
"generationConfig": {"thinkingConfig": {"thinkingLevel": "high", "includeThoughts": true}, "maxOutputTokens": 600}}
Read the summary from output[].type == "reasoning" → summary[].text (OpenAI, xAI Responses), choices[0].message.reasoning_content (xAI Chat), content[].type == "thinking" → thinking (Anthropic), or parts[] with thought: true (Gemini); billed tokens from reasoning_tokens / thinking_tokens / reasoning_tokens / thoughtsTokenCount.
# B5. Caching × reasoning interactions
|
OpenAI |
Anthropic |
xAI |
Gemini |
| changing effort |
invalidates the cache (reasoning_effort_changed); configuration_update on gpt-6-astra |
model-specific; per-message effort (beta) preserves the prefix |
undocumented (automatic cache) |
generationConfig is not part of an explicit cachedContents; implicit prefix unaffected |
| reasoning items in the prefix |
replayed items are cacheable input |
prior thinking blocks are billed input and cache-friendly on models that keep all turns |
replayed reasoning/reasoning_content is required for cache hits and is billed as input |
echoed thoughtSignature counts as input; signatures are part of the replayed history |
| minimum-prefix effects |
effort/verbosity affect the hidden prefix on pre-5.6 models |
thinking config participates in the cache key |
hidden system prefix (~192 tokens) is cached across users |
implicit minimum 4,096 tokens on 3.x Flash / 3.1 Pro |
Related: pricing · state-management · features · docs/openai/prompt-caching.md · docs/anthropic/prompt-caching.md · docs/xai/prompt-caching.md · docs/gemini/context-caching.md · docs/openai/reasoning.md · docs/anthropic/thinking.md · docs/xai/reasoning.md · docs/gemini/thinking.md.