SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
6.6 KB

# OpenAI prompt caching (Responses + Chat Completions)

Status: DOCUMENTED + LIVE_VERIFIED (Chat Completions hit observed: 1408/1556 cached tokens; Responses on gpt-5.4-nano showed no hit within 1.5 s — recorded as-is). GPT-5.6+ explicit breakpoints / prompt_cache_options / diagnostics are DOCUMENTED only (live: gpt-5.4-nano rejects prompt_cache_options). Machine-readable: prompt_cache_* rows in generated/fragments/parameters/openai-responses.json and openai-chat-completions.json; usage/diagnostics objects in generated/fragments/objects/openai-responses-objects.json.

Sources

Last verified: 2026-09-18

# 1. Mechanics

  • Caches the KV state of the rendered prefix (hidden system content, instructions/developer messages, tool definitions, text.format schema, conversation history incl. images/files/audio). Reuse needs an exact prefix match up to an eligible breakpoint, same model, same service tier, same prompt_cache_key.
  • Minimum cacheable prefix: 1,024 visible input tokens on GPT-5.6+; model-/settings-dependent before (tools, images, schemas, reasoning effort, verbosity affect it).
  • Cached tokens pay the cached-input rate (e.g. gpt-5.4-nano $0.02 vs $0.20 per 1M; gpt-4.1-nano $0.025 vs $0.10); cache writes cost extra (1.25×) only on GPT-5.6+. Cached tokens still count toward TPM limits. Caching never changes outputs.
Behaviour GPT-5.6 and later GPT-5.5 / 5.5-pro Earlier (GPT-5.4, 4.1, o-series…)
Implicit breakpoints end of latest eligible message (user msg, last of a tool-output group, last initial developer msg) every 2,048 tokens model-dependent intervals
Explicit breakpoints prompt_cache_breakpoint:{mode:"explicit"} on content parts (≤4 writes/request; not on top-level instructions, not on additional_tools) no no
Mode prompt_cache_options.mode: implicit (default) | explicit — —
Lifetime prompt_cache_options.ttl:"30m" (only value; ≥30 min after last use) prompt_cache_retention:"24h" only prompt_cache_retention: in_memory (5–10 min idle, ≤1 h) | 24h (~30 min, ≤24 h) on gpt-5.4, gpt-5.2, gpt-5.1*, gpt-5, gpt-4.1…
Reporting exact boundary rounded down to ×128, hidden tokens excluded rounded down to ×128
prompt_cache_key optional (separate accounting per customer) routing hint: keep ≈15 req/min per key routing hint
Diagnostics prompt_cache_options.comparison_response_id → prompt_cache_diagnostics no no

Default retention: 24h for non-ZDR orgs, in_memory for ZDR orgs (live echo on gpt-5.4-nano: "prompt_cache_retention":"24h"; "in_memory" when requested).

# 2. Parameters

Parameter Endpoints Type / enum Notes
prompt_cache_key responses, chat, compact string live: accepted and echoed on Responses
prompt_cache_retention responses, chat, compact in_memory | 24h deprecated (still works); use prompt_cache_options.ttl on GPT-5.6+
prompt_cache_options.ttl responses, chat, compact "30m" GPT-5.6+
prompt_cache_options.mode responses, chat, compact implicit | explicit GPT-5.6+
prompt_cache_options.prewarm responses boolean Responses-only (spec ResponsePromptCacheOptionsParam)
prompt_cache_options.comparison_response_id responses string diagnostics baseline; does not load the conversation
…content[].prompt_cache_breakpoint.mode responses (input_text/input_image/input_file, tool outputs), chat (text parts, prediction.content parts) "explicit" GPT-5.6+
user legacy string replaced by safety_identifier + prompt_cache_key

Live errors: prompt_cache_options on gpt-5.4-nano → 400 invalid_parameter "prompt_cache_options is not supported on this model".

# 3. Reading usage

  • Responses: usage.input_tokens_details.cached_tokens, usage.input_tokens_details.cache_write_tokens.
  • Chat: usage.prompt_tokens_details.cached_tokens (no write counter).
  • Cost = (input − cached − written)×p_in + cached×p_cached + written×1.25×p_in (writes only GPT-5.6+).
  • Diagnostics (GPT-5.6+, Responses): prompt_cache_diagnostics.type ∈ cache_hit | cache_miss{reason, comparison_reusable_tokens, cache_missed_tokens} | comparison_response_not_found | unavailable; reason ∈ model_changed, prompt_cache_key_changed, service_tier_changed, tools_changed, text_format_changed, reasoning_effort_changed, verbosity_changed, context_compacted, input_changed. Read it from response.completed.response when streaming. Free; ZDR-compatible.

# 4. Live experiment (2026-09-18)

Prefix: ~1.5k tokens of repeated text, sent twice 1.5 s apart with a stable prompt_cache_key.

Surface Model 1st call 2nd call
Chat system message gpt-4.1-nano prompt_tokens 1556, cached_tokens 0 prompt_tokens 1556, cached_tokens 1408 (=11×128, rounded down as documented; 90 % of prompt)
Responses instructions (store:false, effort:low) gpt-5.4-nano input_tokens 1555, cached 0, cache_write 0 input_tokens 1555, cached_tokens 0, cache_write_tokens 0

Interpretation: the Chat path behaved exactly as documented; the Responses pair on gpt-5.4-nano did not hit (possible causes: routing to another machine, instructions-based prefix, short interval, or hidden-token minimum). Not a documentation error — recorded as an observation (FAILED_VERIFICATION would be wrong: the calls succeeded; the hit simply did not occur). cache_write_tokens is 0 on pre-5.6 models (no write charge).

# 5. Practices (guide digest)

Stable content first (instructions, reference material, tool definitions), dynamic content (timestamps, user data) last or in later messages; append, don't rewrite history; disable tools with tool_choice:"none"/allowed_tools instead of removing definitions; change effort via configuration_update (gpt-6-astra) instead of reasoning.effort; compaction resets the prefix (context_compacted); expand a just-below-minimum prefix with useful stable material when reuse is high (break-even L = M·(r + (w−r)/N)); monitor cached_tokens/cache_write_tokens per user/day and the Prompt Caching dashboard.