OpenAI prompt caching (Responses + Chat Completions)
Status: DOCUMENTED + LIVE_VERIFIED (Chat Completions hit observed: 1408/1556 cached tokens; Responses on gpt-5.4-nano showed no hit within 1.5 s — recorded as-is). GPT-5.6+ explicit breakpoints / prompt_cache_options / diagnostics are DOCUMENTED only (live: gpt-5.4-nano rejects prompt_cache_options). Machine-readable: prompt_cache_* rows in generated/fragments/parameters/openai-responses.json and openai-chat-completions.json; usage/diagnostics objects in generated/fragments/objects/openai-responses-objects.json.
Sources
- https://developers.openai.com/api/docs/guides/prompt-caching · https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics
- https://developers.openai.com/api/docs/pricing (cached-input and cache-write rates)
- OpenAPI
PromptCacheOptionsParam,PromptCacheBreakpointParam,PromptCacheRetentionEnum,PromptCacheDiagnostics,ResponseUsage,CompletionUsage
Last verified: 2026-09-18
1. Mechanics
- Caches the KV state of the rendered prefix (hidden system content,
instructions/developer messages, tool definitions,text.formatschema, conversation history incl. images/files/audio). Reuse needs an exact prefix match up to an eligible breakpoint, same model, same service tier, sameprompt_cache_key. - Minimum cacheable prefix: 1,024 visible input tokens on GPT-5.6+; model-/settings-dependent before (tools, images, schemas, reasoning effort, verbosity affect it).
- Cached tokens pay the cached-input rate (e.g.
gpt-5.4-nano$0.02 vs $0.20 per 1M;gpt-4.1-nano$0.025 vs $0.10); cache writes cost extra (1.25×) only on GPT-5.6+. Cached tokens still count toward TPM limits. Caching never changes outputs.
| Behaviour | GPT-5.6 and later | GPT-5.5 / 5.5-pro | Earlier (GPT-5.4, 4.1, o-series…) |
|---|---|---|---|
| Implicit breakpoints | end of latest eligible message (user msg, last of a tool-output group, last initial developer msg) | every 2,048 tokens | model-dependent intervals |
| Explicit breakpoints | prompt_cache_breakpoint:{mode:"explicit"} on content parts (≤4 writes/request; not on top-level instructions, not on additional_tools) |
no | no |
| Mode | prompt_cache_options.mode: implicit (default) | explicit |
— | — |
| Lifetime | prompt_cache_options.ttl:"30m" (only value; ≥30 min after last use) |
prompt_cache_retention:"24h" only |
prompt_cache_retention: in_memory (5–10 min idle, ≤1 h) | 24h (~30 min, ≤24 h) on gpt-5.4, gpt-5.2, gpt-5.1*, gpt-5, gpt-4.1… |
| Reporting | exact boundary | rounded down to ×128, hidden tokens excluded | rounded down to ×128 |
prompt_cache_key |
optional (separate accounting per customer) | routing hint: keep ≈15 req/min per key | routing hint |
| Diagnostics | prompt_cache_options.comparison_response_id → prompt_cache_diagnostics |
no | no |
Default retention: 24h for non-ZDR orgs, in_memory for ZDR orgs (live echo on gpt-5.4-nano: "prompt_cache_retention":"24h"; "in_memory" when requested).
2. Parameters
| Parameter | Endpoints | Type / enum | Notes |
|---|---|---|---|
prompt_cache_key |
responses, chat, compact | string | live: accepted and echoed on Responses |
prompt_cache_retention |
responses, chat, compact | in_memory | 24h |
deprecated (still works); use prompt_cache_options.ttl on GPT-5.6+ |
prompt_cache_options.ttl |
responses, chat, compact | "30m" |
GPT-5.6+ |
prompt_cache_options.mode |
responses, chat, compact | implicit | explicit |
GPT-5.6+ |
prompt_cache_options.prewarm |
responses | boolean | Responses-only (spec ResponsePromptCacheOptionsParam) |
prompt_cache_options.comparison_response_id |
responses | string | diagnostics baseline; does not load the conversation |
…content[].prompt_cache_breakpoint.mode |
responses (input_text/input_image/input_file, tool outputs), chat (text parts, prediction.content parts) |
"explicit" |
GPT-5.6+ |
user |
legacy | string | replaced by safety_identifier + prompt_cache_key |
Live errors: prompt_cache_options on gpt-5.4-nano → 400 invalid_parameter "prompt_cache_options is not supported on this model".
3. Reading usage
- Responses:
usage.input_tokens_details.cached_tokens,usage.input_tokens_details.cache_write_tokens. - Chat:
usage.prompt_tokens_details.cached_tokens(no write counter). - Cost = (input − cached − written)×p_in + cached×p_cached + written×1.25×p_in (writes only GPT-5.6+).
- Diagnostics (GPT-5.6+, Responses):
prompt_cache_diagnostics.type∈cache_hit|cache_miss{reason, comparison_reusable_tokens, cache_missed_tokens}|comparison_response_not_found|unavailable;reason∈model_changed,prompt_cache_key_changed,service_tier_changed,tools_changed,text_format_changed,reasoning_effort_changed,verbosity_changed,context_compacted,input_changed. Read it fromresponse.completed.responsewhen streaming. Free; ZDR-compatible.
4. Live experiment (2026-09-18)
Prefix: ~1.5k tokens of repeated text, sent twice 1.5 s apart with a stable prompt_cache_key.
| Surface | Model | 1st call | 2nd call |
|---|---|---|---|
Chat system message |
gpt-4.1-nano |
prompt_tokens 1556, cached_tokens 0 |
prompt_tokens 1556, cached_tokens 1408 (=11×128, rounded down as documented; 90 % of prompt) |
Responses instructions (store:false, effort:low) |
gpt-5.4-nano |
input_tokens 1555, cached 0, cache_write 0 |
input_tokens 1555, cached_tokens 0, cache_write_tokens 0 |
Interpretation: the Chat path behaved exactly as documented; the Responses pair on gpt-5.4-nano did not hit (possible causes: routing to another machine, instructions-based prefix, short interval, or hidden-token minimum). Not a documentation error — recorded as an observation (FAILED_VERIFICATION would be wrong: the calls succeeded; the hit simply did not occur). cache_write_tokens is 0 on pre-5.6 models (no write charge).
5. Practices (guide digest)
Stable content first (instructions, reference material, tool definitions), dynamic content (timestamps, user data) last or in later messages; append, don't rewrite history; disable tools with tool_choice:"none"/allowed_tools instead of removing definitions; change effort via configuration_update (gpt-6-astra) instead of reasoning.effort; compaction resets the prefix (context_compacted); expand a just-below-minimum prefix with useful stable material when reuse is high (break-even L = M·(r + (w−r)/N)); monitor cached_tokens/cache_write_tokens per user/day and the Prompt Caching dashboard.