Gemini context caching — implicit and explicit (cachedContents)
Status: DOCUMENTED + BETA (explicit caching is v1beta) + ACCOUNT_RESTRICTED on this key: cache creation on gemini-3.5-flash and gemini-3.5-flash-lite → 429 RESOURCE_EXHAUSTED TotalCachedContentStorageTokensPerModelFreeTier limit exceeded … limit=0 (free tier). Size validation (min_total_token_count=1024) and request-shape errors were LIVE_VERIFIED; CRUD on an existing cache could not be exercised (2026-09-18).
Sources: https://ai.google.dev/gemini-api/docs/generate-content/caching · https://ai.google.dev/api/caching · https://ai.google.dev/gemini-api/docs/pricing · discovery CachedContent
Machine-readable: generated/fragments/parameters/gemini-cached-contents.json, generated/fragments/endpoints/gemini-core.json (5 cachedContents endpoints), generated/fragments/objects/gemini-core-objects.json#CachedContent
Last verified: 2026-09-18
1. Implicit caching (automatic)
Enabled by default on Gemini 2.5+; no guarantee of savings; hits are reported in usageMetadata.cachedContentTokenCount. Minimum prompt size for a hit (docs table):
| Model | Min tokens |
|---|---|
| Gemini 3.8 / 3.7 / 3.6 / 3.5 Flash, 3.1 Pro Preview | 4,096 |
| Gemini 2.5 Flash / Pro | 2,048 |
Tips: put the large, shared prefix first; send similar-prefix requests close in time. /v1/cachedContents does not exist (404, empty body) — explicit caching is v1beta only.
2. Explicit caching — lifecycle
| Step | Call | Notes |
|---|---|---|
| Create | POST /v1beta/cachedContents body {model: "models/<id>", contents[], systemInstruction?, tools?, toolConfig?, displayName? (≤128), ttl? | expireTime?} |
model required in models/{id} form (else 400 Model name must be in the format 'models/{model_name}'). Default TTL 1 h. Contents/systemInstruction/tools are input-only (never returned). Too small → 400 Cached content is too small. total_token_count=6, min_total_token_count=1024 (both 3.5 Flash and Flash-Lite; the docs' 4,096 figure is for implicit caching). Free tier → 429 limit=0. |
| Use | generateContent with "cachedContent": "cachedContents/<id>" |
Same model required; do not repeat the cached systemInstruction/tools. usageMetadata.cachedContentTokenCount (and cacheTokensDetails[]) report the prefix; promptTokenCount still includes it. countTokens with generateContentRequest.cachedContent reports cachedContentTokenCount too. |
| List | GET /v1beta/cachedContents?pageSize=&pageToken= (max 1000) |
metadata only |
| Get | GET /v1beta/cachedContents/{id} |
{name, model, displayName, createTime, updateTime, expireTime, usageMetadata:{totalTokenCount}} |
| Update | PATCH /v1beta/cachedContents/{id} body {ttl: "600s"} or {expireTime: "<RFC3339>"} (+ optional ?updateMask=ttl) |
only expiration is mutable |
| Delete | DELETE /v1beta/cachedContents/{id} → {} |
stops storage billing |
Response object: CachedContent {name, model, displayName, createTime, updateTime, expireTime (always), usageMetadata.totalTokenCount}; ttl is input-only.
3. Pricing (pricing page, 2026-09)
| Model | Cached input / 1M | Storage / 1M tokens / hour | Free tier |
|---|---|---|---|
| gemini-3.5-flash | $0.15 (standard), $0.075 flex | $1.00 | "Free of charge" on the page, but quota limit=0 observed |
| gemini-3.5-flash-lite | $0.03 | $1.00 | Not available (verified 429) |
| gemini-3.8-flash | $0.075 (→ $0.15 in 2027) | $0.50 (→ $1.00 in 2027) | Free of charge (page) |
Billing components: cached tokens at the reduced rate on each use + storage per token-hour for the TTL + normal non-cached input and output. No min/max TTL documented. Cached tokens count toward the model's context window and standard rate limits.
4. SDK
- Python:
cache = client.caches.create(model=, config=types.CreateCachedContentConfig(contents=[...], system_instruction=, ttl="300s", display_name=))→client.models.generate_content(model=, contents=, config=types.GenerateContentConfig(cached_content=cache.name));client.caches.list()/get(name)/update(name, config=types.UpdateCachedContentConfig(ttl=|expire_time=))/delete(name). - Node:
ai.caches.create({model, config:{contents, systemInstruction, ttl}}),ai.caches.update({name, config:{ttl}}),ai.caches.delete({name}). - OpenAI-compat:
extra_body: {cached_content: name}(compat domain).
Examples (examples/gemini/context-caching/) are written to run end-to-end but exit cleanly with ACCOUNT_RESTRICTED on the free tier. Tests gated: tests/gemini/test_caching.py verifies the error shapes and skips CRUD when creation is quota-blocked.