SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
4.6 KB

# Gemini context caching — implicit and explicit (cachedContents)

Status: DOCUMENTED + BETA (explicit caching is v1beta) + ACCOUNT_RESTRICTED on this key: cache creation on gemini-3.5-flash and gemini-3.5-flash-lite → 429 RESOURCE_EXHAUSTED TotalCachedContentStorageTokensPerModelFreeTier limit exceeded … limit=0 (free tier). Size validation (min_total_token_count=1024) and request-shape errors were LIVE_VERIFIED; CRUD on an existing cache could not be exercised (2026-09-18). Sources: https://ai.google.dev/gemini-api/docs/generate-content/caching · https://ai.google.dev/api/caching · https://ai.google.dev/gemini-api/docs/pricing · discovery CachedContent Machine-readable: generated/fragments/parameters/gemini-cached-contents.json, generated/fragments/endpoints/gemini-core.json (5 cachedContents endpoints), generated/fragments/objects/gemini-core-objects.json#CachedContent Last verified: 2026-09-18

# 1. Implicit caching (automatic)

Enabled by default on Gemini 2.5+; no guarantee of savings; hits are reported in usageMetadata.cachedContentTokenCount. Minimum prompt size for a hit (docs table):

Model Min tokens
Gemini 3.8 / 3.7 / 3.6 / 3.5 Flash, 3.1 Pro Preview 4,096
Gemini 2.5 Flash / Pro 2,048

Tips: put the large, shared prefix first; send similar-prefix requests close in time. /v1/cachedContents does not exist (404, empty body) — explicit caching is v1beta only.

# 2. Explicit caching — lifecycle

Step Call Notes
Create POST /v1beta/cachedContents body {model: "models/<id>", contents[], systemInstruction?, tools?, toolConfig?, displayName? (≤128), ttl? | expireTime?} model required in models/{id} form (else 400 Model name must be in the format 'models/{model_name}'). Default TTL 1 h. Contents/systemInstruction/tools are input-only (never returned). Too small → 400 Cached content is too small. total_token_count=6, min_total_token_count=1024 (both 3.5 Flash and Flash-Lite; the docs' 4,096 figure is for implicit caching). Free tier → 429 limit=0.
Use generateContent with "cachedContent": "cachedContents/<id>" Same model required; do not repeat the cached systemInstruction/tools. usageMetadata.cachedContentTokenCount (and cacheTokensDetails[]) report the prefix; promptTokenCount still includes it. countTokens with generateContentRequest.cachedContent reports cachedContentTokenCount too.
List GET /v1beta/cachedContents?pageSize=&pageToken= (max 1000) metadata only
Get GET /v1beta/cachedContents/{id} {name, model, displayName, createTime, updateTime, expireTime, usageMetadata:{totalTokenCount}}
Update PATCH /v1beta/cachedContents/{id} body {ttl: "600s"} or {expireTime: "<RFC3339>"} (+ optional ?updateMask=ttl) only expiration is mutable
Delete DELETE /v1beta/cachedContents/{id} → {} stops storage billing

Response object: CachedContent {name, model, displayName, createTime, updateTime, expireTime (always), usageMetadata.totalTokenCount}; ttl is input-only.

# 3. Pricing (pricing page, 2026-09)

Model Cached input / 1M Storage / 1M tokens / hour Free tier
gemini-3.5-flash $0.15 (standard), $0.075 flex $1.00 "Free of charge" on the page, but quota limit=0 observed
gemini-3.5-flash-lite $0.03 $1.00 Not available (verified 429)
gemini-3.8-flash $0.075 (→ $0.15 in 2027) $0.50 (→ $1.00 in 2027) Free of charge (page)

Billing components: cached tokens at the reduced rate on each use + storage per token-hour for the TTL + normal non-cached input and output. No min/max TTL documented. Cached tokens count toward the model's context window and standard rate limits.

# 4. SDK

  • Python: cache = client.caches.create(model=, config=types.CreateCachedContentConfig(contents=[...], system_instruction=, ttl="300s", display_name=)) → client.models.generate_content(model=, contents=, config=types.GenerateContentConfig(cached_content=cache.name)); client.caches.list()/get(name)/update(name, config=types.UpdateCachedContentConfig(ttl=|expire_time=))/delete(name).
  • Node: ai.caches.create({model, config:{contents, systemInstruction, ttl}}), ai.caches.update({name, config:{ttl}}), ai.caches.delete({name}).
  • OpenAI-compat: extra_body: {cached_content: name} (compat domain).

Examples (examples/gemini/context-caching/) are written to run end-to-end but exit cleanly with ACCOUNT_RESTRICTED on the free tier. Tests gated: tests/gemini/test_caching.py verifies the error shapes and skips CRUD when creation is quota-blocked.