SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
2.2 KB

# xAI prompt caching — automatic prefix cache, x-grok-conv-id / prompt_cache_key

Status: DOCUMENTED + LIVE_VERIFIED (2026-09-19: two identical 920-token prompts with the same key → cached_tokens 192 → 896).

Sources: https://docs.x.ai/developers/advanced-api-usage/prompt-caching (how-it-works, maximizing-cache-hits, multi-turn, usage-and-pricing, best-practices) · https://docs.x.ai/developers/pricing Last verified: 2026-09-19

# How it works

  • Automatic, no cache_control markers. The prefix of the messages array that exactly matches a previous request is served from cache; only the tail is computed.
  • Cache is per server: use sticky routing — header x-grok-conv-id: <id> (Chat Completions, gRPC metadata) or body prompt_cache_key (Responses; also accepted on Chat Completions, "plumbed to x-grok-conv-id"). Not guaranteed (eviction, re-routing).
  • Reasoning models: include previous reasoning_content / encrypted reasoning items, or use previous_response_id; omitting reasoning is "the top cause of cache misses". Never edit/remove/reorder earlier messages.
  • Live: even a fresh minimal request shows cached_tokens: 192 of 196 prompt tokens — the hidden system prefix is cached across all users/keys; a 920-token prompt sent twice with the same prompt_cache_key + x-grok-conv-id went from 192 to 896 cached tokens (~1 s apart).

# Where to read it

API Field
Chat Completions / legacy usage.prompt_tokens_details.cached_tokens (also text_tokens = cached + uncached)
Responses usage.input_tokens_details.cached_tokens
/v1/messages usage.cache_read_input_tokens (cache_creation_input_tokens always 0; input_tokens = non-cached only)
xai-sdk response.usage.cached_prompt_text_tokens

# Pricing

Cached input billed at the cached rate (grok-4.6 $0.50 vs $2.00/M; grok-4.5 $0.30 vs $2.00; grok-4.3 $0.20 vs $1.25; grok-4.20* $0.20 vs $1.25; grok-build-0.1 $0.20 vs $1.00 — per 1M tokens, <200k prompt; long-context tier doubles). Long-context threshold counts cached + uncached tokens. Priority tier (2×) and US endpoint (1.1×) multipliers apply after the cache discount. Batch discount applies to cached tokens too.