SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
46.7 KB · 352 lines markdown
Rendered Raw Blame History
1# Pricing — side-by-side and cost models (OpenAI · Anthropic · xAI · Gemini)23**Status:** every number below is read from `generated/models.json` (`pricing` block of each model, sourced from the vendor pricing/model pages on 2026-09-18) and `generated/pricing.json` (tool/service/rule rows). Figures are USD per 1M tokens unless stated. Statuses next to model ids are the model record statuses. No live billing was reconciled; treat as list prices. Gemini 3.6–3.8 Flash rows are **introductory prices through 2026-12-31** (the records carry a `from_2027_01_01` block at 2×); xAI `x_search` switches from per-call to per-post/per-profile billing on **2026-09-21**.4**Sources:** https://developers.openai.com/api/docs/pricing · https://platform.claude.com/docs/en/about-claude/pricing · https://docs.x.ai/developers/pricing (+ live `GET /v1/language-models` price ticks) · https://ai.google.dev/gemini-api/docs/pricing · docs/openai/pricing.md · docs/anthropic/pricing.md · docs/xai/pricing.md · docs/gemini/pricing.md5**Last verified:** 2026-09-1867## 1. Pricing dimensions — how the four price lists are structured89| Dimension | OpenAI | Anthropic | xAI | Gemini |10|---|---|---|---|---|11| Base tokens | `input`, `cached_input`, `output` per model; separate `*_long_context` table for >272k-token prompts (2× input, 1.5× output) on GPT-5.4/5.5/5.6/6 Astra | `input`, `output`; 1M context at **standard** price on Claude 4.6+ (no long-context premium) | `input`, `cached_input`, `image_input` (= text input price), `output`; **long-context tier** at 2× for prompts **≥ 200,000 tokens — billed on ALL tokens of that request** (`live_raw_ticks.long_context_threshold`); reasoning tokens billed as output on every Grok call | `input`, `cached_input`, `output` per model; modality-specific input rows (`audio_input`, `image_input`, `video_input`); **>200k tier** (`input_over_200k` 2×, `output_over_200k` 1.5×) on Pro models only; thinking tokens billed as output (`output_includes_thinking_tokens`) |12| Cache write | free on models before GPT-5.6; **1.25× input** on GPT-5.6+ (`cache_write` row) | **1.25× input** for 5-minute TTL, **2× input** for 1-hour TTL | none — automatic prefix cache, no write charge | implicit cache: none; **explicit `cachedContents`: storage per 1M tokens per hour** (`cache_storage_hour`: $0.50 Flash 3.6–3.8, $1.00 most Flash, $4.50 Pro; 1.8× on priority) |13| Cache read | model-specific `cached_input` (0.1× on GPT-5.6+, 0.25× on o3, 0.5× on gpt-4o…) | **0.1× input** (0.025× on Fable 5.1 / Mythos 5.1) | `cached_input` per model: 0.25× (grok-4.6 $0.50), 0.15× (grok-4.5 $0.30), 0.16× (grok-4.3 / 4.20 $0.20), 0.2× (grok-build $0.20) | `cached_input` = **0.1× input** on every priced model (e.g. gemini-3.8-flash $0.075, gemini-3.1-pro-preview $0.20) |14| Batch | `batch` tier = 50 % of standard (own `cached_input`) | `batch_input`/`batch_output` = 50 % of standard; cache multipliers stack | **20 % off** (`batch_discount: 0.2`) on grok-4.3 / 4.20 / 4.20-multi-agent only; **not supported** on grok-4.6, grok-4.5, grok-build-0.1 (`batch_discount: null`); Imagine models accepted in Batch but billed at standard rates | `batch` tier = **50 %** of standard (`service:batch_api` multiplier 0.5) on text, image, TTS and embedding models; not on the free tier |15| Cheaper async / best-effort tier | `flex` = batch rates on synchronous calls (slower, may 429) | none | none (Batch API is the only discount) | **`flex`** = 50 % of standard (`service:flex_inference`), 1–15 min target latency, sheddable (429 when capacity is shed); ten text models listed |16| Faster / priority tier | `fast` = 2× standard (renamed from priority 2026-07-30; `fast_long_context` 4×/3×) | `speed: fast` (Opus 5 / 4.8 only) = 2× standard ($10 / $50) | **`service_tier: priority`** = **2×** on all token types (caching discount applied before the multiplier; billed only when the response echoes `service_tier: priority`) | **`priority`** tier = **1.8×** standard (`service:priority_inference`, docs: 75–100 % more); 0.3× the standard rate limit; graceful downgrade to standard when exceeded |17| Reasoning tokens | billed as output (`usage.output_tokens_details.reasoning_tokens`) | billed as output (`usage.output_tokens_details.thinking_tokens`) | billed as output (`usage.completion_tokens_details.reasoning_tokens`); every Grok reasoning model spends them even on `reasoning_effort: low` | billed as output (`usageMetadata.thoughtsTokenCount`); Gemini 3.x thinking cannot be fully disabled on Flash/Pro (`thinkingLevel: minimal` where supported) |18| Regional inference | +10 % on regional hosts for models released ≥ 2026-03-05 | `inference_geo: us` ×1.1 on all token dimensions (Claude 4.6+); Bedrock/Vertex regional +10 % | **`https://us.api.x.ai/v1`** = **1.1×** global token rates (grok-4.6 only in the live catalogue: $2.20 / $0.55 / $6.60); `eu-west-1.api.x.ai` LIVE_DISCOVERED, no separate price row | no regional price row on the Gemini Developer API (data residency is a Vertex AI feature); AI Studio usage free |19| Free tier | none (Free usage tier = rate-limit tier, still billed) | none | none (prepaid credits; Tier 0 = $0 spend) | **yes** — `free_tier` rows on Flash / Flash-Lite / Live / TTS / transcribe / embedding / Gemma models: input & output $0, restrictive limits, content may be used to improve Google products; **Pro models and paid-only media models are not available on the free tier** (observed `limit: 0`) |20| Tool-definition overhead | none published (definitions are input tokens) | published per model: tool-use system prompt 286–804 tokens, toolsets 4,500–6,600 tokens | none published (definitions are input tokens; server-side tool results count as input) | none published; `toolUsePromptTokenCount` reports tokens consumed by URL context / grounding results |21| Per-image / per-second media | image models per token (≈ per image), Sora per second | — | Imagine **per image** ($0.02 / $0.04–$0.08 / $0.05) and **per second** of video ($0.05 / $0.08); voice per minute; TTS per 1M characters; STT per hour | image models per token with a **per-image** equivalent ($0.034–$0.24), Veo **per second** ($0.05–$0.60), Lyria **per song** ($0.04 / $0.08), Live/TTS/transcribe per 1M audio tokens (≈ per minute) |2223## 2. OpenAI current text models (USD / 1M tokens)2425| Model | Status | Input | Cached in | Cache write | Output | Batch in / out | Flex in / out | Fast in / out | Long-context in / out (>272k) |26|---|---|---|---|---|---|---|---|---|---|27| `gpt-6-astra` | `DOCUMENTED` · `LIVE_VERIFIED` | 10 | 1 | 12.5 | 50 | 5 / 25 | 5 / 25 | 20 / 100 | 20 / 75 |28| `gpt-5.6-sol` | `DOCUMENTED` · `LIVE_VERIFIED` | 4 | 0.4 | 5 | 20 | 2 / 10 | 2 / 10 | 8 / 40 | 8 / 30 |29| `gpt-5.6-terra` | `DOCUMENTED` · `LIVE_VERIFIED` | 2 | 0.2 | 2.5 | 12 | 1 / 6 | 1 / 6 | 4 / 24 | 4 / 18 |30| `gpt-5.6-luna` | `DOCUMENTED` · `LIVE_VERIFIED` | 0.2 | 0.02 | 0.25 | 1.2 | 0.1 / 0.6 | 0.1 / 0.6 | 0.4 / 2.4 | 0.4 / 1.8 |31| `gpt-5.5` | `DOCUMENTED` · `LIVE_VERIFIED` | 5 | 0.5 | — | 30 | 2.5 / 15 | 2.5 / 15 | 12.5 / 75 | 10 / 45 |32| `gpt-5.5-pro` | `DOCUMENTED` · `LIVE_VERIFIED` | 30 | — | — | 180 | 15 / 90 | 15 / 90 | — / — | 60 / 270 |33| `gpt-5.4` | `DOCUMENTED` · `LIVE_VERIFIED` | 2.5 | 0.25 | — | 15 | 1.25 / 7.5 | 1.25 / 7.5 | 5 / 30 | 5 / 22.5 |34| `gpt-5.4-pro` | `DOCUMENTED` · `LIVE_VERIFIED` | 30 | — | — | 180 | 15 / 90 | 15 / 90 | — / — | 60 / 270 |35| `gpt-5.4-mini` | `DOCUMENTED` · `LIVE_VERIFIED` | 0.75 | 0.075 | — | 4.5 | 0.375 / 2.25 | 0.375 / 2.25 | 1.5 / 9 | — / — |36| `gpt-5.4-nano` | `DOCUMENTED` · `LIVE_VERIFIED` | 0.2 | 0.02 | — | 1.25 | 0.1 / 0.625 | 0.1 / 0.625 | — / — | — / — |37| `gpt-5.3-codex` | `DOCUMENTED` · `LIVE_VERIFIED` | 1.75 | 0.175 | — | 14 | — / — | — / — | 3.5 / 28 | — / — |38| `gpt-5.2` | `DOCUMENTED` · `LIVE_VERIFIED` | 1.75 | 0.175 | — | 14 | 0.875 / 7 | 0.875 / 7 | 3.5 / 28 | — / — |39| `gpt-5.2-pro` | `DOCUMENTED` · `LIVE_VERIFIED` | 21 | — | — | 168 | 10.5 / 84 | — / — | — / — | — / — |40| `gpt-5.1` | `DOCUMENTED` · `LIVE_VERIFIED` | 1.25 | 0.125 | — | 10 | 0.625 / 5 | 0.625 / 5 | 2.5 / 20 | — / — |41| `gpt-5` | `DOCUMENTED` · `LIVE_VERIFIED` | 1.25 | 0.125 | — | 10 | 0.625 / 5 | 0.625 / 5 | 2.5 / 20 | — / — |42| `gpt-5-mini` | `DOCUMENTED` · `LIVE_VERIFIED` | 0.25 | 0.025 | — | 2 | 0.125 / 1 | 0.125 / 1 | 0.45 / 3.6 | — / — |43| `gpt-5-nano` | `DOCUMENTED` · `LIVE_VERIFIED` | 0.05 | 0.005 | — | 0.4 | 0.025 / 0.2 | 0.025 / 0.2 | — / — | — / — |44| `gpt-5-pro` | `DOCUMENTED` · `LIVE_VERIFIED` | 15 | — | — | 120 | 7.5 / 60 | — / — | — / — | — / — |45| `o3` | `DOCUMENTED` · `LIVE_VERIFIED` | 2 | 0.5 | — | 8 | 1 / 4 | 1 / 4 | 3.5 / 14 | — / — |46| `o3-pro` | `DOCUMENTED` · `LIVE_VERIFIED` | 20 | — | — | 80 | 10 / 40 | — / — | — / — | — / — |47| `o4-mini` | `DOCUMENTED` · `LIVE_VERIFIED` · `DEPRECATED` | 1.1 | 0.275 | — | 4.4 | 0.55 / 2.2 | 0.55 / 2.2 | 2 / 8 | — / — |48| `o3-mini` | `DOCUMENTED` · `LIVE_VERIFIED` · `DEPRECATED` | 1.1 | 0.55 | — | 4.4 | 0.55 / 2.2 | — / — | — / — | — / — |49| `gpt-4.1` | `DOCUMENTED` · `LIVE_VERIFIED` | 2 | 0.5 | — | 8 | 1 / 4 | — / — | 3.5 / 14 | — / — |50| `gpt-4.1-mini` | `DOCUMENTED` · `LIVE_VERIFIED` | 0.4 | 0.1 | — | 1.6 | 0.2 / 0.8 | — / — | 0.7 / 2.8 | — / — |51| `gpt-4.1-nano` | `DOCUMENTED` · `LIVE_VERIFIED` · `DEPRECATED` | 0.1 | 0.025 | — | 0.4 | 0.05 / 0.2 | — / — | 0.2 / 0.8 | — / — |52| `gpt-4o` | `DOCUMENTED` · `LIVE_VERIFIED` | 2.5 | 1.25 | — | 10 | 1.25 / 5 | — / — | 4.25 / 17 | — / — |53| `gpt-4o-mini` | `DOCUMENTED` · `LIVE_VERIFIED` | 0.15 | 0.075 | — | 0.6 | 0.075 / 0.3 | — / — | 0.25 / 1 | — / — |54| `chat-latest` | `DOCUMENTED` · `LIVE_VERIFIED` | 5 | 0.5 | — | 30 | — / — | — / — | — / — | — / — |55| `gpt-5.6-cyber` | `DOCUMENTED` · `ACCOUNT_RESTRICTED` | 12.5 | 1.25 | 15.625 | 75 | — / — | — / — | — / — | — / — |5657Notes from the records: `gpt-5.6` is an alias of `gpt-5.6-sol` (promotional $4/$20 at least through 2026-11-21); pro models have no cached-input price (no caching); `chat-latest` has no batch/flex/fast tiers; `gpt-5.6-cyber` is ACCOUNT_RESTRICTED and priced at short-context rates. The `o3` record carries a model-page table ($1 / $0.25 / $4) that disagrees with its pricing-page standard row ($2 / $0.5 / $8) — flagged as a data inconsistency.5859## 3. Anthropic current models (USD / 1M tokens)6061| Model | Status | Input | Cache write 5m | Cache write 1h | Cache read | Output | Batch in / out | Fast in / out | Min cacheable tokens |62|---|---|---|---|---|---|---|---|---|---|63| `claude-fable-5-1` | `DOCUMENTED` · `LIVE_VERIFIED` | 10 | 12.5 | 20 | 0.25 | 50 | 5 / 25 | — / — | 512 |64| `claude-fable-5` | `DOCUMENTED` · `LIVE_VERIFIED` · `LEGACY` | 10 | 12.5 | 20 | 1 | 50 | 5 / 25 | — / — | 512 |65| `claude-mythos-5-1` | `DOCUMENTED` · `ACCOUNT_RESTRICTED` · `PREVIEW` | 10 | 12.5 | 20 | 0.25 | 50 | 5 / 25 | — / — | 512 |66| `claude-mythos-5` | `DOCUMENTED` · `ACCOUNT_RESTRICTED` · `PREVIEW` | 10 | 12.5 | 20 | 1 | 50 | 5 / 25 | — / — | 512 |67| `claude-opus-5` | `DOCUMENTED` · `LIVE_VERIFIED` | 5 | 6.25 | 10 | 0.5 | 25 | 2.5 / 12.5 | 10 / 50 | 512 |68| `claude-opus-4-8` | `DOCUMENTED` · `LIVE_VERIFIED` · `LEGACY` | 5 | 6.25 | 10 | 0.5 | 25 | 2.5 / 12.5 | 10 / 50 | 1024 |69| `claude-opus-4-7` | `DOCUMENTED` · `LIVE_VERIFIED` · `LEGACY` | 5 | 6.25 | 10 | 0.5 | 25 | 2.5 / 12.5 | — / — | 2048 |70| `claude-opus-4-6` | `DOCUMENTED` · `LIVE_VERIFIED` · `LEGACY` | 5 | 6.25 | 10 | 0.5 | 25 | 2.5 / 12.5 | — / — | 4096 |71| `claude-opus-4-5-20251101` | `DOCUMENTED` · `LIVE_VERIFIED` · `LEGACY` | 5 | 6.25 | 10 | 0.5 | 25 | 2.5 / 12.5 | — / — | 4096 |72| `claude-sonnet-5` | `DOCUMENTED` · `LIVE_VERIFIED` | 2 | 2.5 | 4 | 0.2 | 10 | 1 / 5 | — / — | 1024 |73| `claude-sonnet-4-6` | `DOCUMENTED` · `LIVE_VERIFIED` · `LEGACY` | 3 | 3.75 | 6 | 0.3 | 15 | 1.5 / 7.5 | — / — | 1024 |74| `claude-sonnet-4-5-20250929` | `DOCUMENTED` · `LIVE_VERIFIED` · `LEGACY` | 3 | 3.75 | 6 | 0.3 | 15 | 1.5 / 7.5 | — / — | 1024 |75| `claude-haiku-4-5-20251001` | `DOCUMENTED` · `LIVE_VERIFIED` | 1 | 1.25 | 2 | 0.1 | 5 | 0.5 / 2.5 | — / — | 4096 |7677Notes: Sonnet 5's introductory $2 / $10 was made permanent on 2026-08-10. Fable 5.1 / Mythos 5.1 cache reads are 0.025× ($0.25). Retired models (Opus 4.1 $15/$75, Sonnet 4 $3/$15, Haiku 3.5 $0.80/$4) remain priced on Bedrock/Vertex only. Priority Tier is no longer sold; Claude Platform on AWS and Foundry convert the same USD rates to CCUs at $0.01.7879## 4. xAI current Grok text models (USD / 1M tokens)8081Standard = prompt < 200,000 tokens. **Long context = prompt ≥ 200,000 tokens: the higher rate applies to every token of that request** (input, cached, output). `image_input` equals the text input price on every model. Batch = 20 % off standard where supported. Priority = `service_tier: priority`, 2×. US regional = `https://us.api.x.ai/v1`, 1.1×.8283| Model | Status | Input | Cached in | Output | Long-ctx in / cached / out (≥200k) | Batch in / cached / out | Priority × | US regional × | Context |84|---|---|---|---|---|---|---|---|---|---|85| `grok-4.6` | `DOCUMENTED` · `LIVE_VERIFIED` | 2 | 0.5 | 6 | 4 / 1 / 12 | not supported | 2.0× | 1.1× | 500000 |86| `grok-4.5` | `DOCUMENTED` · `LIVE_VERIFIED` | 2 | 0.3 | 6 | 4 / 0.6 / 12 | not supported | 2.0× | — | 500000 |87| `grok-4.3` | `DOCUMENTED` · `LIVE_VERIFIED` | 1.25 | 0.2 | 2.5 | 2.5 / 0.4 / 5 | 1 / 0.16 / 2 (−20 %) | 2.0× | — | 1000000 |88| `grok-4.20-0309-reasoning` | `DOCUMENTED` · `LIVE_VERIFIED` | 1.25 | 0.2 | 2.5 | 2.5 / 0.4 / 5 | 1 / 0.16 / 2 (−20 %) | 2.0× | — | 1000000 |89| `grok-4.20-0309-non-reasoning` | `DOCUMENTED` · `LIVE_VERIFIED` | 1.25 | 0.2 | 2.5 | 2.5 / 0.4 / 5 | 1 / 0.16 / 2 (−20 %) | 2.0× | — | 1000000 |90| `grok-4.20-multi-agent-0309` | `DOCUMENTED` · `BETA` · `LIVE_VERIFIED` | 1.25 | 0.2 | 2.5 | 2.5 / 0.4 / 5 | 1 / 0.16 / 2 (−20 %) | 2.0 (Chat Completions/Responses; not documented per model) | — | 1000000 |91| `grok-build-0.1` | `DOCUMENTED` · `PREVIEW` · `LIVE_VERIFIED` | 1 | 0.2 | 2 | 2 / 0.4 / 4 | not supported | 2.0× | — | 256000 |9293Notes from the records: `grok-4.20-multi-agent-0309` is BETA and only reachable on `/v1/responses`; its `reasoning.effort` selects the agent count (low/medium = 4, high/xhigh = 16) and all agents' tokens are billed. `grok-build-0.1` (PREVIEW) is the coding model behind the `grok-code-fast-1` redirect. The retired ids `grok-3`, `grok-4-0709`, `grok-4-fast-*`, `grok-4-1-fast-*` **redirect to `grok-4.3` and are billed at grok-4.3 rates**. `grok-embedding-small` has no published price (ACCOUNT_RESTRICTED). Max output is not documented per model (`max_output: null` in every Grok record; Responses `max_output_tokens` defaults to 128,000).9495### 4b. xAI media and voice prices9697| Model / service | Status | Price | Unit / notes |98|---|---|---|---|99| `grok-imagine-image` | `DOCUMENTED` · `LIVE_VERIFIED` | $0.02 per image | single price; batch discount 0.0 (standard rates in Batch) |100| `grok-imagine-image-2.0` | `DOCUMENTED` · `LIVE_VERIFIED` | $0.06 per image | low/1k $0.04; low/2k $0.06; low/1.5k $0.05; medium/1k $0.06; medium/2k $0.08; medium/1.5k $0.07; batch discount 0.0 (standard rates in Batch) |101| `grok-imagine-image-quality` | `DOCUMENTED` · `DEPRECATED` · `LIVE_VERIFIED` | $0.05 per image | single price; batch discount 0.0 (standard rates in Batch) |102| `grok-imagine-video` | `DOCUMENTED` · `LIVE_VERIFIED` | $0.05 per second of generated video | video URLs in batch results expire after 1 hour |103| `grok-imagine-video-1.5` | `DOCUMENTED` · `LIVE_VERIFIED` | $0.08 per second of generated video | video URLs in batch results expire after 1 hour |104| `grok-voice-think-fast-2.0` (speech-to-speech) | `DOCUMENTED` | $0.08 per minute ($4.8 / h) + $0.004 per text item | audio sent or received billed per minute; each conversation.item.create text item billed $0.004 except function_call_output and audio items |105| `grok-voice-transcribe-2.0` (speech-to-text) | `DOCUMENTED` | $0.1 / hour REST (`POST /v1/stt`), $0.2 / hour streaming (`wss://api.x.ai/v1/stt`) | per hour of audio |106| text-to-speech (`POST /v1/tts`, `wss://api.x.ai/v1/tts`) | `DOCUMENTED` | $15 per 1M characters | POST /v1/tts and streaming |107108## 5. Gemini current text models (USD / 1M tokens)109110Standard tier for prompts ≤ 200k tokens; **>200k** columns only exist on Pro models (Flash models have one price for any prompt length). `Cached in` = implicit or explicit cache hit (0.1×). Storage = explicit `cachedContents` storage per 1M tokens per hour. Batch = 50 %, Flex = 50 %, Priority = 1.8×. Free = free-tier availability recorded on the pricing page.111112| Model | Status | Input | Cached in | Output | >200k in / cached / out | Batch in / out | Flex in / out | Priority in / out | Cache storage $/1M/h | Free tier |113|---|---|---|---|---|---|---|---|---|---|---|114| `gemini-3.8-flash` | `DOCUMENTED` · `LIVE_DISCOVERED` · `LIVE_VERIFIED` | 0.75 | 0.075 | 3.75 | — (flat) | 0.375 / 1.875 | 0.375 / 1.875 | 1.35 / 6.75 | 0.5 | input/output/caching free of charge (standard & priority rows); Bat… |115| `gemini-3.7-flash` | `DOCUMENTED` · `LIVE_DISCOVERED` | 0.75 | 0.075 | 3.75 | — (flat) | 0.375 / 1.875 | 0.375 / 1.875 | 1.35 / 6.75 | 0.5 | input/output/caching free of charge (standard & priority rows); Bat… |116| `gemini-3.6-flash` | `DOCUMENTED` · `LIVE_DISCOVERED` | 0.75 | 0.075 | 3.75 | — (flat) | 0.375 / 1.875 | 0.375 / 1.875 | 1.35 / 6.75 | 0.5 | input/output/caching free of charge (standard & priority rows); Bat… |117| `gemini-3.5-flash` | `DOCUMENTED` · `LIVE_DISCOVERED` · `LIVE_VERIFIED` | 1.5 | 0.15 | 9 | — (flat) | 0.75 / 4.5 | 0.75 / 4.5 | 2.7 / 16.2 | 1 | input/output/caching free of charge; Batch/Flex not available |118| `gemini-3.5-flash-lite` | `DOCUMENTED` · `LIVE_DISCOVERED` · `LIVE_VERIFIED` | 0.3 | 0.03 | 2.5 | — (flat) | 0.15 / 1.25 | 0.15 / 1.25 | 0.54 / 4.5 | 1 | input/output free of charge on all four rows (page lists Batch/Flex… |119| `gemini-3.1-pro-preview` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` · `ACCOUNT_RESTRICTED` | 2 | 0.2 | 12 | 4 / 0.4 / 18 | 1 / 6 | 1 / 6 | 3.6 / 21.6 | 4.5 | Not available (paid tier only) |120| `gemini-3.1-pro-preview-customtools` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` | 2 | 0.2 | 12 | 4 / 0.4 / 18 | 1 / 6 | 1 / 6 | 3.6 / 21.6 | 4.5 | Not available (paid tier only) |121| `gemini-3.1-flash-lite` | `DOCUMENTED` · `LIVE_DISCOVERED` · `DEPRECATED` | 0.25 | 0.025 | 1.5 | — (flat) | 0.125 / 0.75 | 0.125 / 0.75 | 0.45 / 2.7 | 1 | input/output free of charge; caching not available on free tier |122| `gemini-3-flash-preview` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` | 0.5 | 0.05 | 3 | — (flat) | 0.25 / 1.5 | 0.25 / 1.5 | 0.9 / 5.4 | 1 | input/output/caching free of charge; Batch/Flex not available |123| `gemini-2.5-pro` | `DOCUMENTED` · `LIVE_DISCOVERED` | 1.25 | 0.125 | 10 | 2.5 / 0.25 / 15 | 0.625 / 5 | 0.625 / 5 | 2.25 / 18 | 4.5 | input/output free of charge (standard/priority rows); caching, Batc… |124| `gemini-2.5-flash` | `DOCUMENTED` · `LIVE_DISCOVERED` | 0.3 | 0.03 | 2.5 | — (flat) | 0.15 / 1.25 | 0.15 / 1.25 | 0.54 / 4.5 | 1 | input/output free of charge; caching not available; Google Search g… |125| `gemini-2.5-flash-lite` | `DOCUMENTED` · `LIVE_DISCOVERED` · `ACCOUNT_RESTRICTED` | 0.1 | 0.01 | 0.4 | — (flat) | 0.05 / 0.2 | 0.05 / 0.2 | 0.18 / 0.72 | 1 | input/output free of charge; caching not available; Google Search g… |126| `gemini-robotics-er-2-preview` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` | 1 | 0.1 | 5 | — (flat) | 0.5 / 2.5 | — / — | — / — | 0.5 | input/output free of charge; caching not available |127| `gemini-2.5-computer-use-preview-10-2025` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` | 1 | — | 5 | — (flat) | — / — | — / — | — / — | — | input/output free of charge |128| `gemma-4-31b-it` | `DOCUMENTED` · `LIVE_DISCOVERED` | free | — | free | — | — | — | — | — | free of charge (input, output, caching, storage) |129| `gemma-4-26b-a4b-it` | `DOCUMENTED` · `LIVE_DISCOVERED` · `LIVE_VERIFIED` | free | — | free | — | — | — | — | — | free of charge (input, output, caching, storage) |130131Notes from the records: gemini-3.8 / 3.7 / 3.6 Flash share one introductory price list ($0.75 / $0.075 / $3.75) **through 2026-12-31**, doubling on 2027-01-01 ($1.50 / $0.15 / $7.50, storage $1.00); gemini-3.5-flash is the higher-priced Flash ($1.50 / $9.00). Pro models (`gemini-3.1-pro-preview`, `gemini-2.5-pro`) are **paid-tier only** (ACCOUNT_RESTRICTED on this free-tier key). Gemini 2.5 models are 'no longer available to new users' (404 on generateContent). Gemma 4 models are free of charge with no paid tier. `gemini-2.5-flash` / `-pro` have separate `audio_input` rows ($1.00 / 1M); Gemini 3.x prices text, image, video and audio input alike. Google Search grounding: Gemini 3.x 5,000 free requests/month then $14 / 1k **search queries**; Gemini 2.5: 1,500 RPD free then $35 / 1k **grounded prompts**.132133### 5b. Gemini media, voice and embedding prices134135| Model | Status | Price | Notes |136|---|---|---|---|137| `gemini-3.1-flash-image` | `DOCUMENTED` · `LIVE_DISCOVERED` | text in $0.5, text out $3, image out $60 / 1M → per image 0.5K (747 tok) $0.045, 1K (1120 tok) $0.067, 2K (1680 tok) $0.101, 4K (2520 tok) $0.151; batch image out $30 | not available |138| `gemini-3.1-flash-lite-image` | `DOCUMENTED` · `LIVE_DISCOVERED` | text in $0.25, image out $30 / 1M → 1K (1120 tok) $0.0336 | not available |139| `gemini-3-pro-image` | `DOCUMENTED` · `LIVE_DISCOVERED` | text in $2, image in $0.0011 per image, image out $120 / 1M → 1K/2K (1120 tok) $0.134, 4K (2000 tok) $0.24; priority image out $216 | not available |140| `gemini-2.5-flash-image` | `DOCUMENTED` · `LIVE_DISCOVERED` · `DEPRECATED` | $0.039 per image ($30 / 1M, 1,290 tok/image); batch $0.0195 | 1290 tokens per image up to 1024x1024 |141| `veo-3.1-generate-preview` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` | per second: 720p $0.4, 1080p $0.4, 4k $0.6 | charged only if the video is successfully generated |142| `veo-3.1-fast-generate-preview` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` | per second: 720p $0.1, 1080p $0.12, 4k $0.3 | not available |143| `veo-3.1-lite-generate-preview` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` · `LIVE_VERIFIED` | per second: 720p $0.05, 1080p $0.08, 4k not supported | not available |144| `gemini-omni-1.1-flash` | `DOCUMENTED` · `LIVE_DISCOVERED` | text in $1.5, text out $9, video out $17.5 / 1M | 5,792 tokens per second of 720p video => ~$0.10 per second |145| `lyria-3.5` | `DOCUMENTED` · `LIVE_DISCOVERED` · `LIVE_VERIFIED` | $0.08 per song (full length) | not available |146| `lyria-3-clip-preview` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` | $0.04 per song (30 s clip) | not available |147| `gemini-3.1-flash-tts-preview` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` | text in $1, audio out $20 / 1M (25 audio tok/s ≈ $0.03 / min); batch $0.5 / $10 | 25 audio tokens per second |148| `gemini-2.5-flash-preview-tts` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` | text in $0.5, audio out $10 / 1M | free of charge (standard) |149| `gemini-3.8-live` | `DOCUMENTED` · `LIVE_DISCOVERED` | text in $0.75, audio in $3 (≈ $0.005 / min), image/video in $1, text out $4.5, audio out $12 (≈ $0.018 / min) | free of charge |150| `gemini-2.5-flash-native-audio-preview-12-2025` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` | text in $0.5, audio/video in $3, text out $2, audio out $12 | free of charge |151| `gemini-3.5-transcribe` | `DOCUMENTED` · `LIVE_DISCOVERED` | audio in $2 (≈ $0.003 / min), text out $12 (≈ $0.002 / min) | blended ~$0.005/min |152| `gemini-3.5-transcribe-live` | `DOCUMENTED` · `LIVE_DISCOVERED` | audio in $3.5 (≈ $0.005 / min), text out $21 | 25 audio tokens/s input, 175 text tokens/min output; blended ~$0.009/min |153| `gemini-3.5-live-translate-preview` | `DOCUMENTED` · `LIVE_DISCOVERED` · `PREVIEW` | audio in $3.5, audio out $21 / 1M | 25 audio tokens/second; effective ~$0.0368 per minute |154| `gemini-embedding-2` | `DOCUMENTED` · `LIVE_DISCOVERED` · `LIVE_VERIFIED` | text $0.2, image $0.45 ($0.0001 per image), audio $6.5 ($0.0002 / s), video $12 ($0.0008 / frame) / 1M; batch text $0.1 | free of charge (standard); batch not available |155| `gemini-embedding-001` | `DOCUMENTED` · `LIVE_DISCOVERED` · `DEPRECATED` | — | not listed on pricing.md (2026-09-18); file-search.md bills indexing embeddings at $0.15 per 1M tokens |156157## 6. Cost model A — 1M input tokens + 100k output tokens, no cache158159Formula: `cost = 1.0 × input_price + 0.1 × output_price`, treating the 1M input tokens as **aggregate volume at the standard (short-prompt) rate**. The last column shows what a **single request whose prompt crosses the long-context threshold** costs instead (OpenAI >272k, xAI ≥200k applied to all tokens, Gemini Pro >200k; Anthropic and Gemini Flash have no threshold). Tiers shown where the model record lists them.160161| Provider | Model | Standard | Batch | Flex | Fast / Priority | Long-context request |162|---|---|---|---|---|---|---|163| openai | `gpt-6-astra` | $15.00 | $7.50 | $7.50 | $30.00 | $27.50 |164| openai | `gpt-5.6-sol` | $6.00 | $3.00 | $3.00 | $12.00 | $11.00 |165| openai | `gpt-5.6-terra` | $3.20 | $1.60 | $1.60 | $6.40 | $5.80 |166| openai | `gpt-5.6-luna` | $0.32 | $0.16 | $0.16 | $0.64 | $0.58 |167| openai | `gpt-5.5` | $8.00 | $4.00 | $4.00 | $20.00 | $14.50 |168| openai | `gpt-5.5-pro` | $48.00 | $24.00 | $24.00 | — | $87.00 |169| openai | `gpt-5.4` | $4.00 | $2.00 | $2.00 | $8.00 | $7.25 |170| openai | `gpt-5.4-pro` | $48.00 | $24.00 | $24.00 | — | $87.00 |171| openai | `gpt-5.4-mini` | $1.20 | $0.60 | $0.60 | $2.40 | — |172| openai | `gpt-5.4-nano` | $0.33 | $0.16 | $0.16 | — | — |173| openai | `gpt-5.3-codex` | $3.15 | — | — | $6.30 | — |174| openai | `gpt-5.2` | $3.15 | $1.58 | $1.58 | $6.30 | — |175| openai | `gpt-5.2-pro` | $37.80 | $18.90 | — | — | — |176| openai | `gpt-5.1` | $2.25 | $1.12 | $1.12 | $4.50 | — |177| openai | `gpt-5` | $2.25 | $1.12 | $1.12 | $4.50 | — |178| openai | `gpt-5-mini` | $0.45 | $0.23 | $0.23 | $0.81 | — |179| openai | `gpt-5-nano` | $0.0900 | $0.0450 | $0.0450 | — | — |180| openai | `gpt-5-pro` | $27.00 | $13.50 | — | — | — |181| openai | `o3` | $2.80 | $1.40 | $1.40 | $4.90 | — |182| openai | `o3-pro` | $28.00 | $14.00 | — | — | — |183| openai | `o4-mini` | $1.54 | $0.77 | $0.77 | $2.80 | — |184| openai | `o3-mini` | $1.54 | $0.77 | — | — | — |185| openai | `gpt-4.1` | $2.80 | $1.40 | — | $4.90 | — |186| openai | `gpt-4.1-mini` | $0.56 | $0.28 | — | $0.98 | — |187| openai | `gpt-4.1-nano` | $0.14 | $0.0700 | — | $0.28 | — |188| openai | `gpt-4o` | $3.50 | $1.75 | — | $5.95 | — |189| openai | `gpt-4o-mini` | $0.21 | $0.10 | — | $0.35 | — |190| openai | `chat-latest` | $8.00 | — | — | — | — |191| openai | `gpt-5.6-cyber` | $20.00 | — | — | — | — |192| anthropic | `claude-fable-5-1` | $15.00 | $7.50 | — | — | same as standard (no premium) |193| anthropic | `claude-fable-5` | $15.00 | $7.50 | — | — | same as standard (no premium) |194| anthropic | `claude-mythos-5-1` | $15.00 | $7.50 | — | — | same as standard (no premium) |195| anthropic | `claude-mythos-5` | $15.00 | $7.50 | — | — | same as standard (no premium) |196| anthropic | `claude-opus-5` | $7.50 | $3.75 | — | $15.00 | same as standard (no premium) |197| anthropic | `claude-opus-4-8` | $7.50 | $3.75 | — | $15.00 | same as standard (no premium) |198| anthropic | `claude-opus-4-7` | $7.50 | $3.75 | — | — | same as standard (no premium) |199| anthropic | `claude-opus-4-6` | $7.50 | $3.75 | — | — | same as standard (no premium) |200| anthropic | `claude-opus-4-5-20251101` | $7.50 | $3.75 | — | — | same as standard (no premium) |201| anthropic | `claude-sonnet-5` | $3.00 | $1.50 | — | — | same as standard (no premium) |202| anthropic | `claude-sonnet-4-6` | $4.50 | $2.25 | — | — | same as standard (no premium) |203| anthropic | `claude-sonnet-4-5-20250929` | $4.50 | $2.25 | — | — | same as standard (no premium) |204| anthropic | `claude-haiku-4-5-20251001` | $1.50 | $0.75 | — | — | same as standard (no premium) |205| xai | `grok-4.6` | $2.60 | not supported | — | $5.20 (priority) | $5.20 |206| xai | `grok-4.5` | $2.60 | not supported | — | $5.20 (priority) | $5.20 |207| xai | `grok-4.3` | $1.50 | $1.20 | — | $3.00 (priority) | $3.00 |208| xai | `grok-4.20-0309-reasoning` | $1.50 | $1.20 | — | $3.00 (priority) | $3.00 |209| xai | `grok-4.20-0309-non-reasoning` | $1.50 | $1.20 | — | $3.00 (priority) | $3.00 |210| xai | `grok-4.20-multi-agent-0309` | $1.50 | $1.20 | — | $3.00 (priority) | $3.00 |211| xai | `grok-build-0.1` | $1.20 | not supported | — | $2.40 (priority) | $2.40 |212| gemini | `gemini-3.8-flash` | $1.12 | $0.56 | $0.56 | $2.03 (priority) | same as standard (flat) |213| gemini | `gemini-3.7-flash` | $1.12 | $0.56 | $0.56 | $2.03 (priority) | same as standard (flat) |214| gemini | `gemini-3.6-flash` | $1.12 | $0.56 | $0.56 | $2.03 (priority) | same as standard (flat) |215| gemini | `gemini-3.5-flash` | $2.40 | $1.20 | $1.20 | $4.32 (priority) | same as standard (flat) |216| gemini | `gemini-3.5-flash-lite` | $0.55 | $0.28 | $0.28 | $0.99 (priority) | same as standard (flat) |217| gemini | `gemini-3.1-pro-preview` | $3.20 | $1.60 | $1.60 | $5.76 (priority) | $5.80 |218| gemini | `gemini-3.1-pro-preview-customtools` | $3.20 | $1.60 | $1.60 | $5.76 (priority) | $5.80 |219| gemini | `gemini-3.1-flash-lite` | $0.40 | $0.20 | $0.20 | $0.72 (priority) | same as standard (flat) |220| gemini | `gemini-3-flash-preview` | $0.80 | $0.40 | $0.40 | $1.44 (priority) | same as standard (flat) |221| gemini | `gemini-2.5-pro` | $2.25 | $1.12 | $1.12 | $4.05 (priority) | $4.00 |222| gemini | `gemini-2.5-flash` | $0.55 | $0.28 | $0.28 | $0.99 (priority) | same as standard (flat) |223| gemini | `gemini-2.5-flash-lite` | $0.14 | $0.0700 | $0.0700 | $0.25 (priority) | same as standard (flat) |224| gemini | `gemini-robotics-er-2-preview` | $1.50 | $0.75 | — | — (priority) | same as standard (flat) |225| gemini | `gemini-2.5-computer-use-preview-10-2025` | $1.50 | — | — | — (priority) | same as standard (flat) |226227Gemma 4 (`gemma-4-31b-it`, `gemma-4-26b-a4b-it`) and every Gemini free-tier row cost $0 at list price but carry free-tier limits and data-use terms; they are omitted from the arithmetic.228229## 7. Cost model B — the same workload with a 900k-token cached prefix (steady state)230231Assumptions: each request = 900k tokens **read** from cache + 100k fresh input + 100k output; the initial cache **write** is amortised over N = 10 requests. Formula per request: `0.9 × cache_read + 0.1 × input + 0.1 × output + write_cost / N`. Write cost: OpenAI pre-5.6 free, GPT-5.6+ `cache_write` × 0.9; Anthropic 5-minute write (1.25×) × 0.9; xAI none (automatic cache); Gemini implicit caching none, **explicit `cachedContents` = storage of 0.9M tokens for one hour** (`cache_storage_hour` × 0.9) amortised over N. A 900k prefix exceeds xAI's 200k threshold, so xAI is shown at **long-context** rates (the only rates that apply to such a request); Gemini Pro at its >200k rates, Gemini Flash flat.232233| Provider | Model | Uncached (A, same rates) | Cached steady-state (N=10) | Saving | Write / storage cost used |234|---|---|---|---|---|---|235| openai | `gpt-6-astra` | $15.00 | $8.03 | 46 % | 12.5 |236| openai | `gpt-5.6-sol` | $6.00 | $3.21 | 46 % | 5 |237| openai | `gpt-5.6-terra` | $3.20 | $1.81 | 44 % | 2.5 |238| openai | `gpt-5.6-luna` | $0.32 | $0.18 | 44 % | 0.25 |239| openai | `gpt-5.5` | $8.00 | $3.95 | 51 % | free |240| openai | `gpt-5.4` | $4.00 | $1.98 | 51 % | free |241| openai | `gpt-5.4-mini` | $1.20 | $0.59 | 51 % | free |242| openai | `gpt-5.4-nano` | $0.33 | $0.16 | 50 % | free |243| openai | `gpt-5.3-codex` | $3.15 | $1.73 | 45 % | free |244| openai | `gpt-5.2` | $3.15 | $1.73 | 45 % | free |245| openai | `gpt-5.1` | $2.25 | $1.24 | 45 % | free |246| openai | `gpt-5` | $2.25 | $1.24 | 45 % | free |247| openai | `gpt-5-mini` | $0.45 | $0.25 | 45 % | free |248| openai | `gpt-5-nano` | $0.0900 | $0.0495 | 45 % | free |249| openai | `o3` | $2.80 | $1.45 | 48 % | free |250| openai | `o4-mini` | $1.54 | $0.80 | 48 % | free |251| openai | `o3-mini` | $1.54 | $1.05 | 32 % | free |252| openai | `gpt-4.1` | $2.80 | $1.45 | 48 % | free |253| openai | `gpt-4.1-mini` | $0.56 | $0.29 | 48 % | free |254| openai | `gpt-4.1-nano` | $0.14 | $0.0725 | 48 % | free |255| openai | `gpt-4o` | $3.50 | $2.38 | 32 % | free |256| openai | `gpt-4o-mini` | $0.21 | $0.14 | 32 % | free |257| openai | `chat-latest` | $8.00 | $3.95 | 51 % | free |258| openai | `gpt-5.6-cyber` | $20.00 | $11.28 | 44 % | 15.625 |259| anthropic | `claude-fable-5-1` | $15.00 | $7.35 | 51 % | 12.5 (5m write) |260| anthropic | `claude-fable-5` | $15.00 | $8.03 | 46 % | 12.5 (5m write) |261| anthropic | `claude-mythos-5-1` | $15.00 | $7.35 | 51 % | 12.5 (5m write) |262| anthropic | `claude-mythos-5` | $15.00 | $8.03 | 46 % | 12.5 (5m write) |263| anthropic | `claude-opus-5` | $7.50 | $4.01 | 46 % | 6.25 (5m write) |264| anthropic | `claude-opus-4-8` | $7.50 | $4.01 | 46 % | 6.25 (5m write) |265| anthropic | `claude-opus-4-7` | $7.50 | $4.01 | 46 % | 6.25 (5m write) |266| anthropic | `claude-opus-4-6` | $7.50 | $4.01 | 46 % | 6.25 (5m write) |267| anthropic | `claude-opus-4-5-20251101` | $7.50 | $4.01 | 46 % | 6.25 (5m write) |268| anthropic | `claude-sonnet-5` | $3.00 | $1.60 | 46 % | 2.5 (5m write) |269| anthropic | `claude-sonnet-4-6` | $4.50 | $2.41 | 46 % | 3.75 (5m write) |270| anthropic | `claude-sonnet-4-5-20250929` | $4.50 | $2.41 | 46 % | 3.75 (5m write) |271| anthropic | `claude-haiku-4-5-20251001` | $1.50 | $0.80 | 46 % | 1.25 (5m write) |272| xai | `grok-4.6` | $5.20 (long-context rates) | $2.50 | 52 % | none (automatic; cached rate 1) |273| xai | `grok-4.5` | $5.20 (long-context rates) | $2.14 | 59 % | none (automatic; cached rate 0.6) |274| xai | `grok-4.3` | $3.00 (long-context rates) | $1.11 | 63 % | none (automatic; cached rate 0.4) |275| xai | `grok-4.20-0309-reasoning` | $3.00 (long-context rates) | $1.11 | 63 % | none (automatic; cached rate 0.4) |276| xai | `grok-4.20-0309-non-reasoning` | $3.00 (long-context rates) | $1.11 | 63 % | none (automatic; cached rate 0.4) |277| xai | `grok-4.20-multi-agent-0309` | $3.00 (long-context rates) | $1.11 | 63 % | none (automatic; cached rate 0.4) |278| xai | `grok-build-0.1` | $2.40 (long-context rates) | $0.96 | 60 % | none (automatic; cached rate 0.4) |279| gemini | `gemini-3.8-flash` | $1.12 | $0.56 | 50 % | implicit: none; explicit: storage 0.5 /1M/h × 0.9 |280| gemini | `gemini-3.7-flash` | $1.12 | $0.56 | 50 % | implicit: none; explicit: storage 0.5 /1M/h × 0.9 |281| gemini | `gemini-3.6-flash` | $1.12 | $0.56 | 50 % | implicit: none; explicit: storage 0.5 /1M/h × 0.9 |282| gemini | `gemini-3.5-flash` | $2.40 | $1.28 | 47 % | implicit: none; explicit: storage 1 /1M/h × 0.9 |283| gemini | `gemini-3.5-flash-lite` | $0.55 | $0.40 | 28 % | implicit: none; explicit: storage 1 /1M/h × 0.9 |284| gemini | `gemini-3.1-pro-preview` | $5.80 (>200k rates) | $2.96 | 49 % | implicit: none; explicit: storage 4.5 /1M/h × 0.9 |285| gemini | `gemini-3.1-pro-preview-customtools` | $5.80 (>200k rates) | $2.96 | 49 % | implicit: none; explicit: storage 4.5 /1M/h × 0.9 |286| gemini | `gemini-3.1-flash-lite` | $0.40 | $0.29 | 28 % | implicit: none; explicit: storage 1 /1M/h × 0.9 |287| gemini | `gemini-3-flash-preview` | $0.80 | $0.48 | 39 % | implicit: none; explicit: storage 1 /1M/h × 0.9 |288| gemini | `gemini-2.5-pro` | $4.00 (>200k rates) | $2.38 | 40 % | implicit: none; explicit: storage 4.5 /1M/h × 0.9 |289| gemini | `gemini-2.5-flash` | $0.55 | $0.40 | 28 % | implicit: none; explicit: storage 1 /1M/h × 0.9 |290| gemini | `gemini-2.5-flash-lite` | $0.14 | $0.15 | -6 % | implicit: none; explicit: storage 1 /1M/h × 0.9 |291| gemini | `gemini-robotics-er-2-preview` | $1.50 | $0.73 | 51 % | implicit: none; explicit: storage 0.5 /1M/h × 0.9 |292293### Caching break-even by provider294295| Provider | Write cost | Read multiplier | Pays for itself on… | Notes |296|---|---|---|---|---|297| OpenAI (pre-5.6) | free | model-specific 0.25×–0.5× | first hit | implicit; `prompt_cache_key` routing; 5–10 min in-memory, `24h` retention option |298| OpenAI (GPT-5.6+, GPT-6) | 1.25× | 0.1× | second use (1.25 + 0.1 = 1.35 < 2.0) | explicit `prompt_cache_breakpoint` optional; `ttl: 30m` |299| Anthropic | 1.25× (5 m) / 2× (1 h) | 0.1× (0.025× Fable 5.1 / Mythos 5.1) | second use (5 m); 1 h only if the gap between calls exceeds 5 min | explicit `cache_control`; hits refresh TTL for free |300| xAI | none | 0.15×–0.25× | first hit | automatic; `prompt_cache_key` / `x-grok-conv-id` for sticky routing; no TTL or minimum documented; cached tokens still count toward TPM |301| Gemini | implicit: none; explicit: storage per token-hour | 0.1× | implicit: first hit (≥ 4,096-token prompts on 3.x, 2,048 on 2.5); explicit: when `0.9 × read + storage_hour × hours / N < 0.9 × input` — for gemini-3.8-flash one hour of storage ($0.45 per 0.9M) is recovered after a single hit ($0.675 − $0.0675 saved) | explicit cache TTL default 1 h (`ttl` / `expireTime`), min 1,024 tokens (live), free tier `limit: 0` |302303## 8. Batch vs synchronous — the same 100k-request job304305| | OpenAI Batch API | Anthropic Message Batches | xAI Batch API | Gemini Batch API |306|---|---|---|---|---|307| Discount | 50 % on input, cached input and output (`batch` tier) | 50 % on input, output, cache writes and cache reads (stacks with caching) | **20 %** on input, output, cached and reasoning tokens — grok-4.3 / 4.20 / 4.20-multi-agent only; grok-4.6, grok-4.5, grok-build: not supported; Imagine models: standard rates | **50 %** (`service:batch_api`) on text, image, TTS, embedding models; Flex gives the same 50 % synchronously (best-effort) |308| Window | 24 h (`completion_window: "24h"`, only value); output retained 30 days | 24 h expiry (`expires_at`); results 29 days | no completion window documented; results listed via `GET /v1/batches/{id}/results`; batch requests do not count toward rate limits; video URLs in results expire after 1 h | target 24 h ('usually much faster'); poll `GET /v1beta/batches/{id}`; concurrent batch jobs 100; enqueued-token caps per model and tier (e.g. Tier 1: 3M gemini-3.8-flash, 5M gemini-3.1-pro-preview) |309| How requests are supplied | JSONL file (`custom_id`, `method`, `url`, `body`) → `POST /v1/batches {input_file_id, endpoint}`; 50,000 requests / 200 MB; one endpoint + one model per file | inline `requests[] {custom_id, params}` → `POST /v1/messages/batches`; 100,000 requests / 256 MB; models mixed freely | `POST /v1/batches` then `POST /v1/batches/{id}/requests` (inline chat/responses/image/video bodies) or JSONL upload; `:cancel`; no DELETE (405) | `POST /v1beta/models/{model}:batchGenerateContent` with inline `requests[]` or a JSONL file from the Files API (2 GB / 20 GB storage); `:asyncBatchEmbedContent` for embeddings; `/v1beta/openai/batches` compatibility route |310| Example: 100k requests × (2k in + 300 out) on the flagship mid-tier | gpt-5.6-sol batch: 200M × $2 + 30M × $10 = **$700** (vs $1,400 standard) | claude-sonnet-5 batch: 200M × $1 + 30M × $5 = **$350** (vs $700 standard) | grok-4.3 batch: 200M × $1.00 + 30M × $2.00 = **$260** (vs $325 standard); grok-4.6 has no batch: **$430** standard | gemini-3.8-flash batch: 200M × $0.375 + 30M × $1.875 = **$131.25** (vs $262.50 standard); gemini-3.1-pro-preview batch: 200M × $1 + 30M × $6 = **$380** (vs $760) |311| Example: same job on the small models | gpt-5.4-mini batch: 200M × $0.375 + 30M × $2.25 = **$142.50** | claude-haiku-4-5 batch: 200M × $0.5 + 30M × $2.5 = **$175** | grok-build-0.1: no batch, **$260** standard | gemini-3.5-flash-lite batch: 200M × $0.15 + 30M × $1.25 = **$67.50** (vs $135) |312313## 9. Tool and platform prices (from `generated/pricing.json`)314315| Item | OpenAI | Anthropic | xAI | Gemini |316|---|---|---|---|---|317| Web search | $10 / 1k calls (`web_search`, image search); preview $25 / 1k on non-reasoning models | $10 / 1k searches (errors not billed) | **$5 / 1k successful calls** (`web_search`, image search included; `view_image` results billed as image tokens) | Google Search grounding: Gemini 3.x **5,000 free queries / month** (shared) then **$14 / 1k search queries**; Gemini 2.5: 1,500 RPD free then **$35 / 1k grounded prompts**; free tier 500 RPD (not Pro) |318| Social / vertical search | — | — | **`x_search`** $5 / 1k calls until 2026-09-21 12:00 PT, then **$5 / 1k posts fetched + $10 / 1k user profiles fetched** | **Google Maps grounding**: Gemini 3.x 5,000 free prompts / month then $14 / 1k queries; Gemini 2.5: $25 / 1k grounded prompts (1,500 RPD free; 10,000 for Pro) |319| Web fetch / URL context | — | $0 per fetch (content tokens only) | — (web search `open_page` action) | `urlContext` free of charge; retrieved content billed as input tokens (`toolUsePromptTokenCount`); ≤ 20 URLs / request |320| Code execution | container session: 1 GB $0.03 · 4 GB $0.12 · 16 GB $0.48 · 64 GB $1.92 per 20 min (per-minute, 5-min minimum since 2026-06-02); shared with hosted shell | 1,550 free container-hours / org / month, then $0.05 / container-hour (5-min minimum); free when `web_search`/`web_fetch` 20260209+ is in the request | **$5 / 1k calls** (`code_interpreter` / `code_execution`) + tokens | **no per-call fee** (`codeExecution`); generated code/results billed as output then as input when re-read; 30 s per execution |321| File search / RAG | $2.50 / 1k calls + $0.10 / GB / day storage (1 GB free) | — | **`file_search` / `collections_search` $2.50 / 1k calls** + tokens; collections storage **$0.10 / GiB / day**, downloads $0.20 / GiB; `attachment_search` (implicit, files attached to messages) **$10 / 1k calls** | **`fileSearch`**: indexing embeddings **$0.15 / 1M tokens** once; storage and query-time embeddings free; retrieved chunks billed as input tokens; store size by tier (Free 1 GB … Tier 3 1 TB) |322| MCP | tokens only (`mcp` tool) | tokens only (`mcp_servers`, beta) | tokens only (`mcp` tool; tool outputs count as input) | not documented (`mcpServers` UNVERIFIED on generateContent; Interactions `mcp_server` type) |323| Computer use | tokens (`computer` tool) | tokens (toolsets ≈ 4,500 tokens of definitions) | — | tokens only (`computerUse`, PREVIEW; screenshots are image input tokens); not on the free tier |324| Files API storage | not priced separately (`purpose=batch` files expire 30 d) | not priced on the pricing page | **$0.025 / GiB / day** storage, $0.20 / GiB downloads; TTL via `expires_after` | free (2 GB / file, 20 GB / project, 48 h TTL) |325| Managed agents runtime | model tokens at Responses rates + container rates (no session fee) | $0.08 per session-hour (`usage.active_seconds`) + model tokens + $10 / 1k web searches | xAI Responses agentic loop: tokens + per-call tool fees; `usage-guideline-violation` fee **$0.05 / request** when a violation is caught before generation | Interactions agents (Deep Research, Antigravity, custom managed agents): model inference at list rates incl. intermediate/reasoning tokens + tool fees; **sandbox compute not billed during preview** |326| Realtime / voice | gpt-realtime-2.1: audio $32 in / $64 out, text $4 / $24, image $5 per 1M; mini $10 / $20 audio; gpt-live-1 $0.05 / min | — | **grok-voice-think-fast-2.0 $0.08 / min** ($4.80 / h) audio sent or received + $0.004 per text item; STT $0.10 / h REST, $0.20 / h streaming; TTS **$15 / 1M characters** | **gemini-3.8-live**: audio in $3 / 1M (≈ $0.005 / min), audio out $12 / 1M (≈ $0.018 / min), text $0.75 / $4.50; TTS `gemini-3.1-flash-tts-preview` $1 text in / $20 audio out per 1M (≈ $0.03 / min); `gemini-3.5-transcribe` ≈ $0.005 / min blended; free tier available |327| Image generation | gpt-image-2 / 2.5: text in $5, image in $8, image out $30 per 1M tokens (≈ $0.006–$0.21 per image); batch 50 % | — | **grok-imagine-image $0.02 / image**, **grok-imagine-image-2.0 $0.04–$0.08** (quality × resolution), grok-imagine-image-quality $0.05 (DEPRECATED → 2026-11-02); `image_generation` tool billed at these rates | **gemini-3.1-flash-image** $0.045 (0.5K) – $0.151 (4K) per image (image out $60 / 1M), **gemini-3.1-flash-lite-image** $0.034 (1K), **gemini-3-pro-image** $0.134 (1K/2K) / $0.24 (4K); batch 50 % |328| Video | sora-2 $0.10 / s (720p), sora-2-pro $0.30–$0.70 / s — shutting down 2026-09-24 | — | **grok-imagine-video $0.05 / s**, **grok-imagine-video-1.5 $0.08 / s** (1080p, reference-to-video) | **Veo 3.1** $0.40 / s (720p/1080p), $0.60 / s (4K); **Veo 3.1 Fast** $0.10 / $0.12 / $0.30; **Veo 3.1 Lite** $0.05 / $0.08 (no 4K); `gemini-omni-1.1-flash` video out $17.50 / 1M tokens (≈ $0.10 / s) |329| Music | — | — | — | **lyria-3.5 $0.08 / song**, lyria-3-clip-preview $0.04 / 30-s clip, lyria-3-pro-preview $0.08; Lyria RealTime (experimental) unpriced |330| Embeddings | text-embedding-3-small $0.02, -large $0.13 per 1M | — | `grok-embedding-small`: no published price (ACCOUNT_RESTRICTED) | **gemini-embedding-2** text $0.20 / 1M (batch $0.10), image $0.45, audio $6.50, video $12 per 1M; free tier available; gemini-embedding-001 unpriced (DEPRECATED → 2028-05-14) |331| Moderation | free (`omni-moderation-latest`) | — (built-in refusals) | — (built-in; usage-guideline violations still billed + $0.05 fee) | — (`safetySettings` thresholds, free) |332| Token counting | free (`POST /v1/responses/input_tokens`) | free (`POST /v1/messages/count_tokens`, own RPM bucket) | free (`POST /v1/tokenize-text`; no `usage`/cost ticks returned) | free (`POST /v1beta/models/{model}:countTokens`) |333| Data residency | +10 % regional hosts (models ≥ 2026-03-05) | ×1.1 `inference_geo: us`; +10 % Bedrock/Vertex regional | ×1.1 on `us.api.x.ai` (grok-4.6) | not priced on the Developer API (Vertex AI feature) |334| Per-request overhead | none published | tool-use system prompt: e.g. Opus 5 406 tokens (any/tool) / 286 (auto); `bash_20250124` +244–325; `text_editor` +700; `computer_toolset_20260801` ≈4,500; `browser_toolset_20260801` ≈6,600 | none published; `cost_in_usd_ticks` in every `usage` object (1 tick = $1e-10) makes the effective cost observable per call | none published; `usageMetadata` breaks tokens down by modality (`promptTokensDetails[]`), `thoughtsTokenCount`, `toolUsePromptTokenCount`, `cachedContentTokenCount` |335336## 10. Reading prices programmatically337338```bash339# OpenAI standard row for one model340jq '.[] | select(.id=="gpt-5.6-sol") | .pricing.standard' generated/models.json341# Anthropic per-dimension rows342jq '[.[] | select(.provider=="anthropic" and .model_or_service=="claude-sonnet-5")]' generated/pricing.json343# xAI standard + long-context blocks and the live price ticks (1 tick = USD 1e-10 per token)344jq '.[] | select(.id=="grok-4.6") | .pricing | {standard, long_context, batch_discount, priority_multiplier, us_regional_multiplier}' generated/models.json345# Gemini tiers (standard/batch/flex/priority) and free-tier note346jq '.[] | select(.id=="gemini-3.8-flash") | .pricing | {tiers, free_tier, from_2027_01_01}' generated/models.json347# every tool / service price row across the four providers348jq '[.[] | select(.model_or_service|test("^(tool:|service:|agent:|rule:)"))]' generated/pricing.json349```350351Related: [models](models.md) · [caching and reasoning](caching-and-reasoning.md) · [features](features.md) · [realtime and media](realtime-and-media.md) · [FAQ](../faq.md).352