# Pricing — side-by-side and cost models (OpenAI · Anthropic · xAI · Gemini)
Status: every number below is read from generated/models.json (pricing block of each model, sourced from the vendor pricing/model pages on 2026-09-18) and generated/pricing.json (tool/service/rule rows). Figures are USD per 1M tokens unless stated. Statuses next to model ids are the model record statuses. No live billing was reconciled; treat as list prices. Gemini 3.6–3.8 Flash rows are introductory prices through 2026-12-31 (the records carry a from_2027_01_01 block at 2×); xAI x_search switches from per-call to per-post/per-profile billing on 2026-09-21.
Sources: https://developers.openai.com/api/docs/pricing · https://platform.claude.com/docs/en/about-claude/pricing · https://docs.x.ai/developers/pricing (+ live GET /v1/language-models price ticks) · https://ai.google.dev/gemini-api/docs/pricing · docs/openai/pricing.md · docs/anthropic/pricing.md · docs/xai/pricing.md · docs/gemini/pricing.md
Last verified: 2026-09-18
# 1. Pricing dimensions — how the four price lists are structured
| Dimension |
OpenAI |
Anthropic |
xAI |
Gemini |
| Base tokens |
input, cached_input, output per model; separate *_long_context table for >272k-token prompts (2× input, 1.5× output) on GPT-5.4/5.5/5.6/6 Astra |
input, output; 1M context at standard price on Claude 4.6+ (no long-context premium) |
input, cached_input, image_input (= text input price), output; long-context tier at 2× for prompts ≥ 200,000 tokens — billed on ALL tokens of that request (live_raw_ticks.long_context_threshold); reasoning tokens billed as output on every Grok call |
input, cached_input, output per model; modality-specific input rows (audio_input, image_input, video_input); >200k tier (input_over_200k 2×, output_over_200k 1.5×) on Pro models only; thinking tokens billed as output (output_includes_thinking_tokens) |
| Cache write |
free on models before GPT-5.6; 1.25× input on GPT-5.6+ (cache_write row) |
1.25× input for 5-minute TTL, 2× input for 1-hour TTL |
none — automatic prefix cache, no write charge |
implicit cache: none; explicit cachedContents: storage per 1M tokens per hour (cache_storage_hour: $0.50 Flash 3.6–3.8, $1.00 most Flash, $4.50 Pro; 1.8× on priority) |
| Cache read |
model-specific cached_input (0.1× on GPT-5.6+, 0.25× on o3, 0.5× on gpt-4o…) |
0.1× input (0.025× on Fable 5.1 / Mythos 5.1) |
cached_input per model: 0.25× (grok-4.6 $0.50), 0.15× (grok-4.5 $0.30), 0.16× (grok-4.3 / 4.20 $0.20), 0.2× (grok-build $0.20) |
cached_input = 0.1× input on every priced model (e.g. gemini-3.8-flash $0.075, gemini-3.1-pro-preview $0.20) |
| Batch |
batch tier = 50 % of standard (own cached_input) |
batch_input/batch_output = 50 % of standard; cache multipliers stack |
20 % off (batch_discount: 0.2) on grok-4.3 / 4.20 / 4.20-multi-agent only; not supported on grok-4.6, grok-4.5, grok-build-0.1 (batch_discount: null); Imagine models accepted in Batch but billed at standard rates |
batch tier = 50 % of standard (service:batch_api multiplier 0.5) on text, image, TTS and embedding models; not on the free tier |
| Cheaper async / best-effort tier |
flex = batch rates on synchronous calls (slower, may 429) |
none |
none (Batch API is the only discount) |
flex = 50 % of standard (service:flex_inference), 1–15 min target latency, sheddable (429 when capacity is shed); ten text models listed |
| Faster / priority tier |
fast = 2× standard (renamed from priority 2026-07-30; fast_long_context 4×/3×) |
speed: fast (Opus 5 / 4.8 only) = 2× standard ($10 / $50) |
service_tier: priority = 2× on all token types (caching discount applied before the multiplier; billed only when the response echoes service_tier: priority) |
priority tier = 1.8× standard (service:priority_inference, docs: 75–100 % more); 0.3× the standard rate limit; graceful downgrade to standard when exceeded |
| Reasoning tokens |
billed as output (usage.output_tokens_details.reasoning_tokens) |
billed as output (usage.output_tokens_details.thinking_tokens) |
billed as output (usage.completion_tokens_details.reasoning_tokens); every Grok reasoning model spends them even on reasoning_effort: low |
billed as output (usageMetadata.thoughtsTokenCount); Gemini 3.x thinking cannot be fully disabled on Flash/Pro (thinkingLevel: minimal where supported) |
| Regional inference |
+10 % on regional hosts for models released ≥ 2026-03-05 |
inference_geo: us ×1.1 on all token dimensions (Claude 4.6+); Bedrock/Vertex regional +10 % |
https://us.api.x.ai/v1 = 1.1× global token rates (grok-4.6 only in the live catalogue: $2.20 / $0.55 / $6.60); eu-west-1.api.x.ai LIVE_DISCOVERED, no separate price row |
no regional price row on the Gemini Developer API (data residency is a Vertex AI feature); AI Studio usage free |
| Free tier |
none (Free usage tier = rate-limit tier, still billed) |
none |
none (prepaid credits; Tier 0 = $0 spend) |
yes — free_tier rows on Flash / Flash-Lite / Live / TTS / transcribe / embedding / Gemma models: input & output $0, restrictive limits, content may be used to improve Google products; Pro models and paid-only media models are not available on the free tier (observed limit: 0) |
| Tool-definition overhead |
none published (definitions are input tokens) |
published per model: tool-use system prompt 286–804 tokens, toolsets 4,500–6,600 tokens |
none published (definitions are input tokens; server-side tool results count as input) |
none published; toolUsePromptTokenCount reports tokens consumed by URL context / grounding results |
| Per-image / per-second media |
image models per token (≈ per image), Sora per second |
— |
Imagine per image ($0.02 / $0.04–$0.08 / $0.05) and per second of video ($0.05 / $0.08); voice per minute; TTS per 1M characters; STT per hour |
image models per token with a per-image equivalent ($0.034–$0.24), Veo per second ($0.05–$0.60), Lyria per song ($0.04 / $0.08), Live/TTS/transcribe per 1M audio tokens (≈ per minute) |
# 2. OpenAI current text models (USD / 1M tokens)
| Model |
Status |
Input |
Cached in |
Cache write |
Output |
Batch in / out |
Flex in / out |
Fast in / out |
Long-context in / out (>272k) |
gpt-6-astra |
DOCUMENTED · LIVE_VERIFIED |
10 |
1 |
12.5 |
50 |
5 / 25 |
5 / 25 |
20 / 100 |
20 / 75 |
gpt-5.6-sol |
DOCUMENTED · LIVE_VERIFIED |
4 |
0.4 |
5 |
20 |
2 / 10 |
2 / 10 |
8 / 40 |
8 / 30 |
gpt-5.6-terra |
DOCUMENTED · LIVE_VERIFIED |
2 |
0.2 |
2.5 |
12 |
1 / 6 |
1 / 6 |
4 / 24 |
4 / 18 |
gpt-5.6-luna |
DOCUMENTED · LIVE_VERIFIED |
0.2 |
0.02 |
0.25 |
1.2 |
0.1 / 0.6 |
0.1 / 0.6 |
0.4 / 2.4 |
0.4 / 1.8 |
gpt-5.5 |
DOCUMENTED · LIVE_VERIFIED |
5 |
0.5 |
— |
30 |
2.5 / 15 |
2.5 / 15 |
12.5 / 75 |
10 / 45 |
gpt-5.5-pro |
DOCUMENTED · LIVE_VERIFIED |
30 |
— |
— |
180 |
15 / 90 |
15 / 90 |
— / — |
60 / 270 |
gpt-5.4 |
DOCUMENTED · LIVE_VERIFIED |
2.5 |
0.25 |
— |
15 |
1.25 / 7.5 |
1.25 / 7.5 |
5 / 30 |
5 / 22.5 |
gpt-5.4-pro |
DOCUMENTED · LIVE_VERIFIED |
30 |
— |
— |
180 |
15 / 90 |
15 / 90 |
— / — |
60 / 270 |
gpt-5.4-mini |
DOCUMENTED · LIVE_VERIFIED |
0.75 |
0.075 |
— |
4.5 |
0.375 / 2.25 |
0.375 / 2.25 |
1.5 / 9 |
— / — |
gpt-5.4-nano |
DOCUMENTED · LIVE_VERIFIED |
0.2 |
0.02 |
— |
1.25 |
0.1 / 0.625 |
0.1 / 0.625 |
— / — |
— / — |
gpt-5.3-codex |
DOCUMENTED · LIVE_VERIFIED |
1.75 |
0.175 |
— |
14 |
— / — |
— / — |
3.5 / 28 |
— / — |
gpt-5.2 |
DOCUMENTED · LIVE_VERIFIED |
1.75 |
0.175 |
— |
14 |
0.875 / 7 |
0.875 / 7 |
3.5 / 28 |
— / — |
gpt-5.2-pro |
DOCUMENTED · LIVE_VERIFIED |
21 |
— |
— |
168 |
10.5 / 84 |
— / — |
— / — |
— / — |
gpt-5.1 |
DOCUMENTED · LIVE_VERIFIED |
1.25 |
0.125 |
— |
10 |
0.625 / 5 |
0.625 / 5 |
2.5 / 20 |
— / — |
gpt-5 |
DOCUMENTED · LIVE_VERIFIED |
1.25 |
0.125 |
— |
10 |
0.625 / 5 |
0.625 / 5 |
2.5 / 20 |
— / — |
gpt-5-mini |
DOCUMENTED · LIVE_VERIFIED |
0.25 |
0.025 |
— |
2 |
0.125 / 1 |
0.125 / 1 |
0.45 / 3.6 |
— / — |
gpt-5-nano |
DOCUMENTED · LIVE_VERIFIED |
0.05 |
0.005 |
— |
0.4 |
0.025 / 0.2 |
0.025 / 0.2 |
— / — |
— / — |
gpt-5-pro |
DOCUMENTED · LIVE_VERIFIED |
15 |
— |
— |
120 |
7.5 / 60 |
— / — |
— / — |
— / — |
o3 |
DOCUMENTED · LIVE_VERIFIED |
2 |
0.5 |
— |
8 |
1 / 4 |
1 / 4 |
3.5 / 14 |
— / — |
o3-pro |
DOCUMENTED · LIVE_VERIFIED |
20 |
— |
— |
80 |
10 / 40 |
— / — |
— / — |
— / — |
o4-mini |
DOCUMENTED · LIVE_VERIFIED · DEPRECATED |
1.1 |
0.275 |
— |
4.4 |
0.55 / 2.2 |
0.55 / 2.2 |
2 / 8 |
— / — |
o3-mini |
DOCUMENTED · LIVE_VERIFIED · DEPRECATED |
1.1 |
0.55 |
— |
4.4 |
0.55 / 2.2 |
— / — |
— / — |
— / — |
gpt-4.1 |
DOCUMENTED · LIVE_VERIFIED |
2 |
0.5 |
— |
8 |
1 / 4 |
— / — |
3.5 / 14 |
— / — |
gpt-4.1-mini |
DOCUMENTED · LIVE_VERIFIED |
0.4 |
0.1 |
— |
1.6 |
0.2 / 0.8 |
— / — |
0.7 / 2.8 |
— / — |
gpt-4.1-nano |
DOCUMENTED · LIVE_VERIFIED · DEPRECATED |
0.1 |
0.025 |
— |
0.4 |
0.05 / 0.2 |
— / — |
0.2 / 0.8 |
— / — |
gpt-4o |
DOCUMENTED · LIVE_VERIFIED |
2.5 |
1.25 |
— |
10 |
1.25 / 5 |
— / — |
4.25 / 17 |
— / — |
gpt-4o-mini |
DOCUMENTED · LIVE_VERIFIED |
0.15 |
0.075 |
— |
0.6 |
0.075 / 0.3 |
— / — |
0.25 / 1 |
— / — |
chat-latest |
DOCUMENTED · LIVE_VERIFIED |
5 |
0.5 |
— |
30 |
— / — |
— / — |
— / — |
— / — |
gpt-5.6-cyber |
DOCUMENTED · ACCOUNT_RESTRICTED |
12.5 |
1.25 |
15.625 |
75 |
— / — |
— / — |
— / — |
— / — |
Notes from the records: gpt-5.6 is an alias of gpt-5.6-sol (promotional $4/$20 at least through 2026-11-21); pro models have no cached-input price (no caching); chat-latest has no batch/flex/fast tiers; gpt-5.6-cyber is ACCOUNT_RESTRICTED and priced at short-context rates. The o3 record carries a model-page table ($1 / $0.25 / $4) that disagrees with its pricing-page standard row ($2 / $0.5 / $8) — flagged as a data inconsistency.
# 3. Anthropic current models (USD / 1M tokens)
| Model |
Status |
Input |
Cache write 5m |
Cache write 1h |
Cache read |
Output |
Batch in / out |
Fast in / out |
Min cacheable tokens |
claude-fable-5-1 |
DOCUMENTED · LIVE_VERIFIED |
10 |
12.5 |
20 |
0.25 |
50 |
5 / 25 |
— / — |
512 |
claude-fable-5 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
10 |
12.5 |
20 |
1 |
50 |
5 / 25 |
— / — |
512 |
claude-mythos-5-1 |
DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW |
10 |
12.5 |
20 |
0.25 |
50 |
5 / 25 |
— / — |
512 |
claude-mythos-5 |
DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW |
10 |
12.5 |
20 |
1 |
50 |
5 / 25 |
— / — |
512 |
claude-opus-5 |
DOCUMENTED · LIVE_VERIFIED |
5 |
6.25 |
10 |
0.5 |
25 |
2.5 / 12.5 |
10 / 50 |
512 |
claude-opus-4-8 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
5 |
6.25 |
10 |
0.5 |
25 |
2.5 / 12.5 |
10 / 50 |
1024 |
claude-opus-4-7 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
5 |
6.25 |
10 |
0.5 |
25 |
2.5 / 12.5 |
— / — |
2048 |
claude-opus-4-6 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
5 |
6.25 |
10 |
0.5 |
25 |
2.5 / 12.5 |
— / — |
4096 |
claude-opus-4-5-20251101 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
5 |
6.25 |
10 |
0.5 |
25 |
2.5 / 12.5 |
— / — |
4096 |
claude-sonnet-5 |
DOCUMENTED · LIVE_VERIFIED |
2 |
2.5 |
4 |
0.2 |
10 |
1 / 5 |
— / — |
1024 |
claude-sonnet-4-6 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
3 |
3.75 |
6 |
0.3 |
15 |
1.5 / 7.5 |
— / — |
1024 |
claude-sonnet-4-5-20250929 |
DOCUMENTED · LIVE_VERIFIED · LEGACY |
3 |
3.75 |
6 |
0.3 |
15 |
1.5 / 7.5 |
— / — |
1024 |
claude-haiku-4-5-20251001 |
DOCUMENTED · LIVE_VERIFIED |
1 |
1.25 |
2 |
0.1 |
5 |
0.5 / 2.5 |
— / — |
4096 |
Notes: Sonnet 5's introductory $2 / $10 was made permanent on 2026-08-10. Fable 5.1 / Mythos 5.1 cache reads are 0.025× ($0.25). Retired models (Opus 4.1 $15/$75, Sonnet 4 $3/$15, Haiku 3.5 $0.80/$4) remain priced on Bedrock/Vertex only. Priority Tier is no longer sold; Claude Platform on AWS and Foundry convert the same USD rates to CCUs at $0.01.
# 4. xAI current Grok text models (USD / 1M tokens)
Standard = prompt < 200,000 tokens. Long context = prompt ≥ 200,000 tokens: the higher rate applies to every token of that request (input, cached, output). image_input equals the text input price on every model. Batch = 20 % off standard where supported. Priority = service_tier: priority, 2×. US regional = https://us.api.x.ai/v1, 1.1×.
| Model |
Status |
Input |
Cached in |
Output |
Long-ctx in / cached / out (≥200k) |
Batch in / cached / out |
Priority × |
US regional × |
Context |
grok-4.6 |
DOCUMENTED · LIVE_VERIFIED |
2 |
0.5 |
6 |
4 / 1 / 12 |
not supported |
2.0× |
1.1× |
500000 |
grok-4.5 |
DOCUMENTED · LIVE_VERIFIED |
2 |
0.3 |
6 |
4 / 0.6 / 12 |
not supported |
2.0× |
— |
500000 |
grok-4.3 |
DOCUMENTED · LIVE_VERIFIED |
1.25 |
0.2 |
2.5 |
2.5 / 0.4 / 5 |
1 / 0.16 / 2 (−20 %) |
2.0× |
— |
1000000 |
grok-4.20-0309-reasoning |
DOCUMENTED · LIVE_VERIFIED |
1.25 |
0.2 |
2.5 |
2.5 / 0.4 / 5 |
1 / 0.16 / 2 (−20 %) |
2.0× |
— |
1000000 |
grok-4.20-0309-non-reasoning |
DOCUMENTED · LIVE_VERIFIED |
1.25 |
0.2 |
2.5 |
2.5 / 0.4 / 5 |
1 / 0.16 / 2 (−20 %) |
2.0× |
— |
1000000 |
grok-4.20-multi-agent-0309 |
DOCUMENTED · BETA · LIVE_VERIFIED |
1.25 |
0.2 |
2.5 |
2.5 / 0.4 / 5 |
1 / 0.16 / 2 (−20 %) |
2.0 (Chat Completions/Responses; not documented per model) |
— |
1000000 |
grok-build-0.1 |
DOCUMENTED · PREVIEW · LIVE_VERIFIED |
1 |
0.2 |
2 |
2 / 0.4 / 4 |
not supported |
2.0× |
— |
256000 |
Notes from the records: grok-4.20-multi-agent-0309 is BETA and only reachable on /v1/responses; its reasoning.effort selects the agent count (low/medium = 4, high/xhigh = 16) and all agents' tokens are billed. grok-build-0.1 (PREVIEW) is the coding model behind the grok-code-fast-1 redirect. The retired ids grok-3, grok-4-0709, grok-4-fast-*, grok-4-1-fast-* redirect to grok-4.3 and are billed at grok-4.3 rates. grok-embedding-small has no published price (ACCOUNT_RESTRICTED). Max output is not documented per model (max_output: null in every Grok record; Responses max_output_tokens defaults to 128,000).
| Model / service |
Status |
Price |
Unit / notes |
grok-imagine-image |
DOCUMENTED · LIVE_VERIFIED |
$0.02 per image |
single price; batch discount 0.0 (standard rates in Batch) |
grok-imagine-image-2.0 |
DOCUMENTED · LIVE_VERIFIED |
$0.06 per image |
low/1k $0.04; low/2k $0.06; low/1.5k $0.05; medium/1k $0.06; medium/2k $0.08; medium/1.5k $0.07; batch discount 0.0 (standard rates in Batch) |
grok-imagine-image-quality |
DOCUMENTED · DEPRECATED · LIVE_VERIFIED |
$0.05 per image |
single price; batch discount 0.0 (standard rates in Batch) |
grok-imagine-video |
DOCUMENTED · LIVE_VERIFIED |
$0.05 per second of generated video |
video URLs in batch results expire after 1 hour |
grok-imagine-video-1.5 |
DOCUMENTED · LIVE_VERIFIED |
$0.08 per second of generated video |
video URLs in batch results expire after 1 hour |
grok-voice-think-fast-2.0 (speech-to-speech) |
DOCUMENTED |
$0.08 per minute ($4.8 / h) + $0.004 per text item |
audio sent or received billed per minute; each conversation.item.create text item billed $0.004 except function_call_output and audio items |
grok-voice-transcribe-2.0 (speech-to-text) |
DOCUMENTED |
$0.1 / hour REST (POST /v1/stt), $0.2 / hour streaming (wss://api.x.ai/v1/stt) |
per hour of audio |
text-to-speech (POST /v1/tts, wss://api.x.ai/v1/tts) |
DOCUMENTED |
$15 per 1M characters |
POST /v1/tts and streaming |
# 5. Gemini current text models (USD / 1M tokens)
Standard tier for prompts ≤ 200k tokens; >200k columns only exist on Pro models (Flash models have one price for any prompt length). Cached in = implicit or explicit cache hit (0.1×). Storage = explicit cachedContents storage per 1M tokens per hour. Batch = 50 %, Flex = 50 %, Priority = 1.8×. Free = free-tier availability recorded on the pricing page.
| Model |
Status |
Input |
Cached in |
Output |
>200k in / cached / out |
Batch in / out |
Flex in / out |
Priority in / out |
Cache storage $/1M/h |
Free tier |
gemini-3.8-flash |
DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED |
0.75 |
0.075 |
3.75 |
— (flat) |
0.375 / 1.875 |
0.375 / 1.875 |
1.35 / 6.75 |
0.5 |
input/output/caching free of charge (standard & priority rows); Bat… |
gemini-3.7-flash |
DOCUMENTED · LIVE_DISCOVERED |
0.75 |
0.075 |
3.75 |
— (flat) |
0.375 / 1.875 |
0.375 / 1.875 |
1.35 / 6.75 |
0.5 |
input/output/caching free of charge (standard & priority rows); Bat… |
gemini-3.6-flash |
DOCUMENTED · LIVE_DISCOVERED |
0.75 |
0.075 |
3.75 |
— (flat) |
0.375 / 1.875 |
0.375 / 1.875 |
1.35 / 6.75 |
0.5 |
input/output/caching free of charge (standard & priority rows); Bat… |
gemini-3.5-flash |
DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED |
1.5 |
0.15 |
9 |
— (flat) |
0.75 / 4.5 |
0.75 / 4.5 |
2.7 / 16.2 |
1 |
input/output/caching free of charge; Batch/Flex not available |
gemini-3.5-flash-lite |
DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED |
0.3 |
0.03 |
2.5 |
— (flat) |
0.15 / 1.25 |
0.15 / 1.25 |
0.54 / 4.5 |
1 |
input/output free of charge on all four rows (page lists Batch/Flex… |
gemini-3.1-pro-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW · ACCOUNT_RESTRICTED |
2 |
0.2 |
12 |
4 / 0.4 / 18 |
1 / 6 |
1 / 6 |
3.6 / 21.6 |
4.5 |
Not available (paid tier only) |
gemini-3.1-pro-preview-customtools |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW |
2 |
0.2 |
12 |
4 / 0.4 / 18 |
1 / 6 |
1 / 6 |
3.6 / 21.6 |
4.5 |
Not available (paid tier only) |
gemini-3.1-flash-lite |
DOCUMENTED · LIVE_DISCOVERED · DEPRECATED |
0.25 |
0.025 |
1.5 |
— (flat) |
0.125 / 0.75 |
0.125 / 0.75 |
0.45 / 2.7 |
1 |
input/output free of charge; caching not available on free tier |
gemini-3-flash-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW |
0.5 |
0.05 |
3 |
— (flat) |
0.25 / 1.5 |
0.25 / 1.5 |
0.9 / 5.4 |
1 |
input/output/caching free of charge; Batch/Flex not available |
gemini-2.5-pro |
DOCUMENTED · LIVE_DISCOVERED |
1.25 |
0.125 |
10 |
2.5 / 0.25 / 15 |
0.625 / 5 |
0.625 / 5 |
2.25 / 18 |
4.5 |
input/output free of charge (standard/priority rows); caching, Batc… |
gemini-2.5-flash |
DOCUMENTED · LIVE_DISCOVERED |
0.3 |
0.03 |
2.5 |
— (flat) |
0.15 / 1.25 |
0.15 / 1.25 |
0.54 / 4.5 |
1 |
input/output free of charge; caching not available; Google Search g… |
gemini-2.5-flash-lite |
DOCUMENTED · LIVE_DISCOVERED · ACCOUNT_RESTRICTED |
0.1 |
0.01 |
0.4 |
— (flat) |
0.05 / 0.2 |
0.05 / 0.2 |
0.18 / 0.72 |
1 |
input/output free of charge; caching not available; Google Search g… |
gemini-robotics-er-2-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW |
1 |
0.1 |
5 |
— (flat) |
0.5 / 2.5 |
— / — |
— / — |
0.5 |
input/output free of charge; caching not available |
gemini-2.5-computer-use-preview-10-2025 |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW |
1 |
— |
5 |
— (flat) |
— / — |
— / — |
— / — |
— |
input/output free of charge |
gemma-4-31b-it |
DOCUMENTED · LIVE_DISCOVERED |
free |
— |
free |
— |
— |
— |
— |
— |
free of charge (input, output, caching, storage) |
gemma-4-26b-a4b-it |
DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED |
free |
— |
free |
— |
— |
— |
— |
— |
free of charge (input, output, caching, storage) |
Notes from the records: gemini-3.8 / 3.7 / 3.6 Flash share one introductory price list ($0.75 / $0.075 / $3.75) through 2026-12-31, doubling on 2027-01-01 ($1.50 / $0.15 / $7.50, storage $1.00); gemini-3.5-flash is the higher-priced Flash ($1.50 / $9.00). Pro models (gemini-3.1-pro-preview, gemini-2.5-pro) are paid-tier only (ACCOUNT_RESTRICTED on this free-tier key). Gemini 2.5 models are 'no longer available to new users' (404 on generateContent). Gemma 4 models are free of charge with no paid tier. gemini-2.5-flash / -pro have separate audio_input rows ($1.00 / 1M); Gemini 3.x prices text, image, video and audio input alike. Google Search grounding: Gemini 3.x 5,000 free requests/month then $14 / 1k search queries; Gemini 2.5: 1,500 RPD free then $35 / 1k grounded prompts.
| Model |
Status |
Price |
Notes |
gemini-3.1-flash-image |
DOCUMENTED · LIVE_DISCOVERED |
text in $0.5, text out $3, image out $60 / 1M → per image 0.5K (747 tok) $0.045, 1K (1120 tok) $0.067, 2K (1680 tok) $0.101, 4K (2520 tok) $0.151; batch image out $30 |
not available |
gemini-3.1-flash-lite-image |
DOCUMENTED · LIVE_DISCOVERED |
text in $0.25, image out $30 / 1M → 1K (1120 tok) $0.0336 |
not available |
gemini-3-pro-image |
DOCUMENTED · LIVE_DISCOVERED |
text in $2, image in $0.0011 per image, image out $120 / 1M → 1K/2K (1120 tok) $0.134, 4K (2000 tok) $0.24; priority image out $216 |
not available |
gemini-2.5-flash-image |
DOCUMENTED · LIVE_DISCOVERED · DEPRECATED |
$0.039 per image ($30 / 1M, 1,290 tok/image); batch $0.0195 |
1290 tokens per image up to 1024x1024 |
veo-3.1-generate-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW |
per second: 720p $0.4, 1080p $0.4, 4k $0.6 |
charged only if the video is successfully generated |
veo-3.1-fast-generate-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW |
per second: 720p $0.1, 1080p $0.12, 4k $0.3 |
not available |
veo-3.1-lite-generate-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW · LIVE_VERIFIED |
per second: 720p $0.05, 1080p $0.08, 4k not supported |
not available |
gemini-omni-1.1-flash |
DOCUMENTED · LIVE_DISCOVERED |
text in $1.5, text out $9, video out $17.5 / 1M |
5,792 tokens per second of 720p video => ~$0.10 per second |
lyria-3.5 |
DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED |
$0.08 per song (full length) |
not available |
lyria-3-clip-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW |
$0.04 per song (30 s clip) |
not available |
gemini-3.1-flash-tts-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW |
text in $1, audio out $20 / 1M (25 audio tok/s ≈ $0.03 / min); batch $0.5 / $10 |
25 audio tokens per second |
gemini-2.5-flash-preview-tts |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW |
text in $0.5, audio out $10 / 1M |
free of charge (standard) |
gemini-3.8-live |
DOCUMENTED · LIVE_DISCOVERED |
text in $0.75, audio in $3 (≈ $0.005 / min), image/video in $1, text out $4.5, audio out $12 (≈ $0.018 / min) |
free of charge |
gemini-2.5-flash-native-audio-preview-12-2025 |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW |
text in $0.5, audio/video in $3, text out $2, audio out $12 |
free of charge |
gemini-3.5-transcribe |
DOCUMENTED · LIVE_DISCOVERED |
audio in $2 (≈ $0.003 / min), text out $12 (≈ $0.002 / min) |
blended ~$0.005/min |
gemini-3.5-transcribe-live |
DOCUMENTED · LIVE_DISCOVERED |
audio in $3.5 (≈ $0.005 / min), text out $21 |
25 audio tokens/s input, 175 text tokens/min output; blended ~$0.009/min |
gemini-3.5-live-translate-preview |
DOCUMENTED · LIVE_DISCOVERED · PREVIEW |
audio in $3.5, audio out $21 / 1M |
25 audio tokens/second; effective ~$0.0368 per minute |
gemini-embedding-2 |
DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED |
text $0.2, image $0.45 ($0.0001 per image), audio $6.5 ($0.0002 / s), video $12 ($0.0008 / frame) / 1M; batch text $0.1 |
free of charge (standard); batch not available |
gemini-embedding-001 |
DOCUMENTED · LIVE_DISCOVERED · DEPRECATED |
— |
not listed on pricing.md (2026-09-18); file-search.md bills indexing embeddings at $0.15 per 1M tokens |
Formula: cost = 1.0 × input_price + 0.1 × output_price, treating the 1M input tokens as aggregate volume at the standard (short-prompt) rate. The last column shows what a single request whose prompt crosses the long-context threshold costs instead (OpenAI >272k, xAI ≥200k applied to all tokens, Gemini Pro >200k; Anthropic and Gemini Flash have no threshold). Tiers shown where the model record lists them.
| Provider |
Model |
Standard |
Batch |
Flex |
Fast / Priority |
Long-context request |
| openai |
gpt-6-astra |
$15.00 |
$7.50 |
$7.50 |
$30.00 |
$27.50 |
| openai |
gpt-5.6-sol |
$6.00 |
$3.00 |
$3.00 |
$12.00 |
$11.00 |
| openai |
gpt-5.6-terra |
$3.20 |
$1.60 |
$1.60 |
$6.40 |
$5.80 |
| openai |
gpt-5.6-luna |
$0.32 |
$0.16 |
$0.16 |
$0.64 |
$0.58 |
| openai |
gpt-5.5 |
$8.00 |
$4.00 |
$4.00 |
$20.00 |
$14.50 |
| openai |
gpt-5.5-pro |
$48.00 |
$24.00 |
$24.00 |
— |
$87.00 |
| openai |
gpt-5.4 |
$4.00 |
$2.00 |
$2.00 |
$8.00 |
$7.25 |
| openai |
gpt-5.4-pro |
$48.00 |
$24.00 |
$24.00 |
— |
$87.00 |
| openai |
gpt-5.4-mini |
$1.20 |
$0.60 |
$0.60 |
$2.40 |
— |
| openai |
gpt-5.4-nano |
$0.33 |
$0.16 |
$0.16 |
— |
— |
| openai |
gpt-5.3-codex |
$3.15 |
— |
— |
$6.30 |
— |
| openai |
gpt-5.2 |
$3.15 |
$1.58 |
$1.58 |
$6.30 |
— |
| openai |
gpt-5.2-pro |
$37.80 |
$18.90 |
— |
— |
— |
| openai |
gpt-5.1 |
$2.25 |
$1.12 |
$1.12 |
$4.50 |
— |
| openai |
gpt-5 |
$2.25 |
$1.12 |
$1.12 |
$4.50 |
— |
| openai |
gpt-5-mini |
$0.45 |
$0.23 |
$0.23 |
$0.81 |
— |
| openai |
gpt-5-nano |
$0.0900 |
$0.0450 |
$0.0450 |
— |
— |
| openai |
gpt-5-pro |
$27.00 |
$13.50 |
— |
— |
— |
| openai |
o3 |
$2.80 |
$1.40 |
$1.40 |
$4.90 |
— |
| openai |
o3-pro |
$28.00 |
$14.00 |
— |
— |
— |
| openai |
o4-mini |
$1.54 |
$0.77 |
$0.77 |
$2.80 |
— |
| openai |
o3-mini |
$1.54 |
$0.77 |
— |
— |
— |
| openai |
gpt-4.1 |
$2.80 |
$1.40 |
— |
$4.90 |
— |
| openai |
gpt-4.1-mini |
$0.56 |
$0.28 |
— |
$0.98 |
— |
| openai |
gpt-4.1-nano |
$0.14 |
$0.0700 |
— |
$0.28 |
— |
| openai |
gpt-4o |
$3.50 |
$1.75 |
— |
$5.95 |
— |
| openai |
gpt-4o-mini |
$0.21 |
$0.10 |
— |
$0.35 |
— |
| openai |
chat-latest |
$8.00 |
— |
— |
— |
— |
| openai |
gpt-5.6-cyber |
$20.00 |
— |
— |
— |
— |
| anthropic |
claude-fable-5-1 |
$15.00 |
$7.50 |
— |
— |
same as standard (no premium) |
| anthropic |
claude-fable-5 |
$15.00 |
$7.50 |
— |
— |
same as standard (no premium) |
| anthropic |
claude-mythos-5-1 |
$15.00 |
$7.50 |
— |
— |
same as standard (no premium) |
| anthropic |
claude-mythos-5 |
$15.00 |
$7.50 |
— |
— |
same as standard (no premium) |
| anthropic |
claude-opus-5 |
$7.50 |
$3.75 |
— |
$15.00 |
same as standard (no premium) |
| anthropic |
claude-opus-4-8 |
$7.50 |
$3.75 |
— |
$15.00 |
same as standard (no premium) |
| anthropic |
claude-opus-4-7 |
$7.50 |
$3.75 |
— |
— |
same as standard (no premium) |
| anthropic |
claude-opus-4-6 |
$7.50 |
$3.75 |
— |
— |
same as standard (no premium) |
| anthropic |
claude-opus-4-5-20251101 |
$7.50 |
$3.75 |
— |
— |
same as standard (no premium) |
| anthropic |
claude-sonnet-5 |
$3.00 |
$1.50 |
— |
— |
same as standard (no premium) |
| anthropic |
claude-sonnet-4-6 |
$4.50 |
$2.25 |
— |
— |
same as standard (no premium) |
| anthropic |
claude-sonnet-4-5-20250929 |
$4.50 |
$2.25 |
— |
— |
same as standard (no premium) |
| anthropic |
claude-haiku-4-5-20251001 |
$1.50 |
$0.75 |
— |
— |
same as standard (no premium) |
| xai |
grok-4.6 |
$2.60 |
not supported |
— |
$5.20 (priority) |
$5.20 |
| xai |
grok-4.5 |
$2.60 |
not supported |
— |
$5.20 (priority) |
$5.20 |
| xai |
grok-4.3 |
$1.50 |
$1.20 |
— |
$3.00 (priority) |
$3.00 |
| xai |
grok-4.20-0309-reasoning |
$1.50 |
$1.20 |
— |
$3.00 (priority) |
$3.00 |
| xai |
grok-4.20-0309-non-reasoning |
$1.50 |
$1.20 |
— |
$3.00 (priority) |
$3.00 |
| xai |
grok-4.20-multi-agent-0309 |
$1.50 |
$1.20 |
— |
$3.00 (priority) |
$3.00 |
| xai |
grok-build-0.1 |
$1.20 |
not supported |
— |
$2.40 (priority) |
$2.40 |
| gemini |
gemini-3.8-flash |
$1.12 |
$0.56 |
$0.56 |
$2.03 (priority) |
same as standard (flat) |
| gemini |
gemini-3.7-flash |
$1.12 |
$0.56 |
$0.56 |
$2.03 (priority) |
same as standard (flat) |
| gemini |
gemini-3.6-flash |
$1.12 |
$0.56 |
$0.56 |
$2.03 (priority) |
same as standard (flat) |
| gemini |
gemini-3.5-flash |
$2.40 |
$1.20 |
$1.20 |
$4.32 (priority) |
same as standard (flat) |
| gemini |
gemini-3.5-flash-lite |
$0.55 |
$0.28 |
$0.28 |
$0.99 (priority) |
same as standard (flat) |
| gemini |
gemini-3.1-pro-preview |
$3.20 |
$1.60 |
$1.60 |
$5.76 (priority) |
$5.80 |
| gemini |
gemini-3.1-pro-preview-customtools |
$3.20 |
$1.60 |
$1.60 |
$5.76 (priority) |
$5.80 |
| gemini |
gemini-3.1-flash-lite |
$0.40 |
$0.20 |
$0.20 |
$0.72 (priority) |
same as standard (flat) |
| gemini |
gemini-3-flash-preview |
$0.80 |
$0.40 |
$0.40 |
$1.44 (priority) |
same as standard (flat) |
| gemini |
gemini-2.5-pro |
$2.25 |
$1.12 |
$1.12 |
$4.05 (priority) |
$4.00 |
| gemini |
gemini-2.5-flash |
$0.55 |
$0.28 |
$0.28 |
$0.99 (priority) |
same as standard (flat) |
| gemini |
gemini-2.5-flash-lite |
$0.14 |
$0.0700 |
$0.0700 |
$0.25 (priority) |
same as standard (flat) |
| gemini |
gemini-robotics-er-2-preview |
$1.50 |
$0.75 |
— |
— (priority) |
same as standard (flat) |
| gemini |
gemini-2.5-computer-use-preview-10-2025 |
$1.50 |
— |
— |
— (priority) |
same as standard (flat) |
Gemma 4 (gemma-4-31b-it, gemma-4-26b-a4b-it) and every Gemini free-tier row cost $0 at list price but carry free-tier limits and data-use terms; they are omitted from the arithmetic.
# 7. Cost model B — the same workload with a 900k-token cached prefix (steady state)
Assumptions: each request = 900k tokens read from cache + 100k fresh input + 100k output; the initial cache write is amortised over N = 10 requests. Formula per request: 0.9 × cache_read + 0.1 × input + 0.1 × output + write_cost / N. Write cost: OpenAI pre-5.6 free, GPT-5.6+ cache_write × 0.9; Anthropic 5-minute write (1.25×) × 0.9; xAI none (automatic cache); Gemini implicit caching none, explicit cachedContents = storage of 0.9M tokens for one hour (cache_storage_hour × 0.9) amortised over N. A 900k prefix exceeds xAI's 200k threshold, so xAI is shown at long-context rates (the only rates that apply to such a request); Gemini Pro at its >200k rates, Gemini Flash flat.
| Provider |
Model |
Uncached (A, same rates) |
Cached steady-state (N=10) |
Saving |
Write / storage cost used |
| openai |
gpt-6-astra |
$15.00 |
$8.03 |
46 % |
12.5 |
| openai |
gpt-5.6-sol |
$6.00 |
$3.21 |
46 % |
5 |
| openai |
gpt-5.6-terra |
$3.20 |
$1.81 |
44 % |
2.5 |
| openai |
gpt-5.6-luna |
$0.32 |
$0.18 |
44 % |
0.25 |
| openai |
gpt-5.5 |
$8.00 |
$3.95 |
51 % |
free |
| openai |
gpt-5.4 |
$4.00 |
$1.98 |
51 % |
free |
| openai |
gpt-5.4-mini |
$1.20 |
$0.59 |
51 % |
free |
| openai |
gpt-5.4-nano |
$0.33 |
$0.16 |
50 % |
free |
| openai |
gpt-5.3-codex |
$3.15 |
$1.73 |
45 % |
free |
| openai |
gpt-5.2 |
$3.15 |
$1.73 |
45 % |
free |
| openai |
gpt-5.1 |
$2.25 |
$1.24 |
45 % |
free |
| openai |
gpt-5 |
$2.25 |
$1.24 |
45 % |
free |
| openai |
gpt-5-mini |
$0.45 |
$0.25 |
45 % |
free |
| openai |
gpt-5-nano |
$0.0900 |
$0.0495 |
45 % |
free |
| openai |
o3 |
$2.80 |
$1.45 |
48 % |
free |
| openai |
o4-mini |
$1.54 |
$0.80 |
48 % |
free |
| openai |
o3-mini |
$1.54 |
$1.05 |
32 % |
free |
| openai |
gpt-4.1 |
$2.80 |
$1.45 |
48 % |
free |
| openai |
gpt-4.1-mini |
$0.56 |
$0.29 |
48 % |
free |
| openai |
gpt-4.1-nano |
$0.14 |
$0.0725 |
48 % |
free |
| openai |
gpt-4o |
$3.50 |
$2.38 |
32 % |
free |
| openai |
gpt-4o-mini |
$0.21 |
$0.14 |
32 % |
free |
| openai |
chat-latest |
$8.00 |
$3.95 |
51 % |
free |
| openai |
gpt-5.6-cyber |
$20.00 |
$11.28 |
44 % |
15.625 |
| anthropic |
claude-fable-5-1 |
$15.00 |
$7.35 |
51 % |
12.5 (5m write) |
| anthropic |
claude-fable-5 |
$15.00 |
$8.03 |
46 % |
12.5 (5m write) |
| anthropic |
claude-mythos-5-1 |
$15.00 |
$7.35 |
51 % |
12.5 (5m write) |
| anthropic |
claude-mythos-5 |
$15.00 |
$8.03 |
46 % |
12.5 (5m write) |
| anthropic |
claude-opus-5 |
$7.50 |
$4.01 |
46 % |
6.25 (5m write) |
| anthropic |
claude-opus-4-8 |
$7.50 |
$4.01 |
46 % |
6.25 (5m write) |
| anthropic |
claude-opus-4-7 |
$7.50 |
$4.01 |
46 % |
6.25 (5m write) |
| anthropic |
claude-opus-4-6 |
$7.50 |
$4.01 |
46 % |
6.25 (5m write) |
| anthropic |
claude-opus-4-5-20251101 |
$7.50 |
$4.01 |
46 % |
6.25 (5m write) |
| anthropic |
claude-sonnet-5 |
$3.00 |
$1.60 |
46 % |
2.5 (5m write) |
| anthropic |
claude-sonnet-4-6 |
$4.50 |
$2.41 |
46 % |
3.75 (5m write) |
| anthropic |
claude-sonnet-4-5-20250929 |
$4.50 |
$2.41 |
46 % |
3.75 (5m write) |
| anthropic |
claude-haiku-4-5-20251001 |
$1.50 |
$0.80 |
46 % |
1.25 (5m write) |
| xai |
grok-4.6 |
$5.20 (long-context rates) |
$2.50 |
52 % |
none (automatic; cached rate 1) |
| xai |
grok-4.5 |
$5.20 (long-context rates) |
$2.14 |
59 % |
none (automatic; cached rate 0.6) |
| xai |
grok-4.3 |
$3.00 (long-context rates) |
$1.11 |
63 % |
none (automatic; cached rate 0.4) |
| xai |
grok-4.20-0309-reasoning |
$3.00 (long-context rates) |
$1.11 |
63 % |
none (automatic; cached rate 0.4) |
| xai |
grok-4.20-0309-non-reasoning |
$3.00 (long-context rates) |
$1.11 |
63 % |
none (automatic; cached rate 0.4) |
| xai |
grok-4.20-multi-agent-0309 |
$3.00 (long-context rates) |
$1.11 |
63 % |
none (automatic; cached rate 0.4) |
| xai |
grok-build-0.1 |
$2.40 (long-context rates) |
$0.96 |
60 % |
none (automatic; cached rate 0.4) |
| gemini |
gemini-3.8-flash |
$1.12 |
$0.56 |
50 % |
implicit: none; explicit: storage 0.5 /1M/h × 0.9 |
| gemini |
gemini-3.7-flash |
$1.12 |
$0.56 |
50 % |
implicit: none; explicit: storage 0.5 /1M/h × 0.9 |
| gemini |
gemini-3.6-flash |
$1.12 |
$0.56 |
50 % |
implicit: none; explicit: storage 0.5 /1M/h × 0.9 |
| gemini |
gemini-3.5-flash |
$2.40 |
$1.28 |
47 % |
implicit: none; explicit: storage 1 /1M/h × 0.9 |
| gemini |
gemini-3.5-flash-lite |
$0.55 |
$0.40 |
28 % |
implicit: none; explicit: storage 1 /1M/h × 0.9 |
| gemini |
gemini-3.1-pro-preview |
$5.80 (>200k rates) |
$2.96 |
49 % |
implicit: none; explicit: storage 4.5 /1M/h × 0.9 |
| gemini |
gemini-3.1-pro-preview-customtools |
$5.80 (>200k rates) |
$2.96 |
49 % |
implicit: none; explicit: storage 4.5 /1M/h × 0.9 |
| gemini |
gemini-3.1-flash-lite |
$0.40 |
$0.29 |
28 % |
implicit: none; explicit: storage 1 /1M/h × 0.9 |
| gemini |
gemini-3-flash-preview |
$0.80 |
$0.48 |
39 % |
implicit: none; explicit: storage 1 /1M/h × 0.9 |
| gemini |
gemini-2.5-pro |
$4.00 (>200k rates) |
$2.38 |
40 % |
implicit: none; explicit: storage 4.5 /1M/h × 0.9 |
| gemini |
gemini-2.5-flash |
$0.55 |
$0.40 |
28 % |
implicit: none; explicit: storage 1 /1M/h × 0.9 |
| gemini |
gemini-2.5-flash-lite |
$0.14 |
$0.15 |
-6 % |
implicit: none; explicit: storage 1 /1M/h × 0.9 |
| gemini |
gemini-robotics-er-2-preview |
$1.50 |
$0.73 |
51 % |
implicit: none; explicit: storage 0.5 /1M/h × 0.9 |
# Caching break-even by provider
| Provider |
Write cost |
Read multiplier |
Pays for itself on… |
Notes |
| OpenAI (pre-5.6) |
free |
model-specific 0.25×–0.5× |
first hit |
implicit; prompt_cache_key routing; 5–10 min in-memory, 24h retention option |
| OpenAI (GPT-5.6+, GPT-6) |
1.25× |
0.1× |
second use (1.25 + 0.1 = 1.35 < 2.0) |
explicit prompt_cache_breakpoint optional; ttl: 30m |
| Anthropic |
1.25× (5 m) / 2× (1 h) |
0.1× (0.025× Fable 5.1 / Mythos 5.1) |
second use (5 m); 1 h only if the gap between calls exceeds 5 min |
explicit cache_control; hits refresh TTL for free |
| xAI |
none |
0.15×–0.25× |
first hit |
automatic; prompt_cache_key / x-grok-conv-id for sticky routing; no TTL or minimum documented; cached tokens still count toward TPM |
| Gemini |
implicit: none; explicit: storage per token-hour |
0.1× |
implicit: first hit (≥ 4,096-token prompts on 3.x, 2,048 on 2.5); explicit: when 0.9 × read + storage_hour × hours / N < 0.9 × input — for gemini-3.8-flash one hour of storage ($0.45 per 0.9M) is recovered after a single hit ($0.675 − $0.0675 saved) |
explicit cache TTL default 1 h (ttl / expireTime), min 1,024 tokens (live), free tier limit: 0 |
# 8. Batch vs synchronous — the same 100k-request job
|
OpenAI Batch API |
Anthropic Message Batches |
xAI Batch API |
Gemini Batch API |
| Discount |
50 % on input, cached input and output (batch tier) |
50 % on input, output, cache writes and cache reads (stacks with caching) |
20 % on input, output, cached and reasoning tokens — grok-4.3 / 4.20 / 4.20-multi-agent only; grok-4.6, grok-4.5, grok-build: not supported; Imagine models: standard rates |
50 % (service:batch_api) on text, image, TTS, embedding models; Flex gives the same 50 % synchronously (best-effort) |
| Window |
24 h (completion_window: "24h", only value); output retained 30 days |
24 h expiry (expires_at); results 29 days |
no completion window documented; results listed via GET /v1/batches/{id}/results; batch requests do not count toward rate limits; video URLs in results expire after 1 h |
target 24 h ('usually much faster'); poll GET /v1beta/batches/{id}; concurrent batch jobs 100; enqueued-token caps per model and tier (e.g. Tier 1: 3M gemini-3.8-flash, 5M gemini-3.1-pro-preview) |
| How requests are supplied |
JSONL file (custom_id, method, url, body) → POST /v1/batches {input_file_id, endpoint}; 50,000 requests / 200 MB; one endpoint + one model per file |
inline requests[] {custom_id, params} → POST /v1/messages/batches; 100,000 requests / 256 MB; models mixed freely |
POST /v1/batches then POST /v1/batches/{id}/requests (inline chat/responses/image/video bodies) or JSONL upload; :cancel; no DELETE (405) |
POST /v1beta/models/{model}:batchGenerateContent with inline requests[] or a JSONL file from the Files API (2 GB / 20 GB storage); :asyncBatchEmbedContent for embeddings; /v1beta/openai/batches compatibility route |
| Example: 100k requests × (2k in + 300 out) on the flagship mid-tier |
gpt-5.6-sol batch: 200M × $2 + 30M × $10 = $700 (vs $1,400 standard) |
claude-sonnet-5 batch: 200M × $1 + 30M × $5 = $350 (vs $700 standard) |
grok-4.3 batch: 200M × $1.00 + 30M × $2.00 = $260 (vs $325 standard); grok-4.6 has no batch: $430 standard |
gemini-3.8-flash batch: 200M × $0.375 + 30M × $1.875 = $131.25 (vs $262.50 standard); gemini-3.1-pro-preview batch: 200M × $1 + 30M × $6 = $380 (vs $760) |
| Example: same job on the small models |
gpt-5.4-mini batch: 200M × $0.375 + 30M × $2.25 = $142.50 |
claude-haiku-4-5 batch: 200M × $0.5 + 30M × $2.5 = $175 |
grok-build-0.1: no batch, $260 standard |
gemini-3.5-flash-lite batch: 200M × $0.15 + 30M × $1.25 = $67.50 (vs $135) |
| Item |
OpenAI |
Anthropic |
xAI |
Gemini |
| Web search |
$10 / 1k calls (web_search, image search); preview $25 / 1k on non-reasoning models |
$10 / 1k searches (errors not billed) |
$5 / 1k successful calls (web_search, image search included; view_image results billed as image tokens) |
Google Search grounding: Gemini 3.x 5,000 free queries / month (shared) then $14 / 1k search queries; Gemini 2.5: 1,500 RPD free then $35 / 1k grounded prompts; free tier 500 RPD (not Pro) |
| Social / vertical search |
— |
— |
x_search $5 / 1k calls until 2026-09-21 12:00 PT, then $5 / 1k posts fetched + $10 / 1k user profiles fetched |
Google Maps grounding: Gemini 3.x 5,000 free prompts / month then $14 / 1k queries; Gemini 2.5: $25 / 1k grounded prompts (1,500 RPD free; 10,000 for Pro) |
| Web fetch / URL context |
— |
$0 per fetch (content tokens only) |
— (web search open_page action) |
urlContext free of charge; retrieved content billed as input tokens (toolUsePromptTokenCount); ≤ 20 URLs / request |
| Code execution |
container session: 1 GB $0.03 · 4 GB $0.12 · 16 GB $0.48 · 64 GB $1.92 per 20 min (per-minute, 5-min minimum since 2026-06-02); shared with hosted shell |
1,550 free container-hours / org / month, then $0.05 / container-hour (5-min minimum); free when web_search/web_fetch 20260209+ is in the request |
$5 / 1k calls (code_interpreter / code_execution) + tokens |
no per-call fee (codeExecution); generated code/results billed as output then as input when re-read; 30 s per execution |
| File search / RAG |
$2.50 / 1k calls + $0.10 / GB / day storage (1 GB free) |
— |
file_search / collections_search $2.50 / 1k calls + tokens; collections storage $0.10 / GiB / day, downloads $0.20 / GiB; attachment_search (implicit, files attached to messages) $10 / 1k calls |
fileSearch: indexing embeddings $0.15 / 1M tokens once; storage and query-time embeddings free; retrieved chunks billed as input tokens; store size by tier (Free 1 GB … Tier 3 1 TB) |
| MCP |
tokens only (mcp tool) |
tokens only (mcp_servers, beta) |
tokens only (mcp tool; tool outputs count as input) |
not documented (mcpServers UNVERIFIED on generateContent; Interactions mcp_server type) |
| Computer use |
tokens (computer tool) |
tokens (toolsets ≈ 4,500 tokens of definitions) |
— |
tokens only (computerUse, PREVIEW; screenshots are image input tokens); not on the free tier |
| Files API storage |
not priced separately (purpose=batch files expire 30 d) |
not priced on the pricing page |
$0.025 / GiB / day storage, $0.20 / GiB downloads; TTL via expires_after |
free (2 GB / file, 20 GB / project, 48 h TTL) |
| Managed agents runtime |
model tokens at Responses rates + container rates (no session fee) |
$0.08 per session-hour (usage.active_seconds) + model tokens + $10 / 1k web searches |
xAI Responses agentic loop: tokens + per-call tool fees; usage-guideline-violation fee $0.05 / request when a violation is caught before generation |
Interactions agents (Deep Research, Antigravity, custom managed agents): model inference at list rates incl. intermediate/reasoning tokens + tool fees; sandbox compute not billed during preview |
| Realtime / voice |
gpt-realtime-2.1: audio $32 in / $64 out, text $4 / $24, image $5 per 1M; mini $10 / $20 audio; gpt-live-1 $0.05 / min |
— |
grok-voice-think-fast-2.0 $0.08 / min ($4.80 / h) audio sent or received + $0.004 per text item; STT $0.10 / h REST, $0.20 / h streaming; TTS $15 / 1M characters |
gemini-3.8-live: audio in $3 / 1M (≈ $0.005 / min), audio out $12 / 1M (≈ $0.018 / min), text $0.75 / $4.50; TTS gemini-3.1-flash-tts-preview $1 text in / $20 audio out per 1M (≈ $0.03 / min); gemini-3.5-transcribe ≈ $0.005 / min blended; free tier available |
| Image generation |
gpt-image-2 / 2.5: text in $5, image in $8, image out $30 per 1M tokens (≈ $0.006–$0.21 per image); batch 50 % |
— |
grok-imagine-image $0.02 / image, grok-imagine-image-2.0 $0.04–$0.08 (quality × resolution), grok-imagine-image-quality $0.05 (DEPRECATED → 2026-11-02); image_generation tool billed at these rates |
gemini-3.1-flash-image $0.045 (0.5K) – $0.151 (4K) per image (image out $60 / 1M), gemini-3.1-flash-lite-image $0.034 (1K), gemini-3-pro-image $0.134 (1K/2K) / $0.24 (4K); batch 50 % |
| Video |
sora-2 $0.10 / s (720p), sora-2-pro $0.30–$0.70 / s — shutting down 2026-09-24 |
— |
grok-imagine-video $0.05 / s, grok-imagine-video-1.5 $0.08 / s (1080p, reference-to-video) |
Veo 3.1 $0.40 / s (720p/1080p), $0.60 / s (4K); Veo 3.1 Fast $0.10 / $0.12 / $0.30; Veo 3.1 Lite $0.05 / $0.08 (no 4K); gemini-omni-1.1-flash video out $17.50 / 1M tokens (≈ $0.10 / s) |
| Music |
— |
— |
— |
lyria-3.5 $0.08 / song, lyria-3-clip-preview $0.04 / 30-s clip, lyria-3-pro-preview $0.08; Lyria RealTime (experimental) unpriced |
| Embeddings |
text-embedding-3-small $0.02, -large $0.13 per 1M |
— |
grok-embedding-small: no published price (ACCOUNT_RESTRICTED) |
gemini-embedding-2 text $0.20 / 1M (batch $0.10), image $0.45, audio $6.50, video $12 per 1M; free tier available; gemini-embedding-001 unpriced (DEPRECATED → 2028-05-14) |
| Moderation |
free (omni-moderation-latest) |
— (built-in refusals) |
— (built-in; usage-guideline violations still billed + $0.05 fee) |
— (safetySettings thresholds, free) |
| Token counting |
free (POST /v1/responses/input_tokens) |
free (POST /v1/messages/count_tokens, own RPM bucket) |
free (POST /v1/tokenize-text; no usage/cost ticks returned) |
free (POST /v1beta/models/{model}:countTokens) |
| Data residency |
+10 % regional hosts (models ≥ 2026-03-05) |
×1.1 inference_geo: us; +10 % Bedrock/Vertex regional |
×1.1 on us.api.x.ai (grok-4.6) |
not priced on the Developer API (Vertex AI feature) |
| Per-request overhead |
none published |
tool-use system prompt: e.g. Opus 5 406 tokens (any/tool) / 286 (auto); bash_20250124 +244–325; text_editor +700; computer_toolset_20260801 ≈4,500; browser_toolset_20260801 ≈6,600 |
none published; cost_in_usd_ticks in every usage object (1 tick = $1e-10) makes the effective cost observable per call |
none published; usageMetadata breaks tokens down by modality (promptTokensDetails[]), thoughtsTokenCount, toolUsePromptTokenCount, cachedContentTokenCount |
# 10. Reading prices programmatically
# OpenAI standard row for one model
jq '.[] | select(.id=="gpt-5.6-sol") | .pricing.standard' generated/models.json
# Anthropic per-dimension rows
jq '[.[] | select(.provider=="anthropic" and .model_or_service=="claude-sonnet-5")]' generated/pricing.json
# xAI standard + long-context blocks and the live price ticks (1 tick = USD 1e-10 per token)
jq '.[] | select(.id=="grok-4.6") | .pricing | {standard, long_context, batch_discount, priority_multiplier, us_regional_multiplier}' generated/models.json
# Gemini tiers (standard/batch/flex/priority) and free-tier note
jq '.[] | select(.id=="gemini-3.8-flash") | .pricing | {tiers, free_tier, from_2027_01_01}' generated/models.json
# every tool / service price row across the four providers
jq '[.[] | select(.model_or_service|test("^(tool:|service:|agent:|rule:)"))]' generated/pricing.json
Related: models · caching and reasoning · features · realtime and media · FAQ.