# Gemini API — pricing
Status: DOCUMENTED (every price on the pricing page transcribed; nothing here is account-specific). Machine-readable twin: generated/fragments/pricing/gemini-pricing.json (527 price records, tier ∈ standard | batch | flex | priority | free).
Sources: https://ai.google.dev/gemini-api/docs/pricing · https://ai.google.dev/gemini-api/docs/billing · https://ai.google.dev/gemini-api/docs/priority-inference · https://ai.google.dev/gemini-api/docs/flex-inference · https://ai.google.dev/gemini-api/docs/google-search#pricing · https://ai.google.dev/gemini-api/docs/file-search · https://ai.google.dev/gemini-api/docs/latest-model#pricing
Last verified: 2026-09-18 (docs). Prices in USD per 1M tokens unless stated. "Output" always includes thinking tokens.
# 1. Plans and tiers
| Plan |
Who |
Key facts |
| Free |
developers, small projects |
free input/output on many models; limited model access (no Pro/image/video/music/agents); content used to improve Google products (Unpaid Services); Batch/Flex "Not available" |
| Paid (Tier 1–3) |
production |
higher limits; context caching; Batch API (−50 %); Flex (−50 %); Priority (+75–100 %); content not used for training; prepay or postpay billing (since 2026-03-23); spend caps $250 / $2,000 / $20,000+ per tier |
| Enterprise |
Gemini Enterprise Agent Platform (Vertex AI) |
provisioned throughput, volume discounts, compliance; different price list |
Service tiers (service_tier request field, Interactions API & OpenAI-compat; serviceTier echoed in usageMetadata and in the X-Gemini-Service-Tier header):
| Tier |
Price vs Standard |
Latency |
Reliability |
| standard |
1× |
seconds–minutes |
high |
| flex |
0.5× |
1–15 min target |
best-effort, sheddable (429 when shed) |
| priority |
1.8× (docs: "75-100 % more") |
seconds |
non-sheddable; overflow gracefully downgraded to standard |
| batch (async API) |
0.5× |
≤ 24 h |
high throughput |
# 2. Text / multimodal models (per 1M tokens)
Introductory pricing marked ★ applies through 2026-12-31; the 2027 price is in parentheses.
| Model |
Input |
Output |
Cached input |
Cache storage /h |
Batch & Flex (in / out) |
Priority (in / out) |
Free tier |
| gemini-3.8-flash ★ |
0.75 (1.50) |
3.75 (7.50) |
0.075 (0.15) |
0.50 (1.00) |
0.375 / 1.875 (0.75 / 3.75) |
1.35 / 6.75 (2.70 / 13.50) |
free |
| gemini-3.7-flash ★ |
0.75 (1.50) |
3.75 (7.50) |
0.075 (0.15) |
0.50 (1.00) |
0.375 / 1.875 |
1.35 / 6.75 |
free |
| gemini-3.6-flash ★ |
0.75 (1.50) |
3.75 (7.50) |
0.075 (0.15) |
0.50 (1.00) |
0.375 / 1.875 |
1.35 / 6.75 |
free |
| gemini-3.5-flash |
1.50 |
9.00 |
0.15 |
1.00 |
0.75 / 4.50 (flex cache 0.08) |
2.70 / 16.20 |
free |
| gemini-3.5-flash-lite |
0.30 (all modalities) |
2.50 |
0.03 |
1.00 |
0.15 / 1.25 |
0.54 / 4.50 |
free (caching n/a) |
| gemini-3.1-flash-lite |
0.25 text-image-video / 0.50 audio |
1.50 |
0.025 / 0.05 audio |
1.00 (batch 0.50, prio 1.80) |
0.125 (0.25 audio) / 0.75 |
0.45 (0.90 audio) / 2.70 |
free |
| gemini-3-flash-preview |
0.50 / 1.00 audio |
3.00 |
0.05 / 0.10 audio |
1.00 (prio 1.80) |
0.25 (0.50 audio) / 1.50 |
0.90 (1.80 audio) / 5.40 |
free |
gemini-3.1-pro-preview (+ -customtools) |
2.00 ≤200k / 4.00 >200k |
12.00 / 18.00 |
0.20 / 0.40 |
4.50 (prio 8.10) |
1.00–2.00 / 6.00–9.00 |
3.60–7.20 / 21.60–32.40 |
not available |
| gemini-2.5-pro |
1.25 / 2.50 >200k |
10.00 / 15.00 |
0.125 / 0.25 |
4.50 (prio 8.10) |
0.625–1.25 / 5.00–7.50 |
2.25–4.50 / 18–27 |
free (no caching) |
| gemini-2.5-flash |
0.30 / 1.00 audio |
2.50 |
0.03 / 0.10 audio |
1.00 (prio 1.80) |
0.15 (0.50 audio) / 1.25 |
0.54 (1.80 audio) / 4.50 |
free |
| gemini-2.5-flash-lite |
0.10 / 0.30 audio |
0.40 |
0.01 / 0.03 audio |
1.00 (prio 1.80) |
0.05 (0.15 audio) / 0.20 |
0.18 (0.54 audio) / 0.72 |
free |
| gemini-robotics-er-2-preview ★ |
1.00 (2.00) |
5.00 (10.00) |
0.10 (0.20) |
0.50 (1.00) |
0.50 / 2.50 |
— |
free |
| gemini-2.5-computer-use-preview-10-2025 ★ |
1.00 (2.00) |
5.00 (10.00) |
— |
— |
— |
— |
free |
| gemma-4-26b-a4b-it, gemma-4-31b-it |
free |
free |
free |
free |
n/a |
n/a |
free only (paid tier "Not available") |
Notes: long-context tier threshold is 200k prompt tokens (Pro models only). Batch/Flex context-cache reads keep standard price on Pro ("Same as Standard"). DOCUMENT (PDF) tokens are billed at the image token rate. The 2.5 family is "no longer available to new users" (live 404) despite still being priced.
# 3. Live API, speech, transcription (per 1M tokens; audio = 25 tokens/s)
| Model |
Input |
Output |
Free tier |
| gemini-3.8-live, gemini-3.8-live-extended-thinking, gemini-3.1-flash-live-preview |
text 0.75 · audio 3.00 (≈$0.005/min) · image/video 1.00 (≈$0.002/min) |
text 4.50 · audio 12.00 (≈$0.018/min) |
free; Google Search grounding supported |
| gemini-2.5-flash-native-audio-preview-12-2025 |
text 0.50 · audio/video 3.00 |
text 2.00 · audio 12.00 |
free |
| gemini-3.5-live-translate-preview |
audio 3.50 (≈$0.0053/min) |
audio 21.00 (≈$0.0315/min); ≈$0.0368/min blended |
free |
| gemini-3.5-transcribe-live |
audio 3.50 (≈$0.005/min) |
text 21.00 (≈$0.004/min; 175 text tokens/min); ≈$0.009/min blended |
free |
| gemini-3.5-transcribe |
audio 2.00 (≈$0.003/min) |
text 12.00 (≈$0.002/min); ≈$0.005/min blended |
free |
| gemini-3.1-flash-tts-preview |
text 1.00 (batch 0.50) |
audio 20.00 (batch 10.00) |
free (standard) |
| gemini-2.5-flash-preview-tts |
text 0.50 (batch 0.25) |
audio 10.00 (batch 5.00) |
free (standard) |
| gemini-2.5-pro-preview-tts |
text 1.00 (batch 0.50) |
audio 20.00 (batch 10.00) |
not available |
# 4. Image generation
| Model |
Input |
Text output |
Image output |
Per-image equivalents |
Batch |
| gemini-3.1-flash-image (Nano Banana 2; preview id same) |
0.50 |
3.00 |
60.00 /1M image tokens |
0.5K 747 tok $0.045 · 1K 1120 tok $0.067 · 2K 1680 tok $0.101 · 4K 2520 tok $0.151 |
0.25 / 1.50 / 30.00 → $0.022–0.076 |
| gemini-3.1-flash-lite-image (Nano Banana 2 Lite) |
0.25 |
1.50 |
30.00 |
1K only: $0.0336 |
0.125 / 0.75 / 15.00 → $0.0168 |
gemini-3-pro-image (Nano Banana Pro; -preview, nano-banana-pro-preview) |
2.00 (image in = 560 tok = $0.0011) |
12.00 |
120.00 |
1K/2K 1120 tok $0.134 · 4K 2000 tok $0.24 |
batch & flex: 1.00 text, $0.0006/image in, 6.00 text out, $0.067 / $0.12 per image; priority 3.60 / 21.60 / 216.00 |
| gemini-2.5-flash-image (Nano Banana, shutdown 2026-10-02) |
0.30 |
— |
30.00 (1290 tok/image) |
$0.039 per image |
batch/flex 0.15 + $0.0195; priority 0.54 + $0.0702 |
| imagen-4.0-* |
retired 2026-08-17 |
|
|
|
|
Google Search grounding on image models: 5,000 free requests/month shared across Gemini 3.x, then $14 per 1,000 (web + image search); retrieved context not charged as input.
# 5. Video, music, embeddings
| Model |
Price |
Unit / notes |
| gemini-omni-1.1-flash / gemini-omni-flash-preview |
input 1.50; output text 9.00, video 17.50 per 1M tokens |
5,792 tokens per second of 720p ⇒ ≈ $0.10 per second; paid tier only |
| veo-3.1-generate-preview |
$0.40/s (720p & 1080p), $0.60/s (4K) |
video with audio; charged only when generation succeeds |
| veo-3.1-fast-generate-preview |
$0.10/s (720p), $0.12/s (1080p), $0.30/s (4K) |
|
| veo-3.1-lite-generate-preview |
$0.05/s (720p), $0.08/s (1080p); no 4K |
|
| lyria-3.5 |
$0.08 per song |
full length |
| lyria-3-clip-preview / lyria-3-pro-preview |
$0.04 (30 s clip) / $0.08 per song |
legacy |
| lyria-realtime-exp |
not priced (experimental) |
|
gemini-embedding-2 (and -preview) |
text 0.20 · image 0.45 ($0.00012/image) · audio 6.50 ($0.00016/s) · video 12.00 ($0.00079/frame) |
batch: 0.10 / 0.225 / 3.25 / 6.00; free tier free (standard) |
| gemini-embedding-001 |
not on the pricing page (2026-09-18); File Search indexing uses $0.15/1M |
shutdown 2028-05-14 |
| Item |
Free tier |
Paid tier |
| Grounding with Google Search |
500 RPD free (Flash/Flash-Lite 2.5; not Pro) |
Gemini 3.x: 5,000 free search requests/month shared, then $14 / 1,000 requests (billed per executed query). Gemini 2.5: 1,500 RPD free, then $35 / 1,000 grounded prompts |
| Grounding with Google Maps |
500 RPD (not Pro) |
Gemini 3: 5,000 prompts/month free then $14 / 1,000 queries; 2.5: 1,500 RPD (10,000 for Pro) then $25 / 1,000 |
| Code execution |
free |
token rates only (generated code + results = output tokens, re-read = input tokens); no runtime charge |
| URL context |
free |
retrieved content billed as input tokens |
| Computer use |
n/a |
token rates of the model |
| File Search |
free |
indexing embeddings $0.15 / 1M tokens; storage and query-time embeddings free; retrieved doc tokens billed as input |
| Custom Tools endpoint |
n/a |
same as gemini-3.1-pro-preview |
| Deep Research agent |
n/a (paid key required) |
list-rate tokens incl. intermediate reasoning; tool fees per tool |
| Managed agents / Antigravity |
n/a |
list-rate tokens; sandbox compute not billed during preview |
| Context-cache storage |
free on Flash free tier |
$0.50–$8.10 per 1M tokens per hour (model/tier dependent, table §2) |
| Google AI Studio |
free in all regions |
separate from API billing; Google AI Pro/Ultra plans apply to the AI Studio UI only |
# 7. Gotchas
- Thinking tokens are billed as output (
usageMetadata.thoughtsTokenCount); with tiny maxOutputTokens the whole budget can go to thoughts (observed MAX_TOKENS with empty text on 3.5/3.8 Flash and Gemma 4).
- Batch/Flex are "Not available" on the free tier for most models but listed "Free of charge" for gemini-3.5-flash-lite (page inconsistency).
- Priority pricing tables show exactly 1.8× standard, while the prose says 75–100 % more.
- Failed requests are not billed; Veo is charged only for successfully generated videos.
- Prepay accounts stop serving at $0 balance; Google Cloud credits are applied first.