SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
9.9 KB

# Gemini API — pricing

Status: DOCUMENTED (every price on the pricing page transcribed; nothing here is account-specific). Machine-readable twin: generated/fragments/pricing/gemini-pricing.json (527 price records, tier ∈ standard | batch | flex | priority | free). Sources: https://ai.google.dev/gemini-api/docs/pricing · https://ai.google.dev/gemini-api/docs/billing · https://ai.google.dev/gemini-api/docs/priority-inference · https://ai.google.dev/gemini-api/docs/flex-inference · https://ai.google.dev/gemini-api/docs/google-search#pricing · https://ai.google.dev/gemini-api/docs/file-search · https://ai.google.dev/gemini-api/docs/latest-model#pricing Last verified: 2026-09-18 (docs). Prices in USD per 1M tokens unless stated. "Output" always includes thinking tokens.

# 1. Plans and tiers

Plan Who Key facts
Free developers, small projects free input/output on many models; limited model access (no Pro/image/video/music/agents); content used to improve Google products (Unpaid Services); Batch/Flex "Not available"
Paid (Tier 1–3) production higher limits; context caching; Batch API (−50 %); Flex (−50 %); Priority (+75–100 %); content not used for training; prepay or postpay billing (since 2026-03-23); spend caps $250 / $2,000 / $20,000+ per tier
Enterprise Gemini Enterprise Agent Platform (Vertex AI) provisioned throughput, volume discounts, compliance; different price list

Service tiers (service_tier request field, Interactions API & OpenAI-compat; serviceTier echoed in usageMetadata and in the X-Gemini-Service-Tier header):

Tier Price vs Standard Latency Reliability
standard 1× seconds–minutes high
flex 0.5× 1–15 min target best-effort, sheddable (429 when shed)
priority 1.8× (docs: "75-100 % more") seconds non-sheddable; overflow gracefully downgraded to standard
batch (async API) 0.5× ≤ 24 h high throughput

# 2. Text / multimodal models (per 1M tokens)

Introductory pricing marked ★ applies through 2026-12-31; the 2027 price is in parentheses.

Model Input Output Cached input Cache storage /h Batch & Flex (in / out) Priority (in / out) Free tier
gemini-3.8-flash ★ 0.75 (1.50) 3.75 (7.50) 0.075 (0.15) 0.50 (1.00) 0.375 / 1.875 (0.75 / 3.75) 1.35 / 6.75 (2.70 / 13.50) free
gemini-3.7-flash ★ 0.75 (1.50) 3.75 (7.50) 0.075 (0.15) 0.50 (1.00) 0.375 / 1.875 1.35 / 6.75 free
gemini-3.6-flash ★ 0.75 (1.50) 3.75 (7.50) 0.075 (0.15) 0.50 (1.00) 0.375 / 1.875 1.35 / 6.75 free
gemini-3.5-flash 1.50 9.00 0.15 1.00 0.75 / 4.50 (flex cache 0.08) 2.70 / 16.20 free
gemini-3.5-flash-lite 0.30 (all modalities) 2.50 0.03 1.00 0.15 / 1.25 0.54 / 4.50 free (caching n/a)
gemini-3.1-flash-lite 0.25 text-image-video / 0.50 audio 1.50 0.025 / 0.05 audio 1.00 (batch 0.50, prio 1.80) 0.125 (0.25 audio) / 0.75 0.45 (0.90 audio) / 2.70 free
gemini-3-flash-preview 0.50 / 1.00 audio 3.00 0.05 / 0.10 audio 1.00 (prio 1.80) 0.25 (0.50 audio) / 1.50 0.90 (1.80 audio) / 5.40 free
gemini-3.1-pro-preview (+ -customtools) 2.00 ≤200k / 4.00 >200k 12.00 / 18.00 0.20 / 0.40 4.50 (prio 8.10) 1.00–2.00 / 6.00–9.00 3.60–7.20 / 21.60–32.40 not available
gemini-2.5-pro 1.25 / 2.50 >200k 10.00 / 15.00 0.125 / 0.25 4.50 (prio 8.10) 0.625–1.25 / 5.00–7.50 2.25–4.50 / 18–27 free (no caching)
gemini-2.5-flash 0.30 / 1.00 audio 2.50 0.03 / 0.10 audio 1.00 (prio 1.80) 0.15 (0.50 audio) / 1.25 0.54 (1.80 audio) / 4.50 free
gemini-2.5-flash-lite 0.10 / 0.30 audio 0.40 0.01 / 0.03 audio 1.00 (prio 1.80) 0.05 (0.15 audio) / 0.20 0.18 (0.54 audio) / 0.72 free
gemini-robotics-er-2-preview ★ 1.00 (2.00) 5.00 (10.00) 0.10 (0.20) 0.50 (1.00) 0.50 / 2.50 — free
gemini-2.5-computer-use-preview-10-2025 ★ 1.00 (2.00) 5.00 (10.00) — — — — free
gemma-4-26b-a4b-it, gemma-4-31b-it free free free free n/a n/a free only (paid tier "Not available")

Notes: long-context tier threshold is 200k prompt tokens (Pro models only). Batch/Flex context-cache reads keep standard price on Pro ("Same as Standard"). DOCUMENT (PDF) tokens are billed at the image token rate. The 2.5 family is "no longer available to new users" (live 404) despite still being priced.

# 3. Live API, speech, transcription (per 1M tokens; audio = 25 tokens/s)

Model Input Output Free tier
gemini-3.8-live, gemini-3.8-live-extended-thinking, gemini-3.1-flash-live-preview text 0.75 · audio 3.00 (≈$0.005/min) · image/video 1.00 (≈$0.002/min) text 4.50 · audio 12.00 (≈$0.018/min) free; Google Search grounding supported
gemini-2.5-flash-native-audio-preview-12-2025 text 0.50 · audio/video 3.00 text 2.00 · audio 12.00 free
gemini-3.5-live-translate-preview audio 3.50 (≈$0.0053/min) audio 21.00 (≈$0.0315/min); ≈$0.0368/min blended free
gemini-3.5-transcribe-live audio 3.50 (≈$0.005/min) text 21.00 (≈$0.004/min; 175 text tokens/min); ≈$0.009/min blended free
gemini-3.5-transcribe audio 2.00 (≈$0.003/min) text 12.00 (≈$0.002/min); ≈$0.005/min blended free
gemini-3.1-flash-tts-preview text 1.00 (batch 0.50) audio 20.00 (batch 10.00) free (standard)
gemini-2.5-flash-preview-tts text 0.50 (batch 0.25) audio 10.00 (batch 5.00) free (standard)
gemini-2.5-pro-preview-tts text 1.00 (batch 0.50) audio 20.00 (batch 10.00) not available

# 4. Image generation

Model Input Text output Image output Per-image equivalents Batch
gemini-3.1-flash-image (Nano Banana 2; preview id same) 0.50 3.00 60.00 /1M image tokens 0.5K 747 tok $0.045 · 1K 1120 tok $0.067 · 2K 1680 tok $0.101 · 4K 2520 tok $0.151 0.25 / 1.50 / 30.00 → $0.022–0.076
gemini-3.1-flash-lite-image (Nano Banana 2 Lite) 0.25 1.50 30.00 1K only: $0.0336 0.125 / 0.75 / 15.00 → $0.0168
gemini-3-pro-image (Nano Banana Pro; -preview, nano-banana-pro-preview) 2.00 (image in = 560 tok = $0.0011) 12.00 120.00 1K/2K 1120 tok $0.134 · 4K 2000 tok $0.24 batch & flex: 1.00 text, $0.0006/image in, 6.00 text out, $0.067 / $0.12 per image; priority 3.60 / 21.60 / 216.00
gemini-2.5-flash-image (Nano Banana, shutdown 2026-10-02) 0.30 — 30.00 (1290 tok/image) $0.039 per image batch/flex 0.15 + $0.0195; priority 0.54 + $0.0702
imagen-4.0-* retired 2026-08-17

Google Search grounding on image models: 5,000 free requests/month shared across Gemini 3.x, then $14 per 1,000 (web + image search); retrieved context not charged as input.

# 5. Video, music, embeddings

Model Price Unit / notes
gemini-omni-1.1-flash / gemini-omni-flash-preview input 1.50; output text 9.00, video 17.50 per 1M tokens 5,792 tokens per second of 720p ⇒ ≈ $0.10 per second; paid tier only
veo-3.1-generate-preview $0.40/s (720p & 1080p), $0.60/s (4K) video with audio; charged only when generation succeeds
veo-3.1-fast-generate-preview $0.10/s (720p), $0.12/s (1080p), $0.30/s (4K)
veo-3.1-lite-generate-preview $0.05/s (720p), $0.08/s (1080p); no 4K
lyria-3.5 $0.08 per song full length
lyria-3-clip-preview / lyria-3-pro-preview $0.04 (30 s clip) / $0.08 per song legacy
lyria-realtime-exp not priced (experimental)
gemini-embedding-2 (and -preview) text 0.20 · image 0.45 ($0.00012/image) · audio 6.50 ($0.00016/s) · video 12.00 ($0.00079/frame) batch: 0.10 / 0.225 / 3.25 / 6.00; free tier free (standard)
gemini-embedding-001 not on the pricing page (2026-09-18); File Search indexing uses $0.15/1M shutdown 2028-05-14

# 6. Tools, agents, services

Item Free tier Paid tier
Grounding with Google Search 500 RPD free (Flash/Flash-Lite 2.5; not Pro) Gemini 3.x: 5,000 free search requests/month shared, then $14 / 1,000 requests (billed per executed query). Gemini 2.5: 1,500 RPD free, then $35 / 1,000 grounded prompts
Grounding with Google Maps 500 RPD (not Pro) Gemini 3: 5,000 prompts/month free then $14 / 1,000 queries; 2.5: 1,500 RPD (10,000 for Pro) then $25 / 1,000
Code execution free token rates only (generated code + results = output tokens, re-read = input tokens); no runtime charge
URL context free retrieved content billed as input tokens
Computer use n/a token rates of the model
File Search free indexing embeddings $0.15 / 1M tokens; storage and query-time embeddings free; retrieved doc tokens billed as input
Custom Tools endpoint n/a same as gemini-3.1-pro-preview
Deep Research agent n/a (paid key required) list-rate tokens incl. intermediate reasoning; tool fees per tool
Managed agents / Antigravity n/a list-rate tokens; sandbox compute not billed during preview
Context-cache storage free on Flash free tier $0.50–$8.10 per 1M tokens per hour (model/tier dependent, table §2)
Google AI Studio free in all regions separate from API billing; Google AI Pro/Ultra plans apply to the AI Studio UI only

# 7. Gotchas

  • Thinking tokens are billed as output (usageMetadata.thoughtsTokenCount); with tiny maxOutputTokens the whole budget can go to thoughts (observed MAX_TOKENS with empty text on 3.5/3.8 Flash and Gemma 4).
  • Batch/Flex are "Not available" on the free tier for most models but listed "Free of charge" for gemini-3.5-flash-lite (page inconsistency).
  • Priority pricing tables show exactly 1.8× standard, while the prose says 75–100 % more.
  • Failed requests are not billed; Veo is charged only for successfully generated videos.
  • Prepay accounts stop serving at $0 balance; Google Cloud credits are applied first.