SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
6.4 KB

# Gemini API — rate limits, usage tiers and 429 semantics

Status: DOCUMENTED + LIVE_DISCOVERED (headers and 429 bodies observed with our Free-tier key on 2026-09-19). Machine-readable twin: generated/fragments/rate-limits/gemini-rate-limits.json. Sources: https://ai.google.dev/gemini-api/docs/rate-limits · https://ai.google.dev/gemini-api/docs/billing · https://ai.google.dev/gemini-api/docs/priority-inference · https://ai.google.dev/gemini-api/docs/flex-inference · https://ai.google.dev/gemini-api/docs/batch-api · https://ai.google.dev/gemini-api/docs/live-api/session · https://ai.google.dev/gemini-api/docs/troubleshooting Last verified: 2026-09-18 (docs) / 2026-09-19 (live).

# 1. How limits work

  • Dimensions: RPM (requests/min), TPM (input tokens/min), RPD (requests/day, reset at midnight Pacific); model-specific extras IPM (images/min, image models) and TPD.
  • Applied per Google Cloud project, not per key. Exceeding any dimension → 429 RESOURCE_EXHAUSTED.
  • Preview/experimental models have tighter limits. "Specified rate limits are not guaranteed and actual capacity may vary."
  • The per-model RPM/TPM/RPD matrix is no longer published in the docs — it lives in AI Studio (https://aistudio.google.com/rate-limit, account-specific). This atlas therefore records tiers, spend limits and batch/Live limits only; per-model numbers are intentionally not invented.

# 2. Usage tiers

Tier Qualification Billing-tier spend cap Spend-based rate limit (rolling 10 min)
Free active project or free trial n/a n/a
Tier 1 link an active Cloud Billing account $250 $10
Tier 2 $100 cumulative Google Cloud spend + 3 days since first payment $2,000 $50
Tier 3 $1,000 cumulative + 30 days $20,000 – $100,000+ $200

Free → Tier 1 is instant; later upgrades within ~10 min; upgrades may be denied "based on other factors". Free tier excludes Pro models (observed: limit: 0), image/video/music generation and agents. Project-level spend caps and billing-account tier caps (2026-03) stop traffic instead of overspending.

# 3. Service-tier limits

Tier Rate limit Behaviour on overflow
standard model/tier limit (AI Studio) 429
priority 0.3× standard (own bucket, also counted against interactive traffic) graceful downgrade to standard, no error
flex own limits; sheddable 429 when capacity is shed → retry (docs provide retry loops and per-request timeouts)

# 4. Batch API limits

Concurrent batch jobs 100; input file 2 GB; file storage 20 GB; turnaround target 24 h; price 50 %. Enqueued-token ceilings per model (all active jobs):

Model Tier 1 Tier 2 Tier 3
gemini-3.1-pro-preview 5M 500M 1B
gemini-3.5-flash-lite, gemini-3.1-flash-lite (+preview), gemini-2.5-flash-lite (+preview) 10M 500M 1B
gemini-3.8-flash, 3.7-flash, 3.6-flash, 3.5-flash, 2.5-flash (+preview, image-preview) 3M 400M 1B
gemini-2.5-pro 5M 500M 1B
gemini-2.5-pro-preview-tts 25k 100k 1M
gemini-2.5-flash-preview-tts 100k 100k 4M
gemini-2.0-flash / flash-lite (shut down) 10M 1B 5B
gemini-3.1-flash-image-preview 1M 250M 750M
gemini-3.1-flash-lite-image, gemini-3-pro-image-preview 2M 270M 1B
gemini-embedding (001 / 2) 500k 5M 10M

The table still names shut-down 2.0/2.5 previews and omits the GA image ids (gemini-3.1-flash-image, gemini-3-pro-image) — assume the preview rows apply.

# 5. Live API session limits

Audio-only session 15 min, audio+video 2 min without contextWindowCompression; WebSocket connection ≈ 10 min (server sends goAway); sessionResumption handles valid 2 h after termination. Ephemeral tokens (POST /v1beta/auth_tokens) for client-side connections. No published concurrent-session limit.

# 6. Grounding quotas (free allowances)

Gemini 3.x: Google Search 5,000 requests/month, Maps 5,000 prompts/month (shared across 3.x). Gemini 2.5: Search 1,500 RPD (free tier 500 RPD, shared Flash/Flash-Lite), Maps 1,500 RPD (10,000 for Pro).

# 7. 429 semantics (observed)

json
{"error":{"code":429,"status":"RESOURCE_EXHAUSTED",
 "message":"You exceeded your current quota, please check your plan and billing details. ...\n* Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_requests, limit: 0, model: gemini-3.1-pro\n...Please retry in 50.868302469s.",
 "details":[{"@type":"type.googleapis.com/google.rpc.Help","links":[{"description":"Learn more about Gemini API quotas","url":"https://ai.google.dev/gemini-api/docs/rate-limits"}]},
            {"@type":"type.googleapis.com/google.rpc.QuotaFailure","violations":[{"quotaMetric":"generativelanguage.googleapis.com/generate_content_free_tier_requests","quotaId":"GenerateRequestsPerDayPerProjectPerModel-FreeTier","quotaDimensions":{"model":"gemini-3.1-pro","location":"global"}}]}]}}
  • Quota ids seen: GenerateContentInputTokensPerModelPerDay-FreeTier, GenerateRequestsPerDayPerProjectPerModel-FreeTier, GenerateRequestsPerMinutePerProjectPerModel-FreeTier. quotaDimensions.model is the base model (gemini-3.1-pro for gemini-3.1-pro-preview and for gemini-pro-latest).
  • Retry delay only in the message text; no Retry-After header.
  • Retry guidance (troubleshooting.md): exponential backoff + jitter on 429/408/5xx only; SDKs retry 429/5xx automatically (Python: 4 attempts, ~1 s initial, 60 s max).

# 8. Headers: documented vs observed

Result
Documented rate-limit headers none
Observed on generateContent (200 and 429) Server-Timing: gfet4t7; dur=87, X-Gemini-Service-Tier: standard, Alt-Svc, Vary, X-Content-Type-Options, X-Frame-Options, X-XSS-Protection, Accept-Ranges, Transfer-Encoding, Content-Type, Date, Server
Observed on /v1beta/openai/* Server-Timing, Set-Cookie, Cache-Control, Expires, P3P (no X-Gemini-Service-Tier)
x-ratelimit-*, Retry-After, request-id absent — remaining quota is not exposed; use AI Studio dashboards; correlate with responseId in the body

# 9. Observed for our key (account-specific, not documentation)

Free tier. OK: gemini-3.5-flash-lite, 3.5-flash, 3.8-flash, gemma-4-26b-a4b-it, gemini-flash-latest, OpenAI-compat chat. Restricted: gemini-3.1-pro-preview and gemini-pro-latest → 429 limit: 0; gemini-2.5-flash-lite → 404 "no longer available to new users".