Gemini API — rate limits, usage tiers and 429 semantics
Status: DOCUMENTED + LIVE_DISCOVERED (headers and 429 bodies observed with our Free-tier key on 2026-09-19). Machine-readable twin: generated/fragments/rate-limits/gemini-rate-limits.json.
Sources: https://ai.google.dev/gemini-api/docs/rate-limits · https://ai.google.dev/gemini-api/docs/billing · https://ai.google.dev/gemini-api/docs/priority-inference · https://ai.google.dev/gemini-api/docs/flex-inference · https://ai.google.dev/gemini-api/docs/batch-api · https://ai.google.dev/gemini-api/docs/live-api/session · https://ai.google.dev/gemini-api/docs/troubleshooting
Last verified: 2026-09-18 (docs) / 2026-09-19 (live).
1. How limits work
- Dimensions: RPM (requests/min), TPM (input tokens/min), RPD (requests/day, reset at midnight Pacific); model-specific extras IPM (images/min, image models) and TPD.
- Applied per Google Cloud project, not per key. Exceeding any dimension →
429 RESOURCE_EXHAUSTED. - Preview/experimental models have tighter limits. "Specified rate limits are not guaranteed and actual capacity may vary."
- The per-model RPM/TPM/RPD matrix is no longer published in the docs — it lives in AI Studio (https://aistudio.google.com/rate-limit, account-specific). This atlas therefore records tiers, spend limits and batch/Live limits only; per-model numbers are intentionally not invented.
2. Usage tiers
| Tier | Qualification | Billing-tier spend cap | Spend-based rate limit (rolling 10 min) |
|---|---|---|---|
| Free | active project or free trial | n/a | n/a |
| Tier 1 | link an active Cloud Billing account | $250 | $10 |
| Tier 2 | $100 cumulative Google Cloud spend + 3 days since first payment | $2,000 | $50 |
| Tier 3 | $1,000 cumulative + 30 days | $20,000 – $100,000+ | $200 |
Free → Tier 1 is instant; later upgrades within ~10 min; upgrades may be denied "based on other factors". Free tier excludes Pro models (observed: limit: 0), image/video/music generation and agents. Project-level spend caps and billing-account tier caps (2026-03) stop traffic instead of overspending.
3. Service-tier limits
| Tier | Rate limit | Behaviour on overflow |
|---|---|---|
| standard | model/tier limit (AI Studio) | 429 |
| priority | 0.3× standard (own bucket, also counted against interactive traffic) | graceful downgrade to standard, no error |
| flex | own limits; sheddable | 429 when capacity is shed → retry (docs provide retry loops and per-request timeouts) |
4. Batch API limits
Concurrent batch jobs 100; input file 2 GB; file storage 20 GB; turnaround target 24 h; price 50 %. Enqueued-token ceilings per model (all active jobs):
| Model | Tier 1 | Tier 2 | Tier 3 |
|---|---|---|---|
| gemini-3.1-pro-preview | 5M | 500M | 1B |
| gemini-3.5-flash-lite, gemini-3.1-flash-lite (+preview), gemini-2.5-flash-lite (+preview) | 10M | 500M | 1B |
| gemini-3.8-flash, 3.7-flash, 3.6-flash, 3.5-flash, 2.5-flash (+preview, image-preview) | 3M | 400M | 1B |
| gemini-2.5-pro | 5M | 500M | 1B |
| gemini-2.5-pro-preview-tts | 25k | 100k | 1M |
| gemini-2.5-flash-preview-tts | 100k | 100k | 4M |
| gemini-2.0-flash / flash-lite (shut down) | 10M | 1B | 5B |
| gemini-3.1-flash-image-preview | 1M | 250M | 750M |
| gemini-3.1-flash-lite-image, gemini-3-pro-image-preview | 2M | 270M | 1B |
| gemini-embedding (001 / 2) | 500k | 5M | 10M |
The table still names shut-down 2.0/2.5 previews and omits the GA image ids (gemini-3.1-flash-image, gemini-3-pro-image) — assume the preview rows apply.
5. Live API session limits
Audio-only session 15 min, audio+video 2 min without contextWindowCompression; WebSocket connection ≈ 10 min (server sends goAway); sessionResumption handles valid 2 h after termination. Ephemeral tokens (POST /v1beta/auth_tokens) for client-side connections. No published concurrent-session limit.
6. Grounding quotas (free allowances)
Gemini 3.x: Google Search 5,000 requests/month, Maps 5,000 prompts/month (shared across 3.x). Gemini 2.5: Search 1,500 RPD (free tier 500 RPD, shared Flash/Flash-Lite), Maps 1,500 RPD (10,000 for Pro).
7. 429 semantics (observed)
{"error":{"code":429,"status":"RESOURCE_EXHAUSTED",
"message":"You exceeded your current quota, please check your plan and billing details. ...\n* Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_requests, limit: 0, model: gemini-3.1-pro\n...Please retry in 50.868302469s.",
"details":[{"@type":"type.googleapis.com/google.rpc.Help","links":[{"description":"Learn more about Gemini API quotas","url":"https://ai.google.dev/gemini-api/docs/rate-limits"}]},
{"@type":"type.googleapis.com/google.rpc.QuotaFailure","violations":[{"quotaMetric":"generativelanguage.googleapis.com/generate_content_free_tier_requests","quotaId":"GenerateRequestsPerDayPerProjectPerModel-FreeTier","quotaDimensions":{"model":"gemini-3.1-pro","location":"global"}}]}]}}- Quota ids seen:
GenerateContentInputTokensPerModelPerDay-FreeTier,GenerateRequestsPerDayPerProjectPerModel-FreeTier,GenerateRequestsPerMinutePerProjectPerModel-FreeTier.quotaDimensions.modelis the base model (gemini-3.1-proforgemini-3.1-pro-previewand forgemini-pro-latest). - Retry delay only in the message text; no
Retry-Afterheader. - Retry guidance (troubleshooting.md): exponential backoff + jitter on 429/408/5xx only; SDKs retry 429/5xx automatically (Python: 4 attempts, ~1 s initial, 60 s max).
8. Headers: documented vs observed
| Result | |
|---|---|
| Documented rate-limit headers | none |
Observed on generateContent (200 and 429) |
Server-Timing: gfet4t7; dur=87, X-Gemini-Service-Tier: standard, Alt-Svc, Vary, X-Content-Type-Options, X-Frame-Options, X-XSS-Protection, Accept-Ranges, Transfer-Encoding, Content-Type, Date, Server |
Observed on /v1beta/openai/* |
Server-Timing, Set-Cookie, Cache-Control, Expires, P3P (no X-Gemini-Service-Tier) |
x-ratelimit-*, Retry-After, request-id |
absent — remaining quota is not exposed; use AI Studio dashboards; correlate with responseId in the body |
9. Observed for our key (account-specific, not documentation)
Free tier. OK: gemini-3.5-flash-lite, 3.5-flash, 3.8-flash, gemma-4-26b-a4b-it, gemini-flash-latest, OpenAI-compat chat. Restricted: gemini-3.1-pro-preview and gemini-pro-latest → 429 limit: 0; gemini-2.5-flash-lite → 404 "no longer available to new users".