# Gemini API — rate limits, usage tiers and 429 semantics **Status:** DOCUMENTED + LIVE_DISCOVERED (headers and 429 bodies observed with our Free-tier key on 2026-09-19). Machine-readable twin: `generated/fragments/rate-limits/gemini-rate-limits.json`. **Sources:** https://ai.google.dev/gemini-api/docs/rate-limits · https://ai.google.dev/gemini-api/docs/billing · https://ai.google.dev/gemini-api/docs/priority-inference · https://ai.google.dev/gemini-api/docs/flex-inference · https://ai.google.dev/gemini-api/docs/batch-api · https://ai.google.dev/gemini-api/docs/live-api/session · https://ai.google.dev/gemini-api/docs/troubleshooting **Last verified:** 2026-09-18 (docs) / 2026-09-19 (live). ## 1. How limits work - Dimensions: **RPM** (requests/min), **TPM** (input tokens/min), **RPD** (requests/day, reset at midnight Pacific); model-specific extras **IPM** (images/min, image models) and **TPD**. - Applied **per Google Cloud project**, not per key. Exceeding any dimension → `429 RESOURCE_EXHAUSTED`. - Preview/experimental models have tighter limits. "Specified rate limits are not guaranteed and actual capacity may vary." - **The per-model RPM/TPM/RPD matrix is no longer published in the docs** — it lives in AI Studio (https://aistudio.google.com/rate-limit, account-specific). This atlas therefore records tiers, spend limits and batch/Live limits only; per-model numbers are intentionally not invented. ## 2. Usage tiers | Tier | Qualification | Billing-tier spend cap | Spend-based rate limit (rolling 10 min) | |---|---|---|---| | Free | active project or free trial | n/a | n/a | | Tier 1 | link an active Cloud Billing account | $250 | $10 | | Tier 2 | $100 cumulative Google Cloud spend + 3 days since first payment | $2,000 | $50 | | Tier 3 | $1,000 cumulative + 30 days | $20,000 – $100,000+ | $200 | Free → Tier 1 is instant; later upgrades within ~10 min; upgrades may be denied "based on other factors". Free tier excludes Pro models (observed: `limit: 0`), image/video/music generation and agents. Project-level spend caps and billing-account tier caps (2026-03) stop traffic instead of overspending. ## 3. Service-tier limits | Tier | Rate limit | Behaviour on overflow | |---|---|---| | standard | model/tier limit (AI Studio) | 429 | | priority | **0.3× standard** (own bucket, also counted against interactive traffic) | graceful downgrade to standard, no error | | flex | own limits; sheddable | 429 when capacity is shed → retry (docs provide retry loops and per-request timeouts) | ## 4. Batch API limits Concurrent batch jobs **100**; input file **2 GB**; file storage **20 GB**; turnaround target 24 h; price 50 %. Enqueued-token ceilings per model (all active jobs): | Model | Tier 1 | Tier 2 | Tier 3 | |---|---|---|---| | gemini-3.1-pro-preview | 5M | 500M | 1B | | gemini-3.5-flash-lite, gemini-3.1-flash-lite (+preview), gemini-2.5-flash-lite (+preview) | 10M | 500M | 1B | | gemini-3.8-flash, 3.7-flash, 3.6-flash, 3.5-flash, 2.5-flash (+preview, image-preview) | 3M | 400M | 1B | | gemini-2.5-pro | 5M | 500M | 1B | | gemini-2.5-pro-preview-tts | 25k | 100k | 1M | | gemini-2.5-flash-preview-tts | 100k | 100k | 4M | | gemini-2.0-flash / flash-lite (shut down) | 10M | 1B | 5B | | gemini-3.1-flash-image-preview | 1M | 250M | 750M | | gemini-3.1-flash-lite-image, gemini-3-pro-image-preview | 2M | 270M | 1B | | gemini-embedding (001 / 2) | 500k | 5M | 10M | The table still names shut-down 2.0/2.5 previews and omits the GA image ids (`gemini-3.1-flash-image`, `gemini-3-pro-image`) — assume the preview rows apply. ## 5. Live API session limits Audio-only session **15 min**, audio+video **2 min** without `contextWindowCompression`; WebSocket connection ≈ **10 min** (server sends `goAway`); `sessionResumption` handles valid **2 h** after termination. Ephemeral tokens (`POST /v1beta/auth_tokens`) for client-side connections. No published concurrent-session limit. ## 6. Grounding quotas (free allowances) Gemini 3.x: Google Search 5,000 requests/month, Maps 5,000 prompts/month (shared across 3.x). Gemini 2.5: Search 1,500 RPD (free tier 500 RPD, shared Flash/Flash-Lite), Maps 1,500 RPD (10,000 for Pro). ## 7. 429 semantics (observed) ```json {"error":{"code":429,"status":"RESOURCE_EXHAUSTED", "message":"You exceeded your current quota, please check your plan and billing details. ...\n* Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_requests, limit: 0, model: gemini-3.1-pro\n...Please retry in 50.868302469s.", "details":[{"@type":"type.googleapis.com/google.rpc.Help","links":[{"description":"Learn more about Gemini API quotas","url":"https://ai.google.dev/gemini-api/docs/rate-limits"}]}, {"@type":"type.googleapis.com/google.rpc.QuotaFailure","violations":[{"quotaMetric":"generativelanguage.googleapis.com/generate_content_free_tier_requests","quotaId":"GenerateRequestsPerDayPerProjectPerModel-FreeTier","quotaDimensions":{"model":"gemini-3.1-pro","location":"global"}}]}]}} ``` - Quota ids seen: `GenerateContentInputTokensPerModelPerDay-FreeTier`, `GenerateRequestsPerDayPerProjectPerModel-FreeTier`, `GenerateRequestsPerMinutePerProjectPerModel-FreeTier`. `quotaDimensions.model` is the **base model** (`gemini-3.1-pro` for `gemini-3.1-pro-preview` and for `gemini-pro-latest`). - Retry delay only in the message text; **no `Retry-After` header**. - Retry guidance (troubleshooting.md): exponential backoff + jitter on 429/408/5xx only; SDKs retry 429/5xx automatically (Python: 4 attempts, ~1 s initial, 60 s max). ## 8. Headers: documented vs observed | | Result | |---|---| | Documented rate-limit headers | none | | Observed on `generateContent` (200 and 429) | `Server-Timing: gfet4t7; dur=87`, `X-Gemini-Service-Tier: standard`, `Alt-Svc`, `Vary`, `X-Content-Type-Options`, `X-Frame-Options`, `X-XSS-Protection`, `Accept-Ranges`, `Transfer-Encoding`, `Content-Type`, `Date`, `Server` | | Observed on `/v1beta/openai/*` | `Server-Timing`, `Set-Cookie`, `Cache-Control`, `Expires`, `P3P` (no `X-Gemini-Service-Tier`) | | `x-ratelimit-*`, `Retry-After`, request-id | **absent** — remaining quota is not exposed; use AI Studio dashboards; correlate with `responseId` in the body | ## 9. Observed for our key (account-specific, not documentation) Free tier. OK: gemini-3.5-flash-lite, 3.5-flash, 3.8-flash, gemma-4-26b-a4b-it, gemini-flash-latest, OpenAI-compat chat. Restricted: gemini-3.1-pro-preview and gemini-pro-latest → 429 `limit: 0`; gemini-2.5-flash-lite → 404 "no longer available to new users".