Python 88.3%
TypeScript 7.6%
Shell 4.1%
1# Gemini API — rate limits, usage tiers and 429 semantics23**Status:** DOCUMENTED + LIVE_DISCOVERED (headers and 429 bodies observed with our Free-tier key on 2026-09-19). Machine-readable twin: `generated/fragments/rate-limits/gemini-rate-limits.json`.4**Sources:** https://ai.google.dev/gemini-api/docs/rate-limits · https://ai.google.dev/gemini-api/docs/billing · https://ai.google.dev/gemini-api/docs/priority-inference · https://ai.google.dev/gemini-api/docs/flex-inference · https://ai.google.dev/gemini-api/docs/batch-api · https://ai.google.dev/gemini-api/docs/live-api/session · https://ai.google.dev/gemini-api/docs/troubleshooting5**Last verified:** 2026-09-18 (docs) / 2026-09-19 (live).67## 1. How limits work8- Dimensions: **RPM** (requests/min), **TPM** (input tokens/min), **RPD** (requests/day, reset at midnight Pacific); model-specific extras **IPM** (images/min, image models) and **TPD**.9- Applied **per Google Cloud project**, not per key. Exceeding any dimension → `429 RESOURCE_EXHAUSTED`.10- Preview/experimental models have tighter limits. "Specified rate limits are not guaranteed and actual capacity may vary."11- **The per-model RPM/TPM/RPD matrix is no longer published in the docs** — it lives in AI Studio (https://aistudio.google.com/rate-limit, account-specific). This atlas therefore records tiers, spend limits and batch/Live limits only; per-model numbers are intentionally not invented.1213## 2. Usage tiers1415| Tier | Qualification | Billing-tier spend cap | Spend-based rate limit (rolling 10 min) |16|---|---|---|---|17| Free | active project or free trial | n/a | n/a |18| Tier 1 | link an active Cloud Billing account | $250 | $10 |19| Tier 2 | $100 cumulative Google Cloud spend + 3 days since first payment | $2,000 | $50 |20| Tier 3 | $1,000 cumulative + 30 days | $20,000 – $100,000+ | $200 |2122Free → Tier 1 is instant; later upgrades within ~10 min; upgrades may be denied "based on other factors". Free tier excludes Pro models (observed: `limit: 0`), image/video/music generation and agents. Project-level spend caps and billing-account tier caps (2026-03) stop traffic instead of overspending.2324## 3. Service-tier limits25| Tier | Rate limit | Behaviour on overflow |26|---|---|---|27| standard | model/tier limit (AI Studio) | 429 |28| priority | **0.3× standard** (own bucket, also counted against interactive traffic) | graceful downgrade to standard, no error |29| flex | own limits; sheddable | 429 when capacity is shed → retry (docs provide retry loops and per-request timeouts) |3031## 4. Batch API limits32Concurrent batch jobs **100**; input file **2 GB**; file storage **20 GB**; turnaround target 24 h; price 50 %. Enqueued-token ceilings per model (all active jobs):3334| Model | Tier 1 | Tier 2 | Tier 3 |35|---|---|---|---|36| gemini-3.1-pro-preview | 5M | 500M | 1B |37| gemini-3.5-flash-lite, gemini-3.1-flash-lite (+preview), gemini-2.5-flash-lite (+preview) | 10M | 500M | 1B |38| gemini-3.8-flash, 3.7-flash, 3.6-flash, 3.5-flash, 2.5-flash (+preview, image-preview) | 3M | 400M | 1B |39| gemini-2.5-pro | 5M | 500M | 1B |40| gemini-2.5-pro-preview-tts | 25k | 100k | 1M |41| gemini-2.5-flash-preview-tts | 100k | 100k | 4M |42| gemini-2.0-flash / flash-lite (shut down) | 10M | 1B | 5B |43| gemini-3.1-flash-image-preview | 1M | 250M | 750M |44| gemini-3.1-flash-lite-image, gemini-3-pro-image-preview | 2M | 270M | 1B |45| gemini-embedding (001 / 2) | 500k | 5M | 10M |4647The table still names shut-down 2.0/2.5 previews and omits the GA image ids (`gemini-3.1-flash-image`, `gemini-3-pro-image`) — assume the preview rows apply.4849## 5. Live API session limits50Audio-only session **15 min**, audio+video **2 min** without `contextWindowCompression`; WebSocket connection ≈ **10 min** (server sends `goAway`); `sessionResumption` handles valid **2 h** after termination. Ephemeral tokens (`POST /v1beta/auth_tokens`) for client-side connections. No published concurrent-session limit.5152## 6. Grounding quotas (free allowances)53Gemini 3.x: Google Search 5,000 requests/month, Maps 5,000 prompts/month (shared across 3.x). Gemini 2.5: Search 1,500 RPD (free tier 500 RPD, shared Flash/Flash-Lite), Maps 1,500 RPD (10,000 for Pro).5455## 7. 429 semantics (observed)56```json57{"error":{"code":429,"status":"RESOURCE_EXHAUSTED",58 "message":"You exceeded your current quota, please check your plan and billing details. ...\n* Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_requests, limit: 0, model: gemini-3.1-pro\n...Please retry in 50.868302469s.",59 "details":[{"@type":"type.googleapis.com/google.rpc.Help","links":[{"description":"Learn more about Gemini API quotas","url":"https://ai.google.dev/gemini-api/docs/rate-limits"}]},60 {"@type":"type.googleapis.com/google.rpc.QuotaFailure","violations":[{"quotaMetric":"generativelanguage.googleapis.com/generate_content_free_tier_requests","quotaId":"GenerateRequestsPerDayPerProjectPerModel-FreeTier","quotaDimensions":{"model":"gemini-3.1-pro","location":"global"}}]}]}}61```62- Quota ids seen: `GenerateContentInputTokensPerModelPerDay-FreeTier`, `GenerateRequestsPerDayPerProjectPerModel-FreeTier`, `GenerateRequestsPerMinutePerProjectPerModel-FreeTier`. `quotaDimensions.model` is the **base model** (`gemini-3.1-pro` for `gemini-3.1-pro-preview` and for `gemini-pro-latest`).63- Retry delay only in the message text; **no `Retry-After` header**.64- Retry guidance (troubleshooting.md): exponential backoff + jitter on 429/408/5xx only; SDKs retry 429/5xx automatically (Python: 4 attempts, ~1 s initial, 60 s max).6566## 8. Headers: documented vs observed67| | Result |68|---|---|69| Documented rate-limit headers | none |70| Observed on `generateContent` (200 and 429) | `Server-Timing: gfet4t7; dur=87`, `X-Gemini-Service-Tier: standard`, `Alt-Svc`, `Vary`, `X-Content-Type-Options`, `X-Frame-Options`, `X-XSS-Protection`, `Accept-Ranges`, `Transfer-Encoding`, `Content-Type`, `Date`, `Server` |71| Observed on `/v1beta/openai/*` | `Server-Timing`, `Set-Cookie`, `Cache-Control`, `Expires`, `P3P` (no `X-Gemini-Service-Tier`) |72| `x-ratelimit-*`, `Retry-After`, request-id | **absent** — remaining quota is not exposed; use AI Studio dashboards; correlate with `responseId` in the body |7374## 9. Observed for our key (account-specific, not documentation)75Free tier. OK: gemini-3.5-flash-lite, 3.5-flash, 3.8-flash, gemma-4-26b-a4b-it, gemini-flash-latest, OpenAI-compat chat. Restricted: gemini-3.1-pro-preview and gemini-pro-latest → 429 `limit: 0`; gemini-2.5-flash-lite → 404 "no longer available to new users".76