# OpenAI — Rate limits (API Atlas) **Status:** DOCUMENTED (rate-limits guide + per-model "Rate limits" tables) · observed headers LIVE_VERIFIED for our key only (2026-09-18). Machine-readable twin: `generated/fragments/rate-limits/openai-rate-limits.json` (`per_model_documented_tiers` holds every model page table verbatim). **Sources:** https://developers.openai.com/api/docs/guides/rate-limits · https://developers.openai.com/api/docs/models/ (Rate limits section) · https://developers.openai.com/api/docs/guides/batch · https://developers.openai.com/api/docs/guides/fast-mode · https://developers.openai.com/api/docs/guides/error-codes · https://developers.openai.com/api/docs/changelog (2026-09-02 error update) **Last verified:** 2026-09-18 > Documented limits and observed headers are kept strictly apart. Nothing below is an account-specific promise; the console (Settings → Limits) is the source of truth for a given org/project. ## 1. Metrics (documented concepts) | Metric | Meaning | Where it applies | |---|---|---| | **RPM** | requests per minute | every model | | **RPD** | requests per day | some models; Free tier rows | | **TPM** | tokens per minute — each request counts `max(max_tokens, estimated prompt tokens)` | text/audio/image-token models | | **TPD** | tokens per day | some models | | **IPM** | images per minute | `gpt-image-*` (the pages show `TPM | IPM`; older `gpt-image-1`/`dall-e` pages show `img/min`) | | **Minutes-of-audio per minute** | admitted audio duration per minute | `gpt-realtime-whisper`, `gpt-realtime-translate` | | **Concurrent sessions** | simultaneous voice sessions | `gpt-live-1` (`v1/live/sessions`); Free tier unsupported | | **Batch queue limit** | total *input tokens* queued in pending batch jobs per model; freed when a batch completes | Batch API, per model | | **Long-context limit** | separate RPM/TPM/queue table for requests > 272K input tokens | GPT-5.4 / 5.5 / 5.6 / GPT-6 Astra (pages show "### Long Context — > 272K input tokens"; GPT-5.4 pages show "> 128k input tokens") | | **Shared limits** | several models can share one pool ("shared limit" list in the console) | organization | | **Project-scoped token limit** | optional project ceiling, surfaced by `x-ratelimit-*-project-tokens` | project | | **Monthly usage limit** | approved monthly spend per organization — distinct from user-configured spend limits and hard spend caps (which return 429 when hit) | organization | | **Vector store ingestion** | `/vector_stores/{id}/files` + `/file_batches` share **300 RPM per vector store** | File search | | **Ramp rate** | traffic may not grow too fast even below RPM/TPM; above ≈1M TPM grow ≤ +50 % every 15 min | every model (Fast mode requests may be downgraded to standard instead) | Limits are enforced at organization **and** project level (never per end-user) and vary per model. Unsuccessful requests still count toward RPM. ## 2. Usage tiers (documented) | Tier | Qualification | Monthly usage limit | |---|---|---| | Free | allowed geography (see supported-countries) | $100 | | Tier 1 | $5 paid | $100 | | Tier 2 | $50 paid | $500 | | Tier 3 | $100 paid | $1,000 | | Tier 4 | $250 paid | $5,000 | | Tier 5 | $1,000 paid | $200,000 | Graduation is automatic with cumulative spend. Beyond Tier 5: **Scale Tier** (predictable capacity for eligible models), **Reserved Tier** (GPT-5.6 and later), **Ultrafast mode** (limited preview for GPT-5.6 Sol, announced 2026-08-13). Enterprise agreements may add latency SLAs for Fast mode / Scale Tier (not for GPT-6 Astra Fast mode). ## 3. Per-model documented tier tables (representative excerpts) Every model page carries its own table; all 101 are in the JSON twin. Column sets vary by model type. ### GPT-6 Astra, GPT-5.6 Sol/Terra/Luna, GPT-5.6 Cyber, Daybreak, GPT-5.5 — "Standard" | Tier | RPM | TPM | Batch queue limit | |---|---|---|---| | Tier 1 | 500 | 500,000 | 1,500,000 | | Tier 2 | 5,000 | 1,000,000 | 3,000,000 | | Tier 3 | 5,000 | 2,000,000 | 100,000,000 | | Tier 4 | 10,000 | 4,000,000 | 200,000,000 | | Tier 5 | 15,000 | 40,000,000 | 15,000,000,000 | ### GPT-5.5 — "Long Context" (> 272K input tokens) | Tier | RPM | TPM | Batch queue limit | |---|---|---|---| | Tier 1 | 200 | 400,000 | 5,000,000 | | Tier 2 | 500 | 1,000,000 | 40,000,000 | | Tier 3 | 1,000 | 2,000,000 | 80,000,000 | | Tier 4 | 2,000 | 10,000,000 | 200,000,000 | | Tier 5 | 8,000 | 20,000,000 | 2,000,000,000 | ### gpt-realtime-2.1 (Realtime) | Tier | RPM | RPD | TPM | |---|---|---|---| | Tier 1 | 200 | 1,000 | 40,000 | | Tier 2 | 400 | — | 200,000 | | Tier 3 | 5,000 | — | 800,000 | | Tier 4 | 10,000 | — | 4,000,000 | | Tier 5 | 20,000 | — | 15,000,000 | ### gpt-image-2.5-flare / sunburst / gpt-image-2 (Images) | Tier | TPM | IPM | |---|---|---| | Tier 1 | 100,000 | 5 | | Tier 2 | 250,000 | 20 | | Tier 3 | 800,000 | 50 | | Tier 4 | 3,000,000 | 150 | | Tier 5 | 8,000,000 | 250 | ### gpt-live-1 (concurrent sessions; Free tier unsupported) | Tier | Concurrent sessions | |---|---| | Tier 1 | 25 | | Tier 2 | 50 | | Tier 3 | 200 | | Tier 4 | 300 | | Tier 5 | 500 | ### whisper-1 (audio, RPM/RPD) | Tier | RPM | RPD | |---|---|---| | free | 3 | 200 | | Tier 1 | 500 | — | | Tier 2 | 2,500 | — | | Tier 3 | 5,000 | — | | Tier 4 | 7,500 | — | | Tier 5 | 10,000 | — | ### text-embedding-3-large | Tier | RPM | RPD | TPM | Batch queue limit | |---|---|---|---|---| | free | 100 | 2,000 | 40,000 | — | | Tier 1 | 3,000 | — | 1,000,000 | 3,000,000 | | Tier 2 | 5,000 | — | 1,000,000 | 20,000,000 | | Tier 3 | 5,000 | — | 5,000,000 | 100,000,000 | | Tier 4 | 10,000 | — | 5,000,000 | 500,000,000 | | Tier 5 | 10,000 | — | 10,000,000 | 4,000,000,000 | ### sora-2 / sora-2-pro (RPM only) Tier 1 25 · Tier 2 50 · Tier 3 125 · Tier 4 200 · Tier 5 375 (Videos API shuts down 2026-09-24). ### gpt-oss-120b / gpt-oss-20b All tiers **0 / 0 / 0** — the open-weight models are documented with an endpoint table but no hosted quota (GET /v1/models/gpt-oss-120b → 404 with our key). Treat as *not served by the hosted API*. ## 4. Response headers | Header | Sample | Meaning | |---|---|---| | `Retry-After` | `56` | minimum seconds to wait; present on temporary 429 (`slow_down`, rate limit) and 503 (`server_is_overloaded`) — **not** on quota/billing errors | | `x-ratelimit-limit-requests` | `60` | RPM ceiling | | `x-ratelimit-limit-tokens` | `150000` | TPM ceiling | | `x-ratelimit-remaining-requests` | `59` | remaining requests | | `x-ratelimit-remaining-tokens` | `149984` | remaining tokens | | `x-ratelimit-reset-requests` | `1s` | time until request budget resets | | `x-ratelimit-reset-tokens` | `6m0s` | time until token budget resets | | `x-ratelimit-limit-project-tokens` | `60000` | project token limit (only when a project-scoped limit applies) | | `x-ratelimit-remaining-project-tokens` | `57000` | remaining project tokens | | `x-ratelimit-reset-project-tokens` | `3s` | project token reset | ## 5. Errors (since the 2026-09-02 update) | HTTP | `error.type` | `error.code` | Meaning | Action | |---|---|---|---|---| | 429 | `rate_limit_error` | `slow_down` | traffic ramped too quickly (can happen below RPM/TPM) | honour `Retry-After`, reduce, ramp gradually | | 429 | `rate_limit_error` | `rate_limit_exceeded` | RPM/TPM/RPD/TPD/IPM exhausted | exponential backoff + jitter; batch; lower `max_tokens` | | 429 | `insufficient_quota` | `insufficient_quota` | monthly usage / hard spend limit reached | do not retry; raise limits | | 503 | `service_unavailable_error` | `server_is_overloaded` | model temporarily overloaded | honour `Retry-After`, then retry with growing delay | Notes: endpoints that used to return 503 `slow_down` for both conditions now return 429 `slow_down` for ramp and 503 `server_is_overloaded` for overload. Video requests that used to return 429 `invalid_request_error`/`rate_limit_exceeded` follow the same new split. Official SDKs retry eligible 429/503 automatically; check how your SDK version honours long `Retry-After` values. For streaming, HTTP errors occur before the stream starts; errors after that arrive as stream events — never replay automatically after consuming output. ## 6. Mitigations documented by OpenAI - Exponential backoff with jitter (Tenacity / backoff / manual); treat `Retry-After` as a minimum. - Set `max_tokens`/`max_output_tokens` close to the expected output (TPM counts the max). - Use the **Batch API** (separate, much larger queue limits; 50 % price; 24 h window) for non-interactive work. - Batch several tasks into one request when RPM-bound but TPM-rich. - Fine-tuning limits: `GET /v1/fine_tuning/model_limits`. - Per-user caps in your own product to avoid abuse-driven exhaustion. ## 7. Observed (our key, 2026-09-18) — do not generalize | Request | Header | Value | |---|---|---| | `POST /v1/responses` (earlier probe) | `x-ratelimit-limit-requests` | 30000 | | `POST /v1/responses` (earlier probe) | `x-ratelimit-limit-tokens` | 180000000 | | `POST /v1/responses` model probes (gpt-5.6-luna … gpt-5.4-mini) | see `observed[]` in the JSON twin | per-model values captured by `scripts/probe_openai_models.py` | These values reflect our organization's tier and any custom limits; they are recorded only to show the header shapes. Documented Tier 5 values for GPT-5.6 are 15,000 RPM / 40M TPM, i.e. the observed figures do not match any published tier row — another reason not to infer tiers from headers. ## 8. Caveats - The models overview says free-tier rows exist for a few models (whisper-1, embeddings, gpt-4o-mini …); most new models start at Tier 1. - Tables on model pages are snapshots of the docs on 2026-09-18; OpenAI adjusts them without changelog entries. - Realtime, Live and Sora limits use different units (RPD, concurrent sessions, RPM only); compare like with like.