# OpenAI — Rate limits (API Atlas)
Status: DOCUMENTED (rate-limits guide + per-model "Rate limits" tables) · observed headers LIVE_VERIFIED for our key only (2026-09-18). Machine-readable twin: generated/fragments/rate-limits/openai-rate-limits.json (per_model_documented_tiers holds every model page table verbatim).
Sources: https://developers.openai.com/api/docs/guides/rate-limits · https://developers.openai.com/api/docs/models/ (Rate limits section) · https://developers.openai.com/api/docs/guides/batch · https://developers.openai.com/api/docs/guides/fast-mode · https://developers.openai.com/api/docs/guides/error-codes · https://developers.openai.com/api/docs/changelog (2026-09-02 error update)
Last verified: 2026-09-18
Documented limits and observed headers are kept strictly apart. Nothing below is an account-specific promise; the console (Settings → Limits) is the source of truth for a given org/project.
# 1. Metrics (documented concepts)
| Metric |
Meaning |
Where it applies |
| RPM |
requests per minute |
every model |
| RPD |
requests per day |
some models; Free tier rows |
| TPM |
tokens per minute — each request counts max(max_tokens, estimated prompt tokens) |
text/audio/image-token models |
| TPD |
tokens per day |
some models |
| IPM |
images per minute |
gpt-image-* (the pages show `TPM |
| Minutes-of-audio per minute |
admitted audio duration per minute |
gpt-realtime-whisper, gpt-realtime-translate |
| Concurrent sessions |
simultaneous voice sessions |
gpt-live-1 (v1/live/sessions); Free tier unsupported |
| Batch queue limit |
total input tokens queued in pending batch jobs per model; freed when a batch completes |
Batch API, per model |
| Long-context limit |
separate RPM/TPM/queue table for requests > 272K input tokens |
GPT-5.4 / 5.5 / 5.6 / GPT-6 Astra (pages show "### Long Context — > 272K input tokens"; GPT-5.4 pages show "> 128k input tokens") |
| Shared limits |
several models can share one pool ("shared limit" list in the console) |
organization |
| Project-scoped token limit |
optional project ceiling, surfaced by x-ratelimit-*-project-tokens |
project |
| Monthly usage limit |
approved monthly spend per organization — distinct from user-configured spend limits and hard spend caps (which return 429 when hit) |
organization |
| Vector store ingestion |
/vector_stores/{id}/files + /file_batches share 300 RPM per vector store |
File search |
| Ramp rate |
traffic may not grow too fast even below RPM/TPM; above ≈1M TPM grow ≤ +50 % every 15 min |
every model (Fast mode requests may be downgraded to standard instead) |
Limits are enforced at organization and project level (never per end-user) and vary per model. Unsuccessful requests still count toward RPM.
# 2. Usage tiers (documented)
| Tier |
Qualification |
Monthly usage limit |
| Free |
allowed geography (see supported-countries) |
$100 |
| Tier 1 |
$5 paid |
$100 |
| Tier 2 |
$50 paid |
$500 |
| Tier 3 |
$100 paid |
$1,000 |
| Tier 4 |
$250 paid |
$5,000 |
| Tier 5 |
$1,000 paid |
$200,000 |
Graduation is automatic with cumulative spend. Beyond Tier 5: Scale Tier (predictable capacity for eligible models), Reserved Tier (GPT-5.6 and later), Ultrafast mode (limited preview for GPT-5.6 Sol, announced 2026-08-13). Enterprise agreements may add latency SLAs for Fast mode / Scale Tier (not for GPT-6 Astra Fast mode).
# 3. Per-model documented tier tables (representative excerpts)
Every model page carries its own table; all 101 are in the JSON twin. Column sets vary by model type.
# GPT-6 Astra, GPT-5.6 Sol/Terra/Luna, GPT-5.6 Cyber, Daybreak, GPT-5.5 — "Standard"
| Tier |
RPM |
TPM |
Batch queue limit |
| Tier 1 |
500 |
500,000 |
1,500,000 |
| Tier 2 |
5,000 |
1,000,000 |
3,000,000 |
| Tier 3 |
5,000 |
2,000,000 |
100,000,000 |
| Tier 4 |
10,000 |
4,000,000 |
200,000,000 |
| Tier 5 |
15,000 |
40,000,000 |
15,000,000,000 |
# GPT-5.5 — "Long Context" (> 272K input tokens)
| Tier |
RPM |
TPM |
Batch queue limit |
| Tier 1 |
200 |
400,000 |
5,000,000 |
| Tier 2 |
500 |
1,000,000 |
40,000,000 |
| Tier 3 |
1,000 |
2,000,000 |
80,000,000 |
| Tier 4 |
2,000 |
10,000,000 |
200,000,000 |
| Tier 5 |
8,000 |
20,000,000 |
2,000,000,000 |
# gpt-realtime-2.1 (Realtime)
| Tier |
RPM |
RPD |
TPM |
| Tier 1 |
200 |
1,000 |
40,000 |
| Tier 2 |
400 |
— |
200,000 |
| Tier 3 |
5,000 |
— |
800,000 |
| Tier 4 |
10,000 |
— |
4,000,000 |
| Tier 5 |
20,000 |
— |
15,000,000 |
# gpt-image-2.5-flare / sunburst / gpt-image-2 (Images)
| Tier |
TPM |
IPM |
| Tier 1 |
100,000 |
5 |
| Tier 2 |
250,000 |
20 |
| Tier 3 |
800,000 |
50 |
| Tier 4 |
3,000,000 |
150 |
| Tier 5 |
8,000,000 |
250 |
# gpt-live-1 (concurrent sessions; Free tier unsupported)
| Tier |
Concurrent sessions |
| Tier 1 |
25 |
| Tier 2 |
50 |
| Tier 3 |
200 |
| Tier 4 |
300 |
| Tier 5 |
500 |
# whisper-1 (audio, RPM/RPD)
| Tier |
RPM |
RPD |
| free |
3 |
200 |
| Tier 1 |
500 |
— |
| Tier 2 |
2,500 |
— |
| Tier 3 |
5,000 |
— |
| Tier 4 |
7,500 |
— |
| Tier 5 |
10,000 |
— |
# text-embedding-3-large
| Tier |
RPM |
RPD |
TPM |
Batch queue limit |
| free |
100 |
2,000 |
40,000 |
— |
| Tier 1 |
3,000 |
— |
1,000,000 |
3,000,000 |
| Tier 2 |
5,000 |
— |
1,000,000 |
20,000,000 |
| Tier 3 |
5,000 |
— |
5,000,000 |
100,000,000 |
| Tier 4 |
10,000 |
— |
5,000,000 |
500,000,000 |
| Tier 5 |
10,000 |
— |
10,000,000 |
4,000,000,000 |
# sora-2 / sora-2-pro (RPM only)
Tier 1 25 · Tier 2 50 · Tier 3 125 · Tier 4 200 · Tier 5 375 (Videos API shuts down 2026-09-24).
# gpt-oss-120b / gpt-oss-20b
All tiers 0 / 0 / 0 — the open-weight models are documented with an endpoint table but no hosted quota (GET /v1/models/gpt-oss-120b → 404 with our key). Treat as not served by the hosted API.
| Header |
Sample |
Meaning |
Retry-After |
56 |
minimum seconds to wait; present on temporary 429 (slow_down, rate limit) and 503 (server_is_overloaded) — not on quota/billing errors |
x-ratelimit-limit-requests |
60 |
RPM ceiling |
x-ratelimit-limit-tokens |
150000 |
TPM ceiling |
x-ratelimit-remaining-requests |
59 |
remaining requests |
x-ratelimit-remaining-tokens |
149984 |
remaining tokens |
x-ratelimit-reset-requests |
1s |
time until request budget resets |
x-ratelimit-reset-tokens |
6m0s |
time until token budget resets |
x-ratelimit-limit-project-tokens |
60000 |
project token limit (only when a project-scoped limit applies) |
x-ratelimit-remaining-project-tokens |
57000 |
remaining project tokens |
x-ratelimit-reset-project-tokens |
3s |
project token reset |
# 5. Errors (since the 2026-09-02 update)
| HTTP |
error.type |
error.code |
Meaning |
Action |
| 429 |
rate_limit_error |
slow_down |
traffic ramped too quickly (can happen below RPM/TPM) |
honour Retry-After, reduce, ramp gradually |
| 429 |
rate_limit_error |
rate_limit_exceeded |
RPM/TPM/RPD/TPD/IPM exhausted |
exponential backoff + jitter; batch; lower max_tokens |
| 429 |
insufficient_quota |
insufficient_quota |
monthly usage / hard spend limit reached |
do not retry; raise limits |
| 503 |
service_unavailable_error |
server_is_overloaded |
model temporarily overloaded |
honour Retry-After, then retry with growing delay |
Notes: endpoints that used to return 503 slow_down for both conditions now return 429 slow_down for ramp and 503 server_is_overloaded for overload. Video requests that used to return 429 invalid_request_error/rate_limit_exceeded follow the same new split. Official SDKs retry eligible 429/503 automatically; check how your SDK version honours long Retry-After values. For streaming, HTTP errors occur before the stream starts; errors after that arrive as stream events — never replay automatically after consuming output.
# 6. Mitigations documented by OpenAI
- Exponential backoff with jitter (Tenacity / backoff / manual); treat
Retry-After as a minimum.
- Set
max_tokens/max_output_tokens close to the expected output (TPM counts the max).
- Use the Batch API (separate, much larger queue limits; 50 % price; 24 h window) for non-interactive work.
- Batch several tasks into one request when RPM-bound but TPM-rich.
- Fine-tuning limits:
GET /v1/fine_tuning/model_limits.
- Per-user caps in your own product to avoid abuse-driven exhaustion.
# 7. Observed (our key, 2026-09-18) — do not generalize
| Request |
Header |
Value |
POST /v1/responses (earlier probe) |
x-ratelimit-limit-requests |
30000 |
POST /v1/responses (earlier probe) |
x-ratelimit-limit-tokens |
180000000 |
POST /v1/responses model probes (gpt-5.6-luna … gpt-5.4-mini) |
see observed[] in the JSON twin |
per-model values captured by scripts/probe_openai_models.py |
These values reflect our organization's tier and any custom limits; they are recorded only to show the header shapes. Documented Tier 5 values for GPT-5.6 are 15,000 RPM / 40M TPM, i.e. the observed figures do not match any published tier row — another reason not to infer tiers from headers.
# 8. Caveats
- The models overview says free-tier rows exist for a few models (whisper-1, embeddings, gpt-4o-mini …); most new models start at Tier 1.
- Tables on model pages are snapshots of the docs on 2026-09-18; OpenAI adjusts them without changelog entries.
- Realtime, Live and Sora limits use different units (RPD, concurrent sessions, RPM only); compare like with like.