SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
14 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
9.7 KB · 179 lines markdown
Rendered Raw Blame History
1# OpenAI — Rate limits (API Atlas)23**Status:** DOCUMENTED (rate-limits guide + per-model "Rate limits" tables) · observed headers LIVE_VERIFIED for our key only (2026-09-18). Machine-readable twin: `generated/fragments/rate-limits/openai-rate-limits.json` (`per_model_documented_tiers` holds every model page table verbatim).45**Sources:** https://developers.openai.com/api/docs/guides/rate-limits · https://developers.openai.com/api/docs/models/<id> (Rate limits section) · https://developers.openai.com/api/docs/guides/batch · https://developers.openai.com/api/docs/guides/fast-mode · https://developers.openai.com/api/docs/guides/error-codes · https://developers.openai.com/api/docs/changelog (2026-09-02 error update)67**Last verified:** 2026-09-1889> Documented limits and observed headers are kept strictly apart. Nothing below is an account-specific promise; the console (Settings → Limits) is the source of truth for a given org/project.1011## 1. Metrics (documented concepts)1213| Metric | Meaning | Where it applies |14|---|---|---|15| **RPM** | requests per minute | every model |16| **RPD** | requests per day | some models; Free tier rows |17| **TPM** | tokens per minute — each request counts `max(max_tokens, estimated prompt tokens)` | text/audio/image-token models |18| **TPD** | tokens per day | some models |19| **IPM** | images per minute | `gpt-image-*` (the pages show `TPM | IPM`; older `gpt-image-1`/`dall-e` pages show `img/min`) |20| **Minutes-of-audio per minute** | admitted audio duration per minute | `gpt-realtime-whisper`, `gpt-realtime-translate` |21| **Concurrent sessions** | simultaneous voice sessions | `gpt-live-1` (`v1/live/sessions`); Free tier unsupported |22| **Batch queue limit** | total *input tokens* queued in pending batch jobs per model; freed when a batch completes | Batch API, per model |23| **Long-context limit** | separate RPM/TPM/queue table for requests > 272K input tokens | GPT-5.4 / 5.5 / 5.6 / GPT-6 Astra (pages show "### Long Context — > 272K input tokens"; GPT-5.4 pages show "> 128k input tokens") |24| **Shared limits** | several models can share one pool ("shared limit" list in the console) | organization |25| **Project-scoped token limit** | optional project ceiling, surfaced by `x-ratelimit-*-project-tokens` | project |26| **Monthly usage limit** | approved monthly spend per organization — distinct from user-configured spend limits and hard spend caps (which return 429 when hit) | organization |27| **Vector store ingestion** | `/vector_stores/{id}/files` + `/file_batches` share **300 RPM per vector store** | File search |28| **Ramp rate** | traffic may not grow too fast even below RPM/TPM; above ≈1M TPM grow ≤ +50 % every 15 min | every model (Fast mode requests may be downgraded to standard instead) |2930Limits are enforced at organization **and** project level (never per end-user) and vary per model. Unsuccessful requests still count toward RPM.3132## 2. Usage tiers (documented)3334| Tier | Qualification | Monthly usage limit |35|---|---|---|36| Free | allowed geography (see supported-countries) | $100 |37| Tier 1 | $5 paid | $100 |38| Tier 2 | $50 paid | $500 |39| Tier 3 | $100 paid | $1,000 |40| Tier 4 | $250 paid | $5,000 |41| Tier 5 | $1,000 paid | $200,000 |4243Graduation is automatic with cumulative spend. Beyond Tier 5: **Scale Tier** (predictable capacity for eligible models), **Reserved Tier** (GPT-5.6 and later), **Ultrafast mode** (limited preview for GPT-5.6 Sol, announced 2026-08-13). Enterprise agreements may add latency SLAs for Fast mode / Scale Tier (not for GPT-6 Astra Fast mode).4445## 3. Per-model documented tier tables (representative excerpts)4647Every model page carries its own table; all 101 are in the JSON twin. Column sets vary by model type.4849### GPT-6 Astra, GPT-5.6 Sol/Terra/Luna, GPT-5.6 Cyber, Daybreak, GPT-5.5 — "Standard"5051| Tier | RPM | TPM | Batch queue limit |52|---|---|---|---|53| Tier 1 | 500 | 500,000 | 1,500,000 |54| Tier 2 | 5,000 | 1,000,000 | 3,000,000 |55| Tier 3 | 5,000 | 2,000,000 | 100,000,000 |56| Tier 4 | 10,000 | 4,000,000 | 200,000,000 |57| Tier 5 | 15,000 | 40,000,000 | 15,000,000,000 |5859### GPT-5.5 — "Long Context" (> 272K input tokens)6061| Tier | RPM | TPM | Batch queue limit |62|---|---|---|---|63| Tier 1 | 200 | 400,000 | 5,000,000 |64| Tier 2 | 500 | 1,000,000 | 40,000,000 |65| Tier 3 | 1,000 | 2,000,000 | 80,000,000 |66| Tier 4 | 2,000 | 10,000,000 | 200,000,000 |67| Tier 5 | 8,000 | 20,000,000 | 2,000,000,000 |6869### gpt-realtime-2.1 (Realtime)7071| Tier | RPM | RPD | TPM |72|---|---|---|---|73| Tier 1 | 200 | 1,000 | 40,000 |74| Tier 2 | 400 | — | 200,000 |75| Tier 3 | 5,000 | — | 800,000 |76| Tier 4 | 10,000 | — | 4,000,000 |77| Tier 5 | 20,000 | — | 15,000,000 |7879### gpt-image-2.5-flare / sunburst / gpt-image-2 (Images)8081| Tier | TPM | IPM |82|---|---|---|83| Tier 1 | 100,000 | 5 |84| Tier 2 | 250,000 | 20 |85| Tier 3 | 800,000 | 50 |86| Tier 4 | 3,000,000 | 150 |87| Tier 5 | 8,000,000 | 250 |8889### gpt-live-1 (concurrent sessions; Free tier unsupported)9091| Tier | Concurrent sessions |92|---|---|93| Tier 1 | 25 |94| Tier 2 | 50 |95| Tier 3 | 200 |96| Tier 4 | 300 |97| Tier 5 | 500 |9899### whisper-1 (audio, RPM/RPD)100101| Tier | RPM | RPD |102|---|---|---|103| free | 3 | 200 |104| Tier 1 | 500 | — |105| Tier 2 | 2,500 | — |106| Tier 3 | 5,000 | — |107| Tier 4 | 7,500 | — |108| Tier 5 | 10,000 | — |109110### text-embedding-3-large111112| Tier | RPM | RPD | TPM | Batch queue limit |113|---|---|---|---|---|114| free | 100 | 2,000 | 40,000 | — |115| Tier 1 | 3,000 | — | 1,000,000 | 3,000,000 |116| Tier 2 | 5,000 | — | 1,000,000 | 20,000,000 |117| Tier 3 | 5,000 | — | 5,000,000 | 100,000,000 |118| Tier 4 | 10,000 | — | 5,000,000 | 500,000,000 |119| Tier 5 | 10,000 | — | 10,000,000 | 4,000,000,000 |120121### sora-2 / sora-2-pro (RPM only)122123Tier 1 25 · Tier 2 50 · Tier 3 125 · Tier 4 200 · Tier 5 375 (Videos API shuts down 2026-09-24).124125### gpt-oss-120b / gpt-oss-20b126127All tiers **0 / 0 / 0** — the open-weight models are documented with an endpoint table but no hosted quota (GET /v1/models/gpt-oss-120b → 404 with our key). Treat as *not served by the hosted API*.128129## 4. Response headers130131| Header | Sample | Meaning |132|---|---|---|133| `Retry-After` | `56` | minimum seconds to wait; present on temporary 429 (`slow_down`, rate limit) and 503 (`server_is_overloaded`) — **not** on quota/billing errors |134| `x-ratelimit-limit-requests` | `60` | RPM ceiling |135| `x-ratelimit-limit-tokens` | `150000` | TPM ceiling |136| `x-ratelimit-remaining-requests` | `59` | remaining requests |137| `x-ratelimit-remaining-tokens` | `149984` | remaining tokens |138| `x-ratelimit-reset-requests` | `1s` | time until request budget resets |139| `x-ratelimit-reset-tokens` | `6m0s` | time until token budget resets |140| `x-ratelimit-limit-project-tokens` | `60000` | project token limit (only when a project-scoped limit applies) |141| `x-ratelimit-remaining-project-tokens` | `57000` | remaining project tokens |142| `x-ratelimit-reset-project-tokens` | `3s` | project token reset |143144## 5. Errors (since the 2026-09-02 update)145146| HTTP | `error.type` | `error.code` | Meaning | Action |147|---|---|---|---|---|148| 429 | `rate_limit_error` | `slow_down` | traffic ramped too quickly (can happen below RPM/TPM) | honour `Retry-After`, reduce, ramp gradually |149| 429 | `rate_limit_error` | `rate_limit_exceeded` | RPM/TPM/RPD/TPD/IPM exhausted | exponential backoff + jitter; batch; lower `max_tokens` |150| 429 | `insufficient_quota` | `insufficient_quota` | monthly usage / hard spend limit reached | do not retry; raise limits |151| 503 | `service_unavailable_error` | `server_is_overloaded` | model temporarily overloaded | honour `Retry-After`, then retry with growing delay |152153Notes: endpoints that used to return 503 `slow_down` for both conditions now return 429 `slow_down` for ramp and 503 `server_is_overloaded` for overload. Video requests that used to return 429 `invalid_request_error`/`rate_limit_exceeded` follow the same new split. Official SDKs retry eligible 429/503 automatically; check how your SDK version honours long `Retry-After` values. For streaming, HTTP errors occur before the stream starts; errors after that arrive as stream events — never replay automatically after consuming output.154155## 6. Mitigations documented by OpenAI156157- Exponential backoff with jitter (Tenacity / backoff / manual); treat `Retry-After` as a minimum.158- Set `max_tokens`/`max_output_tokens` close to the expected output (TPM counts the max).159- Use the **Batch API** (separate, much larger queue limits; 50 % price; 24 h window) for non-interactive work.160- Batch several tasks into one request when RPM-bound but TPM-rich.161- Fine-tuning limits: `GET /v1/fine_tuning/model_limits`.162- Per-user caps in your own product to avoid abuse-driven exhaustion.163164## 7. Observed (our key, 2026-09-18) — do not generalize165166| Request | Header | Value |167|---|---|---|168| `POST /v1/responses` (earlier probe) | `x-ratelimit-limit-requests` | 30000 |169| `POST /v1/responses` (earlier probe) | `x-ratelimit-limit-tokens` | 180000000 |170| `POST /v1/responses` model probes (gpt-5.6-luna … gpt-5.4-mini) | see `observed[]` in the JSON twin | per-model values captured by `scripts/probe_openai_models.py` |171172These values reflect our organization's tier and any custom limits; they are recorded only to show the header shapes. Documented Tier 5 values for GPT-5.6 are 15,000 RPM / 40M TPM, i.e. the observed figures do not match any published tier row — another reason not to infer tiers from headers.173174## 8. Caveats175176- The models overview says free-tier rows exist for a few models (whisper-1, embeddings, gpt-4o-mini …); most new models start at Tier 1.177- Tables on model pages are snapshots of the docs on 2026-09-18; OpenAI adjusts them without changelog entries.178- Realtime, Live and Sora limits use different units (RPD, concurrent sessions, RPM only); compare like with like.179