# Anthropic rate limits and spend limits **Status:** `DOCUMENTED` (published tables) + `LIVE_DISCOVERED` (headers observed with our key on 2026-09-18; account-specific). **Sources:** https://platform.claude.com/docs/en/api/rate-limits · api/service-tiers · build-with-claude/fast-mode · build-with-claude/token-counting · build-with-claude/files · api/overview · release-notes/overview (retrieved 2026-09-18). **Last verified:** 2026-09-18 · Machine-readable: `generated/fragments/rate-limits/anthropic-rate-limits.json`. ## 1. Model - Two kinds of limits: **spend limits** (monthly USD cap per usage tier) and **rate limits** (RPM / ITPM / OTPM per model class). - Organization-level, tier assigned automatically from usage history; new orgs may start in an **Evaluation** tier below the published numbers. Token-bucket enforcement (bursts can 429). Limits are maxima, not guarantees. - Separate buckets per model class; `inference_geo` values share one pool. - **Acceleration limits**: a sharp traffic increase can 429 (`rate_limit_error`) since 2025-08-11 (previously 529). - History: TPM → ITPM+OTPM (2024-11-20); dedicated 1M-context limits removed (2026-03-13); tiers 1–4 consolidated into **Start / Build / Scale** with Sonnet/Haiku raised to Opus levels (2026-06-26). ## 2. Usage tiers and spend caps | Tier | Monthly spend cap | Notes | |---|---|---| | Start | $500 | Claude Platform on AWS orgs start here | | Build | $1,000 | | | Scale | $200,000 | | | Custom | none | negotiated; contact sales via Console → Rate limits | Reaching the tier cap → HTTP 429 `rate_limit_error` with `error.details.error_code = "enforced_spend_limit_reached"`, **no `retry-after`**, access resumes 00:00 UTC on the 1st or on tier increase. A self-set (lower) limit → HTTP 400 `invalid_request_error` "You have reached your specified API usage limits…" (workspace variant exists). Claude Code workspace limits are checked separately and can 429 with `retry-after`. ## 3. Messages API limits per model class (RPM / ITPM / OTPM) | Model class | Start | Build | Scale | |---|---|---|---| | Fable 5.x ¹ | 1,000 / 500k / 100k | 2,000 / 1.5M / 300k | 4,000 / 4M / 800k | | Opus 5 | 1,000 / 2M / 400k | 5,000 / 5M / 1M | 10,000 / 10M / 2M | | Opus 4.x ² | 1,000 / 2M / 400k | 5,000 / 5M / 1M | 10,000 / 10M / 2M | | Sonnet 5 | 1,000 / 2M / 400k | 5,000 / 5M / 1M | 10,000 / 10M / 2M | | Sonnet 4.x ³ | 1,000 / 2M / 400k | 5,000 / 5M / 1M | 10,000 / 10M / 2M | | Haiku 4.5 | 1,000 / 2M / 400k | 5,000 / 5M / 1M | 10,000 / 10M / 2M | | Haiku 3.5 (retired on API) ⁴ | 1,000 / 100k / 20k | 2,000 / 200k / 40k | 4,000 / 400k / 80k | ¹ combined Fable 5.1 + Fable 5; Mythos 5.1 + Mythos 5 share a separate combined limit on the same terms. ² combined Opus 4.8/4.7/4.6/4.5 (Opus 5 separate). ³ combined Sonnet 4.6/4.5 (Sonnet 5 separate). ⁴ counts `cache_read_input_tokens` toward ITPM. ### Cache-aware ITPM - Counted: `input_tokens` (after the last cache breakpoint) + `cache_creation_input_tokens`. **Not counted**: `cache_read_input_tokens` (except Haiku 3.5). - ITPM estimated at request start then corrected; OTPM counted on actual generated tokens (`max_tokens` irrelevant). - Example: 2M ITPM at 80% cache hits ≈ 10M total input tokens/min. ## 4. Other limit families | Family | Start | Build | Scale | Notes | |---|---|---|---|---| | Message Batches API (all models shared) | 1,000 RPM · 200k queued · 100k/batch | 2,000 · 300k · 100k | 4,000 · 500k · 100k | queued = batch requests not yet processed | | Token counting | 5,000 RPM | 10,000 | 20,000 | free | | Files API | ~500 RPM (all tiers, per org, all file ops) | | | contact sales to raise | | Managed Agents | create endpoints 300 RPM · read endpoints 1,200 RPM | | | per org, separate from Messages | | Fast mode (`speed: fast`, Opus 5 / 4.8) | dedicated limits, separate from Opus | | | 429 + `retry-after`; headers `anthropic-fast-{input,output}-tokens-{limit,remaining,reset}` | | Programmatic tool calling | same limits; each call from code execution = one invocation | | | | | Priority Tier (legacy commitments) | committed ITPM/OTPM per model; burndown cache read 0.1, 5m write 1.25, 1h write 2.0, `inference_geo: us` 1.1 | | | `service_tier: auto|standard_only`; headers `anthropic-priority-*` | | Workspaces | per-workspace spend/rate limits below org limit (not on default workspace); org limits always apply | | | Rate Limits API to read configured limits | ## 5. Response headers (documented) | Header | Meaning | |---|---| | `retry-after` | seconds to wait; earlier retries fail; absent on the spend-cap 429 | | `anthropic-ratelimit-requests-{limit,remaining,reset}` | requests in the period; reset RFC 3339 | | `anthropic-ratelimit-tokens-{limit,remaining,reset}` | most restrictive token limit in effect (workspace limit if exceeded, else total input+output); remaining rounded to nearest thousand | | `anthropic-ratelimit-input-tokens-{limit,remaining,reset}` | input tokens | | `anthropic-ratelimit-output-tokens-{limit,remaining,reset}` | output tokens | | `anthropic-priority-{input,output}-tokens-{limit,remaining,reset}` | Priority Tier only | | `anthropic-fast-{input,output}-tokens-{limit,remaining,reset}` | fast mode only | | `request-id`, `anthropic-organization-id`, `anthropic-workspace-id` | identification (workspace header since 2026-08-11; absent on Admin API) | ## 6. Observed with our key (2026-09-18/19 UTC) — account-specific, not documentation | Endpoint | Observed | |---|---| | `POST /v1/messages` (opus-5, sonnet-5, fable-5-1, opus-4-8, sonnet-4-6) | `input-tokens-limit 10000000`, `output-tokens-limit 2000000`, `requests-limit 10000`, `tokens-limit 12000000` (= ITPM+OTPM); `*-remaining` and `*-reset` present; `request-id`, `anthropic-organization-id`, `CF-RAY` present | | `GET /v1/models`, `GET /v1/models/{id}` | no `anthropic-ratelimit-*` headers | | `POST /v1/messages/count_tokens` | no `anthropic-ratelimit-*` headers | | `GET /v1/organizations/me` (Admin API, observed by another agent this run) | only `anthropic-ratelimit-requests-limit: 100` | | `retry-after` | not observed (no 429 triggered) | Interpretation: the Opus/Sonnet numbers match the documented **Scale** tier. `claude-fable-5-1` returned the same 10M/2M/10k values although the documented Fable Scale limits are 4,000 / 4M / 800k — recorded as observed for this key, not generalized. ## 7. 429 / error semantics summary | Situation | HTTP | `error.type` | `retry-after` | |---|---|---|---| | RPM/ITPM/OTPM exceeded | 429 | `rate_limit_error` | yes | | Acceleration limit | 429 | `rate_limit_error` | yes | | Tier spend cap | 429 | `rate_limit_error` + `details.error_code=enforced_spend_limit_reached` | **no** | | Self-set spend limit | 400 | `invalid_request_error` | — | | Overloaded | 529 | `overloaded_error` | — (see api/errors) |