# Anthropic rate limits and spend limits
Status: DOCUMENTED (published tables) + LIVE_DISCOVERED (headers observed with our key on 2026-09-18; account-specific).
Sources: https://platform.claude.com/docs/en/api/rate-limits · api/service-tiers · build-with-claude/fast-mode · build-with-claude/token-counting · build-with-claude/files · api/overview · release-notes/overview (retrieved 2026-09-18).
Last verified: 2026-09-18 · Machine-readable: generated/fragments/rate-limits/anthropic-rate-limits.json.
# 1. Model
- Two kinds of limits: spend limits (monthly USD cap per usage tier) and rate limits (RPM / ITPM / OTPM per model class).
- Organization-level, tier assigned automatically from usage history; new orgs may start in an Evaluation tier below the published numbers. Token-bucket enforcement (bursts can 429). Limits are maxima, not guarantees.
- Separate buckets per model class;
inference_geo values share one pool.
- Acceleration limits: a sharp traffic increase can 429 (
rate_limit_error) since 2025-08-11 (previously 529).
- History: TPM → ITPM+OTPM (2024-11-20); dedicated 1M-context limits removed (2026-03-13); tiers 1–4 consolidated into Start / Build / Scale with Sonnet/Haiku raised to Opus levels (2026-06-26).
# 2. Usage tiers and spend caps
| Tier |
Monthly spend cap |
Notes |
| Start |
$500 |
Claude Platform on AWS orgs start here |
| Build |
$1,000 |
|
| Scale |
$200,000 |
|
| Custom |
none |
negotiated; contact sales via Console → Rate limits |
Reaching the tier cap → HTTP 429 rate_limit_error with error.details.error_code = "enforced_spend_limit_reached", no retry-after, access resumes 00:00 UTC on the 1st or on tier increase. A self-set (lower) limit → HTTP 400 invalid_request_error "You have reached your specified API usage limits…" (workspace variant exists). Claude Code workspace limits are checked separately and can 429 with retry-after.
# 3. Messages API limits per model class (RPM / ITPM / OTPM)
| Model class |
Start |
Build |
Scale |
| Fable 5.x ¹ |
1,000 / 500k / 100k |
2,000 / 1.5M / 300k |
4,000 / 4M / 800k |
| Opus 5 |
1,000 / 2M / 400k |
5,000 / 5M / 1M |
10,000 / 10M / 2M |
| Opus 4.x ² |
1,000 / 2M / 400k |
5,000 / 5M / 1M |
10,000 / 10M / 2M |
| Sonnet 5 |
1,000 / 2M / 400k |
5,000 / 5M / 1M |
10,000 / 10M / 2M |
| Sonnet 4.x ³ |
1,000 / 2M / 400k |
5,000 / 5M / 1M |
10,000 / 10M / 2M |
| Haiku 4.5 |
1,000 / 2M / 400k |
5,000 / 5M / 1M |
10,000 / 10M / 2M |
| Haiku 3.5 (retired on API) ⁴ |
1,000 / 100k / 20k |
2,000 / 200k / 40k |
4,000 / 400k / 80k |
¹ combined Fable 5.1 + Fable 5; Mythos 5.1 + Mythos 5 share a separate combined limit on the same terms. ² combined Opus 4.8/4.7/4.6/4.5 (Opus 5 separate). ³ combined Sonnet 4.6/4.5 (Sonnet 5 separate). ⁴ counts cache_read_input_tokens toward ITPM.
# Cache-aware ITPM
- Counted:
input_tokens (after the last cache breakpoint) + cache_creation_input_tokens. Not counted: cache_read_input_tokens (except Haiku 3.5).
- ITPM estimated at request start then corrected; OTPM counted on actual generated tokens (
max_tokens irrelevant).
- Example: 2M ITPM at 80% cache hits ≈ 10M total input tokens/min.
# 4. Other limit families
| Family |
Start |
Build |
Scale |
Notes |
| Message Batches API (all models shared) |
1,000 RPM · 200k queued · 100k/batch |
2,000 · 300k · 100k |
4,000 · 500k · 100k |
queued = batch requests not yet processed |
| Token counting |
5,000 RPM |
10,000 |
20,000 |
free |
| Files API |
~500 RPM (all tiers, per org, all file ops) |
|
|
contact sales to raise |
| Managed Agents |
create endpoints 300 RPM · read endpoints 1,200 RPM |
|
|
per org, separate from Messages |
Fast mode (speed: fast, Opus 5 / 4.8) |
dedicated limits, separate from Opus |
|
|
429 + retry-after; headers anthropic-fast-{input,output}-tokens-{limit,remaining,reset} |
| Programmatic tool calling |
same limits; each call from code execution = one invocation |
|
|
|
| Priority Tier (legacy commitments) |
committed ITPM/OTPM per model; burndown cache read 0.1, 5m write 1.25, 1h write 2.0, inference_geo: us 1.1 |
|
|
`service_tier: auto |
| Workspaces |
per-workspace spend/rate limits below org limit (not on default workspace); org limits always apply |
|
|
Rate Limits API to read configured limits |
| Header |
Meaning |
retry-after |
seconds to wait; earlier retries fail; absent on the spend-cap 429 |
anthropic-ratelimit-requests-{limit,remaining,reset} |
requests in the period; reset RFC 3339 |
anthropic-ratelimit-tokens-{limit,remaining,reset} |
most restrictive token limit in effect (workspace limit if exceeded, else total input+output); remaining rounded to nearest thousand |
anthropic-ratelimit-input-tokens-{limit,remaining,reset} |
input tokens |
anthropic-ratelimit-output-tokens-{limit,remaining,reset} |
output tokens |
anthropic-priority-{input,output}-tokens-{limit,remaining,reset} |
Priority Tier only |
anthropic-fast-{input,output}-tokens-{limit,remaining,reset} |
fast mode only |
request-id, anthropic-organization-id, anthropic-workspace-id |
identification (workspace header since 2026-08-11; absent on Admin API) |
# 6. Observed with our key (2026-09-18/19 UTC) — account-specific, not documentation
| Endpoint |
Observed |
POST /v1/messages (opus-5, sonnet-5, fable-5-1, opus-4-8, sonnet-4-6) |
input-tokens-limit 10000000, output-tokens-limit 2000000, requests-limit 10000, tokens-limit 12000000 (= ITPM+OTPM); *-remaining and *-reset present; request-id, anthropic-organization-id, CF-RAY present |
GET /v1/models, GET /v1/models/{id} |
no anthropic-ratelimit-* headers |
POST /v1/messages/count_tokens |
no anthropic-ratelimit-* headers |
GET /v1/organizations/me (Admin API, observed by another agent this run) |
only anthropic-ratelimit-requests-limit: 100 |
retry-after |
not observed (no 429 triggered) |
Interpretation: the Opus/Sonnet numbers match the documented Scale tier. claude-fable-5-1 returned the same 10M/2M/10k values although the documented Fable Scale limits are 4,000 / 4M / 800k — recorded as observed for this key, not generalized.
# 7. 429 / error semantics summary
| Situation |
HTTP |
error.type |
retry-after |
| RPM/ITPM/OTPM exceeded |
429 |
rate_limit_error |
yes |
| Acceleration limit |
429 |
rate_limit_error |
yes |
| Tier spend cap |
429 |
rate_limit_error + details.error_code=enforced_spend_limit_reached |
no |
| Self-set spend limit |
400 |
invalid_request_error |
— |
| Overloaded |
529 |
overloaded_error |
— (see api/errors) |