SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
6.7 KB

# Anthropic rate limits and spend limits

Status: DOCUMENTED (published tables) + LIVE_DISCOVERED (headers observed with our key on 2026-09-18; account-specific).
Sources: https://platform.claude.com/docs/en/api/rate-limits · api/service-tiers · build-with-claude/fast-mode · build-with-claude/token-counting · build-with-claude/files · api/overview · release-notes/overview (retrieved 2026-09-18).
Last verified: 2026-09-18 · Machine-readable: generated/fragments/rate-limits/anthropic-rate-limits.json.

# 1. Model

  • Two kinds of limits: spend limits (monthly USD cap per usage tier) and rate limits (RPM / ITPM / OTPM per model class).
  • Organization-level, tier assigned automatically from usage history; new orgs may start in an Evaluation tier below the published numbers. Token-bucket enforcement (bursts can 429). Limits are maxima, not guarantees.
  • Separate buckets per model class; inference_geo values share one pool.
  • Acceleration limits: a sharp traffic increase can 429 (rate_limit_error) since 2025-08-11 (previously 529).
  • History: TPM → ITPM+OTPM (2024-11-20); dedicated 1M-context limits removed (2026-03-13); tiers 1–4 consolidated into Start / Build / Scale with Sonnet/Haiku raised to Opus levels (2026-06-26).

# 2. Usage tiers and spend caps

Tier Monthly spend cap Notes
Start $500 Claude Platform on AWS orgs start here
Build $1,000
Scale $200,000
Custom none negotiated; contact sales via Console → Rate limits

Reaching the tier cap → HTTP 429 rate_limit_error with error.details.error_code = "enforced_spend_limit_reached", no retry-after, access resumes 00:00 UTC on the 1st or on tier increase. A self-set (lower) limit → HTTP 400 invalid_request_error "You have reached your specified API usage limits…" (workspace variant exists). Claude Code workspace limits are checked separately and can 429 with retry-after.

# 3. Messages API limits per model class (RPM / ITPM / OTPM)

Model class Start Build Scale
Fable 5.x ¹ 1,000 / 500k / 100k 2,000 / 1.5M / 300k 4,000 / 4M / 800k
Opus 5 1,000 / 2M / 400k 5,000 / 5M / 1M 10,000 / 10M / 2M
Opus 4.x ² 1,000 / 2M / 400k 5,000 / 5M / 1M 10,000 / 10M / 2M
Sonnet 5 1,000 / 2M / 400k 5,000 / 5M / 1M 10,000 / 10M / 2M
Sonnet 4.x ³ 1,000 / 2M / 400k 5,000 / 5M / 1M 10,000 / 10M / 2M
Haiku 4.5 1,000 / 2M / 400k 5,000 / 5M / 1M 10,000 / 10M / 2M
Haiku 3.5 (retired on API) ⁴ 1,000 / 100k / 20k 2,000 / 200k / 40k 4,000 / 400k / 80k

¹ combined Fable 5.1 + Fable 5; Mythos 5.1 + Mythos 5 share a separate combined limit on the same terms. ² combined Opus 4.8/4.7/4.6/4.5 (Opus 5 separate). ³ combined Sonnet 4.6/4.5 (Sonnet 5 separate). ⁴ counts cache_read_input_tokens toward ITPM.

# Cache-aware ITPM

  • Counted: input_tokens (after the last cache breakpoint) + cache_creation_input_tokens. Not counted: cache_read_input_tokens (except Haiku 3.5).
  • ITPM estimated at request start then corrected; OTPM counted on actual generated tokens (max_tokens irrelevant).
  • Example: 2M ITPM at 80% cache hits ≈ 10M total input tokens/min.

# 4. Other limit families

Family Start Build Scale Notes
Message Batches API (all models shared) 1,000 RPM · 200k queued · 100k/batch 2,000 · 300k · 100k 4,000 · 500k · 100k queued = batch requests not yet processed
Token counting 5,000 RPM 10,000 20,000 free
Files API ~500 RPM (all tiers, per org, all file ops) contact sales to raise
Managed Agents create endpoints 300 RPM · read endpoints 1,200 RPM per org, separate from Messages
Fast mode (speed: fast, Opus 5 / 4.8) dedicated limits, separate from Opus 429 + retry-after; headers anthropic-fast-{input,output}-tokens-{limit,remaining,reset}
Programmatic tool calling same limits; each call from code execution = one invocation
Priority Tier (legacy commitments) committed ITPM/OTPM per model; burndown cache read 0.1, 5m write 1.25, 1h write 2.0, inference_geo: us 1.1 `service_tier: auto
Workspaces per-workspace spend/rate limits below org limit (not on default workspace); org limits always apply Rate Limits API to read configured limits

# 5. Response headers (documented)

Header Meaning
retry-after seconds to wait; earlier retries fail; absent on the spend-cap 429
anthropic-ratelimit-requests-{limit,remaining,reset} requests in the period; reset RFC 3339
anthropic-ratelimit-tokens-{limit,remaining,reset} most restrictive token limit in effect (workspace limit if exceeded, else total input+output); remaining rounded to nearest thousand
anthropic-ratelimit-input-tokens-{limit,remaining,reset} input tokens
anthropic-ratelimit-output-tokens-{limit,remaining,reset} output tokens
anthropic-priority-{input,output}-tokens-{limit,remaining,reset} Priority Tier only
anthropic-fast-{input,output}-tokens-{limit,remaining,reset} fast mode only
request-id, anthropic-organization-id, anthropic-workspace-id identification (workspace header since 2026-08-11; absent on Admin API)

# 6. Observed with our key (2026-09-18/19 UTC) — account-specific, not documentation

Endpoint Observed
POST /v1/messages (opus-5, sonnet-5, fable-5-1, opus-4-8, sonnet-4-6) input-tokens-limit 10000000, output-tokens-limit 2000000, requests-limit 10000, tokens-limit 12000000 (= ITPM+OTPM); *-remaining and *-reset present; request-id, anthropic-organization-id, CF-RAY present
GET /v1/models, GET /v1/models/{id} no anthropic-ratelimit-* headers
POST /v1/messages/count_tokens no anthropic-ratelimit-* headers
GET /v1/organizations/me (Admin API, observed by another agent this run) only anthropic-ratelimit-requests-limit: 100
retry-after not observed (no 429 triggered)

Interpretation: the Opus/Sonnet numbers match the documented Scale tier. claude-fable-5-1 returned the same 10M/2M/10k values although the documented Fable Scale limits are 4,000 / 4M / 800k — recorded as observed for this key, not generalized.

# 7. 429 / error semantics summary

Situation HTTP error.type retry-after
RPM/ITPM/OTPM exceeded 429 rate_limit_error yes
Acceleration limit 429 rate_limit_error yes
Tier spend cap 429 rate_limit_error + details.error_code=enforced_spend_limit_reached no
Self-set spend limit 400 invalid_request_error —
Overloaded 529 overloaded_error — (see api/errors)