Anthropic service tiers (standard / priority / batch / fast)
Status: DOCUMENTED; service_tier semantics LIVE_VERIFIED only insofar as 5 minimal messages returned usage.service_tier: "standard" without sending the parameter.
Sources: https://platform.claude.com/docs/en/api/service-tiers · build-with-claude/fast-mode · build-with-claude/batch-processing · about-claude/pricing · api/rate-limits (retrieved 2026-09-18).
Last verified: 2026-09-18.
1. The tiers
| Tier | How to select | Price | Availability semantics | Models |
|---|---|---|---|---|
| Standard | default | list price | best-effort alongside all other traffic | all |
| Priority Tier | service_tier: "auto" (default) with an existing capacity commitment |
list price; commitment contract (1/3/6/12 months, ITPM + OTPM, specific model) | prioritized over all other requests, 99.5% uptime target, fewer 529 overloaded_error; overflow falls back to standard |
all except Fable 5.1, Mythos 5.1, Mythos 5, Mythos Preview, Opus 5, Sonnet 5. No longer sold (existing contracts honoured). |
| Batch | POST /v1/messages/batches |
50% off input and output | asynchronous, results within 24h; outside your normal capacity; own rate limits (see rate-limits doc); up to 300k output with output-300k-2026-03-24 on Opus 5/4.8/4.7/4.6, Sonnet 5/4.6 |
all active models (capabilities.batch.supported true on all 11 live ids) |
| Fast mode | speed: "fast" + anthropic-beta: fast-mode-2026-02-01 |
premium $10 / $50 per MTok | up to 2.5x output tokens/s; dedicated rate limits; research preview; Claude API only (not Batch, not clouds) | Opus 5, Opus 4.8 (Opus 4.7 → error since 2026-07-24; Opus 4.6 → silently standard since 2026-06-29) |
2. service_tier request parameter (POST /v1/messages)
| Value | Meaning |
|---|---|
"auto" (default) |
use Priority Tier capacity if the org has enough committed input and output TPM for the request, otherwise standard |
"standard_only" |
never consume Priority Tier capacity |
Priority assignment burns committed capacity at: cache reads 0.1 token/token, 5m cache writes 1.25, 1h cache writes 2.0, inference_geo: "us" (4.6+) 1.1 on input and output, everything else 1.0. Priority requests also draw from the regular rate limits; if those would be exceeded the request is declined.
3. Response fields and headers
usage.service_tier:"standard"|"priority"(also"batch"on batch results per the Messages reference). Observed live:"standard"on all 5 probes (no parameter sent).usage.speed:"fast"|"standard"— present only when fast mode is involved (absent on our standard probes).usage.inference_geo:"global"observed (4.6+ models echo the geo).- Priority headers when eligible (even if over limit):
anthropic-priority-input-tokens-{limit,remaining,reset},anthropic-priority-output-tokens-{limit,remaining,reset}. - Fast headers:
anthropic-fast-input-tokens-{limit,remaining,reset},anthropic-fast-output-tokens-{limit,remaining,reset}.
4. Interactions
- Batch × prompt caching stack (batch results carry cache fields); fast mode × caching ×
inference_geostack; fast mode is not available on Batch. - Managed Agents sessions: no batch mode; fast mode premium applies when
model.speed: "fast". - Claude Platform on AWS: same rate limits, no fast mode, no per-workspace limits.