SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
3.4 KB

# Anthropic service tiers (standard / priority / batch / fast)

Status: DOCUMENTED; service_tier semantics LIVE_VERIFIED only insofar as 5 minimal messages returned usage.service_tier: "standard" without sending the parameter.
Sources: https://platform.claude.com/docs/en/api/service-tiers · build-with-claude/fast-mode · build-with-claude/batch-processing · about-claude/pricing · api/rate-limits (retrieved 2026-09-18).
Last verified: 2026-09-18.

# 1. The tiers

Tier How to select Price Availability semantics Models
Standard default list price best-effort alongside all other traffic all
Priority Tier service_tier: "auto" (default) with an existing capacity commitment list price; commitment contract (1/3/6/12 months, ITPM + OTPM, specific model) prioritized over all other requests, 99.5% uptime target, fewer 529 overloaded_error; overflow falls back to standard all except Fable 5.1, Mythos 5.1, Mythos 5, Mythos Preview, Opus 5, Sonnet 5. No longer sold (existing contracts honoured).
Batch POST /v1/messages/batches 50% off input and output asynchronous, results within 24h; outside your normal capacity; own rate limits (see rate-limits doc); up to 300k output with output-300k-2026-03-24 on Opus 5/4.8/4.7/4.6, Sonnet 5/4.6 all active models (capabilities.batch.supported true on all 11 live ids)
Fast mode speed: "fast" + anthropic-beta: fast-mode-2026-02-01 premium $10 / $50 per MTok up to 2.5x output tokens/s; dedicated rate limits; research preview; Claude API only (not Batch, not clouds) Opus 5, Opus 4.8 (Opus 4.7 → error since 2026-07-24; Opus 4.6 → silently standard since 2026-06-29)

# 2. service_tier request parameter (POST /v1/messages)

Value Meaning
"auto" (default) use Priority Tier capacity if the org has enough committed input and output TPM for the request, otherwise standard
"standard_only" never consume Priority Tier capacity

Priority assignment burns committed capacity at: cache reads 0.1 token/token, 5m cache writes 1.25, 1h cache writes 2.0, inference_geo: "us" (4.6+) 1.1 on input and output, everything else 1.0. Priority requests also draw from the regular rate limits; if those would be exceeded the request is declined.

# 3. Response fields and headers

  • usage.service_tier: "standard" | "priority" (also "batch" on batch results per the Messages reference). Observed live: "standard" on all 5 probes (no parameter sent).
  • usage.speed: "fast" | "standard" — present only when fast mode is involved (absent on our standard probes).
  • usage.inference_geo: "global" observed (4.6+ models echo the geo).
  • Priority headers when eligible (even if over limit): anthropic-priority-input-tokens-{limit,remaining,reset}, anthropic-priority-output-tokens-{limit,remaining,reset}.
  • Fast headers: anthropic-fast-input-tokens-{limit,remaining,reset}, anthropic-fast-output-tokens-{limit,remaining,reset}.

# 4. Interactions

  • Batch × prompt caching stack (batch results carry cache fields); fast mode × caching × inference_geo stack; fast mode is not available on Batch.
  • Managed Agents sessions: no batch mode; fast mode premium applies when model.speed: "fast".
  • Claude Platform on AWS: same rate limits, no fast mode, no per-workspace limits.