SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
3.4 KB

# Anthropic — Fast mode (speed: "fast")

Status: DOCUMENTED · PREVIEW (research preview, waitlist / account manager) · BETA header fast-mode-2026-02-01 · ACCOUNT_RESTRICTED for our key (429 rate limit of 0 fast mode input tokens per minute on claude-opus-5, 2026-09-18). Sources: Fast mode · Beta Messages API (speed) · Release notes (2026-02-05 launch Opus 4.6; 2026-04 Opus 4.7; 2026-05-28 Opus 4.8; 2026-06-26 removed 4.6; 2026-07-24 removed 4.7 / Opus 5) Last verified: 2026-09-18

# Request

json
{"model": "claude-opus-5", "max_tokens": 1024, "speed": "fast", "messages": [...]}

Header: anthropic-beta: fast-mode-2026-02-01. speed enum: "standard" (default) | "fast". Available only through the beta surface: without the header the field is rejected (400 speed: Extra inputs are not permitted, live).

# Supported models

Model speed: "fast" Notes
Opus 5 (claude-opus-5) ✔ research preview Claude API + Managed Agents only
Opus 4.8 (claude-opus-4-8) ✔ research preview idem
Opus 4.7 ✘ error (removed 2026-07-24) does not fall back
Opus 4.6 ✘ silently runs standard, standard billing, usage.speed: "standard" unique fallback behaviour
All others (Sonnet, Haiku, Fable…) ✘ 400 '<model>' does not support the speed parameter. This feature is only available on supported models. (live, Haiku)

Not available on Bedrock, Claude Platform on AWS, Vertex, Foundry, Batch API, Priority Tier.

# What you get

  • Up to 2.5× output tokens/s; TTFT unchanged; same weights and behaviour; best seen with streaming.
  • Pricing: $10 input / $50 output per MTok (2× standard) on the whole context, incl. > 200k; cache and data-residency multipliers stack on top.
  • Dedicated rate limit with headers anthropic-fast-input-tokens-{limit,remaining,reset} and anthropic-fast-output-tokens-{limit,remaining,reset}; 429 + retry-after when exceeded; 529 on capacity.
  • usage.speed reports "fast" or "standard" (live: speed: "standard" returned when we sent speed: "standard" with the header on Opus 5).
  • Switching speed invalidates the prompt cache (system + messages); fallback to standard = cache miss.
  • SDK retries 429 twice by default (max_retries); for a client-side fallback set max_retries: 0, catch the rate-limit error and resend without speed.

# Live observations (2026-09-18)

Call Result
Opus 5, speed: fast, header, thinking: disabled, max_tokens 8 429 rate_limit_error: "This request would exceed your organization's rate limit of 0 fast mode input tokens per minute … model: claude-opus-5". Headers: x-should-retry: true, no anthropic-fast-* headers returned on the 429. → ACCOUNT_RESTRICTED (feature exists; org not enrolled).
Opus 5, speed: fast, no header 400 speed: Extra inputs are not permitted
Haiku 4.5, speed: fast, header 400 does not support the speed parameter
Opus 5, speed: standard, header 200, usage.speed: "standard", $0.00016

No example directory: the feature cannot be exercised with this account; the request shape above is in generated/fragments/parameters/anthropic-advanced.json (speed).