SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
7.0 KB

# Anthropic pricing (Claude API)

Status: DOCUMENTED (prices are not verifiable by API; the 5 minimal live calls of this run are consistent with the documented usage shape). All prices USD.
Sources: https://platform.claude.com/docs/en/about-claude/pricing · models/overview · models/*/overview · build-with-claude/fast-mode · build-with-claude/prompt-caching · manage-claude/data-residency · agents-and-tools/tool-use/{web-search-tool,web-fetch-tool,code-execution-tool} · api/service-tiers · release-notes/overview (retrieved 2026-09-18).
Last verified: 2026-09-18 · Machine-readable: generated/fragments/pricing/anthropic-pricing.json.

# 1. Model token prices (per MTok)

Model Base input 5m cache write (1.25x) 1h cache write (2x) Cache read Output Batch in / out (50%)
Claude Fable 5.1 10 12.50 20 0.25 (0.025x) 50 5 / 25
Claude Mythos 5.1 (invite only) 10 12.50 20 0.25 (0.025x) 50 5 / 25
Claude Fable 5 10 12.50 20 1.00 50 5 / 25
Claude Mythos 5 (invite only) 10 12.50 20 1.00 50 5 / 25
Claude Opus 5 5 6.25 10 0.50 25 2.50 / 12.50
Claude Opus 4.8 5 6.25 10 0.50 25 2.50 / 12.50
Claude Opus 4.7 5 6.25 10 0.50 25 2.50 / 12.50
Claude Opus 4.6 5 6.25 10 0.50 25 2.50 / 12.50
Claude Opus 4.5 5 6.25 10 0.50 25 2.50 / 12.50
Claude Opus 4.1 (retired; Bedrock/Google Cloud only) 15 18.75 30 1.50 75 7.50 / 37.50
Claude Opus 4 (retired; Google Cloud only) 15 18.75 30 1.50 75 7.50 / 37.50
Claude Sonnet 5 2 2.50 4 0.20 10 1 / 5
Claude Sonnet 4.6 3 3.75 6 0.30 15 1.50 / 7.50
Claude Sonnet 4.5 3 3.75 6 0.30 15 1.50 / 7.50
Claude Sonnet 4 (retired; Bedrock/Google Cloud) 3 3.75 6 0.30 15 1.50 / 7.50
Claude Haiku 4.5 1 1.25 2 0.10 5 0.50 / 2.50
Claude Haiku 3.5 (retired; Bedrock/Google Cloud) 0.80 1.00 1.60 0.08 4 0.40 / 2

Notes

  • Sonnet 5's introductory $2/$10 became the standard price on 2026-08-10 (the planned $3/$15 increase on 2026-09-01 was cancelled).
  • Cache reads: 0.1x base input on every model except Fable 5.1 / Mythos 5.1 (0.025x since 2026-09-01). Caching pays off after 1 read (5m) or 2 reads (1h).
  • Claude Mythos Preview and Claude 3.x models are not on the current pricing page (no price recorded).
  • Tokenizer: Claude 4.7+ and Mythos Preview produce ~30% more tokens for the same text than Sonnet 4.6 and earlier — same per-token price, higher token count.

# 2. Long context, data residency, clouds

Item Price Notes
1M context (>200k input) standard per-token price Claude 4.6+ and Mythos Preview; no long-context premium; caching and batch discounts apply across the window. Pre-4.6 models are 200k.
inference_geo: "us" ×1.1 on all token dimensions Claude 4.6+ only (400 on earlier models); "global" = ×1.0. Also Foundry "US Data Zone Standard".
Bedrock / Vertex regional & multi-region endpoints +10% over global endpoints Sonnet 4.5, Haiku 4.5, Opus 4.5 and later; provider invoices.
Claude Platform on AWS / Microsoft Foundry CCU at $0.01 (100 CCU = $1 of list price after discounts) hourly metering to the marketplace; postpaid; same per-model rates.

# 3. Fast mode (research preview, Claude API only)

Model Input Output Notes
Claude Opus 5, Claude Opus 4.8 10 50 speed: "fast" + anthropic-beta: fast-mode-2026-02-01; applies across the full context window; caching (1.25x/2x/0.1x) and inference_geo (1.1x) multipliers stack on top; not available on Batch or on any cloud.
Claude Opus 4.7 — — removed 2026-07-24: speed: "fast" → error
Claude Opus 4.6 — — removed 2026-06-29: runs at standard speed and standard price; usage.speed = "standard"

# 4. Service tiers

Tier Price effect Notes
standard list price default
batch 50% off input and output Message Batches API; 24h window; stacks with caching
priority same token price; capacity commitment no longer sold; burndown of committed TPM: cache read 0.1, 5m write 1.25, 1h write 2.0, inference_geo: us 1.1
fast premium table above Opus 5 / Opus 4.8

# 5. Tools and features

Item Price Unit Notes
Web search 10 per 1,000 searches + tokens for results (counted as input in the turn and later turns); errors not billed; 1 search = 1 use regardless of results
Web fetch 0 — only tokens of fetched content (~2,500 tok per 10 kB page, ~25k per 100 kB, ~125k per 500 kB PDF); cap with max_content_tokens
Code execution with web_search_20260209+ / web_fetch_20260209+ in the request 0 — free (tokens only), since 2026-02-17
Code execution standalone 0.05 per hour per container after 1,550 free hours / org / month; 5-minute minimum; billed even if the tool isn't called when files are preloaded; usage.server_tool_use.code_execution_requests
Managed Agents session runtime 0.08 per session-hour metered per ms while running; replaces container-hour billing; tokens at model rates; no batch discount; fast mode premium if model.speed: fast; ×1.1 if model.inference_geo: us
Token counting 0 — own RPM limits
Files API no storage price documented — ~500 RPM; optional expires_in_seconds
Claude Code not on the platform pricing page —

# Token overheads (input tokens added per request)

Overhead Tokens Applies to
Tool-use system prompt, tool_choice auto/none → any/tool Opus 5 286→406 · Opus 4.8 290→410 · Opus 4.7 675→804 · Opus 4.6 497→589 · Opus 4.5 496→588 · Opus 4.1/Opus 4/Sonnet 4 313→315 · Sonnet 5 354→474 · Sonnet 4.6 497→589 · Sonnet 4.5 496→588 · Haiku 4.5 496→588 · Haiku 3.5 264→355 any request with ≥1 tool (none with no tools = 0)
bash_20250124 definition 325 (Opus 5/4.8/4.7) · 244 (Opus 4.6, Sonnet 4.6 and earlier) + command output tokens
text_editor_20250429 (Claude 4.x) 700
computer_toolset_20260801 (default members) ~4,500 (~4,520 Fable 5/Mythos 5/Opus 5/Opus 4.8; ~4,590 Sonnet 5); −~410 if zoom disabled screenshots/zoom images billed as image input
computer_20251124 / computer_20250124 ~735 per definition + 466–499 system prompt
browser_toolset_20260801 ~6,600 (~6,610 / ~6,670 Sonnet 5); +~880 with all 4 optional members

# 6. Billing facts

  • Billing on actual monthly usage, USD, credit card or invoicing; small free credits for new users; volume/enterprise discounts negotiated.
  • Spend caps per usage tier: Start $500, Build $1,000, Scale $200,000, Custom none (see rate-limits doc).
  • Refusals: since 2026-06-02 a request that returns stop_reason: "refusal" before any output is not billed; fallbacks re-runs are billed at the fallback model's rate.
  • Discounts stack: batch × caching × (fast) × data residency.