Anthropic token counting — POST /v1/messages/count_tokens
Status: DOCUMENTED + LIVE_VERIFIED 2026-09-18 (8 calls on claude-haiku-4-5-20251001, free). No beta header needed (the old token-counting-2024-11-01 header is obsolete).
Sources: https://platform.claude.com/docs/en/api/messages/count_tokens · https://platform.claude.com/docs/en/build-with-claude/token-counting · https://platform.claude.com/docs/en/api/errors#request-size-limits
Machine-readable: generated/fragments/parameters/anthropic-count-tokens.json, generated/fragments/endpoints/anthropic-core.json
Last verified: 2026-09-18
Request
Same headers as Messages. Body limit 32 MB. Body is a Messages request without max_tokens, stream, stop_sequences, metadata, service_tier, sampling params:
| Field | Req. | Counted? | Notes |
|---|---|---|---|
model |
yes | — | tokenizer of this model is used; Claude 4.7+/Fable/Mythos tokenizer ≈ +30 % vs earlier models — recount per target model |
messages[] |
yes | yes | all content block types; image/document must be base64 (url/file sources → invalid_request_error) |
system |
no | yes | +6 tokens for "You are a helpful assistant." |
tools[] |
no | yes | client tools + advisor tool only; other server tools → 400 "Server tools are not supported in the count_tokens endpoint: web_search_20250305. Use the /v1/messages endpoint instead." |
tool_choice |
no | affects count | any/tool rejected on Fable 5.1 / Mythos 5.1 (400) |
thinking |
no | yes | enabled/1024 → 40 vs 11 tokens; previous-turn thinking blocks count only on models that keep all turns; current-turn thinking counts |
output_config |
no | yes | effort, format (json_schema) — small schema ≈ +140 tokens |
cache_control |
no | ignored | accepted; counting never uses the cache |
output_format (beta legacy) |
no | SDK still exposes it on count_tokens; API without beta header → 400 "Extra inputs are not permitted" (observed on /v1/messages/count_tokens) | |
MCP connector (mcp_servers) |
— | rejected (documented) |
Response
{"input_tokens": 11}An estimate; may include system-added tokens which are not billed. Actual usage.input_tokens on the real call matched exactly for our plain prompt (11 = 11).
Live measurements (Haiku 4.5, "Reply with OK.")
| Variant | input_tokens |
|---|---|
| messages only | 11 |
+ system string (7 words) |
17 |
+ one client tool (get_weather, 1 property) |
561 |
| + system + tool (shell example) | 568 |
| + 1×1 PNG image block | 16 |
+ text document ("The sky is blue.", citations.enabled) |
562 |
+ thinking: {type: enabled, budget_tokens: 1024} |
40 |
+ output_config.format json_schema (1 property) |
151 |
+ web_search_20250305 server tool |
400 |
Limits & pricing
Free. Separate RPM bucket from Messages (documented per usage tier: Start 5,000 · Build 10,000 · Scale 20,000 RPM). Rate-limit headers are not returned on count_tokens responses (only request-id, anthropic-organization-id, CF-RAY observed).
SDKs
client.messages.count_tokens(...) / client.messages.countTokens(...) → MessageTokensCount.input_tokens; beta namespace client.beta.messages.count_tokens(betas=[...]). Examples: examples/anthropic/messages/count_tokens.{sh,py,ts}; tests: tests/anthropic/test_count_tokens.py.