Anthropic Messages API — POST /v1/messages (exhaustive reference)
Status: DOCUMENTED + LIVE_VERIFIED (46-call probe on claude-haiku-4-5-20251001, 2026-09-18, est. $0.0023; raws in tmp-live/anthropic-core/). Beta fields marked BETA. Sampling params DEPRECATED.
Sources: https://platform.claude.com/docs/en/api/messages · https://platform.claude.com/docs/en/api/beta/messages · https://platform.claude.com/docs/en/api/errors · https://platform.claude.com/docs/en/api/service-tiers · https://platform.claude.com/docs/en/build-with-claude/streaming
Machine-readable: generated/fragments/endpoints/anthropic-core.json, generated/fragments/parameters/anthropic-messages.json (120 records), generated/fragments/objects/anthropic-messages-objects.json
Last verified: 2026-09-18
1. Request
POST https://api.anthropic.com/v1/messages
x-api-key: <key> (or Authorization: Bearer <key|WIF/OAuth token>)
anthropic-version: 2023-06-01
content-type: application/json
[anthropic-beta: a,b] [anthropic-workspace-id: wrkspc_…] [anthropic-user-profile-id: … (beta)]Body limit 32 MB (413 request_too_large, returned by Cloudflare). The API is stateless: replay the whole history every call. Consecutive same-role turns are merged; max 100,000 messages.
1.1 Top-level body parameters (GA surface)
| Field | Type | Req. | Default / range | Notes (live behaviour in italics) |
|---|---|---|---|---|
model |
string | yes | — | id or alias (claude-haiku-4-5, claude-sonnet-5, claude-fable-5-1…). Unknown → 404 not_found_error "model: …" |
messages[] |
MessageParam[] | yes | ≤100k | {role: user|assistant, content: string | ContentBlockParam[]}; string = one text block. Empty → 400 "at least one message is required" |
max_tokens |
integer ≥0 | yes | model max (Haiku 4.5: 64000) | missing → 400 "max_tokens: Field required"; 1,000,000 → 400 "1000000 > 64000 …"; 0 pre-warms the cache: content [], output_tokens 0, stop_reason max_tokens |
system |
string | TextBlockParam[] | no | — | each block: {type: text, text, cache_control?, citations?}. No system role in messages (the schema lists system in the role enum but docs say use this field). |
stop_sequences |
string[] | no | — | hit → stop_reason: stop_sequence, stop_sequence: "3", matched text excluded ("1 2 ") |
stream |
boolean | no | false | SSE — see docs/anthropic/streaming.md |
metadata.user_id |
string ≤512 | no | — | opaque end-user id, no PII |
service_tier |
auto | standard_only |
no | auto |
both → usage.service_tier: "standard" (no Priority Tier commitment; Priority Tier is no longer sold) |
cache_control (top-level) |
{type: ephemeral, ttl?: 5m|1h} |
no | — | auto-marks the last cacheable block. accepted |
inference_geo |
string | no | workspace default | e.g. "us" (Claude 4.6+, priced 1.1x). Haiku 4.5 → 400 "does not support inference_geo" |
container |
string | {id?, skills?[{type: anthropic|custom, skill_id, version?}]} |
no | — | code-execution container reuse (skills ≤20) |
output_config |
object | no | — | effort: low|medium|high|xhigh|max, format: {type: json_schema, schema} (structured outputs, GA), task_budget (beta) |
thinking |
object | no | — | {type: enabled, budget_tokens ≥1024, display?} | {type: adaptive, display?} | {type: disabled}. Haiku 4.5: adaptive → 400 "adaptive thinking is not supported on this model". Deep-dive: thinking doc. |
tools[] / tool_choice |
see tools doc | no | tool_choice.type: auto |
tool_choice.type ∈ auto|any|tool(name)|none, disable_parallel_tool_use. forced tool → stop_reason: tool_use, block has caller: {type: direct} |
temperature |
number 0..1 | no | 1.0 | DEPRECATED (models after Opus 4.6 accept only 1.0). Haiku 4.5: with top_p → 400 "temperature and top_p cannot both be specified for this model"; 5 → 400 "temperature: range: 0..1" |
top_p |
number 0..1 | no | — | DEPRECATED (only ≥0.99 accepted on newer models); mutually exclusive with temperature (live) |
top_k |
integer ≥0 | no | — | DEPRECATED (rejected on models after Opus 4.6). "abc" → 400 "top_k: Input should be a valid integer" |
anthropic Python SDK 1.7.0 has removed the
temperature/top_p/top_kkeyword arguments frommessages.create/stream/parse(TypeError). Useextra_body={"temperature": 0.2}. The TS SDK 0.126.0 still typestemperature?.
1.2 Beta-only body parameters (need anthropic-beta; SDK client.beta.messages + betas=[…])
| Field | Beta header | Shape | Live |
|---|---|---|---|
context_management.edits[] |
context-management-2025-06-27 |
{type: clear_tool_uses_20250919, trigger?, keep?, clear_at_least?, clear_tool_inputs?, exclude_tools?} · {type: clear_thinking_20251015, keep?} · {type: compact_20260112, trigger?, instructions?, pause_after_compaction?} |
accepted on Haiku 4.5 → response context_management: {applied_edits: []} |
compaction |
compact-2026-09-04 |
{type: summarize, instructions? ≤16384} — compacts the whole conversation, returns one signed compaction block, samples nothing |
not tested |
mcp_servers[] (≤20) |
mcp-client-2025-11-20 (older mcp-client-2025-04-04) |
{type: url, name, url, authorization_token?, tool_configuration?: {enabled?, allowed_tools?}} |
not tested |
speed |
fast-mode-2026-02-01 |
standard | fast |
Haiku 4.5 → 400 "does not support the speed parameter" |
diagnostics.previous_message_id |
cache-diagnosis-2026-04-07 |
string ≤256 → response diagnostics (cache-miss reasons: model/system/tools/messages changed, previous not found, unavailable) |
not tested |
fallbacks |
server-side-fallback-2026-07-01 |
[{model, max_tokens?, output_config?, speed?, thinking?}] or "default" |
not tested |
fallback_credit_token |
fallback-credit-2026-07-01 |
string | {token, mode: strict|best_effort} (redeem within 5 min of a refusal) |
not tested |
output_format |
structured-outputs-2025-11-13 |
legacy location of output_config.format |
without header → 400 "output_format: Extra inputs are not permitted" |
thinking.display: updates |
thinking-display-updates-2026-08-18 |
streams only progress updates | — |
thinking.block_binding.prefix_mismatch_behavior |
thinking-binding-controls-2026-08-01 |
error | drop_block |
— |
output_config.task_budget |
task-budgets-2026-03-13 |
{type: tokens, total ≥1024, remaining?} |
— |
betas |
— | SDK-only (client.beta.messages.create(betas=[…])). As body field → 400 "betas: Extra inputs are not permitted" |
1.3 Message content blocks (request)
messages[].content[] union (16 types): text, image, document, search_result, thinking, redacted_thinking, tool_use, tool_result, server_tool_use, web_search_tool_result, web_fetch_tool_result, code_execution_tool_result, bash_code_execution_tool_result, text_editor_code_execution_tool_result, tool_search_tool_result, container_upload (+ beta: mcp_tool_use, mcp_tool_result, advisor_tool_result, compaction, fallback). Every block accepts cache_control. Full field tables: docs/anthropic/content-blocks.md.
1.4 Prefill (assistant-final message)
Ending messages with an assistant turn makes the model continue it. Haiku 4.5: allowed ("The word is:" → " OK"). claude-sonnet-5: 400 "This model does not support assistant message prefill. The conversation must end with a user message." — Claude 4.6+ and Mythos Preview reject prefill; use output_config.format or system instructions.
2. Response — Message
{"id":"msg_011CfBxdcZj7KMHsf2wSPauQ","type":"message","role":"assistant","model":"claude-haiku-4-5-20251001",
"content":[{"type":"text","text":"OK."}],"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"container":null,
"usage":{"input_tokens":11,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,
"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},
"output_tokens":5,"service_tier":"standard","inference_geo":"not_available"}}| Field | Type | Notes |
|---|---|---|
id |
msg_… |
format may change |
type / role |
message / assistant |
constants |
model |
string | resolved snapshot id (alias in → snapshot out) |
content[] |
ContentBlock[] | response union (12 GA types + beta) — see content-blocks doc; empty for max_tokens: 0 |
stop_reason |
enum | end_turn · max_tokens · stop_sequence · tool_use · pause_turn · refusal · model_context_window_exceeded · compaction (beta). Null only in streaming message_start. See docs/anthropic/stop-reasons.md |
stop_sequence |
string|null | matched custom stop |
stop_details |
{type: refusal, category: cyber|bio|frontier_llm|reasoning_extraction|general_harms|null, explanation} | null |
null on all normal responses (observed live) |
container |
{id, expires_at, skills[]} | null |
code-execution container |
usage |
object | input_tokens, output_tokens, cache_creation_input_tokens, cache_read_input_tokens, cache_creation{ephemeral_5m_input_tokens, ephemeral_1h_input_tokens}, output_tokens_details{thinking_tokens} (present when thinking ran — 39 output / 32 thinking live), server_tool_use{web_search_requests, web_fetch_requests}, service_tier (standard|priority|batch), inference_geo (live value "not_available" — not enumerated in docs); beta: iterations[], speed, fallback_credit |
context_management (beta) |
{applied_edits: []} |
present only when the beta header was sent (observed) |
diagnostics, input_transformations (beta) |
cache-miss diagnosis; dropped/mismatched thinking blocks |
Total billed input = input_tokens + cache_creation_input_tokens + cache_read_input_tokens. output_tokens is non-zero even for empty text (except max_tokens: 0).
3. Response headers (live, 200)
request-id: req_…, anthropic-organization-id, anthropic-ratelimit-{requests,tokens,input-tokens,output-tokens}-{limit,remaining,reset} (observed limits 10,000 RPM / 12M TPM / 10M ITPM / 2M OTPM on this key — account-specific, not documented values), CF-RAY. Errors add x-should-retry: false. Details: docs/anthropic/headers.md.
4. Versioning
Only anthropic-version: 2023-06-01 works: 2020-01-01 → 400 "is not a valid version"; 2023-01-01 → 400 "not allowed for this endpoint". Within a version Anthropic may add optional inputs, new output values and new enum variants (e.g. stream event types) — parse defensively.
5. Long requests
Non-streaming requests should stay under ~10 minutes: SDKs refuse non-streaming calls expected to exceed 10 min (TS scales the timeout by max_tokens/128000 up to 60 min), set TCP keep-alive, and retry twice. Prefer stream: true (+ get_final_message() / finalMessage()) or the Batches API for very large max_tokens. 504 timeout_error otherwise.
6. Related endpoints
| Endpoint | Doc | Live |
|---|---|---|
POST /v1/messages/count_tokens |
docs/anthropic/token-counting.md |
200 |
GET /v1/models, GET /v1/models/{id} |
models doc | 200; alias claude-haiku-4-5 → claude-haiku-4-5-20251001, max_tokens 64000, max_input_tokens 200000 |
POST /v1/complete |
docs/anthropic/text-completions-legacy.md |
400 "endpoint has been deprecated" for every model |
| Batches | batches doc | — |