Anthropic — Message Batches API
Status: DOCUMENTED · LIVE_VERIFIED (all 6 endpoints exercised on 2026-09-18 with a regular API key; two batches created, polled, results downloaded, cancelled and deleted).
Sources:
- https://platform.claude.com/docs/en/build-with-claude/batch-processing (user guide)
- https://platform.claude.com/docs/en/api/messages/batches/create · /retrieve · /list · /cancel · /delete · /results (reference;
api/beta/messages/batches/*are the beta-typed twins) - https://platform.claude.com/docs/en/api/rate-limits#message-batches-api
- Live run:
tmp-live/platform-anthropic/batches-scenario.json(sanitized),reports/live-requests.jsonl
Last verified: 2026-09-18
1. What it is
Asynchronous, bulk submission of Messages API requests. You POST a list of independent requests[] (each = custom_id + full Messages params), the batch is processed in the background, you poll the batch object until processing_status == "ended", then download a .jsonl results file. Each request is processed independently (one failure does not affect the others) and results are not ordered — match them with custom_id.
| Fact | Value (documented) | Observed 2026-09-18 |
|---|---|---|
| Base path | /v1/messages/batches |
— |
| Beta header | none required (GA). The reference examples still show anthropic-beta: message-batches-2024-09-24; output-300k-2026-03-24 raises max_tokens to 300k inside batches (Opus 5/4.8/4.7/4.6, Sonnet 5/4.6) |
worked without any beta header |
| Max size | 100,000 requests or 256 MB, whichever first | — |
| Processing SLA | most batches < 1 h; hard expiry at 24 h (expires_at = created_at + 24h) |
2-request batch: ended after ~102 s; 1-request cancelled batch: ended ~30 s after cancel |
| Results retention | 29 days after created_at (then archived_at is set and results are gone) |
results_url present once ended |
| Scope | Workspace (batches and results only visible inside the creating workspace; anthropic-workspace-id header accepted) |
— |
| Pricing | 50 % of standard price on all tokens (input, output, cache write/read). Stacks with prompt-caching discounts. Batch results carry usage.service_tier: "batch" |
service_tier: "batch" confirmed in the succeeded result |
| Models | all active models | claude-haiku-4-5-20251001 |
| Not allowed inside a batch | stream: true, speed (fast mode), max_tokens: 0 |
— |
| Rate limits (documented, tier-dependent) | Start: 1,000 RPM / 200k queued requests; Build: 2,000 RPM / 300k; higher tiers 4,000 RPM / 500k; 100k requests per batch. Batch usage does not consume Messages API rate limits | headers on batch calls did not expose ratelimit values in our run |
| Spend limits | high concurrency may slightly overshoot the workspace spend limit | — |
What can be batched
Everything the Messages API accepts, mixed freely inside one batch: vision, PDFs, system prompts, multi-turn, extended thinking, tool use incl. all server tools (web search, web fetch, code execution, MCP connector, advisor, tool search), structured outputs, prompt caching, "most beta features" (put the anthropic-beta header on the batch create call; it applies to every request). The batch worker runs the same server-side agentic loop as the synchronous API but with more iterations per turn before returning stop_reason: "pause_turn"; a pause_turn result must be continued in a follow-up request. web_search is throttled per organization inside batches and retried automatically. max_tokens: 0 (cache pre-warm) is rejected. Prompt-cache hits are best-effort (30–98 %); the 1-hour TTL (cache_control: {type: "ephemeral", ttl: "1h"}) is recommended.
2. Endpoints
| Method | Path | Title | Path/query params | Body | Pagination | Beta header | Status |
|---|---|---|---|---|---|---|---|
GET |
/v1/messages/batches |
List Message Batches | after_id, before_id, limit | — | cursor_ids | message-batches-2024-09-24 | LV |
POST |
/v1/messages/batches |
Create a Message Batch | — | json | — | message-batches-2024-09-24 | LV |
DELETE |
/v1/messages/batches/{message_batch_id} |
Delete a Message Batch | — | — | — | message-batches-2024-09-24 | LV |
GET |
/v1/messages/batches/{message_batch_id} |
Retrieve a Message Batch | — | — | — | message-batches-2024-09-24 | LV |
POST |
/v1/messages/batches/{message_batch_id}/cancel |
Cancel a Message Batch | — | — | — | message-batches-2024-09-24 | LV |
GET |
/v1/messages/batches/{message_batch_id}/results |
Retrieve Message Batch results | — | — | — | message-batches-2024-09-24 | LV |
results_url returned on the batch object is the absolute URL of the results endpoint (https://api.anthropic.com/v1/messages/batches/{id}/results).
Pagination (list)
ID-cursor style: limit (1–100, default 20), before_id, after_id. Response { "data": [...], "has_more": bool, "first_id": ..., "last_id": ... }. Most recently created first. Observed: ?limit=1 → one batch, has_more: false (only one batch existed); ?after_id=<that id> → data: [].
3. Request body (create)
{
"requests": [
{ "custom_id": "ok-1",
"params": { "model": "claude-haiku-4-5-20251001", "max_tokens": 8,
"messages": [{"role": "user", "content": "Reply with OK."}] } },
{ "custom_id": "bad-model-1",
"params": { "model": "claude-does-not-exist-1", "max_tokens": 8,
"messages": [{"role": "user", "content": "Reply with OK."}] } }
]
}requests[]— 1..100,000 items. Empty list →400 invalid_request_error: "requests: List should have at least 1 item after validation, not 0"(observed).requests[].custom_id— string, 1–64 chars,^[a-zA-Z0-9_-]{1,64}$, unique within the batch. Duplicate →400 "requests:custom_ids must be unique within a batch. Duplicatecustom_idfound:dup"(observed).requests[].params— the completePOST /v1/messagesbody (model, max_tokens, messages, system, tools, tool_choice, thinking, output_config, container, mcp_servers, context_management, metadata, service_tier, stop_sequences, temperature/top_p/top_k, …). Full flattened list (380 parameters):generated/fragments/parameters/anthropic-batches.json.- Per-request validation that can only be detected at processing time (e.g. unknown model) does not fail the create call: it surfaces as an
erroredresult (observed:not_found_error "model: claude-does-not-exist-1").
4. The MessageBatch object
{
"id": "msgbatch_01F32YGd3yxLE1nLvRTSo8QD", "type": "message_batch",
"processing_status": "in_progress",
"request_counts": {"processing": 2, "succeeded": 0, "errored": 0, "canceled": 0, "expired": 0},
"ended_at": null, "created_at": "2026-09-19T01:49:44.445332+00:00",
"expires_at": "2026-09-20T01:49:44.445332+00:00", "archived_at": null,
"cancel_initiated_at": null, "results_url": null
}| Field | Type | Notes |
|---|---|---|
processing_status |
in_progress · canceling · ended |
lifecycle below |
request_counts |
{processing, succeeded, errored, canceled, expired} |
sums to the number of requests |
created_at, expires_at |
RFC 3339 | expires_at = created_at + 24 h (observed exactly) |
ended_at |
RFC 3339 or null | set when every request is terminal |
cancel_initiated_at |
RFC 3339 or null | set by POST …/cancel (observed 0.3 s after create) |
archived_at |
RFC 3339 or null | set when results are purged (29 days) |
results_url |
string or null | null until ended |
Lifecycle
stateDiagram-v2 [*] --> in_progress : POST /v1/messages/batches in_progress --> ended : all requests succeeded/errored/expired (≤ 24 h) in_progress --> canceling : POST …/cancel canceling --> ended : in-flight requests finish or are canceled ended --> [*] : DELETE (allowed only here)
DELETEwhilein_progress→400 invalid_request_error "In progress batches cannot be deleted. Consider canceling the batch instead."(observed). Whilecanceling→400 "Batch…has not finished canceling. Please wait for cancellation to completed before deleting."(observed, sic). Afterended→200 {"id": "...", "type": "message_batch_deleted"}; a later GET returns404 not_found_error "Message Batch…not found.".POST …/cancelis idempotent: second call returned200with the samecancelingobject. Requests already finished before the cancel keep their result; the rest end ascanceled. Observed for a fresh 1-request batch:request_counts.canceled = 1,ended~30 s later.GET …/resultsbeforeended→404 not_found_error "Message Batch…has no available results."(observed).
5. Results (GET …/results)
- Content type
application/x-jsonl(one JSON object per line, streamed; may be large — read line by line). Order is not the request order (observed:bad-model-1came beforeok-1). - Line shape (
MessageBatchIndividualResponse):{"custom_id": "...", "result": {...}}withresult.type∈:
result.type |
Extra fields | Observed example |
|---|---|---|
succeeded |
message = full Message object (usage.service_tier: "batch") |
{"custom_id":"ok-1","result":{"type":"succeeded","message":{"id":"msg_…","content":[{"type":"text","text":"OK."}],"stop_reason":"end_turn","usage":{"input_tokens":11,"output_tokens":5,"service_tier":"batch","inference_geo":"not_available",…}}}} |
errored |
error = {"type":"error","error":{"type":"<error type>","message":"…","details":null},"request_id":null} |
{"custom_id":"bad-model-1","result":{"type":"errored","error":{"type":"error","error":{"details":null,"type":"not_found_error","message":"model: claude-does-not-exist-1"},"request_id":null}}} |
canceled |
none | {"custom_id":"cancel-1","result":{"type":"canceled"}} |
expired |
none | not observed (needs a 24 h timeout) |
Errored requests are not billed; expired/canceled requests are not billed either (documented). Errors inside errored can be any Messages API error (invalid_request_error, not_found_error, rate_limit_error, overloaded_error, api_error…) — retry the overloaded/api_error ones in a new batch.
6. Errors (batch endpoints themselves)
| HTTP | error.type |
When (observed unless noted) |
|---|---|---|
| 400 | invalid_request_error |
empty requests, duplicate custom_id, delete while in_progress/canceling, unsupported params (stream, speed, max_tokens: 0 — documented) |
| 404 | not_found_error |
unknown/deleted batch id; results not yet available |
| 413 | request_too_large |
body > 256 MB (documented) |
| 429 | rate_limit_error |
batch RPM or queued-requests limit (documented) |
7. SDK surface
| Operation | Python | TypeScript |
|---|---|---|
| create | client.messages.batches.create(requests=[...]) (client.beta.messages.batches.create(betas=[...], …) for beta params) |
client.messages.batches.create({requests}) |
| retrieve | client.messages.batches.retrieve(id) |
client.messages.batches.retrieve(id) |
| list | client.messages.batches.list(limit=…) (auto-paginating iterator) |
for await (const b of client.messages.batches.list()) |
| cancel | client.messages.batches.cancel(id) |
client.messages.batches.cancel(id) |
| delete | client.messages.batches.delete(id) |
client.messages.batches.delete(id) |
| results | for r in client.messages.batches.results(id): r.result.type (JSONL decoder; streams) |
for await (const r of await client.messages.batches.results(id)) |
| CLI | ant messages:batches create <<'YAML' … (documented) |
— |
Python typing helpers (documented): anthropic.types.messages.batch_create_params.Request, anthropic.types.message_create_params.MessageCreateParamsNonStreaming.
8. Live run summary (2026-09-18, claude-haiku-4-5-20251001, cost ≈ $0.0001)
| Step | Call | Status | Timing |
|---|---|---|---|
| create (2 req) | POST /v1/messages/batches |
200 in_progress |
0.34 s |
| retrieve / list / list after_id | GET |
200 | 0.2–0.3 s |
| delete while in_progress | DELETE |
400 | — |
| results while in_progress | GET …/results |
404 | — |
| poll ×6 every 20 s | GET |
200 → ended |
102 s after create (ended_at − created_at = 87 s) |
| results | GET …/results |
200, 2 lines (errored, succeeded) |
0.3 s |
| delete after ended | DELETE |
200 message_batch_deleted |
— |
| retrieve after delete | GET |
404 | — |
| create #2 (1 req) + cancel immediately | POST, POST …/cancel ×2 |
200 canceling (idempotent) |
cancel 0.2 s |
| delete while canceling | DELETE |
400 | — |
| poll → ended | GET |
request_counts.canceled: 1 |
~30 s |
| results #2 | GET …/results |
200, {"type":"canceled"} |
— |
| delete #2 | DELETE |
200 | — |
| empty requests / duplicate custom_id / unknown id | 400 / 400 / 404 | — |
Both batches were deleted; no residue. Examples: examples/anthropic/batch/ · Tests: tests/anthropic/test_batches.py (run_batch_tests marker).
9. Gotchas
max_tokens: 0andstreamare rejected at create time; an invalid model is not — it becomes anerroredresult.- Results are unordered and keyed only by
custom_id; the results endpoint 404s untilended. - You cannot delete an in-progress batch — cancel first, then wait for
ended. - Retention clock is
created_at(29 days), notended_at. - Batch responses report
usage.service_tier: "batch"; Usage/Cost Admin reports can filterservice_tiers[]=batch. - Console downloads of results can be disabled per organization/workspace (privacy control).