SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
13.5 KB

# Anthropic — Message Batches API

Status: DOCUMENTED · LIVE_VERIFIED (all 6 endpoints exercised on 2026-09-18 with a regular API key; two batches created, polled, results downloaded, cancelled and deleted).
Sources:

Last verified: 2026-09-18


# 1. What it is

Asynchronous, bulk submission of Messages API requests. You POST a list of independent requests[] (each = custom_id + full Messages params), the batch is processed in the background, you poll the batch object until processing_status == "ended", then download a .jsonl results file. Each request is processed independently (one failure does not affect the others) and results are not ordered — match them with custom_id.

Fact Value (documented) Observed 2026-09-18
Base path /v1/messages/batches —
Beta header none required (GA). The reference examples still show anthropic-beta: message-batches-2024-09-24; output-300k-2026-03-24 raises max_tokens to 300k inside batches (Opus 5/4.8/4.7/4.6, Sonnet 5/4.6) worked without any beta header
Max size 100,000 requests or 256 MB, whichever first —
Processing SLA most batches < 1 h; hard expiry at 24 h (expires_at = created_at + 24h) 2-request batch: ended after ~102 s; 1-request cancelled batch: ended ~30 s after cancel
Results retention 29 days after created_at (then archived_at is set and results are gone) results_url present once ended
Scope Workspace (batches and results only visible inside the creating workspace; anthropic-workspace-id header accepted) —
Pricing 50 % of standard price on all tokens (input, output, cache write/read). Stacks with prompt-caching discounts. Batch results carry usage.service_tier: "batch" service_tier: "batch" confirmed in the succeeded result
Models all active models claude-haiku-4-5-20251001
Not allowed inside a batch stream: true, speed (fast mode), max_tokens: 0 —
Rate limits (documented, tier-dependent) Start: 1,000 RPM / 200k queued requests; Build: 2,000 RPM / 300k; higher tiers 4,000 RPM / 500k; 100k requests per batch. Batch usage does not consume Messages API rate limits headers on batch calls did not expose ratelimit values in our run
Spend limits high concurrency may slightly overshoot the workspace spend limit —

# What can be batched

Everything the Messages API accepts, mixed freely inside one batch: vision, PDFs, system prompts, multi-turn, extended thinking, tool use incl. all server tools (web search, web fetch, code execution, MCP connector, advisor, tool search), structured outputs, prompt caching, "most beta features" (put the anthropic-beta header on the batch create call; it applies to every request). The batch worker runs the same server-side agentic loop as the synchronous API but with more iterations per turn before returning stop_reason: "pause_turn"; a pause_turn result must be continued in a follow-up request. web_search is throttled per organization inside batches and retried automatically. max_tokens: 0 (cache pre-warm) is rejected. Prompt-cache hits are best-effort (30–98 %); the 1-hour TTL (cache_control: {type: "ephemeral", ttl: "1h"}) is recommended.

# 2. Endpoints

Method Path Title Path/query params Body Pagination Beta header Status
GET /v1/messages/batches List Message Batches after_id, before_id, limit — cursor_ids message-batches-2024-09-24 LV
POST /v1/messages/batches Create a Message Batch — json — message-batches-2024-09-24 LV
DELETE /v1/messages/batches/{message_batch_id} Delete a Message Batch — — — message-batches-2024-09-24 LV
GET /v1/messages/batches/{message_batch_id} Retrieve a Message Batch — — — message-batches-2024-09-24 LV
POST /v1/messages/batches/{message_batch_id}/cancel Cancel a Message Batch — — — message-batches-2024-09-24 LV
GET /v1/messages/batches/{message_batch_id}/results Retrieve Message Batch results — — — message-batches-2024-09-24 LV

results_url returned on the batch object is the absolute URL of the results endpoint (https://api.anthropic.com/v1/messages/batches/{id}/results).

# Pagination (list)

ID-cursor style: limit (1–100, default 20), before_id, after_id. Response { "data": [...], "has_more": bool, "first_id": ..., "last_id": ... }. Most recently created first. Observed: ?limit=1 → one batch, has_more: false (only one batch existed); ?after_id=<that id> → data: [].

# 3. Request body (create)

json
{
  "requests": [
    { "custom_id": "ok-1",
      "params": { "model": "claude-haiku-4-5-20251001", "max_tokens": 8,
                  "messages": [{"role": "user", "content": "Reply with OK."}] } },
    { "custom_id": "bad-model-1",
      "params": { "model": "claude-does-not-exist-1", "max_tokens": 8,
                  "messages": [{"role": "user", "content": "Reply with OK."}] } }
  ]
}
  • requests[] — 1..100,000 items. Empty list → 400 invalid_request_error: "requests: List should have at least 1 item after validation, not 0" (observed).
  • requests[].custom_id — string, 1–64 chars, ^[a-zA-Z0-9_-]{1,64}$, unique within the batch. Duplicate → 400 "requests: custom_ids must be unique within a batch. Duplicate custom_idfound:dup" (observed).
  • requests[].params — the complete POST /v1/messages body (model, max_tokens, messages, system, tools, tool_choice, thinking, output_config, container, mcp_servers, context_management, metadata, service_tier, stop_sequences, temperature/top_p/top_k, …). Full flattened list (380 parameters): generated/fragments/parameters/anthropic-batches.json.
  • Per-request validation that can only be detected at processing time (e.g. unknown model) does not fail the create call: it surfaces as an errored result (observed: not_found_error "model: claude-does-not-exist-1").

# 4. The MessageBatch object

json
{
  "id": "msgbatch_01F32YGd3yxLE1nLvRTSo8QD", "type": "message_batch",
  "processing_status": "in_progress",
  "request_counts": {"processing": 2, "succeeded": 0, "errored": 0, "canceled": 0, "expired": 0},
  "ended_at": null, "created_at": "2026-09-19T01:49:44.445332+00:00",
  "expires_at": "2026-09-20T01:49:44.445332+00:00", "archived_at": null,
  "cancel_initiated_at": null, "results_url": null
}
Field Type Notes
processing_status in_progress · canceling · ended lifecycle below
request_counts {processing, succeeded, errored, canceled, expired} sums to the number of requests
created_at, expires_at RFC 3339 expires_at = created_at + 24 h (observed exactly)
ended_at RFC 3339 or null set when every request is terminal
cancel_initiated_at RFC 3339 or null set by POST …/cancel (observed 0.3 s after create)
archived_at RFC 3339 or null set when results are purged (29 days)
results_url string or null null until ended

# Lifecycle

stateDiagram-v2
  [*] --> in_progress : POST /v1/messages/batches
  in_progress --> ended : all requests succeeded/errored/expired (≤ 24 h)
  in_progress --> canceling : POST …/cancel
  canceling --> ended : in-flight requests finish or are canceled
  ended --> [*] : DELETE (allowed only here)
  • DELETE while in_progress → 400 invalid_request_error "In progress batches cannot be deleted. Consider canceling the batch instead." (observed). While canceling → 400 "Batch … has not finished canceling. Please wait for cancellation to completed before deleting." (observed, sic). After ended → 200 {"id": "...", "type": "message_batch_deleted"}; a later GET returns 404 not_found_error "Message Batch … not found.".
  • POST …/cancel is idempotent: second call returned 200 with the same canceling object. Requests already finished before the cancel keep their result; the rest end as canceled. Observed for a fresh 1-request batch: request_counts.canceled = 1, ended ~30 s later.
  • GET …/results before ended → 404 not_found_error "Message Batch … has no available results." (observed).

# 5. Results (GET …/results)

  • Content type application/x-jsonl (one JSON object per line, streamed; may be large — read line by line). Order is not the request order (observed: bad-model-1 came before ok-1).
  • Line shape (MessageBatchIndividualResponse): {"custom_id": "...", "result": {...}} with result.type ∈:
result.type Extra fields Observed example
succeeded message = full Message object (usage.service_tier: "batch") {"custom_id":"ok-1","result":{"type":"succeeded","message":{"id":"msg_…","content":[{"type":"text","text":"OK."}],"stop_reason":"end_turn","usage":{"input_tokens":11,"output_tokens":5,"service_tier":"batch","inference_geo":"not_available",…}}}}
errored error = {"type":"error","error":{"type":"<error type>","message":"…","details":null},"request_id":null} {"custom_id":"bad-model-1","result":{"type":"errored","error":{"type":"error","error":{"details":null,"type":"not_found_error","message":"model: claude-does-not-exist-1"},"request_id":null}}}
canceled none {"custom_id":"cancel-1","result":{"type":"canceled"}}
expired none not observed (needs a 24 h timeout)

Errored requests are not billed; expired/canceled requests are not billed either (documented). Errors inside errored can be any Messages API error (invalid_request_error, not_found_error, rate_limit_error, overloaded_error, api_error…) — retry the overloaded/api_error ones in a new batch.

# 6. Errors (batch endpoints themselves)

HTTP error.type When (observed unless noted)
400 invalid_request_error empty requests, duplicate custom_id, delete while in_progress/canceling, unsupported params (stream, speed, max_tokens: 0 — documented)
404 not_found_error unknown/deleted batch id; results not yet available
413 request_too_large body > 256 MB (documented)
429 rate_limit_error batch RPM or queued-requests limit (documented)

# 7. SDK surface

Operation Python TypeScript
create client.messages.batches.create(requests=[...]) (client.beta.messages.batches.create(betas=[...], …) for beta params) client.messages.batches.create({requests})
retrieve client.messages.batches.retrieve(id) client.messages.batches.retrieve(id)
list client.messages.batches.list(limit=…) (auto-paginating iterator) for await (const b of client.messages.batches.list())
cancel client.messages.batches.cancel(id) client.messages.batches.cancel(id)
delete client.messages.batches.delete(id) client.messages.batches.delete(id)
results for r in client.messages.batches.results(id): r.result.type (JSONL decoder; streams) for await (const r of await client.messages.batches.results(id))
CLI ant messages:batches create <<'YAML' … (documented) —

Python typing helpers (documented): anthropic.types.messages.batch_create_params.Request, anthropic.types.message_create_params.MessageCreateParamsNonStreaming.

# 8. Live run summary (2026-09-18, claude-haiku-4-5-20251001, cost ≈ $0.0001)

Step Call Status Timing
create (2 req) POST /v1/messages/batches 200 in_progress 0.34 s
retrieve / list / list after_id GET 200 0.2–0.3 s
delete while in_progress DELETE 400 —
results while in_progress GET …/results 404 —
poll ×6 every 20 s GET 200 → ended 102 s after create (ended_at − created_at = 87 s)
results GET …/results 200, 2 lines (errored, succeeded) 0.3 s
delete after ended DELETE 200 message_batch_deleted —
retrieve after delete GET 404 —
create #2 (1 req) + cancel immediately POST, POST …/cancel ×2 200 canceling (idempotent) cancel 0.2 s
delete while canceling DELETE 400 —
poll → ended GET request_counts.canceled: 1 ~30 s
results #2 GET …/results 200, {"type":"canceled"} —
delete #2 DELETE 200 —
empty requests / duplicate custom_id / unknown id 400 / 400 / 404 —

Both batches were deleted; no residue. Examples: examples/anthropic/batch/ · Tests: tests/anthropic/test_batches.py (run_batch_tests marker).

# 9. Gotchas

  • max_tokens: 0 and stream are rejected at create time; an invalid model is not — it becomes an errored result.
  • Results are unordered and keyed only by custom_id; the results endpoint 404s until ended.
  • You cannot delete an in-progress batch — cancel first, then wait for ended.
  • Retention clock is created_at (29 days), not ended_at.
  • Batch responses report usage.service_tier: "batch"; Usage/Cost Admin reports can filter service_tiers[]=batch.
  • Console downloads of results can be disabled per organization/workspace (privacy control).