Anthropic API errors — catalogue, retry matrix, robust handlers
Status: DOCUMENTED; 400/401/404 shapes and messages LIVE_VERIFIED 2026-09-18 (15 deliberate errors on claude-haiku-4-5-20251001, nothing billed; raws tmp-live/anthropic-core/i*_*.json). x-should-retry header LIVE_DISCOVERED.
Sources: https://platform.claude.com/docs/en/api/errors · https://platform.claude.com/docs/en/build-with-claude/streaming#error-events · https://platform.claude.com/docs/en/api/rate-limits · https://platform.claude.com/docs/en/api/beta-headers#error-handling · SDK pages (…/sdks/python#handling-errors etc.)
Machine-readable: generated/fragments/errors/anthropic-errors.json
Last verified: 2026-09-18
Error shape
{"type": "error", "error": {"type": "not_found_error", "message": "model: claude-does-not-exist"}, "request_id": "req_011CfBxeBpDojVzndwYnobEw"}Always JSON; error.type/error.message always present; request_id mirrors the request-id header (quote it to support). Enum of error.type may grow (versioning policy). Mid-stream errors (event: error) have the same error object but no request_id.
Catalogue + retry matrix
| HTTP | error.type |
Meaning | Retry? | Action | Live message(s) |
|---|---|---|---|---|---|
| 400 | invalid_request_error |
malformed/invalid request, unsupported param for the model, invalid version/beta header, spend limit reached (org/workspace) | no | fix the request; the message names the field | max_tokens: Field required · max_tokens: 1000000 > 64000, which is the maximum allowed number of output tokens for claude-haiku-4-5-20251001 · temperature: range: 0..1 · top_k: Input should be a valid integer · `temperature` and `top_p` cannot both be specified for this model. Please use only one. · messages: at least one message is required · The request body is not valid JSON: … · anthropic-version: "2020-01-01" is not a valid version · anthropic-version: "2023-01-01" not allowed for this endpoint · Unexpected value(s) `does-not-exist-2099-01-01` for the `anthropic-beta` header… · This model does not support assistant message prefill. The conversation must end with a user message. (claude-sonnet-5) · adaptive thinking is not supported on this model · messages.2: `tool_use` ids were found without `tool_result` blocks immediately after: … · Server tools are not supported in the count_tokens endpoint: web_search_20250305… · 'claude-haiku-4-5-20251001' does not support inference_geo. · '…' does not support the `speed` parameter… · betas: Extra inputs are not permitted · output_format: Extra inputs are not permitted · The /v1/complete endpoint has been deprecated… |
| 401 | authentication_error |
key malformed/revoked/expired (AWS: bad SigV4) | no | fix credentials | invalid x-api-key |
| 402 | billing_error |
billing problem | no | fix payment in Console / AWS Marketplace | — |
| 403 | permission_error |
key lacks permission (workspace/org settings) | no | check access; not "does not exist" | — |
| 404 | not_found_error |
unknown path, id or model | no | check path/ids; GET /v1/models |
model: claude-does-not-exist · Not found (unknown path) |
| 409 | conflict_error |
state conflict / uniqueness | yes (SDK default) | resolve then retry | — |
| 413 | request_too_large |
> 32 MB Messages/count_tokens, 256 MB batches, 500 MB files (from Cloudflare) | no | shrink / Files API / batches | — |
| 422 | (SDK UnprocessableEntityError) |
not in the HTTP list; SDK class exists | no | treat as 400 | — |
| 429 | rate_limit_error |
RPM/ITPM/OTPM bucket, acceleration limit, monthly tier spend cap (no retry-after, keeps failing), Claude Code workspace spend limit |
yes, unless spend cap | honor retry-after; exponential backoff + jitter; ramp gradually |
— |
| 500 | api_error |
internal error | yes | backoff; support with request_id | — |
| 504 | timeout_error |
processing timeout | yes | stream / batches for long requests | — |
| 529 | overloaded_error |
capacity (all users); also as SSE error event after 200 |
yes | backoff; Priority Tier reduces it | — |
| 200 | stop_reason: refusal |
not an error — classifier refusal with stop_details.category |
n/a | fallback model / fallbacks beta |
— |
Documented validation families (400): prefill unsupported (Claude 4.6+, Mythos Preview); thinking blocks modified (messages.i.content.j: … cannot be modified); thinking.type.enabled unsupported on Claude 4.7+ ("Use thinking.type.adaptive and output_config.effort"); adaptive unsupported on ≤4.5; disabled unsupported on Fable 5.x / Mythos 5.x / Mythos Preview; forced tool_choice any/tool unsupported on Fable 5.1 / Mythos 5.1; thinking block bound to another conversation (Invalid signature in thinking block…, thinking-binding-controls-2026-08-01); block_binding: Extra inputs are not permitted without that header; AWS "Outbound web identity federation is disabled for your account".
Server-tool failures are not HTTP errors: result blocks carry error_code (web_search: invalid_tool_input|unavailable|max_uses_exceeded|too_many_requests|query_too_long|request_too_large; web_fetch adds url_too_long|url_not_allowed|url_not_in_prior_context|url_not_accessible|unsupported_content_type|content_too_large; code execution: execution_time_exceeded, output_file_too_large, file_not_found). Treat too_many_requests/unavailable as transient.
Response headers on errors
request-id (always), x-should-retry: false (every 4xx observed; SDKs honor true/false), anthropic-organization-id (absent on 401 and unknown-path 404), retry-after (429/529 when applicable), CF-RAY. Rate-limit buckets are not sent on error responses.
SDK defaults
Retries: 2, exponential backoff (Python 0.5 s → 8 s, jitter), on connection errors, 408, 409, 429, ≥500, timeouts; retry-after honored. Timeout 10 min; non-streaming requests expected to exceed it are refused (ValueError / AnthropicError("Streaming is required…")). Typed classes: Python anthropic.NotFoundError etc.; TS Anthropic.NotFoundError; Ruby Anthropic::Errors::NotFoundError; Java com.anthropic.errors.NotFoundException; C# AnthropicNotFoundException; Go *anthropic.Error (branch on StatusCode). Catch the most specific class first; never string-match messages.
Robust handler — Python (examples/shared/errors/anthropic_error_handling.py, LIVE_VERIFIED)
import random, time, anthropic
client = anthropic.Anthropic(max_retries=0) # we do our own retry here; default 2 is fine in production
FATAL = (anthropic.BadRequestError, anthropic.AuthenticationError, anthropic.PermissionDeniedError,
anthropic.NotFoundError, anthropic.UnprocessableEntityError)
def call_with_backoff(fn, attempts=4, base=0.5, cap=8.0):
for i in range(attempts):
try:
return fn()
except anthropic.APIStatusError as e: # 4xx/5xx with parsed body
if isinstance(e, FATAL) or e.response.headers.get("x-should-retry") == "false" or i == attempts - 1:
raise # log e.status_code, e.request_id, e.body["error"]
ra = e.response.headers.get("retry-after")
time.sleep(float(ra) if ra else min(cap, base * 2 ** i) * (1 + random.random() / 4))
except (anthropic.APIConnectionError, anthropic.APITimeoutError):
if i == attempts - 1: raise
time.sleep(min(cap, base * 2 ** i))
msg = call_with_backoff(lambda: client.messages.create(model="claude-haiku-4-5-20251001", max_tokens=16,
messages=[{"role": "user", "content": "Reply with OK."}]))
if msg.stop_reason == "refusal": # HTTP 200 but no usable answer
... # retry on a fallback model, read msg.stop_details.categoryStreaming: wrap the with client.messages.stream(...) block in the same handler — a mid-stream error event raises anthropic.APIError; retry the whole request. Note: anthropic 1.7.0 dropped temperature/top_p/top_k kwargs — use extra_body.
Robust handler — TypeScript (examples/shared/errors/anthropic_error_handling.ts, LIVE_VERIFIED)
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({ maxRetries: 0 });
const fatal = (e: Anthropic.APIError) => e instanceof Anthropic.BadRequestError || e instanceof Anthropic.AuthenticationError ||
e instanceof Anthropic.PermissionDeniedError || e instanceof Anthropic.NotFoundError || e instanceof Anthropic.UnprocessableEntityError;
async function withBackoff<T>(fn: () => Promise<T>, attempts = 4, base = 500, cap = 8000): Promise<T> {
for (let i = 0; ; i++) {
try { return await fn(); }
catch (e) {
if (!(e instanceof Anthropic.APIError)) throw e; // APIConnectionError etc. also extend APIError
if (fatal(e) || e.headers?.get?.("x-should-retry") === "false" || i === attempts - 1) throw e; // e.status, e.requestID, e.error
const ra = e.headers?.get?.("retry-after");
await new Promise(r => setTimeout(r, ra ? Number(ra) * 1000 : Math.min(cap, base * 2 ** i) * (1 + Math.random() / 4)));
}
}
}
const msg = await withBackoff(() => client.messages.create({ model: "claude-haiku-4-5-20251001", max_tokens: 16, messages: [{ role: "user", content: "Reply with OK." }] }));
if (msg.stop_reason === "refusal") { /* fallback model; msg.stop_details?.category */ }For client.messages.stream(), attach .on("error", …) and/or await finalMessage() inside the same wrapper.
Errors vs stop reasons
Errors = HTTP 4xx/5xx (or SSE error event), no usable content. Stop reasons = HTTP 200 with content and a stop_reason (refusal, max_tokens, pause_turn need application handling) — see docs/anthropic/stop-reasons.md.