Resilience patterns for OpenAI, Anthropic, xAI and Gemini clients
Status: DOCUMENTED (rules) · LIVE_VERIFIED (reference implementation exercised end-to-end via the four provider adapters: OpenAI/Anthropic 2026-09-18, xAI/Gemini 2026-09-19) · offline unit tests: tests/shared/test_resilient_client.py (46 tests incl. recorded xAI/Gemini error bodies) + resilientClient.ts --selftest
Sources:
- https://developers.openai.com/api/docs/guides/error-codes · …/guides/rate-limits · …/guides/background (stream resume) · …/guides/production-best-practices · OpenAI OpenAPI spec (Idempotency-Key, sequence_number)
- https://platform.claude.com/docs/en/api/errors · …/api/rate-limits · …/build-with-claude/streaming (in-stream
error) · …/manage-claude/compliance-errors (x-should-retry) - xAI: https://docs.x.ai/developers/debugging · https://docs.x.ai/developers/rate-limits (exponential backoff
2**attempt, 5 retries; Batch API exempt) ·generated/fragments/errors/xai-errors.json,headers/xai-headers.json(live 2026-09-19: 400 bad key, 422 string body,x-ratelimit-*headers) - Gemini: https://ai.google.dev/gemini-api/docs/api-errors · …/troubleshooting · …/rate-limits ·
generated/fragments/errors/gemini-errors.json(live 429 body withQuotaFailure+RetryInfo),headers/gemini-headers.jsonLast verified: 2026-09-19
Implementation: examples/shared/resilient-client/resilient_client.py (stdlib) and resilientClient.ts (fetch). Both are transport-pluggable so the whole policy is unit-testable offline. ResilientClient(provider) for openai | anthropic | xai | gemini (env keys *_API_KEY, base URLs https://api.openai.com, https://api.anthropic.com, https://api.x.ai, https://generativelanguage.googleapis.com; auth Bearer / x-api-key / Bearer xai-… / x-goog-api-key).
1. Error envelopes (parse before classifying)
| Provider | Shape | extract_error() → (type, code, message) |
|---|---|---|
| OpenAI / Anthropic | {"error": {"type", "code", "message"}} |
(type, code, message) |
| xAI inference API | {"code": "invalid-argument", "error": "<message>"} — kebab-case code; some 422/400 bodies are a bare JSON string (serde: "Failed to deserialize the JSON body into the target type: …"); 405 is an empty body |
(code, code, error) / (None, None, <string>) |
| xAI Management API | gRPC-style {"code": 16, "message": "Invalid bearer token…", "details": []} |
("grpc", "16", message) |
| Gemini | google.rpc.Status {"error": {"code": 429, "message", "status": "RESOURCE_EXHAUSTED", "details": [{"@type": "…google.rpc.Help"}, {"@type": "…QuotaFailure", "violations": [{quotaMetric, quotaId, quotaDimensions{model, location}}]}, {"@type": "…RetryInfo", "retryDelay": "40s"}]}}; the OpenAI-compat layer may wrap it in a one-element array |
(status, "429", message) |
| Gemini Interactions API | {"error": {"code": "<snake_case>", "message"}} in bodies, event: error in SSE streams |
(None, "<snake_case>", message) |
2. Error classification (what to retry)
| HTTP | OpenAI (error.type / code) |
Anthropic (error.type) |
xAI (code) |
Gemini (error.status) |
Retry? |
|---|---|---|---|---|---|
| 400 | invalid_request_error |
invalid_request_error (also your own spend limit) |
invalid-argument — also the answer to an incorrect API key ("Incorrect API key provided…") and to unsupported params; bare-string body for multi-agent on chat |
INVALID_ARGUMENT (bad field, unknown name, invalid key "API key not valid"); FAILED_PRECONDITION (billing not enabled for a paid-only feature, region not supported) |
No — never retry a 400; on xAI check the message for "API key" before blaming the request |
| 401 | invalid key / IP not allowlisted | authentication_error |
unauthenticated:no-credentials (header missing); Management API {code:16} |
UNAUTHENTICATED (OAuth/ephemeral token) |
No |
| 402 | — | billing_error |
— | — | No |
| 403 | permission_error |
permission_error |
missing ACL (api-key:endpoint:* / api-key:model:*), disabled/blocked key or team |
PERMISSION_DENIED (key restricted to other APIs, leaked key auto-blocked, foreign File/tuned model) |
No |
| 404 | not_found_error / model_not_found |
not_found_error |
not-found — same message for unknown and restricted models ("does not exist or your team … does not have access") |
NOT_FOUND — also "no longer available to new users" (2.5 models) and preview ids on /v1 |
No (model fallback may apply) |
| 405/415 | — | — | wrong method (empty body) / wrong Content-Type | — | No |
| 408 | timeout | — | — | — | Yes |
| 409 | conflict | conflict_error |
— | ABORTED (retry) vs ALREADY_EXISTS (don't) — without a status in the body assume ALREADY_EXISTS |
per status |
| 413 | — | request_too_large |
— | — | No |
| 416 | — | — | — | OUT_OF_RANGE |
No |
| 422 | UnprocessableEntityError |
— | bare-string serde error ("text: invalid type: integer 123, expected a string at line 1 column 33") |
— | OpenAI once; xAI No (fix the field type) |
| 429 | rate_limit_error (slow_down) with Retry-After |
rate_limit_error with retry-after |
too-many-requests / gRPC 8 RESOURCE_EXHAUSTED: team RPS/TPM or per-key qps/qpm/tpm; no Retry-After documented or observed |
RESOURCE_EXHAUSTED — retry delay only in details[] RetryInfo.retryDelay ("40s") and/or the message tail ("Please retry in 54.22098241s."); no Retry-After header |
Yes with backoff (xAI: pure exponential; Gemini: parsed delay treated as the minimum) |
| 429 | insufficient_quota, spend-limit codes |
rate_limit_error without retry-after (tier spend cap) |
message mentioning credits/billing (prepaid exhausted, $0 invoiced limit — status undocumented) | every "Quota exceeded" line says limit: 0 (free-tier key on a paid-only model such as gemini-3.1-pro, caching, Search grounding) |
No — user action (billing / tier) required; gemini_zero_quota_retryable=True overrides |
| 499 | — | — | — | CANCELLED (client closed) |
n/a |
| 500 | server error | api_error |
internal (see https://status.x.ai) |
INTERNAL (also oversized/unusual inputs → reduce context) |
Yes |
| 501 | — | — | — | UNIMPLEMENTED (PaLM methods, unsupported feature) |
No |
| 502/504 | gateway / timeout | timeout_error 504 |
gRPC 4 DEADLINE_EXCEEDED → 504 (SDK default timeout 1620 s) | DEADLINE_EXCEEDED 504 (long thinking, Flex queueing → raise client timeout, stream, background/Batch) |
Yes |
| 503 | service_unavailable_error / server_is_overloaded |
— | gRPC 14 UNAVAILABLE | UNAVAILABLE ("The model is overloaded. Please try again later.", Flex capacity shed) |
Yes, honour Retry-After when present |
| 529 | — | overloaded_error |
— | — | Yes |
| 200 | — | — | usage-guideline violation still billed ($0.05 on Responses) | promptFeedback.blockReason (no candidates) / candidates[].finishReason ≠ STOP (SAFETY, RECITATION, MAX_TOKENS, MISSING_THOUGHT_SIGNATURE…) |
No — not transport errors; handle in the adapter (stop_reason refusal/max_tokens/other) |
| 202 | — | — | deferred completion not ready (GET /v1/chat/deferred-completion/{id}) |
— | poll with backoff |
| network | APIConnectionError, APITimeoutError |
same | openai SDK on base_url=https://api.x.ai/v1 (same classes) / grpc.RpcError in xai-sdk |
google.genai.errors.APIError (code, details) / connection errors |
Yes (bounded) |
Server hints that override everything:
x-should-retry: false|true(Anthropic) — applied whenever present.Retry-After(OpenAI 429/503; Anthropic 429/529; absent on xAI and Gemini). Treat as a minimum, add jitter, cap it (max_retry_after_s, 60 s).- Gemini:
parse_gemini_retry_delay(body)prefersRetryInfo.retryDelay(protobuf Duration string"40s", or{seconds, nanos}), then the message regexPlease retry in (\d+(\.\d+)?)s. Both were observed together in one body (40 s vs 54.22 s) — the structured value wins. - Anthropic in-stream
event: error(overloaded_error) after HTTP 200 → classify like the status. GeminigenerateContentSSE has no documented mid-stream error event (a{"error": …}chunk is still handled); a stream that closes withoutfinishReasonis a failure. Gemini Interactions streams emitevent: error.
3. Backoff
Full jitter: sleep = U(0, min(cap, base·2^(attempt-1))) (defaults base 0.5 s, cap 20 s), bounded by both max_attempts (4) and a wall-clock max_total_s (120 s). All four vendors recommend exponential backoff with jitter: OpenAI adds "limit both the number of attempts and the total time"; xAI's docs show 2**attempt with 5 retries and point bulk work to the Batch API (exempt from rate limits); Gemini's troubleshooting page says 429/408/5xx only and its SDK defaults to 4 attempts, ~1 s initial, 60 s max.
SDK interaction: openai 3.16.2 and anthropic 1.7.0 retry 2× by default (DEFAULT_MAX_RETRIES = 2); google-genai 2.24 retries 429/5xx via http_options.retry_options; the openai SDK against https://api.x.ai/v1 keeps OpenAI's defaults; xai-sdk (gRPC) surfaces grpc.StatusCode and the docs' sample loop catches RESOURCE_EXHAUSTED. When you own the policy, disable SDK retries (max_retries=0, retry_options=None) and feed exceptions to sdk_error_is_retryable(exc, provider=…) (pass provider="xai" explicitly for the OpenAI SDK-on-xAI case; gRPC errors are mapped by e.code()).
4. Timeouts and connection pooling
- Per-request: connect 5 s, read 600 s (OpenAI/Anthropic SDK defaults). xAI reasoning models can be slow: the docs use 3600 s client timeouts and
xai-sdkdefaults to 1620 s; Gemini's SDK sends itshttp_options.timeoutasX-Server-Timeoutand long Flex/Deep-Research jobs should usebackground: true(Interactions) or Batch rather than a long socket. - Prefer streaming for long generations on all four providers (Anthropic's SDKs refuse non-streaming requests expected to exceed 10 min; Gemini 504
DEADLINE_EXCEEDEDguidance says the same). - Python stdlib
urllib(reference client) opens a connection per request; production:http.clientkeep-alive,httpx.Client(limits=...), or the SDKs. Nodefetch(undici) pools by default. xAI prompt caches are per server: keep a conversation on one host and sendx-grok-conv-id(Chat) /prompt_cache_key(Responses) so retries land on the same cache.
5. Rate-limit headers (observed shape, parse defensively)
| OpenAI | Anthropic | xAI (undocumented, observed 2026-09-19) | Gemini | |
|---|---|---|---|---|
| Requests | x-ratelimit-limit-requests, -remaining-requests, -reset-requests (1s, 6m0s) |
anthropic-ratelimit-requests-limit/-remaining/-reset (RFC 3339) |
x-ratelimit-limit-requests / -remaining-requests (per-minute budget: 7200 grok-4.6/4.5, 1800 grok-4.3/4.20/build); no reset header |
none |
| Tokens | x-ratelimit-limit-tokens, -remaining-tokens, -reset-tokens (+ project-scoped) |
anthropic-ratelimit-tokens-*, input-tokens-*, output-tokens-* |
x-ratelimit-limit-tokens / -remaining-tokens (TPM = documented Tier 0: 50M / 10M); absent on /v1/tokenize-text, /v1/models, /v1/api-key |
none — remaining quota lives only in AI Studio; limits are per project, per model (RPM/TPM/RPD) |
| Retry hint | Retry-After |
retry-after (absent on spend cap) |
none | none (message text / RetryInfo) |
| Request id | x-request-id |
request-id |
x-request-id (= chat.completion.id; absent on catalogue GETs) |
none — use body responseId |
| Extras | openai-processing-ms |
anthropic-organization-id |
x-zero-data-retention: true|false, Cloudflare Server-Timing/CF-RAY |
undocumented X-Gemini-Service-Tier (also on 429), Server-Timing: gfet4t7 |
parse_rate_limit_headers() returns a RateLimitInfo; for Gemini it is intentionally empty except raw. Never generalise observed values into "documented limits" (xAI's header values are below the documented Tier 0 RPS and may be a new-team allowance).
6. Idempotency — facts
- OpenAI:
Idempotency-Keydocumented on exactly one operation (POST /v1/agents/sessions/{id}/events). Anthropic: none. xAI: none documented on any endpoint;POST /v1/files/{id}/public-urlis idempotent by behaviour (same URL returned). Gemini: none on generateContent/Interactions;files.namecustom ids and File Search store names give create-once semantics (409ALREADY_EXISTS). Webhook receivers dedupe onwebhook-id(OpenAI, Anthropic, Gemini static webhooks). - Consequence: a retried
POSTafter a timeout may have succeeded server-side. Mitigations: stream; OpenAIbackground: true+ resume; xAIstore: true+GET /v1/responses/{id}(live: retrievable even withstore:false) or deferred completions (deferred: true, result fetchable once within 24 h); Gemini Interactionsbackground: true+GET /v1beta/interactions/{id}; client-side dedupe keyed on an id you also place inmetadata(OpenAI/Anthropic) /labels(Gemini) — xAI Responses rejectsmetadata, so usesafety_identifier/userorprompt_cache_key.
7. Stream reconnection
- OpenAI Responses:
background: true, stream: true+sequence_numbercursor →GET /v1/responses/{id}?stream=true&starting_after=N. - xAI Responses: same event vocabulary with
sequence_number, butbackgroundis rejected (400) → no resume; restart the request (keep the cursor for logs only). WebSocket mode (wss://api.x.ai/v1/responses) is the documented alternative for long agentic runs. - Anthropic Messages: no resume →
anthropic_stream_with_restart(syntheticrestartmarker; also on in-streamerror). - Gemini generateContent: no resume; every chunk carries running
usageMetadataand the last chunkfinishReason— a close without it means restart. Gemini Interactions: resumable SSEGET /v1beta/interactions/{id}?stream=true&last_event_id=…forbackground: trueinteractions; Live API:sessionResumptionhandles (valid 2 h) aftergoAway/connection reset (~10 min). - All: keep the emitted text length to de-duplicate on resume/restart.
8. Circuit breaker
Per provider (optionally per model): closed → open after 5 consecutive failures → half-open after 30 s → one probe. Open circuits raise CircuitOpen immediately so fallback chains skip a degraded provider.
9. Fallbacks
call_with_fallback([Target…]) tries targets in order; Target.path may contain {model} (Gemini: /v1beta/models/{model}:generateContent).
- Same provider, other model on
RetryExhausted,CircuitOpen, 404model_not_found(xAI: 404not-foundfor restricted models; Gemini: 404 "no longer available to new users"), transport errors, and Gemini 429limit: 0(the model is paid-only for this key → try a Flash/Lite model). - Other provider — the request body must be rebuilt (Responses ≠ Messages ≠ generateContent); use the adapters. Only portable parameters survive; reasoning artefacts never do.
- Never fall back on auth/validation errors: OpenAI/Anthropic
authentication_error/permission_error/invalid_request_error, xAIinvalid-argument(incl. incorrect key)/unauthenticated*, GeminiINVALID_ARGUMENT/UNAUTHENTICATED/PERMISSION_DENIED/FAILED_PRECONDITION(AUTH_OR_VALIDATION_ERROR_TYPES) — the same error will recur and you would mask a real bug. Tested chain: Geminilimit: 0→ xAI (bad key stops the chain) / → Anthropic (succeeds).
10. Budget guard
BudgetGuard(max_usd, prices) converts usage into USD from any of the four shapes: OpenAI/xAI Responses (input_tokens, output_tokens, input_tokens_details.cached_tokens), OpenAI/xAI Chat (prompt_tokens, completion_tokens + completion_tokens_details.reasoning_tokens, prompt_tokens_details.cached_tokens), Anthropic (cache_read_input_tokens, cache_creation_input_tokens), Gemini usageMetadata (promptTokenCount, candidatesTokenCount + thoughtsTokenCount, cachedContentTokenCount). xAI also returns usage.cost_in_usd_ticks (1 tick = $1e-8) — the authoritative per-request cost when available. Prices come from generated/pricing.json; unknown models use a deliberately high default.
11. Checklist
- Parse the envelope first (
{error:{}},{code,error}, bare string, google.rpc, array-wrapped); classify by status and type/code/message. - Never retry: quota/spend codes, Anthropic spend-cap 429, xAI 400 (incl. bad key) / 422, Gemini
FAILED_PRECONDITIONandlimit: 0429s. - Honour
Retry-Afterandx-should-retry; for Gemini readRetryInfo.retryDelay/ "Please retry in Ns"; cap the hint; add jitter; xAI gets pure backoff. - Bound retries by attempts and elapsed time; disable SDK retries when you own the loop; long client timeouts for reasoning models.
- Stream long generations; resume only where supported (OpenAI background, Gemini Interactions
last_event_id, LivesessionResumption); restart elsewhere. - Treat POST retries as non-idempotent: dedupe client-side (
metadata/labels/ your own store). - Log
x-request-id(OpenAI, xAI) /request-id(Anthropic) / bodyresponseId(Gemini) with every failure. - Circuit-break per provider; fall back only on capacity or model-availability errors.
- Guard spend with a price-table-driven budget (xAI: cross-check
cost_in_usd_ticks).