# Resilience patterns for OpenAI, Anthropic, xAI and Gemini clients **Status:** DOCUMENTED (rules) · LIVE_VERIFIED (reference implementation exercised end-to-end via the four provider adapters: OpenAI/Anthropic 2026-09-18, xAI/Gemini 2026-09-19) · offline unit tests: `tests/shared/test_resilient_client.py` (46 tests incl. recorded xAI/Gemini error bodies) + `resilientClient.ts --selftest` **Sources:** - https://developers.openai.com/api/docs/guides/error-codes · …/guides/rate-limits · …/guides/background (stream resume) · …/guides/production-best-practices · OpenAI OpenAPI spec (Idempotency-Key, sequence_number) - https://platform.claude.com/docs/en/api/errors · …/api/rate-limits · …/build-with-claude/streaming (in-stream `error`) · …/manage-claude/compliance-errors (`x-should-retry`) - xAI: https://docs.x.ai/developers/debugging · https://docs.x.ai/developers/rate-limits (exponential backoff `2**attempt`, 5 retries; Batch API exempt) · `generated/fragments/errors/xai-errors.json`, `headers/xai-headers.json` (live 2026-09-19: 400 bad key, 422 string body, `x-ratelimit-*` headers) - Gemini: https://ai.google.dev/gemini-api/docs/api-errors · …/troubleshooting · …/rate-limits · `generated/fragments/errors/gemini-errors.json` (live 429 body with `QuotaFailure` + `RetryInfo`), `headers/gemini-headers.json` **Last verified:** 2026-09-19 Implementation: `examples/shared/resilient-client/resilient_client.py` (stdlib) and `resilientClient.ts` (fetch). Both are transport-pluggable so the whole policy is unit-testable offline. `ResilientClient(provider)` for `openai | anthropic | xai | gemini` (env keys `*_API_KEY`, base URLs `https://api.openai.com`, `https://api.anthropic.com`, `https://api.x.ai`, `https://generativelanguage.googleapis.com`; auth `Bearer` / `x-api-key` / `Bearer xai-…` / `x-goog-api-key`). ## 1. Error envelopes (parse before classifying) | Provider | Shape | `extract_error()` → (type, code, message) | |---|---|---| | OpenAI / Anthropic | `{"error": {"type", "code", "message"}}` | `(type, code, message)` | | xAI inference API | `{"code": "invalid-argument", "error": ""}` — kebab-case code; **some 422/400 bodies are a bare JSON string** (serde: "Failed to deserialize the JSON body into the target type: …"); 405 is an empty body | `(code, code, error)` / `(None, None, )` | | xAI Management API | gRPC-style `{"code": 16, "message": "Invalid bearer token…", "details": []}` | `("grpc", "16", message)` | | Gemini | google.rpc.Status `{"error": {"code": 429, "message", "status": "RESOURCE_EXHAUSTED", "details": [{"@type": "…google.rpc.Help"}, {"@type": "…QuotaFailure", "violations": [{quotaMetric, quotaId, quotaDimensions{model, location}}]}, {"@type": "…RetryInfo", "retryDelay": "40s"}]}}`; the OpenAI-compat layer may **wrap it in a one-element array** | `(status, "429", message)` | | Gemini Interactions API | `{"error": {"code": "", "message"}}` in bodies, `event: error` in SSE streams | `(None, "", message)` | ## 2. Error classification (what to retry) | HTTP | OpenAI (`error.type` / `code`) | Anthropic (`error.type`) | xAI (`code`) | Gemini (`error.status`) | Retry? | |---|---|---|---|---|---| | 400 | `invalid_request_error` | `invalid_request_error` (also your own spend limit) | `invalid-argument` — **also the answer to an incorrect API key** ("Incorrect API key provided…") and to unsupported params; bare-string body for multi-agent on chat | `INVALID_ARGUMENT` (bad field, unknown name, invalid key "API key not valid"); **`FAILED_PRECONDITION`** (billing not enabled for a paid-only feature, region not supported) | **No** — never retry a 400; on xAI check the message for "API key" before blaming the request | | 401 | invalid key / IP not allowlisted | `authentication_error` | `unauthenticated:no-credentials` (header missing); Management API `{code:16}` | `UNAUTHENTICATED` (OAuth/ephemeral token) | No | | 402 | — | `billing_error` | — | — | No | | 403 | `permission_error` | `permission_error` | missing ACL (`api-key:endpoint:*` / `api-key:model:*`), disabled/blocked key or team | `PERMISSION_DENIED` (key restricted to other APIs, leaked key auto-blocked, foreign File/tuned model) | No | | 404 | `not_found_error` / `model_not_found` | `not_found_error` | `not-found` — same message for unknown **and** restricted models ("does not exist or your team … does not have access") | `NOT_FOUND` — also "no longer available to new users" (2.5 models) and preview ids on `/v1` | No (model fallback may apply) | | 405/415 | — | — | wrong method (empty body) / wrong Content-Type | — | No | | 408 | timeout | — | — | — | Yes | | 409 | conflict | `conflict_error` | — | `ABORTED` (retry) vs `ALREADY_EXISTS` (don't) — without a status in the body assume `ALREADY_EXISTS` | per status | | 413 | — | `request_too_large` | — | — | No | | 416 | — | — | — | `OUT_OF_RANGE` | No | | 422 | `UnprocessableEntityError` | — | **bare-string serde error** ("text: invalid type: integer `123`, expected a string at line 1 column 33") | — | OpenAI once; xAI **No** (fix the field type) | | 429 | `rate_limit_error` (`slow_down`) with `Retry-After` | `rate_limit_error` **with** `retry-after` | `too-many-requests` / gRPC 8 RESOURCE_EXHAUSTED: team RPS/TPM or per-key qps/qpm/tpm; **no Retry-After documented or observed** | `RESOURCE_EXHAUSTED` — retry delay **only** in `details[]` `RetryInfo.retryDelay` ("40s") and/or the message tail ("Please retry in 54.22098241s."); **no Retry-After header** | **Yes** with backoff (xAI: pure exponential; Gemini: parsed delay treated as the minimum) | | 429 | `insufficient_quota`, spend-limit codes | `rate_limit_error` **without** `retry-after` (tier spend cap) | message mentioning credits/billing (prepaid exhausted, $0 invoiced limit — status undocumented) | every "Quota exceeded" line says **`limit: 0`** (free-tier key on a paid-only model such as `gemini-3.1-pro`, caching, Search grounding) | **No** — user action (billing / tier) required; `gemini_zero_quota_retryable=True` overrides | | 499 | — | — | — | `CANCELLED` (client closed) | n/a | | 500 | server error | `api_error` | `internal` (see https://status.x.ai) | `INTERNAL` (also oversized/unusual inputs → reduce context) | Yes | | 501 | — | — | — | `UNIMPLEMENTED` (PaLM methods, unsupported feature) | No | | 502/504 | gateway / timeout | `timeout_error` 504 | gRPC 4 DEADLINE_EXCEEDED → 504 (SDK default timeout 1620 s) | `DEADLINE_EXCEEDED` 504 (long thinking, Flex queueing → raise client timeout, stream, `background`/Batch) | Yes | | 503 | `service_unavailable_error` / `server_is_overloaded` | — | gRPC 14 UNAVAILABLE | `UNAVAILABLE` ("The model is overloaded. Please try again later.", Flex capacity shed) | Yes, honour `Retry-After` when present | | 529 | — | `overloaded_error` | — | — | Yes | | 200 | — | — | usage-guideline violation still billed ($0.05 on Responses) | `promptFeedback.blockReason` (no candidates) / `candidates[].finishReason` ≠ STOP (`SAFETY`, `RECITATION`, `MAX_TOKENS`, `MISSING_THOUGHT_SIGNATURE`…) | **No** — not transport errors; handle in the adapter (`stop_reason` refusal/max_tokens/other) | | 202 | — | — | deferred completion not ready (`GET /v1/chat/deferred-completion/{id}`) | — | poll with backoff | | network | `APIConnectionError`, `APITimeoutError` | same | `openai` SDK on `base_url=https://api.x.ai/v1` (same classes) / `grpc.RpcError` in xai-sdk | `google.genai.errors.APIError` (`code`, `details`) / connection errors | Yes (bounded) | Server hints that override everything: - `x-should-retry: false|true` (Anthropic) — applied whenever present. - `Retry-After` (OpenAI 429/503; Anthropic 429/529; **absent on xAI and Gemini**). Treat as a **minimum**, add jitter, cap it (`max_retry_after_s`, 60 s). - Gemini: `parse_gemini_retry_delay(body)` prefers `RetryInfo.retryDelay` (protobuf Duration string `"40s"`, or `{seconds, nanos}`), then the message regex `Please retry in (\d+(\.\d+)?)s`. Both were observed together in one body (40 s vs 54.22 s) — the structured value wins. - Anthropic in-stream `event: error` (`overloaded_error`) after HTTP 200 → classify like the status. Gemini `generateContent` SSE has no documented mid-stream error event (a `{"error": …}` chunk is still handled); a stream that closes without `finishReason` is a failure. Gemini Interactions streams emit `event: error`. ## 3. Backoff Full jitter: `sleep = U(0, min(cap, base·2^(attempt-1)))` (defaults base 0.5 s, cap 20 s), bounded by **both** `max_attempts` (4) and a wall-clock `max_total_s` (120 s). All four vendors recommend exponential backoff with jitter: OpenAI adds "limit both the number of attempts and the total time"; xAI's docs show `2**attempt` with 5 retries and point bulk work to the Batch API (exempt from rate limits); Gemini's troubleshooting page says 429/408/5xx only and its SDK defaults to 4 attempts, ~1 s initial, 60 s max. SDK interaction: `openai 3.16.2` and `anthropic 1.7.0` retry 2× by default (`DEFAULT_MAX_RETRIES = 2`); `google-genai 2.24` retries 429/5xx via `http_options.retry_options`; the `openai` SDK against `https://api.x.ai/v1` keeps OpenAI's defaults; `xai-sdk` (gRPC) surfaces `grpc.StatusCode` and the docs' sample loop catches `RESOURCE_EXHAUSTED`. When you own the policy, disable SDK retries (`max_retries=0`, `retry_options=None`) and feed exceptions to `sdk_error_is_retryable(exc, provider=…)` (pass `provider="xai"` explicitly for the OpenAI SDK-on-xAI case; gRPC errors are mapped by `e.code()`). ## 4. Timeouts and connection pooling - Per-request: connect 5 s, read 600 s (OpenAI/Anthropic SDK defaults). xAI reasoning models can be slow: the docs use **3600 s** client timeouts and `xai-sdk` defaults to **1620 s**; Gemini's SDK sends its `http_options.timeout` as `X-Server-Timeout` and long Flex/Deep-Research jobs should use `background: true` (Interactions) or Batch rather than a long socket. - Prefer **streaming** for long generations on all four providers (Anthropic's SDKs refuse non-streaming requests expected to exceed 10 min; Gemini 504 `DEADLINE_EXCEEDED` guidance says the same). - Python stdlib `urllib` (reference client) opens a connection per request; production: `http.client` keep-alive, `httpx.Client(limits=...)`, or the SDKs. Node `fetch` (undici) pools by default. xAI prompt caches are **per server**: keep a conversation on one host and send `x-grok-conv-id` (Chat) / `prompt_cache_key` (Responses) so retries land on the same cache. ## 5. Rate-limit headers (observed shape, parse defensively) | | OpenAI | Anthropic | xAI (undocumented, observed 2026-09-19) | Gemini | |---|---|---|---|---| | Requests | `x-ratelimit-limit-requests`, `-remaining-requests`, `-reset-requests` (`1s`, `6m0s`) | `anthropic-ratelimit-requests-limit/-remaining/-reset` (RFC 3339) | `x-ratelimit-limit-requests` / `-remaining-requests` (per-minute budget: 7200 grok-4.6/4.5, 1800 grok-4.3/4.20/build); **no reset header** | **none** | | Tokens | `x-ratelimit-limit-tokens`, `-remaining-tokens`, `-reset-tokens` (+ project-scoped) | `anthropic-ratelimit-tokens-*`, `input-tokens-*`, `output-tokens-*` | `x-ratelimit-limit-tokens` / `-remaining-tokens` (TPM = documented Tier 0: 50M / 10M); absent on `/v1/tokenize-text`, `/v1/models`, `/v1/api-key` | **none** — remaining quota lives only in AI Studio; limits are per **project**, per model (RPM/TPM/RPD) | | Retry hint | `Retry-After` | `retry-after` (absent on spend cap) | none | none (message text / `RetryInfo`) | | Request id | `x-request-id` | `request-id` | `x-request-id` (= `chat.completion.id`; absent on catalogue GETs) | none — use body `responseId` | | Extras | `openai-processing-ms` | `anthropic-organization-id` | `x-zero-data-retention: true\|false`, Cloudflare `Server-Timing`/`CF-RAY` | undocumented `X-Gemini-Service-Tier` (also on 429), `Server-Timing: gfet4t7` | `parse_rate_limit_headers()` returns a `RateLimitInfo`; for Gemini it is intentionally empty except `raw`. Never generalise observed values into "documented limits" (xAI's header values are below the documented Tier 0 RPS and may be a new-team allowance). ## 6. Idempotency — facts - **OpenAI:** `Idempotency-Key` documented on exactly one operation (`POST /v1/agents/sessions/{id}/events`). **Anthropic:** none. **xAI:** none documented on any endpoint; `POST /v1/files/{id}/public-url` is idempotent by behaviour (same URL returned). **Gemini:** none on generateContent/Interactions; `files.name` custom ids and File Search store names give create-once semantics (409 `ALREADY_EXISTS`). Webhook receivers dedupe on `webhook-id` (OpenAI, Anthropic, Gemini static webhooks). - Consequence: a retried `POST` after a timeout **may have succeeded server-side**. Mitigations: stream; OpenAI `background: true` + resume; xAI `store: true` + `GET /v1/responses/{id}` (live: retrievable even with `store:false`) or deferred completions (`deferred: true`, result fetchable once within 24 h); Gemini Interactions `background: true` + `GET /v1beta/interactions/{id}`; client-side dedupe keyed on an id you also place in `metadata` (OpenAI/Anthropic) / `labels` (Gemini) — xAI Responses rejects `metadata`, so use `safety_identifier`/`user` or `prompt_cache_key`. ## 7. Stream reconnection - **OpenAI Responses:** `background: true, stream: true` + `sequence_number` cursor → `GET /v1/responses/{id}?stream=true&starting_after=N`. - **xAI Responses:** same event vocabulary with `sequence_number`, **but `background` is rejected (400)** → no resume; restart the request (keep the cursor for logs only). WebSocket mode (`wss://api.x.ai/v1/responses`) is the documented alternative for long agentic runs. - **Anthropic Messages:** no resume → `anthropic_stream_with_restart` (synthetic `restart` marker; also on in-stream `error`). - **Gemini generateContent:** no resume; every chunk carries running `usageMetadata` and the last chunk `finishReason` — a close without it means restart. **Gemini Interactions:** resumable SSE `GET /v1beta/interactions/{id}?stream=true&last_event_id=…` for `background: true` interactions; **Live API:** `sessionResumption` handles (valid 2 h) after `goAway`/connection reset (~10 min). - All: keep the emitted text length to de-duplicate on resume/restart. ## 8. Circuit breaker Per provider (optionally per model): closed → open after 5 consecutive failures → half-open after 30 s → one probe. Open circuits raise `CircuitOpen` immediately so fallback chains skip a degraded provider. ## 9. Fallbacks `call_with_fallback([Target…])` tries targets in order; `Target.path` may contain `{model}` (Gemini: `/v1beta/models/{model}:generateContent`). - **Same provider, other model** on `RetryExhausted`, `CircuitOpen`, 404 `model_not_found` (xAI: 404 `not-found` for restricted models; Gemini: 404 "no longer available to new users"), transport errors, **and Gemini 429 `limit: 0`** (the model is paid-only for this key → try a Flash/Lite model). - **Other provider** — the request body must be rebuilt (Responses ≠ Messages ≠ generateContent); use the adapters. Only portable parameters survive; reasoning artefacts never do. - **Never** fall back on auth/validation errors: OpenAI/Anthropic `authentication_error`/`permission_error`/`invalid_request_error`, xAI `invalid-argument` (incl. incorrect key)/`unauthenticated*`, Gemini `INVALID_ARGUMENT`/`UNAUTHENTICATED`/`PERMISSION_DENIED`/`FAILED_PRECONDITION` (`AUTH_OR_VALIDATION_ERROR_TYPES`) — the same error will recur and you would mask a real bug. Tested chain: Gemini `limit: 0` → xAI (bad key stops the chain) / → Anthropic (succeeds). ## 10. Budget guard `BudgetGuard(max_usd, prices)` converts usage into USD from any of the four shapes: OpenAI/xAI Responses (`input_tokens`, `output_tokens`, `input_tokens_details.cached_tokens`), OpenAI/xAI Chat (`prompt_tokens`, `completion_tokens` **+ `completion_tokens_details.reasoning_tokens`**, `prompt_tokens_details.cached_tokens`), Anthropic (`cache_read_input_tokens`, `cache_creation_input_tokens`), Gemini `usageMetadata` (`promptTokenCount`, `candidatesTokenCount` + `thoughtsTokenCount`, `cachedContentTokenCount`). xAI also returns `usage.cost_in_usd_ticks` (1 tick = $1e-8) — the authoritative per-request cost when available. Prices come from `generated/pricing.json`; unknown models use a deliberately high default. ## 11. Checklist - [ ] Parse the envelope first (`{error:{}}`, `{code,error}`, bare string, google.rpc, array-wrapped); classify by status **and** type/code/message. - [ ] Never retry: quota/spend codes, Anthropic spend-cap 429, xAI 400 (incl. bad key) / 422, Gemini `FAILED_PRECONDITION` and `limit: 0` 429s. - [ ] Honour `Retry-After` and `x-should-retry`; for Gemini read `RetryInfo.retryDelay` / "Please retry in Ns"; cap the hint; add jitter; xAI gets pure backoff. - [ ] Bound retries by attempts **and** elapsed time; disable SDK retries when you own the loop; long client timeouts for reasoning models. - [ ] Stream long generations; resume only where supported (OpenAI background, Gemini Interactions `last_event_id`, Live `sessionResumption`); restart elsewhere. - [ ] Treat POST retries as non-idempotent: dedupe client-side (`metadata` / `labels` / your own store). - [ ] Log `x-request-id` (OpenAI, xAI) / `request-id` (Anthropic) / body `responseId` (Gemini) with every failure. - [ ] Circuit-break per provider; fall back only on capacity or model-availability errors. - [ ] Guard spend with a price-table-driven budget (xAI: cross-check `cost_in_usd_ticks`).