SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
17.3 KB

# Resilience patterns for OpenAI, Anthropic, xAI and Gemini clients

Status: DOCUMENTED (rules) · LIVE_VERIFIED (reference implementation exercised end-to-end via the four provider adapters: OpenAI/Anthropic 2026-09-18, xAI/Gemini 2026-09-19) · offline unit tests: tests/shared/test_resilient_client.py (46 tests incl. recorded xAI/Gemini error bodies) + resilientClient.ts --selftest Sources:

Implementation: examples/shared/resilient-client/resilient_client.py (stdlib) and resilientClient.ts (fetch). Both are transport-pluggable so the whole policy is unit-testable offline. ResilientClient(provider) for openai | anthropic | xai | gemini (env keys *_API_KEY, base URLs https://api.openai.com, https://api.anthropic.com, https://api.x.ai, https://generativelanguage.googleapis.com; auth Bearer / x-api-key / Bearer xai-… / x-goog-api-key).

# 1. Error envelopes (parse before classifying)

Provider Shape extract_error() → (type, code, message)
OpenAI / Anthropic {"error": {"type", "code", "message"}} (type, code, message)
xAI inference API {"code": "invalid-argument", "error": "<message>"} — kebab-case code; some 422/400 bodies are a bare JSON string (serde: "Failed to deserialize the JSON body into the target type: …"); 405 is an empty body (code, code, error) / (None, None, <string>)
xAI Management API gRPC-style {"code": 16, "message": "Invalid bearer token…", "details": []} ("grpc", "16", message)
Gemini google.rpc.Status {"error": {"code": 429, "message", "status": "RESOURCE_EXHAUSTED", "details": [{"@type": "…google.rpc.Help"}, {"@type": "…QuotaFailure", "violations": [{quotaMetric, quotaId, quotaDimensions{model, location}}]}, {"@type": "…RetryInfo", "retryDelay": "40s"}]}}; the OpenAI-compat layer may wrap it in a one-element array (status, "429", message)
Gemini Interactions API {"error": {"code": "<snake_case>", "message"}} in bodies, event: error in SSE streams (None, "<snake_case>", message)

# 2. Error classification (what to retry)

HTTP OpenAI (error.type / code) Anthropic (error.type) xAI (code) Gemini (error.status) Retry?
400 invalid_request_error invalid_request_error (also your own spend limit) invalid-argument — also the answer to an incorrect API key ("Incorrect API key provided…") and to unsupported params; bare-string body for multi-agent on chat INVALID_ARGUMENT (bad field, unknown name, invalid key "API key not valid"); FAILED_PRECONDITION (billing not enabled for a paid-only feature, region not supported) No — never retry a 400; on xAI check the message for "API key" before blaming the request
401 invalid key / IP not allowlisted authentication_error unauthenticated:no-credentials (header missing); Management API {code:16} UNAUTHENTICATED (OAuth/ephemeral token) No
402 — billing_error — — No
403 permission_error permission_error missing ACL (api-key:endpoint:* / api-key:model:*), disabled/blocked key or team PERMISSION_DENIED (key restricted to other APIs, leaked key auto-blocked, foreign File/tuned model) No
404 not_found_error / model_not_found not_found_error not-found — same message for unknown and restricted models ("does not exist or your team … does not have access") NOT_FOUND — also "no longer available to new users" (2.5 models) and preview ids on /v1 No (model fallback may apply)
405/415 — — wrong method (empty body) / wrong Content-Type — No
408 timeout — — — Yes
409 conflict conflict_error — ABORTED (retry) vs ALREADY_EXISTS (don't) — without a status in the body assume ALREADY_EXISTS per status
413 — request_too_large — — No
416 — — — OUT_OF_RANGE No
422 UnprocessableEntityError — bare-string serde error ("text: invalid type: integer 123, expected a string at line 1 column 33") — OpenAI once; xAI No (fix the field type)
429 rate_limit_error (slow_down) with Retry-After rate_limit_error with retry-after too-many-requests / gRPC 8 RESOURCE_EXHAUSTED: team RPS/TPM or per-key qps/qpm/tpm; no Retry-After documented or observed RESOURCE_EXHAUSTED — retry delay only in details[] RetryInfo.retryDelay ("40s") and/or the message tail ("Please retry in 54.22098241s."); no Retry-After header Yes with backoff (xAI: pure exponential; Gemini: parsed delay treated as the minimum)
429 insufficient_quota, spend-limit codes rate_limit_error without retry-after (tier spend cap) message mentioning credits/billing (prepaid exhausted, $0 invoiced limit — status undocumented) every "Quota exceeded" line says limit: 0 (free-tier key on a paid-only model such as gemini-3.1-pro, caching, Search grounding) No — user action (billing / tier) required; gemini_zero_quota_retryable=True overrides
499 — — — CANCELLED (client closed) n/a
500 server error api_error internal (see https://status.x.ai) INTERNAL (also oversized/unusual inputs → reduce context) Yes
501 — — — UNIMPLEMENTED (PaLM methods, unsupported feature) No
502/504 gateway / timeout timeout_error 504 gRPC 4 DEADLINE_EXCEEDED → 504 (SDK default timeout 1620 s) DEADLINE_EXCEEDED 504 (long thinking, Flex queueing → raise client timeout, stream, background/Batch) Yes
503 service_unavailable_error / server_is_overloaded — gRPC 14 UNAVAILABLE UNAVAILABLE ("The model is overloaded. Please try again later.", Flex capacity shed) Yes, honour Retry-After when present
529 — overloaded_error — — Yes
200 — — usage-guideline violation still billed ($0.05 on Responses) promptFeedback.blockReason (no candidates) / candidates[].finishReason ≠ STOP (SAFETY, RECITATION, MAX_TOKENS, MISSING_THOUGHT_SIGNATURE…) No — not transport errors; handle in the adapter (stop_reason refusal/max_tokens/other)
202 — — deferred completion not ready (GET /v1/chat/deferred-completion/{id}) — poll with backoff
network APIConnectionError, APITimeoutError same openai SDK on base_url=https://api.x.ai/v1 (same classes) / grpc.RpcError in xai-sdk google.genai.errors.APIError (code, details) / connection errors Yes (bounded)

Server hints that override everything:

  • x-should-retry: false|true (Anthropic) — applied whenever present.
  • Retry-After (OpenAI 429/503; Anthropic 429/529; absent on xAI and Gemini). Treat as a minimum, add jitter, cap it (max_retry_after_s, 60 s).
  • Gemini: parse_gemini_retry_delay(body) prefers RetryInfo.retryDelay (protobuf Duration string "40s", or {seconds, nanos}), then the message regex Please retry in (\d+(\.\d+)?)s. Both were observed together in one body (40 s vs 54.22 s) — the structured value wins.
  • Anthropic in-stream event: error (overloaded_error) after HTTP 200 → classify like the status. Gemini generateContent SSE has no documented mid-stream error event (a {"error": …} chunk is still handled); a stream that closes without finishReason is a failure. Gemini Interactions streams emit event: error.

# 3. Backoff

Full jitter: sleep = U(0, min(cap, base·2^(attempt-1))) (defaults base 0.5 s, cap 20 s), bounded by both max_attempts (4) and a wall-clock max_total_s (120 s). All four vendors recommend exponential backoff with jitter: OpenAI adds "limit both the number of attempts and the total time"; xAI's docs show 2**attempt with 5 retries and point bulk work to the Batch API (exempt from rate limits); Gemini's troubleshooting page says 429/408/5xx only and its SDK defaults to 4 attempts, ~1 s initial, 60 s max.

SDK interaction: openai 3.16.2 and anthropic 1.7.0 retry 2× by default (DEFAULT_MAX_RETRIES = 2); google-genai 2.24 retries 429/5xx via http_options.retry_options; the openai SDK against https://api.x.ai/v1 keeps OpenAI's defaults; xai-sdk (gRPC) surfaces grpc.StatusCode and the docs' sample loop catches RESOURCE_EXHAUSTED. When you own the policy, disable SDK retries (max_retries=0, retry_options=None) and feed exceptions to sdk_error_is_retryable(exc, provider=…) (pass provider="xai" explicitly for the OpenAI SDK-on-xAI case; gRPC errors are mapped by e.code()).

# 4. Timeouts and connection pooling

  • Per-request: connect 5 s, read 600 s (OpenAI/Anthropic SDK defaults). xAI reasoning models can be slow: the docs use 3600 s client timeouts and xai-sdk defaults to 1620 s; Gemini's SDK sends its http_options.timeout as X-Server-Timeout and long Flex/Deep-Research jobs should use background: true (Interactions) or Batch rather than a long socket.
  • Prefer streaming for long generations on all four providers (Anthropic's SDKs refuse non-streaming requests expected to exceed 10 min; Gemini 504 DEADLINE_EXCEEDED guidance says the same).
  • Python stdlib urllib (reference client) opens a connection per request; production: http.client keep-alive, httpx.Client(limits=...), or the SDKs. Node fetch (undici) pools by default. xAI prompt caches are per server: keep a conversation on one host and send x-grok-conv-id (Chat) / prompt_cache_key (Responses) so retries land on the same cache.

# 5. Rate-limit headers (observed shape, parse defensively)

OpenAI Anthropic xAI (undocumented, observed 2026-09-19) Gemini
Requests x-ratelimit-limit-requests, -remaining-requests, -reset-requests (1s, 6m0s) anthropic-ratelimit-requests-limit/-remaining/-reset (RFC 3339) x-ratelimit-limit-requests / -remaining-requests (per-minute budget: 7200 grok-4.6/4.5, 1800 grok-4.3/4.20/build); no reset header none
Tokens x-ratelimit-limit-tokens, -remaining-tokens, -reset-tokens (+ project-scoped) anthropic-ratelimit-tokens-*, input-tokens-*, output-tokens-* x-ratelimit-limit-tokens / -remaining-tokens (TPM = documented Tier 0: 50M / 10M); absent on /v1/tokenize-text, /v1/models, /v1/api-key none — remaining quota lives only in AI Studio; limits are per project, per model (RPM/TPM/RPD)
Retry hint Retry-After retry-after (absent on spend cap) none none (message text / RetryInfo)
Request id x-request-id request-id x-request-id (= chat.completion.id; absent on catalogue GETs) none — use body responseId
Extras openai-processing-ms anthropic-organization-id x-zero-data-retention: true|false, Cloudflare Server-Timing/CF-RAY undocumented X-Gemini-Service-Tier (also on 429), Server-Timing: gfet4t7

parse_rate_limit_headers() returns a RateLimitInfo; for Gemini it is intentionally empty except raw. Never generalise observed values into "documented limits" (xAI's header values are below the documented Tier 0 RPS and may be a new-team allowance).

# 6. Idempotency — facts

  • OpenAI: Idempotency-Key documented on exactly one operation (POST /v1/agents/sessions/{id}/events). Anthropic: none. xAI: none documented on any endpoint; POST /v1/files/{id}/public-url is idempotent by behaviour (same URL returned). Gemini: none on generateContent/Interactions; files.name custom ids and File Search store names give create-once semantics (409 ALREADY_EXISTS). Webhook receivers dedupe on webhook-id (OpenAI, Anthropic, Gemini static webhooks).
  • Consequence: a retried POST after a timeout may have succeeded server-side. Mitigations: stream; OpenAI background: true + resume; xAI store: true + GET /v1/responses/{id} (live: retrievable even with store:false) or deferred completions (deferred: true, result fetchable once within 24 h); Gemini Interactions background: true + GET /v1beta/interactions/{id}; client-side dedupe keyed on an id you also place in metadata (OpenAI/Anthropic) / labels (Gemini) — xAI Responses rejects metadata, so use safety_identifier/user or prompt_cache_key.

# 7. Stream reconnection

  • OpenAI Responses: background: true, stream: true + sequence_number cursor → GET /v1/responses/{id}?stream=true&starting_after=N.
  • xAI Responses: same event vocabulary with sequence_number, but background is rejected (400) → no resume; restart the request (keep the cursor for logs only). WebSocket mode (wss://api.x.ai/v1/responses) is the documented alternative for long agentic runs.
  • Anthropic Messages: no resume → anthropic_stream_with_restart (synthetic restart marker; also on in-stream error).
  • Gemini generateContent: no resume; every chunk carries running usageMetadata and the last chunk finishReason — a close without it means restart. Gemini Interactions: resumable SSE GET /v1beta/interactions/{id}?stream=true&last_event_id=… for background: true interactions; Live API: sessionResumption handles (valid 2 h) after goAway/connection reset (~10 min).
  • All: keep the emitted text length to de-duplicate on resume/restart.

# 8. Circuit breaker

Per provider (optionally per model): closed → open after 5 consecutive failures → half-open after 30 s → one probe. Open circuits raise CircuitOpen immediately so fallback chains skip a degraded provider.

# 9. Fallbacks

call_with_fallback([Target…]) tries targets in order; Target.path may contain {model} (Gemini: /v1beta/models/{model}:generateContent).

  • Same provider, other model on RetryExhausted, CircuitOpen, 404 model_not_found (xAI: 404 not-found for restricted models; Gemini: 404 "no longer available to new users"), transport errors, and Gemini 429 limit: 0 (the model is paid-only for this key → try a Flash/Lite model).
  • Other provider — the request body must be rebuilt (Responses ≠ Messages ≠ generateContent); use the adapters. Only portable parameters survive; reasoning artefacts never do.
  • Never fall back on auth/validation errors: OpenAI/Anthropic authentication_error/permission_error/invalid_request_error, xAI invalid-argument (incl. incorrect key)/unauthenticated*, Gemini INVALID_ARGUMENT/UNAUTHENTICATED/PERMISSION_DENIED/FAILED_PRECONDITION (AUTH_OR_VALIDATION_ERROR_TYPES) — the same error will recur and you would mask a real bug. Tested chain: Gemini limit: 0 → xAI (bad key stops the chain) / → Anthropic (succeeds).

# 10. Budget guard

BudgetGuard(max_usd, prices) converts usage into USD from any of the four shapes: OpenAI/xAI Responses (input_tokens, output_tokens, input_tokens_details.cached_tokens), OpenAI/xAI Chat (prompt_tokens, completion_tokens + completion_tokens_details.reasoning_tokens, prompt_tokens_details.cached_tokens), Anthropic (cache_read_input_tokens, cache_creation_input_tokens), Gemini usageMetadata (promptTokenCount, candidatesTokenCount + thoughtsTokenCount, cachedContentTokenCount). xAI also returns usage.cost_in_usd_ticks (1 tick = $1e-8) — the authoritative per-request cost when available. Prices come from generated/pricing.json; unknown models use a deliberately high default.

# 11. Checklist

  • Parse the envelope first ({error:{}}, {code,error}, bare string, google.rpc, array-wrapped); classify by status and type/code/message.
  • Never retry: quota/spend codes, Anthropic spend-cap 429, xAI 400 (incl. bad key) / 422, Gemini FAILED_PRECONDITION and limit: 0 429s.
  • Honour Retry-After and x-should-retry; for Gemini read RetryInfo.retryDelay / "Please retry in Ns"; cap the hint; add jitter; xAI gets pure backoff.
  • Bound retries by attempts and elapsed time; disable SDK retries when you own the loop; long client timeouts for reasoning models.
  • Stream long generations; resume only where supported (OpenAI background, Gemini Interactions last_event_id, Live sessionResumption); restart elsewhere.
  • Treat POST retries as non-idempotent: dedupe client-side (metadata / labels / your own store).
  • Log x-request-id (OpenAI, xAI) / request-id (Anthropic) / body responseId (Gemini) with every failure.
  • Circuit-break per provider; fall back only on capacity or model-availability errors.
  • Guard spend with a price-table-driven budget (xAI: cross-check cost_in_usd_ticks).