xAI deferred chat completions — deferred: true + GET /v1/chat/deferred-completion/{request_id}
Status: DOCUMENTED + LIVE_VERIFIED (2026-09-19: start → 202 → 200 in ~2 s; second fetch also 200).
Sources: https://docs.x.ai/developers/advanced-api-usage/deferred-chat-completions · https://docs.x.ai/developers/rest-api-reference/inference/chat-completions#get-deferred-chat-completions · OpenAPI StartDeferredChatResponse
Last verified: 2026-09-19
Flow
POST /v1/chat/completionswith"deferred": true(any other chat parameters allowed;streamis meaningless) → 200{"request_id":"3ee7c57a-…"}immediately.GET /v1/chat/deferred-completion/{request_id}→ 202 Accepted with empty body while pending; 200 ChatCompletion (same body as a normal completion, incl.reasoning_content,usage) when ready; 404 for unknown ids.- Docs: result retrievable exactly once within 24 h, then discarded. Live: a second GET a few seconds later returned 200 again (not strictly once).
Rate limit = same as chat completions. Only via REST or xai-sdk (chat.defer(timeout=timedelta(minutes=10), interval=timedelta(seconds=10))); no equivalent on /v1/responses (background:true → 400). Poll with backoff (docs use exponential 1–60 s).
Examples: examples/xai/deferred/ (sh, py). Test: tests/xai/test_chat.py::test_deferred_completion_roundtrip.