xAI authentication, base URLs, headers and errors
Status: DOCUMENTED + LIVE_VERIFIED (auth paths, 400/401/404/405/422 shapes, rate-limit headers and regional hosts probed on 2026-09-19; 403/429/5xx not triggered).
Sources: https://docs.x.ai/developers/quickstart · /developers/debugging · /developers/rate-limits · /developers/advanced-api-usage/{regions,mtls,websocket-mode,deferred-chat-completions,prompt-caching/maximizing-cache-hits} · /developers/faq/security · /developers/management-api-guide · /developers/grpc-api-reference · tmp-live/xai-models/catalogue-probe*.json.
Last verified: 2026-09-18 · Machine-readable: generated/fragments/headers/xai-headers.json, generated/fragments/errors/xai-errors.json.
1. Authentication
| Credential | Where | Notes |
|---|---|---|
API key xai-… |
Authorization: Bearer <XAI_API_KEY> on every REST, WebSocket and gRPC call (gRPC: same header as call metadata) |
created per team in the console (API Keys) or via the Management API; scoped by ACLs (api-key:endpoint:*, api-key:model:*); may carry per-key qps/qpm/tpm caps and expireTime; can be disabled/blocked; team can be blocked |
| Management key | Authorization: Bearer <management key> on https://management-api.x.ai |
distinct key type; inference key → 401 {code:16} |
| OAuth token | GET /v1/me documents "works with both API keys and OAuth tokens" (oauth{client_id} branch) |
OAuth flow itself is not documented publicly |
| mTLS | enterprise, team-level, host https://mtls.api.x.ai |
client certificate in addition to the API key; rotation without downtime; curl -v https://mtls.api.x.ai/v1/api-key as smoke test |
| Ephemeral client secrets | POST /v1/realtime/client_secrets |
short-lived tokens for browser voice clients (voice agent's page) |
There is no version header and no beta header. Introspect a key with GET /v1/api-key (acls, blocked/disabled flags) and GET /v1/me (user, team, zdr_status).
2. Base URLs
| Host | Scope | Status |
|---|---|---|
https://api.x.ai/v1 |
global; region not guaranteed | LIVE_VERIFIED |
https://us.api.x.ai/v1 |
US-only handling, inference, moderation, retained data; grok-4.6 only; tokens ×1.1; no image/video/voice; same keys | LIVE_VERIFIED (GET /v1/models → only grok-4.6 at 22000/5500/66000 ticks) |
https://eu-west-1.api.x.ai/v1 |
undocumented; GET /v1/models → 200 listing only grok-4.3 at global prices (model page lists an eu-west-1 cluster for grok-4.3 and grok-voice-think-fast-2.0) |
LIVE_DISCOVERED — do not rely on for compliance |
https://mtls.api.x.ai/v1 |
mTLS teams (global only) | DOCUMENTED |
wss://api.x.ai/v1/responses |
WebSocket Responses mode | DOCUMENTED |
wss://api.x.ai/v1/realtime, wss://api.x.ai/v1/stt, wss://api.x.ai/v1/tts |
voice | DOCUMENTED |
api.x.ai:443 (gRPC) |
services xai_api.* |
DOCUMENTED |
https://management-api.x.ai |
Management API | LIVE_VERIFIED (401 with inference key) |
| clusters named in docs | us-east-1, us-west-2, us-central-1, eu-west-1, us-saltlake-2; propagation example cloud9.api.x.ai, us-east-1.api.x.ai |
DOCUMENTED |
Prompt caches do not carry over between hosts; keep a conversation on one host and set prompt_cache_key / x-grok-conv-id.
3. Request headers
| Header | Required | Purpose |
|---|---|---|
Authorization: Bearer … |
yes | see §1 |
Content-Type: application/json (multipart for uploads/STT) |
for bodies | wrong type → 415 |
x-grok-conv-id: <conversation id> |
recommended (Chat Completions) | routes the conversation to the same server so the automatic prompt cache hits; Responses API uses body prompt_cache_key |
xai-sdk-version, xai-sdk-language |
auto (gRPC metadata from xai-sdk) | telemetry |
4. Response headers (observed 2026-09-19)
| Header | Where seen | Meaning |
|---|---|---|
x-request-id |
most inference and catalogue-detail responses; equals chat.completion.id |
quote in bug reports (support@x.ai) |
x-ratelimit-limit-requests / x-ratelimit-remaining-requests |
chat completions, tokenize-text | per-minute request budget of the key for the model (7200 for grok-4.6/4.5, 1800 for grok-4.3/4.20/build) — undocumented |
x-ratelimit-limit-tokens / x-ratelimit-remaining-tokens |
chat completions | TPM budget (50M / 10M = documented Tier 0) — undocumented |
x-zero-data-retention: true|false |
documented (FAQ) | ZDR state of the team |
Server-Timing: cfEdge;dur=…,cfOrigin;dur=…, CF-RAY |
all | Cloudflare edge; origin duration ≈ model latency |
Retry-After |
not documented, not observed | use exponential backoff on 429 |
Not present: x-ratelimit-reset-*, version/organization headers.
5. Error format
Inference API (REST): {"code": "<kebab-case code>", "error": "<message>"} — e.g. invalid-argument, not-found, unauthenticated:no-credentials. Some validation failures return a bare JSON string (422 deserialization errors; the multi-agent-on-chat-completions 400). 405 returns an empty body. Management API: gRPC-style {"code": 16, "message": "…", "details": []}. gRPC: canonical status codes (RESOURCE_EXHAUSTED = rate limit).
| HTTP | code (live) | Cause | Verified example |
|---|---|---|---|
| 400 | invalid-argument |
bad body/URL, unsupported parameter, incorrect API key | "Model grok-build-0.1 does not support parameter reasoningEffort." · "Incorrect API key provided. You can obtain an API key from https://console.x.ai." · bare "Multi Agent requests are not allowed on chat completions" |
| 401 | unauthenticated:no-credentials |
no Authorization header (inference); invalid bearer (Management API, {code:16}) |
"No credentials presented." |
| 403 | — | missing ACL / disabled key / blocked team ("ask your team admin") | not triggered |
| 404 | not-found |
unknown route or model not available to the team (same message for both; regional host without the model) | "The model grok-2-image does not exist or your team <id> does not have access to it. If you believe this is a mistake, please contact support and quote your team ID and the model name." |
| 405 | (empty) | wrong method | DELETE /v1/models |
| 415 | — | unsupported Content-Type | not triggered |
| 422 | (string) | field type/format invalid | "Failed to deserialize the JSON body into the target type: text: invalid type: integer 123, expected a string at line 1 column 33" |
| 429 | (gRPC 8) | team RPS/TPM or per-key qps/qpm/tpm exceeded | not triggered; back off exponentially; Batch API bypasses limits |
| 202 | — | deferred completion not ready (GET /v1/chat/deferred-completion/{id}, empty body) |
— |
| 5xx | — | check https://status.x.ai (RSS /feed.xml) |
not triggered |
Billing-side "errors": usage-guideline violations are still charged ($0.05 fee when caught before generation on Responses); exhausted prepaid credits with a $0 invoiced limit → requests rejected.
6. gRPC ↔ HTTP code mapping (for xai-sdk users)
| gRPC | HTTP | SDK handling |
|---|---|---|
| 3 INVALID_ARGUMENT | 400 | grpc.RpcError, e.code() == grpc.StatusCode.INVALID_ARGUMENT |
| 16 UNAUTHENTICATED | 401 | check key / host |
| 7 PERMISSION_DENIED | 403 | ACLs |
| 5 NOT_FOUND | 404 | model/route |
| 8 RESOURCE_EXHAUSTED | 429 | docs' backoff example catches exactly this |
| 4 DEADLINE_EXCEEDED | 504 | SDK default timeout 1620 s |
| 13 INTERNAL / 14 UNAVAILABLE | 500 / 503 | retry |