# xAI authentication, base URLs, headers and errors **Status:** `DOCUMENTED` + `LIVE_VERIFIED` (auth paths, 400/401/404/405/422 shapes, rate-limit headers and regional hosts probed on 2026-09-19; 403/429/5xx not triggered). **Sources:** https://docs.x.ai/developers/quickstart · /developers/debugging · /developers/rate-limits · /developers/advanced-api-usage/{regions,mtls,websocket-mode,deferred-chat-completions,prompt-caching/maximizing-cache-hits} · /developers/faq/security · /developers/management-api-guide · /developers/grpc-api-reference · `tmp-live/xai-models/catalogue-probe*.json`. **Last verified:** 2026-09-18 · Machine-readable: `generated/fragments/headers/xai-headers.json`, `generated/fragments/errors/xai-errors.json`. ## 1. Authentication | Credential | Where | Notes | |---|---|---| | API key `xai-…` | `Authorization: Bearer ` on every REST, WebSocket and gRPC call (gRPC: same header as call metadata) | created per team in the console (API Keys) or via the Management API; scoped by ACLs (`api-key:endpoint:*`, `api-key:model:*`); may carry per-key `qps`/`qpm`/`tpm` caps and `expireTime`; can be disabled/blocked; team can be blocked | | Management key | `Authorization: Bearer ` on `https://management-api.x.ai` | distinct key type; inference key → 401 `{code:16}` | | OAuth token | `GET /v1/me` documents "works with both API keys and OAuth tokens" (`oauth{client_id}` branch) | OAuth flow itself is not documented publicly | | mTLS | enterprise, team-level, host `https://mtls.api.x.ai` | client certificate **in addition to** the API key; rotation without downtime; `curl -v https://mtls.api.x.ai/v1/api-key` as smoke test | | Ephemeral client secrets | `POST /v1/realtime/client_secrets` | short-lived tokens for browser voice clients (voice agent's page) | There is **no version header and no beta header**. Introspect a key with `GET /v1/api-key` (acls, blocked/disabled flags) and `GET /v1/me` (user, team, `zdr_status`). ## 2. Base URLs | Host | Scope | Status | |---|---|---| | `https://api.x.ai/v1` | global; region not guaranteed | LIVE_VERIFIED | | `https://us.api.x.ai/v1` | US-only handling, inference, moderation, retained data; **grok-4.6 only**; tokens ×1.1; no image/video/voice; same keys | LIVE_VERIFIED (`GET /v1/models` → only grok-4.6 at 22000/5500/66000 ticks) | | `https://eu-west-1.api.x.ai/v1` | **undocumented**; `GET /v1/models` → 200 listing only grok-4.3 at global prices (model page lists an `eu-west-1` cluster for grok-4.3 and grok-voice-think-fast-2.0) | LIVE_DISCOVERED — do not rely on for compliance | | `https://mtls.api.x.ai/v1` | mTLS teams (global only) | DOCUMENTED | | `wss://api.x.ai/v1/responses` | WebSocket Responses mode | DOCUMENTED | | `wss://api.x.ai/v1/realtime`, `wss://api.x.ai/v1/stt`, `wss://api.x.ai/v1/tts` | voice | DOCUMENTED | | `api.x.ai:443` (gRPC) | services `xai_api.*` | DOCUMENTED | | `https://management-api.x.ai` | Management API | LIVE_VERIFIED (401 with inference key) | | clusters named in docs | us-east-1, us-west-2, us-central-1, eu-west-1, us-saltlake-2; propagation example `cloud9.api.x.ai`, `us-east-1.api.x.ai` | DOCUMENTED | Prompt caches do not carry over between hosts; keep a conversation on one host and set `prompt_cache_key` / `x-grok-conv-id`. ## 3. Request headers | Header | Required | Purpose | |---|---|---| | `Authorization: Bearer …` | yes | see §1 | | `Content-Type: application/json` (multipart for uploads/STT) | for bodies | wrong type → 415 | | `x-grok-conv-id: ` | recommended (Chat Completions) | routes the conversation to the same server so the automatic prompt cache hits; Responses API uses body `prompt_cache_key` | | `xai-sdk-version`, `xai-sdk-language` | auto (gRPC metadata from xai-sdk) | telemetry | ## 4. Response headers (observed 2026-09-19) | Header | Where seen | Meaning | |---|---|---| | `x-request-id` | most inference and catalogue-detail responses; equals `chat.completion.id` | quote in bug reports (support@x.ai) | | `x-ratelimit-limit-requests` / `x-ratelimit-remaining-requests` | chat completions, tokenize-text | per-minute request budget of the key for the model (7200 for grok-4.6/4.5, 1800 for grok-4.3/4.20/build) — undocumented | | `x-ratelimit-limit-tokens` / `x-ratelimit-remaining-tokens` | chat completions | TPM budget (50M / 10M = documented Tier 0) — undocumented | | `x-zero-data-retention: true\|false` | documented (FAQ) | ZDR state of the team | | `Server-Timing: cfEdge;dur=…,cfOrigin;dur=…`, `CF-RAY` | all | Cloudflare edge; origin duration ≈ model latency | | `Retry-After` | not documented, not observed | use exponential backoff on 429 | Not present: `x-ratelimit-reset-*`, version/organization headers. ## 5. Error format Inference API (REST): `{"code": "", "error": ""}` — e.g. `invalid-argument`, `not-found`, `unauthenticated:no-credentials`. Some validation failures return a **bare JSON string** (422 deserialization errors; the multi-agent-on-chat-completions 400). 405 returns an empty body. Management API: gRPC-style `{"code": 16, "message": "…", "details": []}`. gRPC: canonical status codes (`RESOURCE_EXHAUSTED` = rate limit). | HTTP | code (live) | Cause | Verified example | |---|---|---|---| | 400 | `invalid-argument` | bad body/URL, unsupported parameter, **incorrect API key** | `"Model grok-build-0.1 does not support parameter reasoningEffort."` · `"Incorrect API key provided. You can obtain an API key from https://console.x.ai."` · bare `"Multi Agent requests are not allowed on chat completions"` | | 401 | `unauthenticated:no-credentials` | no Authorization header (inference); invalid bearer (Management API, `{code:16}`) | `"No credentials presented."` | | 403 | — | missing ACL / disabled key / blocked team ("ask your team admin") | not triggered | | 404 | `not-found` | unknown route or model not available to the team (same message for both; regional host without the model) | `"The model grok-2-image does not exist or your team does not have access to it. If you believe this is a mistake, please contact support and quote your team ID and the model name."` | | 405 | (empty) | wrong method | `DELETE /v1/models` | | 415 | — | unsupported Content-Type | not triggered | | 422 | (string) | field type/format invalid | `"Failed to deserialize the JSON body into the target type: text: invalid type: integer `123`, expected a string at line 1 column 33"` | | 429 | (gRPC 8) | team RPS/TPM or per-key qps/qpm/tpm exceeded | not triggered; back off exponentially; Batch API bypasses limits | | 202 | — | deferred completion not ready (`GET /v1/chat/deferred-completion/{id}`, empty body) | — | | 5xx | — | check https://status.x.ai (RSS `/feed.xml`) | not triggered | Billing-side "errors": usage-guideline violations are still charged ($0.05 fee when caught before generation on Responses); exhausted prepaid credits with a $0 invoiced limit → requests rejected. ## 6. gRPC ↔ HTTP code mapping (for xai-sdk users) | gRPC | HTTP | SDK handling | |---|---|---| | 3 INVALID_ARGUMENT | 400 | `grpc.RpcError`, `e.code() == grpc.StatusCode.INVALID_ARGUMENT` | | 16 UNAUTHENTICATED | 401 | check key / host | | 7 PERMISSION_DENIED | 403 | ACLs | | 5 NOT_FOUND | 404 | model/route | | 8 RESOURCE_EXHAUSTED | 429 | docs' backoff example catches exactly this | | 4 DEADLINE_EXCEEDED | 504 | SDK default timeout 1620 s | | 13 INTERNAL / 14 UNAVAILABLE | 500 / 503 | retry |