xAI rate limits
Status: DOCUMENTED (tier tables) + LIVE_DISCOVERED (undocumented x-ratelimit-* response headers observed with our Tier 0 key on 2026-09-19). No 429 was triggered.
Sources: https://docs.x.ai/developers/rate-limits · /developers/models/ (per-model "Rate limits" tables = Tier 0 values) · /developers/management-api-guide (per-key qps/qpm/tpm) · /developers/debugging · /developers/advanced-api-usage/batch-api · tmp-live/xai-models/catalogue-probe.json.
Last verified: 2026-09-18 · Machine-readable: generated/fragments/rate-limits/xai-rate-limits.json.
1. Principles
- Limits are per team, per model, on two dimensions: RPS (requests per second, derived from a per-minute request budget: RPM/60, so a minute's budget cannot be spent in one second) and TPM (tokens per minute).
- Tier = cumulative spend on the xAI API since 2026-01-01 (prepaid purchases or fulfilled invoices). Automatic, permanent (never downgrades).
- TPM counts prompt tokens (text, image, audio), completion tokens, reasoning tokens and cached prompt tokens.
- Exceeding either → HTTP 429 (gRPC
RESOURCE_EXHAUSTED). Docs recommend exponential backoff (2**attempt, 5 retries). NoRetry-Afterdocumented. - Batch API requests do not count towards rate limits. Tiers apply to text, embedding and voice models; Imagine (image/video) limits are RPS-only and increased via sales@x.ai.
- Per-key caps (in addition to team limits):
qps,qpm,tpmset when creating/updating a key through the Management API (POST /auth/teams/{teamId}/api-keys,PUT /auth/api-keys/{id}withfieldMask). The tpm limiter engages when strictly exceeded and does not abort in-flight requests. - Personalised limits are on the console Models page (
https://console.x.ai/team/default/models,?cluster=us-central-1for the US endpoint).
2. Tiers
| Tier | Cumulative spend |
|---|---|
| Tier 0 | $0 (default) |
| Tier 1 | $50 |
| Tier 2 | $250 |
| Tier 3 | $1,000 |
| Tier 4 | $5,000 |
| Enterprise | on request |
3. Per-model limits (RPS / TPM by tier)
| Model | RPS T0 · T1 · T2 · T3 · T4 | TPM T0 · T1 · T2 · T3 · T4 |
|---|---|---|
| grok-4.6 | 150 · 172 · 208 · 312 · 500 | 50M · 53M · 60M · 74M · 100M |
| grok-4.5 | 150 · 172 · 208 · 312 · 500 | 50M · 53M · 60M · 74M · 100M |
| grok-4.3 | 37 · 50 · 75 · 125 · 208 | 10M · 15M · 25M · 45M · 85M |
| grok-4.20-0309-reasoning | 37 · 50 · 75 · 125 · 208 | 10M · 15M · 25M · 45M · 85M |
| grok-4.20-0309-non-reasoning | 37 · 50 · 75 · 125 · 208 | 10M · 15M · 25M · 45M · 85M |
| grok-build-0.1 | 37 · 50 · 75 · 125 · 208 | 10M · 15M · 25M · 45M · 85M |
| grok-4.20-multi-agent-0309 | 9 · 12 · 18 · 31 · 56 | 2.5M · 3.7M · 6.2M · 11M · 21M |
| grok-imagine-image, -quality, -2.0 | 6 · 12 · 25 · 50 · 100 | — |
| grok-imagine-video, -1.5 | 10 · 20 · 39 · 79 · 158 | — |
Voice & audio
| Service | RPS T0 · T4 | Concurrent sessions T0 · T1 · T2 · T3 · T4 | Other |
|---|---|---|---|
| grok-voice-think-fast-2.0 (speech-to-speech) | — | 10 · 20 · 50 · 100 · 200 | max session 120 min |
| Text to Speech | 50 · 50 · 100 · 250 · 500 | 100 · 200 · 200 · 300 · 500 | |
| Speech to Text | 10 · 10 · 20 · 30 · 40 | 100 · 200 · 200 · 300 · 500 (streaming) |
Embedding models: tiers apply, no table published.
4. Observed headers (our key, 2026-09-19 — account-specific, not documentation)
| Request | x-ratelimit-limit-requests |
x-ratelimit-limit-tokens |
remaining-* |
|---|---|---|---|
| POST /v1/chat/completions grok-4.6, grok-4.5 | 7200 | 50000000 | equal to limits after one call |
| POST /v1/chat/completions grok-4.3, grok-4.20-0309-non-reasoning, grok-build-0.1 | 1800 | 10000000 | idem |
| POST /v1/tokenize-text (grok-4.3) | 1800 | (absent) | |
| POST /v1/responses grok-4.20-multi-agent-0309 | (absent) | (absent) | |
| GET /v1/models, /v1/language-models, /v1/api-key, /v1/me | (absent) | (absent) |
Interpretation: x-ratelimit-limit-tokens equals the documented Tier 0 TPM. x-ratelimit-limit-requests looks like a per-minute request budget (7200 → 120/s, 1800 → 30/s), i.e. below the documented Tier 0 RPS (150, 37) — possibly a new-team allowance; the docs' "RPS = RPM/60" statement fits this header being the RPM. No x-ratelimit-reset-* header exists. Every inference response also carries x-request-id, Cloudflare Server-Timing and CF-RAY.
5. Increasing limits
Spend more (automatic), request an increase from the console Models page (also for limits beyond Tier 4), or contact sales@x.ai. For Imagine limits contact sales.
6. Related caps
max_session_minutes120 for speech-to-speech.- Deferred completion results are retrievable once within 24 h; batch completes typically within 24 h.
- Prepaid credits exhausted with a $0 invoiced-billing limit → requests rejected (status code not documented).