xAI pricing
Status: DOCUMENTED (pricing page + per-model pages) and LIVE_VERIFIED for the per-token/per-image amounts, which are also returned by GET /v1/models / /v1/language-models in ticks (1 USD = 10^10 ticks) and match the pages exactly; every live call's usage.cost_in_usd_ticks was consistent. All prices USD.
Sources: https://docs.x.ai/developers/pricing · /developers/models · /developers/models/ · /developers/cost-tracking · /developers/advanced-api-usage/{priority-processing,regions,batch-api,prompt-caching} · /developers/tools/tool-usage-details · /developers/migration/* · sources/xai/language-models-raw.json.
Last verified: 2026-09-18 · Machine-readable: generated/fragments/pricing/xai-pricing.json (120 price records).
1. Text models (per 1M tokens)
| Model | Context | Input | Cached input | Image input | Output | ≥ 200k prompt: input / cached / output | Batch discount | Priority |
|---|---|---|---|---|---|---|---|---|
| grok-4.6 | 500k | 2.00 | 0.50 | 2.00 | 6.00 | 4.00 / 1.00 / 12.00 | not supported | ×2 |
| grok-4.5 | 500k | 2.00 | 0.30 | 2.00 | 6.00 | 4.00 / 0.60 / 12.00 | not supported | ×2 |
| grok-4.3 | 1M | 1.25 | 0.20 | 1.25 | 2.50 | 2.50 / 0.40 / 5.00 | 20 % | ×2 |
| grok-4.20-0309-reasoning | 1M | 1.25 | 0.20 | 1.25 | 2.50 | 2.50 / 0.40 / 5.00 | 20 % | ×2 |
| grok-4.20-0309-non-reasoning | 1M | 1.25 | 0.20 | 1.25 | 2.50 | 2.50 / 0.40 / 5.00 | 20 % | ×2 |
| grok-4.20-multi-agent-0309 | 1M | 1.25 | 0.20 | 1.25 | 2.50 | 2.50 / 0.40 / 5.00 | 20 % | (Responses only; priority not documented for it) |
| grok-build-0.1 | 256k | 1.00 | 0.20 | 1.00 | 2.00 | 2.00 / 0.40 / 4.00 | not supported | ×2 |
Rules:
- Tiered (long-context) pricing: "requests whose prompt reaches the listed token threshold [200,000] are billed at the higher rate for all tokens in the request". Live field
long_context_threshold: 200000; when it is0the model has no long-context tier. - Reasoning tokens are billed as output tokens (
completion_tokens_details.reasoning_tokens;total_tokensincludes them). Reasoning cannot be disabled except on grok-4.3 (reasoning_effort: none). - Cached input: automatic prompt caching;
prompt_tokens_details.cached_tokensbilled at the cached rate (also counts toward TPM). Route withx-grok-conv-id(chat) /prompt_cache_key(Responses) for reliable hits. - Image input tokens are priced like text (
prompt_image_token_price=prompt_text_token_price). - Batch API: 20 % off all token types (input, output, cached, reasoning) for grok-4.3 and the three grok-4.20 ids; other models either reject batch (4.6, 4.5, build) or have no discount. Image/video accepted in batch at standard rates.
- Priority Processing:
service_tier: "priority"→ 2× on all token types, caching discount applied before the multiplier; billed only when the response says"service_tier": "priority"(we always got"default"). Chat Completions and Responses only. - US regional endpoint
https://us.api.x.ai/v1: 1.1× on input, output and cached tokens incl. long-context; grok-4.6 only. Live catalogue there: 22000 / 5500 / 66000 ticks (= 2.20 / 0.55 / 6.60) and 44000 / 11000 / 132000 above 200k. The undocumentedeu-west-1.api.x.aihost lists grok-4.3 at global prices. - Retired slugs (grok-3, grok-4-0709, grok-4-fast-, grok-4-1-fast-) are billed at grok-4.3 rates since 2026-05-15;
grok-code-fast-1at grok-build-0.1 rates.
2. Imagine (images and video)
| Model | Price | Notes |
|---|---|---|
| grok-imagine-image | $0.02 / image | |
| grok-imagine-image-quality | $0.05 / image | retires 2026-11-02 → served as 2.0 low ($0.01 cheaper) |
| grok-imagine-image-2.0 | $0.04 / image (default) | quality × resolution matrix (live pricing[]): low 1k 0.04, 1.5k 0.05, 2k 0.06; medium 1k 0.06, 1.5k 0.07, 2k 0.08. quality: auto (default) bills at the quality served (low for generation, medium for editing). Docs schema note says "medium is the default quality" — superseded by the August release note (auto). |
| grok-imagine-video | $0.050 / second | edits/extensions billed per output second |
| grok-imagine-video-1.5 | $0.080 / second |
Batch: image and video generation supported (where the model allows) at standard rates; URLs in batch results expire after 1 h. The image_generation server-side tool bills at Imagine rates with no per-call fee.
3. Voice
| Mode | Price |
|---|---|
Speech to Speech (grok-voice-think-fast-2.0) |
$0.08 / min ($4.80 / h) of audio sent or received + $0.004 per text conversation.item.create item (except function_call_output and audio items; response.create is free) |
| Speech to Text | $0.10 / h REST (POST /v1/stt), $0.20 / h streaming (wss://api.x.ai/v1/stt) |
| Text to Speech | $15.00 / 1M input characters |
4. Server-side tools (per 1,000 successful calls)
Tool (type) |
Price | Notes |
|---|---|---|
web_search |
$5 | includes image search (enable_image_search) |
x_search |
$5 / 1k calls until 2026-09-21 12:00 PT; then $5 / 1k posts fetched + $10 / 1k user profiles fetched | every post returned (incl. parent/quoted) counts |
code_execution (code_interpreter) |
$5 | |
attachment_search |
$10 | files attached to messages |
collections_search (file_search) |
$2.50 | |
image_generation |
Imagine rates | no call fee |
view_image, view_x_video |
token-based only | on search results only |
| remote MCP | token-based only |
Failed tool attempts are not charged (server_side_tool_usage counts billable calls; the Responses usage.num_server_side_tools_used, num_sources_used). cost_in_usd_ticks in every response already includes tool invocations.
5. Storage, downloads, fees
| Item | Price |
|---|---|
| File storage | $0.025 / GiB / day |
| Collection storage | $0.10 / GiB / day |
| File / collection downloads | $0.20 / GiB |
| Usage-guideline violation caught before generation (Responses API) | $0.05 / request (violating generations are still charged) |
Management API, catalogue/account endpoints, tokenize-text |
no price documented; no usage returned (free in practice) |
6. Cost tracking
usage.cost_in_usd_ticks (integer, 1 USD = 10,000,000,000 ticks) is returned by chat completions, Responses, image and video generation, and in the final streaming chunk (stream_options.include_usage: true for OpenAI-SDK/REST; running total per chunk in xai-sdk, response.cost_usd). It is the actual billed amount after caching discounts and includes tool calls. Batch objects expose cost_breakdown. Observed: grok-4.6 "Reply with OK." → 10,640,000 ticks = $0.001064 (640 prompt of which 512 cached, 1 completion, 91 reasoning: 128×2.00 + 512×0.50 + 92×6.00 per 1M = 0.000256 + 0.000256 + 0.000552 = $0.001064 ✔).
7. Billing model (console)
Prepaid credits (guest checkout, promo codes, auto top-up) or monthly invoiced billing (off by default, $0 invoiced limit → requests rejected when credits run out; raise the limit or contact sales). Usage is deducted immediately per request. Rate-limit tiers are unlocked by cumulative spend since 2026-01-01 ($50 / $250 / $1,000 / $5,000). Model access may vary by geography/account (console Models page). No refunds on prepaid credits unless required by law.