# xAI pricing **Status:** `DOCUMENTED` (pricing page + per-model pages) and `LIVE_VERIFIED` for the per-token/per-image amounts, which are also returned by `GET /v1/models` / `/v1/language-models` in ticks (1 USD = 10^10 ticks) and match the pages exactly; every live call's `usage.cost_in_usd_ticks` was consistent. All prices USD. **Sources:** https://docs.x.ai/developers/pricing · /developers/models · /developers/models/ · /developers/cost-tracking · /developers/advanced-api-usage/{priority-processing,regions,batch-api,prompt-caching} · /developers/tools/tool-usage-details · /developers/migration/* · `sources/xai/language-models-raw.json`. **Last verified:** 2026-09-18 · Machine-readable: `generated/fragments/pricing/xai-pricing.json` (120 price records). ## 1. Text models (per 1M tokens) | Model | Context | Input | Cached input | Image input | Output | ≥ 200k prompt: input / cached / output | Batch discount | Priority | |---|---|---|---|---|---|---|---|---| | grok-4.6 | 500k | 2.00 | 0.50 | 2.00 | 6.00 | 4.00 / 1.00 / 12.00 | not supported | ×2 | | grok-4.5 | 500k | 2.00 | 0.30 | 2.00 | 6.00 | 4.00 / 0.60 / 12.00 | not supported | ×2 | | grok-4.3 | 1M | 1.25 | 0.20 | 1.25 | 2.50 | 2.50 / 0.40 / 5.00 | 20 % | ×2 | | grok-4.20-0309-reasoning | 1M | 1.25 | 0.20 | 1.25 | 2.50 | 2.50 / 0.40 / 5.00 | 20 % | ×2 | | grok-4.20-0309-non-reasoning | 1M | 1.25 | 0.20 | 1.25 | 2.50 | 2.50 / 0.40 / 5.00 | 20 % | ×2 | | grok-4.20-multi-agent-0309 | 1M | 1.25 | 0.20 | 1.25 | 2.50 | 2.50 / 0.40 / 5.00 | 20 % | (Responses only; priority not documented for it) | | grok-build-0.1 | 256k | 1.00 | 0.20 | 1.00 | 2.00 | 2.00 / 0.40 / 4.00 | not supported | ×2 | Rules: - **Tiered (long-context) pricing**: "requests whose prompt reaches the listed token threshold [200,000] are billed at the higher rate for **all** tokens in the request". Live field `long_context_threshold: 200000`; when it is `0` the model has no long-context tier. - **Reasoning tokens** are billed as output tokens (`completion_tokens_details.reasoning_tokens`; `total_tokens` includes them). Reasoning cannot be disabled except on grok-4.3 (`reasoning_effort: none`). - **Cached input**: automatic prompt caching; `prompt_tokens_details.cached_tokens` billed at the cached rate (also counts toward TPM). Route with `x-grok-conv-id` (chat) / `prompt_cache_key` (Responses) for reliable hits. - **Image input tokens** are priced like text (`prompt_image_token_price` = `prompt_text_token_price`). - **Batch API**: 20 % off all token types (input, output, cached, reasoning) for grok-4.3 and the three grok-4.20 ids; other models either reject batch (4.6, 4.5, build) or have no discount. Image/video accepted in batch at standard rates. - **Priority Processing**: `service_tier: "priority"` → 2× on all token types, caching discount applied before the multiplier; billed only when the response says `"service_tier": "priority"` (we always got `"default"`). Chat Completions and Responses only. - **US regional endpoint** `https://us.api.x.ai/v1`: 1.1× on input, output and cached tokens incl. long-context; grok-4.6 only. Live catalogue there: 22000 / 5500 / 66000 ticks (= 2.20 / 0.55 / 6.60) and 44000 / 11000 / 132000 above 200k. The undocumented `eu-west-1.api.x.ai` host lists grok-4.3 at global prices. - **Retired slugs** (grok-3, grok-4-0709, grok-4-fast-*, grok-4-1-fast-*) are billed at grok-4.3 rates since 2026-05-15; `grok-code-fast-1` at grok-build-0.1 rates. ## 2. Imagine (images and video) | Model | Price | Notes | |---|---|---| | grok-imagine-image | $0.02 / image | | | grok-imagine-image-quality | $0.05 / image | retires 2026-11-02 → served as 2.0 `low` ($0.01 cheaper) | | grok-imagine-image-2.0 | $0.04 / image (default) | quality × resolution matrix (live `pricing[]`): low 1k **0.04**, 1.5k 0.05, 2k 0.06; medium 1k 0.06, 1.5k 0.07, 2k 0.08. `quality: auto` (default) bills at the quality served (low for generation, medium for editing). Docs schema note says "medium is the default quality" — superseded by the August release note (auto). | | grok-imagine-video | $0.050 / second | edits/extensions billed per output second | | grok-imagine-video-1.5 | $0.080 / second | | Batch: image and video generation supported (where the model allows) at standard rates; URLs in batch results expire after 1 h. The `image_generation` server-side tool bills at Imagine rates with no per-call fee. ## 3. Voice | Mode | Price | |---|---| | Speech to Speech (`grok-voice-think-fast-2.0`) | $0.08 / min ($4.80 / h) of audio sent or received + $0.004 per text `conversation.item.create` item (except `function_call_output` and audio items; `response.create` is free) | | Speech to Text | $0.10 / h REST (`POST /v1/stt`), $0.20 / h streaming (`wss://api.x.ai/v1/stt`) | | Text to Speech | $15.00 / 1M input characters | ## 4. Server-side tools (per 1,000 **successful** calls) | Tool (`type`) | Price | Notes | |---|---|---| | `web_search` | $5 | includes image search (`enable_image_search`) | | `x_search` | $5 / 1k calls **until 2026-09-21 12:00 PT**; then **$5 / 1k posts fetched + $10 / 1k user profiles fetched** | every post returned (incl. parent/quoted) counts | | `code_execution` (`code_interpreter`) | $5 | | | `attachment_search` | $10 | files attached to messages | | `collections_search` (`file_search`) | $2.50 | | | `image_generation` | Imagine rates | no call fee | | `view_image`, `view_x_video` | token-based only | on search results only | | remote MCP | token-based only | | Failed tool attempts are not charged (`server_side_tool_usage` counts billable calls; the Responses `usage.num_server_side_tools_used`, `num_sources_used`). `cost_in_usd_ticks` in every response already includes tool invocations. ## 5. Storage, downloads, fees | Item | Price | |---|---| | File storage | $0.025 / GiB / day | | Collection storage | $0.10 / GiB / day | | File / collection downloads | $0.20 / GiB | | Usage-guideline violation caught before generation (Responses API) | $0.05 / request (violating generations are still charged) | | Management API, catalogue/account endpoints, `tokenize-text` | no price documented; no `usage` returned (free in practice) | ## 6. Cost tracking `usage.cost_in_usd_ticks` (integer, 1 USD = 10,000,000,000 ticks) is returned by chat completions, Responses, image and video generation, and in the final streaming chunk (`stream_options.include_usage: true` for OpenAI-SDK/REST; running total per chunk in xai-sdk, `response.cost_usd`). It is the actual billed amount after caching discounts and includes tool calls. Batch objects expose `cost_breakdown`. Observed: grok-4.6 "Reply with OK." → 10,640,000 ticks = $0.001064 (640 prompt of which 512 cached, 1 completion, 91 reasoning: 128×2.00 + 512×0.50 + 92×6.00 per 1M = 0.000256 + 0.000256 + 0.000552 = $0.001064 ✔). ## 7. Billing model (console) Prepaid credits (guest checkout, promo codes, auto top-up) or monthly invoiced billing (off by default, $0 invoiced limit → requests rejected when credits run out; raise the limit or contact sales). Usage is deducted immediately per request. Rate-limit tiers are unlocked by cumulative spend since 2026-01-01 ($50 / $250 / $1,000 / $5,000). Model access may vary by geography/account (console Models page). No refunds on prepaid credits unless required by law.