SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
7.2 KB

# xAI pricing

Status: DOCUMENTED (pricing page + per-model pages) and LIVE_VERIFIED for the per-token/per-image amounts, which are also returned by GET /v1/models / /v1/language-models in ticks (1 USD = 10^10 ticks) and match the pages exactly; every live call's usage.cost_in_usd_ticks was consistent. All prices USD.
Sources: https://docs.x.ai/developers/pricing · /developers/models · /developers/models/ · /developers/cost-tracking · /developers/advanced-api-usage/{priority-processing,regions,batch-api,prompt-caching} · /developers/tools/tool-usage-details · /developers/migration/* · sources/xai/language-models-raw.json.
Last verified: 2026-09-18 · Machine-readable: generated/fragments/pricing/xai-pricing.json (120 price records).

# 1. Text models (per 1M tokens)

Model Context Input Cached input Image input Output ≥ 200k prompt: input / cached / output Batch discount Priority
grok-4.6 500k 2.00 0.50 2.00 6.00 4.00 / 1.00 / 12.00 not supported ×2
grok-4.5 500k 2.00 0.30 2.00 6.00 4.00 / 0.60 / 12.00 not supported ×2
grok-4.3 1M 1.25 0.20 1.25 2.50 2.50 / 0.40 / 5.00 20 % ×2
grok-4.20-0309-reasoning 1M 1.25 0.20 1.25 2.50 2.50 / 0.40 / 5.00 20 % ×2
grok-4.20-0309-non-reasoning 1M 1.25 0.20 1.25 2.50 2.50 / 0.40 / 5.00 20 % ×2
grok-4.20-multi-agent-0309 1M 1.25 0.20 1.25 2.50 2.50 / 0.40 / 5.00 20 % (Responses only; priority not documented for it)
grok-build-0.1 256k 1.00 0.20 1.00 2.00 2.00 / 0.40 / 4.00 not supported ×2

Rules:

  • Tiered (long-context) pricing: "requests whose prompt reaches the listed token threshold [200,000] are billed at the higher rate for all tokens in the request". Live field long_context_threshold: 200000; when it is 0 the model has no long-context tier.
  • Reasoning tokens are billed as output tokens (completion_tokens_details.reasoning_tokens; total_tokens includes them). Reasoning cannot be disabled except on grok-4.3 (reasoning_effort: none).
  • Cached input: automatic prompt caching; prompt_tokens_details.cached_tokens billed at the cached rate (also counts toward TPM). Route with x-grok-conv-id (chat) / prompt_cache_key (Responses) for reliable hits.
  • Image input tokens are priced like text (prompt_image_token_price = prompt_text_token_price).
  • Batch API: 20 % off all token types (input, output, cached, reasoning) for grok-4.3 and the three grok-4.20 ids; other models either reject batch (4.6, 4.5, build) or have no discount. Image/video accepted in batch at standard rates.
  • Priority Processing: service_tier: "priority" → 2× on all token types, caching discount applied before the multiplier; billed only when the response says "service_tier": "priority" (we always got "default"). Chat Completions and Responses only.
  • US regional endpoint https://us.api.x.ai/v1: 1.1× on input, output and cached tokens incl. long-context; grok-4.6 only. Live catalogue there: 22000 / 5500 / 66000 ticks (= 2.20 / 0.55 / 6.60) and 44000 / 11000 / 132000 above 200k. The undocumented eu-west-1.api.x.ai host lists grok-4.3 at global prices.
  • Retired slugs (grok-3, grok-4-0709, grok-4-fast-, grok-4-1-fast-) are billed at grok-4.3 rates since 2026-05-15; grok-code-fast-1 at grok-build-0.1 rates.

# 2. Imagine (images and video)

Model Price Notes
grok-imagine-image $0.02 / image
grok-imagine-image-quality $0.05 / image retires 2026-11-02 → served as 2.0 low ($0.01 cheaper)
grok-imagine-image-2.0 $0.04 / image (default) quality × resolution matrix (live pricing[]): low 1k 0.04, 1.5k 0.05, 2k 0.06; medium 1k 0.06, 1.5k 0.07, 2k 0.08. quality: auto (default) bills at the quality served (low for generation, medium for editing). Docs schema note says "medium is the default quality" — superseded by the August release note (auto).
grok-imagine-video $0.050 / second edits/extensions billed per output second
grok-imagine-video-1.5 $0.080 / second

Batch: image and video generation supported (where the model allows) at standard rates; URLs in batch results expire after 1 h. The image_generation server-side tool bills at Imagine rates with no per-call fee.

# 3. Voice

Mode Price
Speech to Speech (grok-voice-think-fast-2.0) $0.08 / min ($4.80 / h) of audio sent or received + $0.004 per text conversation.item.create item (except function_call_output and audio items; response.create is free)
Speech to Text $0.10 / h REST (POST /v1/stt), $0.20 / h streaming (wss://api.x.ai/v1/stt)
Text to Speech $15.00 / 1M input characters

# 4. Server-side tools (per 1,000 successful calls)

Tool (type) Price Notes
web_search $5 includes image search (enable_image_search)
x_search $5 / 1k calls until 2026-09-21 12:00 PT; then $5 / 1k posts fetched + $10 / 1k user profiles fetched every post returned (incl. parent/quoted) counts
code_execution (code_interpreter) $5
attachment_search $10 files attached to messages
collections_search (file_search) $2.50
image_generation Imagine rates no call fee
view_image, view_x_video token-based only on search results only
remote MCP token-based only

Failed tool attempts are not charged (server_side_tool_usage counts billable calls; the Responses usage.num_server_side_tools_used, num_sources_used). cost_in_usd_ticks in every response already includes tool invocations.

# 5. Storage, downloads, fees

Item Price
File storage $0.025 / GiB / day
Collection storage $0.10 / GiB / day
File / collection downloads $0.20 / GiB
Usage-guideline violation caught before generation (Responses API) $0.05 / request (violating generations are still charged)
Management API, catalogue/account endpoints, tokenize-text no price documented; no usage returned (free in practice)

# 6. Cost tracking

usage.cost_in_usd_ticks (integer, 1 USD = 10,000,000,000 ticks) is returned by chat completions, Responses, image and video generation, and in the final streaming chunk (stream_options.include_usage: true for OpenAI-SDK/REST; running total per chunk in xai-sdk, response.cost_usd). It is the actual billed amount after caching discounts and includes tool calls. Batch objects expose cost_breakdown. Observed: grok-4.6 "Reply with OK." → 10,640,000 ticks = $0.001064 (640 prompt of which 512 cached, 1 completion, 91 reasoning: 128×2.00 + 512×0.50 + 92×6.00 per 1M = 0.000256 + 0.000256 + 0.000552 = $0.001064 ✔).

# 7. Billing model (console)

Prepaid credits (guest checkout, promo codes, auto top-up) or monthly invoiced billing (off by default, $0 invoiced limit → requests rejected when credits run out; raise the limit or contact sales). Usage is deducted immediately per request. Rate-limit tiers are unlocked by cumulative spend since 2026-01-01 ($50 / $250 / $1,000 / $5,000). Model access may vary by geography/account (console Models page). No refunds on prepaid credits unless required by law.