SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
46.7 KB

# Pricing — side-by-side and cost models (OpenAI · Anthropic · xAI · Gemini)

Status: every number below is read from generated/models.json (pricing block of each model, sourced from the vendor pricing/model pages on 2026-09-18) and generated/pricing.json (tool/service/rule rows). Figures are USD per 1M tokens unless stated. Statuses next to model ids are the model record statuses. No live billing was reconciled; treat as list prices. Gemini 3.6–3.8 Flash rows are introductory prices through 2026-12-31 (the records carry a from_2027_01_01 block at 2×); xAI x_search switches from per-call to per-post/per-profile billing on 2026-09-21. Sources: https://developers.openai.com/api/docs/pricing · https://platform.claude.com/docs/en/about-claude/pricing · https://docs.x.ai/developers/pricing (+ live GET /v1/language-models price ticks) · https://ai.google.dev/gemini-api/docs/pricing · docs/openai/pricing.md · docs/anthropic/pricing.md · docs/xai/pricing.md · docs/gemini/pricing.md Last verified: 2026-09-18

# 1. Pricing dimensions — how the four price lists are structured

Dimension OpenAI Anthropic xAI Gemini
Base tokens input, cached_input, output per model; separate *_long_context table for >272k-token prompts (2× input, 1.5× output) on GPT-5.4/5.5/5.6/6 Astra input, output; 1M context at standard price on Claude 4.6+ (no long-context premium) input, cached_input, image_input (= text input price), output; long-context tier at 2× for prompts ≥ 200,000 tokens — billed on ALL tokens of that request (live_raw_ticks.long_context_threshold); reasoning tokens billed as output on every Grok call input, cached_input, output per model; modality-specific input rows (audio_input, image_input, video_input); >200k tier (input_over_200k 2×, output_over_200k 1.5×) on Pro models only; thinking tokens billed as output (output_includes_thinking_tokens)
Cache write free on models before GPT-5.6; 1.25× input on GPT-5.6+ (cache_write row) 1.25× input for 5-minute TTL, 2× input for 1-hour TTL none — automatic prefix cache, no write charge implicit cache: none; explicit cachedContents: storage per 1M tokens per hour (cache_storage_hour: $0.50 Flash 3.6–3.8, $1.00 most Flash, $4.50 Pro; 1.8× on priority)
Cache read model-specific cached_input (0.1× on GPT-5.6+, 0.25× on o3, 0.5× on gpt-4o…) 0.1× input (0.025× on Fable 5.1 / Mythos 5.1) cached_input per model: 0.25× (grok-4.6 $0.50), 0.15× (grok-4.5 $0.30), 0.16× (grok-4.3 / 4.20 $0.20), 0.2× (grok-build $0.20) cached_input = 0.1× input on every priced model (e.g. gemini-3.8-flash $0.075, gemini-3.1-pro-preview $0.20)
Batch batch tier = 50 % of standard (own cached_input) batch_input/batch_output = 50 % of standard; cache multipliers stack 20 % off (batch_discount: 0.2) on grok-4.3 / 4.20 / 4.20-multi-agent only; not supported on grok-4.6, grok-4.5, grok-build-0.1 (batch_discount: null); Imagine models accepted in Batch but billed at standard rates batch tier = 50 % of standard (service:batch_api multiplier 0.5) on text, image, TTS and embedding models; not on the free tier
Cheaper async / best-effort tier flex = batch rates on synchronous calls (slower, may 429) none none (Batch API is the only discount) flex = 50 % of standard (service:flex_inference), 1–15 min target latency, sheddable (429 when capacity is shed); ten text models listed
Faster / priority tier fast = 2× standard (renamed from priority 2026-07-30; fast_long_context 4×/3×) speed: fast (Opus 5 / 4.8 only) = 2× standard ($10 / $50) service_tier: priority = 2× on all token types (caching discount applied before the multiplier; billed only when the response echoes service_tier: priority) priority tier = 1.8× standard (service:priority_inference, docs: 75–100 % more); 0.3× the standard rate limit; graceful downgrade to standard when exceeded
Reasoning tokens billed as output (usage.output_tokens_details.reasoning_tokens) billed as output (usage.output_tokens_details.thinking_tokens) billed as output (usage.completion_tokens_details.reasoning_tokens); every Grok reasoning model spends them even on reasoning_effort: low billed as output (usageMetadata.thoughtsTokenCount); Gemini 3.x thinking cannot be fully disabled on Flash/Pro (thinkingLevel: minimal where supported)
Regional inference +10 % on regional hosts for models released ≥ 2026-03-05 inference_geo: us ×1.1 on all token dimensions (Claude 4.6+); Bedrock/Vertex regional +10 % https://us.api.x.ai/v1 = 1.1× global token rates (grok-4.6 only in the live catalogue: $2.20 / $0.55 / $6.60); eu-west-1.api.x.ai LIVE_DISCOVERED, no separate price row no regional price row on the Gemini Developer API (data residency is a Vertex AI feature); AI Studio usage free
Free tier none (Free usage tier = rate-limit tier, still billed) none none (prepaid credits; Tier 0 = $0 spend) yes — free_tier rows on Flash / Flash-Lite / Live / TTS / transcribe / embedding / Gemma models: input & output $0, restrictive limits, content may be used to improve Google products; Pro models and paid-only media models are not available on the free tier (observed limit: 0)
Tool-definition overhead none published (definitions are input tokens) published per model: tool-use system prompt 286–804 tokens, toolsets 4,500–6,600 tokens none published (definitions are input tokens; server-side tool results count as input) none published; toolUsePromptTokenCount reports tokens consumed by URL context / grounding results
Per-image / per-second media image models per token (≈ per image), Sora per second — Imagine per image ($0.02 / $0.04–$0.08 / $0.05) and per second of video ($0.05 / $0.08); voice per minute; TTS per 1M characters; STT per hour image models per token with a per-image equivalent ($0.034–$0.24), Veo per second ($0.05–$0.60), Lyria per song ($0.04 / $0.08), Live/TTS/transcribe per 1M audio tokens (≈ per minute)

# 2. OpenAI current text models (USD / 1M tokens)

Model Status Input Cached in Cache write Output Batch in / out Flex in / out Fast in / out Long-context in / out (>272k)
gpt-6-astra DOCUMENTED · LIVE_VERIFIED 10 1 12.5 50 5 / 25 5 / 25 20 / 100 20 / 75
gpt-5.6-sol DOCUMENTED · LIVE_VERIFIED 4 0.4 5 20 2 / 10 2 / 10 8 / 40 8 / 30
gpt-5.6-terra DOCUMENTED · LIVE_VERIFIED 2 0.2 2.5 12 1 / 6 1 / 6 4 / 24 4 / 18
gpt-5.6-luna DOCUMENTED · LIVE_VERIFIED 0.2 0.02 0.25 1.2 0.1 / 0.6 0.1 / 0.6 0.4 / 2.4 0.4 / 1.8
gpt-5.5 DOCUMENTED · LIVE_VERIFIED 5 0.5 — 30 2.5 / 15 2.5 / 15 12.5 / 75 10 / 45
gpt-5.5-pro DOCUMENTED · LIVE_VERIFIED 30 — — 180 15 / 90 15 / 90 — / — 60 / 270
gpt-5.4 DOCUMENTED · LIVE_VERIFIED 2.5 0.25 — 15 1.25 / 7.5 1.25 / 7.5 5 / 30 5 / 22.5
gpt-5.4-pro DOCUMENTED · LIVE_VERIFIED 30 — — 180 15 / 90 15 / 90 — / — 60 / 270
gpt-5.4-mini DOCUMENTED · LIVE_VERIFIED 0.75 0.075 — 4.5 0.375 / 2.25 0.375 / 2.25 1.5 / 9 — / —
gpt-5.4-nano DOCUMENTED · LIVE_VERIFIED 0.2 0.02 — 1.25 0.1 / 0.625 0.1 / 0.625 — / — — / —
gpt-5.3-codex DOCUMENTED · LIVE_VERIFIED 1.75 0.175 — 14 — / — — / — 3.5 / 28 — / —
gpt-5.2 DOCUMENTED · LIVE_VERIFIED 1.75 0.175 — 14 0.875 / 7 0.875 / 7 3.5 / 28 — / —
gpt-5.2-pro DOCUMENTED · LIVE_VERIFIED 21 — — 168 10.5 / 84 — / — — / — — / —
gpt-5.1 DOCUMENTED · LIVE_VERIFIED 1.25 0.125 — 10 0.625 / 5 0.625 / 5 2.5 / 20 — / —
gpt-5 DOCUMENTED · LIVE_VERIFIED 1.25 0.125 — 10 0.625 / 5 0.625 / 5 2.5 / 20 — / —
gpt-5-mini DOCUMENTED · LIVE_VERIFIED 0.25 0.025 — 2 0.125 / 1 0.125 / 1 0.45 / 3.6 — / —
gpt-5-nano DOCUMENTED · LIVE_VERIFIED 0.05 0.005 — 0.4 0.025 / 0.2 0.025 / 0.2 — / — — / —
gpt-5-pro DOCUMENTED · LIVE_VERIFIED 15 — — 120 7.5 / 60 — / — — / — — / —
o3 DOCUMENTED · LIVE_VERIFIED 2 0.5 — 8 1 / 4 1 / 4 3.5 / 14 — / —
o3-pro DOCUMENTED · LIVE_VERIFIED 20 — — 80 10 / 40 — / — — / — — / —
o4-mini DOCUMENTED · LIVE_VERIFIED · DEPRECATED 1.1 0.275 — 4.4 0.55 / 2.2 0.55 / 2.2 2 / 8 — / —
o3-mini DOCUMENTED · LIVE_VERIFIED · DEPRECATED 1.1 0.55 — 4.4 0.55 / 2.2 — / — — / — — / —
gpt-4.1 DOCUMENTED · LIVE_VERIFIED 2 0.5 — 8 1 / 4 — / — 3.5 / 14 — / —
gpt-4.1-mini DOCUMENTED · LIVE_VERIFIED 0.4 0.1 — 1.6 0.2 / 0.8 — / — 0.7 / 2.8 — / —
gpt-4.1-nano DOCUMENTED · LIVE_VERIFIED · DEPRECATED 0.1 0.025 — 0.4 0.05 / 0.2 — / — 0.2 / 0.8 — / —
gpt-4o DOCUMENTED · LIVE_VERIFIED 2.5 1.25 — 10 1.25 / 5 — / — 4.25 / 17 — / —
gpt-4o-mini DOCUMENTED · LIVE_VERIFIED 0.15 0.075 — 0.6 0.075 / 0.3 — / — 0.25 / 1 — / —
chat-latest DOCUMENTED · LIVE_VERIFIED 5 0.5 — 30 — / — — / — — / — — / —
gpt-5.6-cyber DOCUMENTED · ACCOUNT_RESTRICTED 12.5 1.25 15.625 75 — / — — / — — / — — / —

Notes from the records: gpt-5.6 is an alias of gpt-5.6-sol (promotional $4/$20 at least through 2026-11-21); pro models have no cached-input price (no caching); chat-latest has no batch/flex/fast tiers; gpt-5.6-cyber is ACCOUNT_RESTRICTED and priced at short-context rates. The o3 record carries a model-page table ($1 / $0.25 / $4) that disagrees with its pricing-page standard row ($2 / $0.5 / $8) — flagged as a data inconsistency.

# 3. Anthropic current models (USD / 1M tokens)

Model Status Input Cache write 5m Cache write 1h Cache read Output Batch in / out Fast in / out Min cacheable tokens
claude-fable-5-1 DOCUMENTED · LIVE_VERIFIED 10 12.5 20 0.25 50 5 / 25 — / — 512
claude-fable-5 DOCUMENTED · LIVE_VERIFIED · LEGACY 10 12.5 20 1 50 5 / 25 — / — 512
claude-mythos-5-1 DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW 10 12.5 20 0.25 50 5 / 25 — / — 512
claude-mythos-5 DOCUMENTED · ACCOUNT_RESTRICTED · PREVIEW 10 12.5 20 1 50 5 / 25 — / — 512
claude-opus-5 DOCUMENTED · LIVE_VERIFIED 5 6.25 10 0.5 25 2.5 / 12.5 10 / 50 512
claude-opus-4-8 DOCUMENTED · LIVE_VERIFIED · LEGACY 5 6.25 10 0.5 25 2.5 / 12.5 10 / 50 1024
claude-opus-4-7 DOCUMENTED · LIVE_VERIFIED · LEGACY 5 6.25 10 0.5 25 2.5 / 12.5 — / — 2048
claude-opus-4-6 DOCUMENTED · LIVE_VERIFIED · LEGACY 5 6.25 10 0.5 25 2.5 / 12.5 — / — 4096
claude-opus-4-5-20251101 DOCUMENTED · LIVE_VERIFIED · LEGACY 5 6.25 10 0.5 25 2.5 / 12.5 — / — 4096
claude-sonnet-5 DOCUMENTED · LIVE_VERIFIED 2 2.5 4 0.2 10 1 / 5 — / — 1024
claude-sonnet-4-6 DOCUMENTED · LIVE_VERIFIED · LEGACY 3 3.75 6 0.3 15 1.5 / 7.5 — / — 1024
claude-sonnet-4-5-20250929 DOCUMENTED · LIVE_VERIFIED · LEGACY 3 3.75 6 0.3 15 1.5 / 7.5 — / — 1024
claude-haiku-4-5-20251001 DOCUMENTED · LIVE_VERIFIED 1 1.25 2 0.1 5 0.5 / 2.5 — / — 4096

Notes: Sonnet 5's introductory $2 / $10 was made permanent on 2026-08-10. Fable 5.1 / Mythos 5.1 cache reads are 0.025× ($0.25). Retired models (Opus 4.1 $15/$75, Sonnet 4 $3/$15, Haiku 3.5 $0.80/$4) remain priced on Bedrock/Vertex only. Priority Tier is no longer sold; Claude Platform on AWS and Foundry convert the same USD rates to CCUs at $0.01.

# 4. xAI current Grok text models (USD / 1M tokens)

Standard = prompt < 200,000 tokens. Long context = prompt ≥ 200,000 tokens: the higher rate applies to every token of that request (input, cached, output). image_input equals the text input price on every model. Batch = 20 % off standard where supported. Priority = service_tier: priority, 2×. US regional = https://us.api.x.ai/v1, 1.1×.

Model Status Input Cached in Output Long-ctx in / cached / out (≥200k) Batch in / cached / out Priority × US regional × Context
grok-4.6 DOCUMENTED · LIVE_VERIFIED 2 0.5 6 4 / 1 / 12 not supported 2.0× 1.1× 500000
grok-4.5 DOCUMENTED · LIVE_VERIFIED 2 0.3 6 4 / 0.6 / 12 not supported 2.0× — 500000
grok-4.3 DOCUMENTED · LIVE_VERIFIED 1.25 0.2 2.5 2.5 / 0.4 / 5 1 / 0.16 / 2 (−20 %) 2.0× — 1000000
grok-4.20-0309-reasoning DOCUMENTED · LIVE_VERIFIED 1.25 0.2 2.5 2.5 / 0.4 / 5 1 / 0.16 / 2 (−20 %) 2.0× — 1000000
grok-4.20-0309-non-reasoning DOCUMENTED · LIVE_VERIFIED 1.25 0.2 2.5 2.5 / 0.4 / 5 1 / 0.16 / 2 (−20 %) 2.0× — 1000000
grok-4.20-multi-agent-0309 DOCUMENTED · BETA · LIVE_VERIFIED 1.25 0.2 2.5 2.5 / 0.4 / 5 1 / 0.16 / 2 (−20 %) 2.0 (Chat Completions/Responses; not documented per model) — 1000000
grok-build-0.1 DOCUMENTED · PREVIEW · LIVE_VERIFIED 1 0.2 2 2 / 0.4 / 4 not supported 2.0× — 256000

Notes from the records: grok-4.20-multi-agent-0309 is BETA and only reachable on /v1/responses; its reasoning.effort selects the agent count (low/medium = 4, high/xhigh = 16) and all agents' tokens are billed. grok-build-0.1 (PREVIEW) is the coding model behind the grok-code-fast-1 redirect. The retired ids grok-3, grok-4-0709, grok-4-fast-*, grok-4-1-fast-* redirect to grok-4.3 and are billed at grok-4.3 rates. grok-embedding-small has no published price (ACCOUNT_RESTRICTED). Max output is not documented per model (max_output: null in every Grok record; Responses max_output_tokens defaults to 128,000).

# 4b. xAI media and voice prices

Model / service Status Price Unit / notes
grok-imagine-image DOCUMENTED · LIVE_VERIFIED $0.02 per image single price; batch discount 0.0 (standard rates in Batch)
grok-imagine-image-2.0 DOCUMENTED · LIVE_VERIFIED $0.06 per image low/1k $0.04; low/2k $0.06; low/1.5k $0.05; medium/1k $0.06; medium/2k $0.08; medium/1.5k $0.07; batch discount 0.0 (standard rates in Batch)
grok-imagine-image-quality DOCUMENTED · DEPRECATED · LIVE_VERIFIED $0.05 per image single price; batch discount 0.0 (standard rates in Batch)
grok-imagine-video DOCUMENTED · LIVE_VERIFIED $0.05 per second of generated video video URLs in batch results expire after 1 hour
grok-imagine-video-1.5 DOCUMENTED · LIVE_VERIFIED $0.08 per second of generated video video URLs in batch results expire after 1 hour
grok-voice-think-fast-2.0 (speech-to-speech) DOCUMENTED $0.08 per minute ($4.8 / h) + $0.004 per text item audio sent or received billed per minute; each conversation.item.create text item billed $0.004 except function_call_output and audio items
grok-voice-transcribe-2.0 (speech-to-text) DOCUMENTED $0.1 / hour REST (POST /v1/stt), $0.2 / hour streaming (wss://api.x.ai/v1/stt) per hour of audio
text-to-speech (POST /v1/tts, wss://api.x.ai/v1/tts) DOCUMENTED $15 per 1M characters POST /v1/tts and streaming

# 5. Gemini current text models (USD / 1M tokens)

Standard tier for prompts ≤ 200k tokens; >200k columns only exist on Pro models (Flash models have one price for any prompt length). Cached in = implicit or explicit cache hit (0.1×). Storage = explicit cachedContents storage per 1M tokens per hour. Batch = 50 %, Flex = 50 %, Priority = 1.8×. Free = free-tier availability recorded on the pricing page.

Model Status Input Cached in Output >200k in / cached / out Batch in / out Flex in / out Priority in / out Cache storage $/1M/h Free tier
gemini-3.8-flash DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED 0.75 0.075 3.75 — (flat) 0.375 / 1.875 0.375 / 1.875 1.35 / 6.75 0.5 input/output/caching free of charge (standard & priority rows); Bat…
gemini-3.7-flash DOCUMENTED · LIVE_DISCOVERED 0.75 0.075 3.75 — (flat) 0.375 / 1.875 0.375 / 1.875 1.35 / 6.75 0.5 input/output/caching free of charge (standard & priority rows); Bat…
gemini-3.6-flash DOCUMENTED · LIVE_DISCOVERED 0.75 0.075 3.75 — (flat) 0.375 / 1.875 0.375 / 1.875 1.35 / 6.75 0.5 input/output/caching free of charge (standard & priority rows); Bat…
gemini-3.5-flash DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED 1.5 0.15 9 — (flat) 0.75 / 4.5 0.75 / 4.5 2.7 / 16.2 1 input/output/caching free of charge; Batch/Flex not available
gemini-3.5-flash-lite DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED 0.3 0.03 2.5 — (flat) 0.15 / 1.25 0.15 / 1.25 0.54 / 4.5 1 input/output free of charge on all four rows (page lists Batch/Flex…
gemini-3.1-pro-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW · ACCOUNT_RESTRICTED 2 0.2 12 4 / 0.4 / 18 1 / 6 1 / 6 3.6 / 21.6 4.5 Not available (paid tier only)
gemini-3.1-pro-preview-customtools DOCUMENTED · LIVE_DISCOVERED · PREVIEW 2 0.2 12 4 / 0.4 / 18 1 / 6 1 / 6 3.6 / 21.6 4.5 Not available (paid tier only)
gemini-3.1-flash-lite DOCUMENTED · LIVE_DISCOVERED · DEPRECATED 0.25 0.025 1.5 — (flat) 0.125 / 0.75 0.125 / 0.75 0.45 / 2.7 1 input/output free of charge; caching not available on free tier
gemini-3-flash-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW 0.5 0.05 3 — (flat) 0.25 / 1.5 0.25 / 1.5 0.9 / 5.4 1 input/output/caching free of charge; Batch/Flex not available
gemini-2.5-pro DOCUMENTED · LIVE_DISCOVERED 1.25 0.125 10 2.5 / 0.25 / 15 0.625 / 5 0.625 / 5 2.25 / 18 4.5 input/output free of charge (standard/priority rows); caching, Batc…
gemini-2.5-flash DOCUMENTED · LIVE_DISCOVERED 0.3 0.03 2.5 — (flat) 0.15 / 1.25 0.15 / 1.25 0.54 / 4.5 1 input/output free of charge; caching not available; Google Search g…
gemini-2.5-flash-lite DOCUMENTED · LIVE_DISCOVERED · ACCOUNT_RESTRICTED 0.1 0.01 0.4 — (flat) 0.05 / 0.2 0.05 / 0.2 0.18 / 0.72 1 input/output free of charge; caching not available; Google Search g…
gemini-robotics-er-2-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW 1 0.1 5 — (flat) 0.5 / 2.5 — / — — / — 0.5 input/output free of charge; caching not available
gemini-2.5-computer-use-preview-10-2025 DOCUMENTED · LIVE_DISCOVERED · PREVIEW 1 — 5 — (flat) — / — — / — — / — — input/output free of charge
gemma-4-31b-it DOCUMENTED · LIVE_DISCOVERED free — free — — — — — free of charge (input, output, caching, storage)
gemma-4-26b-a4b-it DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED free — free — — — — — free of charge (input, output, caching, storage)

Notes from the records: gemini-3.8 / 3.7 / 3.6 Flash share one introductory price list ($0.75 / $0.075 / $3.75) through 2026-12-31, doubling on 2027-01-01 ($1.50 / $0.15 / $7.50, storage $1.00); gemini-3.5-flash is the higher-priced Flash ($1.50 / $9.00). Pro models (gemini-3.1-pro-preview, gemini-2.5-pro) are paid-tier only (ACCOUNT_RESTRICTED on this free-tier key). Gemini 2.5 models are 'no longer available to new users' (404 on generateContent). Gemma 4 models are free of charge with no paid tier. gemini-2.5-flash / -pro have separate audio_input rows ($1.00 / 1M); Gemini 3.x prices text, image, video and audio input alike. Google Search grounding: Gemini 3.x 5,000 free requests/month then $14 / 1k search queries; Gemini 2.5: 1,500 RPD free then $35 / 1k grounded prompts.

# 5b. Gemini media, voice and embedding prices

Model Status Price Notes
gemini-3.1-flash-image DOCUMENTED · LIVE_DISCOVERED text in $0.5, text out $3, image out $60 / 1M → per image 0.5K (747 tok) $0.045, 1K (1120 tok) $0.067, 2K (1680 tok) $0.101, 4K (2520 tok) $0.151; batch image out $30 not available
gemini-3.1-flash-lite-image DOCUMENTED · LIVE_DISCOVERED text in $0.25, image out $30 / 1M → 1K (1120 tok) $0.0336 not available
gemini-3-pro-image DOCUMENTED · LIVE_DISCOVERED text in $2, image in $0.0011 per image, image out $120 / 1M → 1K/2K (1120 tok) $0.134, 4K (2000 tok) $0.24; priority image out $216 not available
gemini-2.5-flash-image DOCUMENTED · LIVE_DISCOVERED · DEPRECATED $0.039 per image ($30 / 1M, 1,290 tok/image); batch $0.0195 1290 tokens per image up to 1024x1024
veo-3.1-generate-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW per second: 720p $0.4, 1080p $0.4, 4k $0.6 charged only if the video is successfully generated
veo-3.1-fast-generate-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW per second: 720p $0.1, 1080p $0.12, 4k $0.3 not available
veo-3.1-lite-generate-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW · LIVE_VERIFIED per second: 720p $0.05, 1080p $0.08, 4k not supported not available
gemini-omni-1.1-flash DOCUMENTED · LIVE_DISCOVERED text in $1.5, text out $9, video out $17.5 / 1M 5,792 tokens per second of 720p video => ~$0.10 per second
lyria-3.5 DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED $0.08 per song (full length) not available
lyria-3-clip-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW $0.04 per song (30 s clip) not available
gemini-3.1-flash-tts-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW text in $1, audio out $20 / 1M (25 audio tok/s ≈ $0.03 / min); batch $0.5 / $10 25 audio tokens per second
gemini-2.5-flash-preview-tts DOCUMENTED · LIVE_DISCOVERED · PREVIEW text in $0.5, audio out $10 / 1M free of charge (standard)
gemini-3.8-live DOCUMENTED · LIVE_DISCOVERED text in $0.75, audio in $3 (≈ $0.005 / min), image/video in $1, text out $4.5, audio out $12 (≈ $0.018 / min) free of charge
gemini-2.5-flash-native-audio-preview-12-2025 DOCUMENTED · LIVE_DISCOVERED · PREVIEW text in $0.5, audio/video in $3, text out $2, audio out $12 free of charge
gemini-3.5-transcribe DOCUMENTED · LIVE_DISCOVERED audio in $2 (≈ $0.003 / min), text out $12 (≈ $0.002 / min) blended ~$0.005/min
gemini-3.5-transcribe-live DOCUMENTED · LIVE_DISCOVERED audio in $3.5 (≈ $0.005 / min), text out $21 25 audio tokens/s input, 175 text tokens/min output; blended ~$0.009/min
gemini-3.5-live-translate-preview DOCUMENTED · LIVE_DISCOVERED · PREVIEW audio in $3.5, audio out $21 / 1M 25 audio tokens/second; effective ~$0.0368 per minute
gemini-embedding-2 DOCUMENTED · LIVE_DISCOVERED · LIVE_VERIFIED text $0.2, image $0.45 ($0.0001 per image), audio $6.5 ($0.0002 / s), video $12 ($0.0008 / frame) / 1M; batch text $0.1 free of charge (standard); batch not available
gemini-embedding-001 DOCUMENTED · LIVE_DISCOVERED · DEPRECATED — not listed on pricing.md (2026-09-18); file-search.md bills indexing embeddings at $0.15 per 1M tokens

# 6. Cost model A — 1M input tokens + 100k output tokens, no cache

Formula: cost = 1.0 × input_price + 0.1 × output_price, treating the 1M input tokens as aggregate volume at the standard (short-prompt) rate. The last column shows what a single request whose prompt crosses the long-context threshold costs instead (OpenAI >272k, xAI ≥200k applied to all tokens, Gemini Pro >200k; Anthropic and Gemini Flash have no threshold). Tiers shown where the model record lists them.

Provider Model Standard Batch Flex Fast / Priority Long-context request
openai gpt-6-astra $15.00 $7.50 $7.50 $30.00 $27.50
openai gpt-5.6-sol $6.00 $3.00 $3.00 $12.00 $11.00
openai gpt-5.6-terra $3.20 $1.60 $1.60 $6.40 $5.80
openai gpt-5.6-luna $0.32 $0.16 $0.16 $0.64 $0.58
openai gpt-5.5 $8.00 $4.00 $4.00 $20.00 $14.50
openai gpt-5.5-pro $48.00 $24.00 $24.00 — $87.00
openai gpt-5.4 $4.00 $2.00 $2.00 $8.00 $7.25
openai gpt-5.4-pro $48.00 $24.00 $24.00 — $87.00
openai gpt-5.4-mini $1.20 $0.60 $0.60 $2.40 —
openai gpt-5.4-nano $0.33 $0.16 $0.16 — —
openai gpt-5.3-codex $3.15 — — $6.30 —
openai gpt-5.2 $3.15 $1.58 $1.58 $6.30 —
openai gpt-5.2-pro $37.80 $18.90 — — —
openai gpt-5.1 $2.25 $1.12 $1.12 $4.50 —
openai gpt-5 $2.25 $1.12 $1.12 $4.50 —
openai gpt-5-mini $0.45 $0.23 $0.23 $0.81 —
openai gpt-5-nano $0.0900 $0.0450 $0.0450 — —
openai gpt-5-pro $27.00 $13.50 — — —
openai o3 $2.80 $1.40 $1.40 $4.90 —
openai o3-pro $28.00 $14.00 — — —
openai o4-mini $1.54 $0.77 $0.77 $2.80 —
openai o3-mini $1.54 $0.77 — — —
openai gpt-4.1 $2.80 $1.40 — $4.90 —
openai gpt-4.1-mini $0.56 $0.28 — $0.98 —
openai gpt-4.1-nano $0.14 $0.0700 — $0.28 —
openai gpt-4o $3.50 $1.75 — $5.95 —
openai gpt-4o-mini $0.21 $0.10 — $0.35 —
openai chat-latest $8.00 — — — —
openai gpt-5.6-cyber $20.00 — — — —
anthropic claude-fable-5-1 $15.00 $7.50 — — same as standard (no premium)
anthropic claude-fable-5 $15.00 $7.50 — — same as standard (no premium)
anthropic claude-mythos-5-1 $15.00 $7.50 — — same as standard (no premium)
anthropic claude-mythos-5 $15.00 $7.50 — — same as standard (no premium)
anthropic claude-opus-5 $7.50 $3.75 — $15.00 same as standard (no premium)
anthropic claude-opus-4-8 $7.50 $3.75 — $15.00 same as standard (no premium)
anthropic claude-opus-4-7 $7.50 $3.75 — — same as standard (no premium)
anthropic claude-opus-4-6 $7.50 $3.75 — — same as standard (no premium)
anthropic claude-opus-4-5-20251101 $7.50 $3.75 — — same as standard (no premium)
anthropic claude-sonnet-5 $3.00 $1.50 — — same as standard (no premium)
anthropic claude-sonnet-4-6 $4.50 $2.25 — — same as standard (no premium)
anthropic claude-sonnet-4-5-20250929 $4.50 $2.25 — — same as standard (no premium)
anthropic claude-haiku-4-5-20251001 $1.50 $0.75 — — same as standard (no premium)
xai grok-4.6 $2.60 not supported — $5.20 (priority) $5.20
xai grok-4.5 $2.60 not supported — $5.20 (priority) $5.20
xai grok-4.3 $1.50 $1.20 — $3.00 (priority) $3.00
xai grok-4.20-0309-reasoning $1.50 $1.20 — $3.00 (priority) $3.00
xai grok-4.20-0309-non-reasoning $1.50 $1.20 — $3.00 (priority) $3.00
xai grok-4.20-multi-agent-0309 $1.50 $1.20 — $3.00 (priority) $3.00
xai grok-build-0.1 $1.20 not supported — $2.40 (priority) $2.40
gemini gemini-3.8-flash $1.12 $0.56 $0.56 $2.03 (priority) same as standard (flat)
gemini gemini-3.7-flash $1.12 $0.56 $0.56 $2.03 (priority) same as standard (flat)
gemini gemini-3.6-flash $1.12 $0.56 $0.56 $2.03 (priority) same as standard (flat)
gemini gemini-3.5-flash $2.40 $1.20 $1.20 $4.32 (priority) same as standard (flat)
gemini gemini-3.5-flash-lite $0.55 $0.28 $0.28 $0.99 (priority) same as standard (flat)
gemini gemini-3.1-pro-preview $3.20 $1.60 $1.60 $5.76 (priority) $5.80
gemini gemini-3.1-pro-preview-customtools $3.20 $1.60 $1.60 $5.76 (priority) $5.80
gemini gemini-3.1-flash-lite $0.40 $0.20 $0.20 $0.72 (priority) same as standard (flat)
gemini gemini-3-flash-preview $0.80 $0.40 $0.40 $1.44 (priority) same as standard (flat)
gemini gemini-2.5-pro $2.25 $1.12 $1.12 $4.05 (priority) $4.00
gemini gemini-2.5-flash $0.55 $0.28 $0.28 $0.99 (priority) same as standard (flat)
gemini gemini-2.5-flash-lite $0.14 $0.0700 $0.0700 $0.25 (priority) same as standard (flat)
gemini gemini-robotics-er-2-preview $1.50 $0.75 — — (priority) same as standard (flat)
gemini gemini-2.5-computer-use-preview-10-2025 $1.50 — — — (priority) same as standard (flat)

Gemma 4 (gemma-4-31b-it, gemma-4-26b-a4b-it) and every Gemini free-tier row cost $0 at list price but carry free-tier limits and data-use terms; they are omitted from the arithmetic.

# 7. Cost model B — the same workload with a 900k-token cached prefix (steady state)

Assumptions: each request = 900k tokens read from cache + 100k fresh input + 100k output; the initial cache write is amortised over N = 10 requests. Formula per request: 0.9 × cache_read + 0.1 × input + 0.1 × output + write_cost / N. Write cost: OpenAI pre-5.6 free, GPT-5.6+ cache_write × 0.9; Anthropic 5-minute write (1.25×) × 0.9; xAI none (automatic cache); Gemini implicit caching none, explicit cachedContents = storage of 0.9M tokens for one hour (cache_storage_hour × 0.9) amortised over N. A 900k prefix exceeds xAI's 200k threshold, so xAI is shown at long-context rates (the only rates that apply to such a request); Gemini Pro at its >200k rates, Gemini Flash flat.

Provider Model Uncached (A, same rates) Cached steady-state (N=10) Saving Write / storage cost used
openai gpt-6-astra $15.00 $8.03 46 % 12.5
openai gpt-5.6-sol $6.00 $3.21 46 % 5
openai gpt-5.6-terra $3.20 $1.81 44 % 2.5
openai gpt-5.6-luna $0.32 $0.18 44 % 0.25
openai gpt-5.5 $8.00 $3.95 51 % free
openai gpt-5.4 $4.00 $1.98 51 % free
openai gpt-5.4-mini $1.20 $0.59 51 % free
openai gpt-5.4-nano $0.33 $0.16 50 % free
openai gpt-5.3-codex $3.15 $1.73 45 % free
openai gpt-5.2 $3.15 $1.73 45 % free
openai gpt-5.1 $2.25 $1.24 45 % free
openai gpt-5 $2.25 $1.24 45 % free
openai gpt-5-mini $0.45 $0.25 45 % free
openai gpt-5-nano $0.0900 $0.0495 45 % free
openai o3 $2.80 $1.45 48 % free
openai o4-mini $1.54 $0.80 48 % free
openai o3-mini $1.54 $1.05 32 % free
openai gpt-4.1 $2.80 $1.45 48 % free
openai gpt-4.1-mini $0.56 $0.29 48 % free
openai gpt-4.1-nano $0.14 $0.0725 48 % free
openai gpt-4o $3.50 $2.38 32 % free
openai gpt-4o-mini $0.21 $0.14 32 % free
openai chat-latest $8.00 $3.95 51 % free
openai gpt-5.6-cyber $20.00 $11.28 44 % 15.625
anthropic claude-fable-5-1 $15.00 $7.35 51 % 12.5 (5m write)
anthropic claude-fable-5 $15.00 $8.03 46 % 12.5 (5m write)
anthropic claude-mythos-5-1 $15.00 $7.35 51 % 12.5 (5m write)
anthropic claude-mythos-5 $15.00 $8.03 46 % 12.5 (5m write)
anthropic claude-opus-5 $7.50 $4.01 46 % 6.25 (5m write)
anthropic claude-opus-4-8 $7.50 $4.01 46 % 6.25 (5m write)
anthropic claude-opus-4-7 $7.50 $4.01 46 % 6.25 (5m write)
anthropic claude-opus-4-6 $7.50 $4.01 46 % 6.25 (5m write)
anthropic claude-opus-4-5-20251101 $7.50 $4.01 46 % 6.25 (5m write)
anthropic claude-sonnet-5 $3.00 $1.60 46 % 2.5 (5m write)
anthropic claude-sonnet-4-6 $4.50 $2.41 46 % 3.75 (5m write)
anthropic claude-sonnet-4-5-20250929 $4.50 $2.41 46 % 3.75 (5m write)
anthropic claude-haiku-4-5-20251001 $1.50 $0.80 46 % 1.25 (5m write)
xai grok-4.6 $5.20 (long-context rates) $2.50 52 % none (automatic; cached rate 1)
xai grok-4.5 $5.20 (long-context rates) $2.14 59 % none (automatic; cached rate 0.6)
xai grok-4.3 $3.00 (long-context rates) $1.11 63 % none (automatic; cached rate 0.4)
xai grok-4.20-0309-reasoning $3.00 (long-context rates) $1.11 63 % none (automatic; cached rate 0.4)
xai grok-4.20-0309-non-reasoning $3.00 (long-context rates) $1.11 63 % none (automatic; cached rate 0.4)
xai grok-4.20-multi-agent-0309 $3.00 (long-context rates) $1.11 63 % none (automatic; cached rate 0.4)
xai grok-build-0.1 $2.40 (long-context rates) $0.96 60 % none (automatic; cached rate 0.4)
gemini gemini-3.8-flash $1.12 $0.56 50 % implicit: none; explicit: storage 0.5 /1M/h × 0.9
gemini gemini-3.7-flash $1.12 $0.56 50 % implicit: none; explicit: storage 0.5 /1M/h × 0.9
gemini gemini-3.6-flash $1.12 $0.56 50 % implicit: none; explicit: storage 0.5 /1M/h × 0.9
gemini gemini-3.5-flash $2.40 $1.28 47 % implicit: none; explicit: storage 1 /1M/h × 0.9
gemini gemini-3.5-flash-lite $0.55 $0.40 28 % implicit: none; explicit: storage 1 /1M/h × 0.9
gemini gemini-3.1-pro-preview $5.80 (>200k rates) $2.96 49 % implicit: none; explicit: storage 4.5 /1M/h × 0.9
gemini gemini-3.1-pro-preview-customtools $5.80 (>200k rates) $2.96 49 % implicit: none; explicit: storage 4.5 /1M/h × 0.9
gemini gemini-3.1-flash-lite $0.40 $0.29 28 % implicit: none; explicit: storage 1 /1M/h × 0.9
gemini gemini-3-flash-preview $0.80 $0.48 39 % implicit: none; explicit: storage 1 /1M/h × 0.9
gemini gemini-2.5-pro $4.00 (>200k rates) $2.38 40 % implicit: none; explicit: storage 4.5 /1M/h × 0.9
gemini gemini-2.5-flash $0.55 $0.40 28 % implicit: none; explicit: storage 1 /1M/h × 0.9
gemini gemini-2.5-flash-lite $0.14 $0.15 -6 % implicit: none; explicit: storage 1 /1M/h × 0.9
gemini gemini-robotics-er-2-preview $1.50 $0.73 51 % implicit: none; explicit: storage 0.5 /1M/h × 0.9

# Caching break-even by provider

Provider Write cost Read multiplier Pays for itself on… Notes
OpenAI (pre-5.6) free model-specific 0.25×–0.5× first hit implicit; prompt_cache_key routing; 5–10 min in-memory, 24h retention option
OpenAI (GPT-5.6+, GPT-6) 1.25× 0.1× second use (1.25 + 0.1 = 1.35 < 2.0) explicit prompt_cache_breakpoint optional; ttl: 30m
Anthropic 1.25× (5 m) / 2× (1 h) 0.1× (0.025× Fable 5.1 / Mythos 5.1) second use (5 m); 1 h only if the gap between calls exceeds 5 min explicit cache_control; hits refresh TTL for free
xAI none 0.15×–0.25× first hit automatic; prompt_cache_key / x-grok-conv-id for sticky routing; no TTL or minimum documented; cached tokens still count toward TPM
Gemini implicit: none; explicit: storage per token-hour 0.1× implicit: first hit (≥ 4,096-token prompts on 3.x, 2,048 on 2.5); explicit: when 0.9 × read + storage_hour × hours / N < 0.9 × input — for gemini-3.8-flash one hour of storage ($0.45 per 0.9M) is recovered after a single hit ($0.675 − $0.0675 saved) explicit cache TTL default 1 h (ttl / expireTime), min 1,024 tokens (live), free tier limit: 0

# 8. Batch vs synchronous — the same 100k-request job

OpenAI Batch API Anthropic Message Batches xAI Batch API Gemini Batch API
Discount 50 % on input, cached input and output (batch tier) 50 % on input, output, cache writes and cache reads (stacks with caching) 20 % on input, output, cached and reasoning tokens — grok-4.3 / 4.20 / 4.20-multi-agent only; grok-4.6, grok-4.5, grok-build: not supported; Imagine models: standard rates 50 % (service:batch_api) on text, image, TTS, embedding models; Flex gives the same 50 % synchronously (best-effort)
Window 24 h (completion_window: "24h", only value); output retained 30 days 24 h expiry (expires_at); results 29 days no completion window documented; results listed via GET /v1/batches/{id}/results; batch requests do not count toward rate limits; video URLs in results expire after 1 h target 24 h ('usually much faster'); poll GET /v1beta/batches/{id}; concurrent batch jobs 100; enqueued-token caps per model and tier (e.g. Tier 1: 3M gemini-3.8-flash, 5M gemini-3.1-pro-preview)
How requests are supplied JSONL file (custom_id, method, url, body) → POST /v1/batches {input_file_id, endpoint}; 50,000 requests / 200 MB; one endpoint + one model per file inline requests[] {custom_id, params} → POST /v1/messages/batches; 100,000 requests / 256 MB; models mixed freely POST /v1/batches then POST /v1/batches/{id}/requests (inline chat/responses/image/video bodies) or JSONL upload; :cancel; no DELETE (405) POST /v1beta/models/{model}:batchGenerateContent with inline requests[] or a JSONL file from the Files API (2 GB / 20 GB storage); :asyncBatchEmbedContent for embeddings; /v1beta/openai/batches compatibility route
Example: 100k requests × (2k in + 300 out) on the flagship mid-tier gpt-5.6-sol batch: 200M × $2 + 30M × $10 = $700 (vs $1,400 standard) claude-sonnet-5 batch: 200M × $1 + 30M × $5 = $350 (vs $700 standard) grok-4.3 batch: 200M × $1.00 + 30M × $2.00 = $260 (vs $325 standard); grok-4.6 has no batch: $430 standard gemini-3.8-flash batch: 200M × $0.375 + 30M × $1.875 = $131.25 (vs $262.50 standard); gemini-3.1-pro-preview batch: 200M × $1 + 30M × $6 = $380 (vs $760)
Example: same job on the small models gpt-5.4-mini batch: 200M × $0.375 + 30M × $2.25 = $142.50 claude-haiku-4-5 batch: 200M × $0.5 + 30M × $2.5 = $175 grok-build-0.1: no batch, $260 standard gemini-3.5-flash-lite batch: 200M × $0.15 + 30M × $1.25 = $67.50 (vs $135)

# 9. Tool and platform prices (from generated/pricing.json)

Item OpenAI Anthropic xAI Gemini
Web search $10 / 1k calls (web_search, image search); preview $25 / 1k on non-reasoning models $10 / 1k searches (errors not billed) $5 / 1k successful calls (web_search, image search included; view_image results billed as image tokens) Google Search grounding: Gemini 3.x 5,000 free queries / month (shared) then $14 / 1k search queries; Gemini 2.5: 1,500 RPD free then $35 / 1k grounded prompts; free tier 500 RPD (not Pro)
Social / vertical search — — x_search $5 / 1k calls until 2026-09-21 12:00 PT, then $5 / 1k posts fetched + $10 / 1k user profiles fetched Google Maps grounding: Gemini 3.x 5,000 free prompts / month then $14 / 1k queries; Gemini 2.5: $25 / 1k grounded prompts (1,500 RPD free; 10,000 for Pro)
Web fetch / URL context — $0 per fetch (content tokens only) — (web search open_page action) urlContext free of charge; retrieved content billed as input tokens (toolUsePromptTokenCount); ≤ 20 URLs / request
Code execution container session: 1 GB $0.03 · 4 GB $0.12 · 16 GB $0.48 · 64 GB $1.92 per 20 min (per-minute, 5-min minimum since 2026-06-02); shared with hosted shell 1,550 free container-hours / org / month, then $0.05 / container-hour (5-min minimum); free when web_search/web_fetch 20260209+ is in the request $5 / 1k calls (code_interpreter / code_execution) + tokens no per-call fee (codeExecution); generated code/results billed as output then as input when re-read; 30 s per execution
File search / RAG $2.50 / 1k calls + $0.10 / GB / day storage (1 GB free) — file_search / collections_search $2.50 / 1k calls + tokens; collections storage $0.10 / GiB / day, downloads $0.20 / GiB; attachment_search (implicit, files attached to messages) $10 / 1k calls fileSearch: indexing embeddings $0.15 / 1M tokens once; storage and query-time embeddings free; retrieved chunks billed as input tokens; store size by tier (Free 1 GB … Tier 3 1 TB)
MCP tokens only (mcp tool) tokens only (mcp_servers, beta) tokens only (mcp tool; tool outputs count as input) not documented (mcpServers UNVERIFIED on generateContent; Interactions mcp_server type)
Computer use tokens (computer tool) tokens (toolsets ≈ 4,500 tokens of definitions) — tokens only (computerUse, PREVIEW; screenshots are image input tokens); not on the free tier
Files API storage not priced separately (purpose=batch files expire 30 d) not priced on the pricing page $0.025 / GiB / day storage, $0.20 / GiB downloads; TTL via expires_after free (2 GB / file, 20 GB / project, 48 h TTL)
Managed agents runtime model tokens at Responses rates + container rates (no session fee) $0.08 per session-hour (usage.active_seconds) + model tokens + $10 / 1k web searches xAI Responses agentic loop: tokens + per-call tool fees; usage-guideline-violation fee $0.05 / request when a violation is caught before generation Interactions agents (Deep Research, Antigravity, custom managed agents): model inference at list rates incl. intermediate/reasoning tokens + tool fees; sandbox compute not billed during preview
Realtime / voice gpt-realtime-2.1: audio $32 in / $64 out, text $4 / $24, image $5 per 1M; mini $10 / $20 audio; gpt-live-1 $0.05 / min — grok-voice-think-fast-2.0 $0.08 / min ($4.80 / h) audio sent or received + $0.004 per text item; STT $0.10 / h REST, $0.20 / h streaming; TTS $15 / 1M characters gemini-3.8-live: audio in $3 / 1M (≈ $0.005 / min), audio out $12 / 1M (≈ $0.018 / min), text $0.75 / $4.50; TTS gemini-3.1-flash-tts-preview $1 text in / $20 audio out per 1M (≈ $0.03 / min); gemini-3.5-transcribe ≈ $0.005 / min blended; free tier available
Image generation gpt-image-2 / 2.5: text in $5, image in $8, image out $30 per 1M tokens (≈ $0.006–$0.21 per image); batch 50 % — grok-imagine-image $0.02 / image, grok-imagine-image-2.0 $0.04–$0.08 (quality × resolution), grok-imagine-image-quality $0.05 (DEPRECATED → 2026-11-02); image_generation tool billed at these rates gemini-3.1-flash-image $0.045 (0.5K) – $0.151 (4K) per image (image out $60 / 1M), gemini-3.1-flash-lite-image $0.034 (1K), gemini-3-pro-image $0.134 (1K/2K) / $0.24 (4K); batch 50 %
Video sora-2 $0.10 / s (720p), sora-2-pro $0.30–$0.70 / s — shutting down 2026-09-24 — grok-imagine-video $0.05 / s, grok-imagine-video-1.5 $0.08 / s (1080p, reference-to-video) Veo 3.1 $0.40 / s (720p/1080p), $0.60 / s (4K); Veo 3.1 Fast $0.10 / $0.12 / $0.30; Veo 3.1 Lite $0.05 / $0.08 (no 4K); gemini-omni-1.1-flash video out $17.50 / 1M tokens (≈ $0.10 / s)
Music — — — lyria-3.5 $0.08 / song, lyria-3-clip-preview $0.04 / 30-s clip, lyria-3-pro-preview $0.08; Lyria RealTime (experimental) unpriced
Embeddings text-embedding-3-small $0.02, -large $0.13 per 1M — grok-embedding-small: no published price (ACCOUNT_RESTRICTED) gemini-embedding-2 text $0.20 / 1M (batch $0.10), image $0.45, audio $6.50, video $12 per 1M; free tier available; gemini-embedding-001 unpriced (DEPRECATED → 2028-05-14)
Moderation free (omni-moderation-latest) — (built-in refusals) — (built-in; usage-guideline violations still billed + $0.05 fee) — (safetySettings thresholds, free)
Token counting free (POST /v1/responses/input_tokens) free (POST /v1/messages/count_tokens, own RPM bucket) free (POST /v1/tokenize-text; no usage/cost ticks returned) free (POST /v1beta/models/{model}:countTokens)
Data residency +10 % regional hosts (models ≥ 2026-03-05) ×1.1 inference_geo: us; +10 % Bedrock/Vertex regional ×1.1 on us.api.x.ai (grok-4.6) not priced on the Developer API (Vertex AI feature)
Per-request overhead none published tool-use system prompt: e.g. Opus 5 406 tokens (any/tool) / 286 (auto); bash_20250124 +244–325; text_editor +700; computer_toolset_20260801 ≈4,500; browser_toolset_20260801 ≈6,600 none published; cost_in_usd_ticks in every usage object (1 tick = $1e-10) makes the effective cost observable per call none published; usageMetadata breaks tokens down by modality (promptTokensDetails[]), thoughtsTokenCount, toolUsePromptTokenCount, cachedContentTokenCount

# 10. Reading prices programmatically

bash
# OpenAI standard row for one model
jq '.[] | select(.id=="gpt-5.6-sol") | .pricing.standard' generated/models.json
# Anthropic per-dimension rows
jq '[.[] | select(.provider=="anthropic" and .model_or_service=="claude-sonnet-5")]' generated/pricing.json
# xAI standard + long-context blocks and the live price ticks (1 tick = USD 1e-10 per token)
jq '.[] | select(.id=="grok-4.6") | .pricing | {standard, long_context, batch_discount, priority_multiplier, us_regional_multiplier}' generated/models.json
# Gemini tiers (standard/batch/flex/priority) and free-tier note
jq '.[] | select(.id=="gemini-3.8-flash") | .pricing | {tiers, free_tier, from_2027_01_01}' generated/models.json
# every tool / service price row across the four providers
jq '[.[] | select(.model_or_service|test("^(tool:|service:|agent:|rule:)"))]' generated/pricing.json

Related: models · caching and reasoning · features · realtime and media · FAQ.