# OpenAI — Pricing (API Atlas) **Status:** DOCUMENTED (prices copied from the official pricing page and model pages; not billed-verified beyond the minimal probes). Machine-readable twin: `generated/fragments/pricing/openai-pricing.json`. **Sources:** https://developers.openai.com/api/docs/pricing · https://developers.openai.com/api/docs/models/ · https://developers.openai.com/api/docs/guides/batch · https://developers.openai.com/api/docs/guides/flex-processing · https://developers.openai.com/api/docs/guides/fast-mode · https://developers.openai.com/api/docs/guides/prompt-caching · https://developers.openai.com/api/docs/guides/your-data **Last verified:** 2026-09-18 All prices USD. Token prices are per 1M tokens. `cached_input` = cache read; `cache_write` (GPT-5.6+/GPT-6 only) = 1.25× uncached input. Long context = prompts >272K input tokens (2× input/cache, 1.5× output for the whole request). ## Service tiers | Tier | `service_tier` | Price rule | Notes | |---|---|---|---| | Standard | `default` / omitted | list price | | | Batch | Batch API (`/v1/batches`, `completion_window: 24h`) | 50% of standard | separate, much higher queue limits; 24h turnaround | | Flex | `flex` | Batch rates (50%) | beta, limited models; slower; may return 429 `resource_unavailable`; raise client timeout (15 min recommended) | | Fast (ex-Priority) | `fast` or `priority` | 2× standard for GPT-5.6 Sol / GPT-6 Astra (per-model table) | renamed 2026-07-30; up to 2.5× faster; downgraded requests return `service_tier: default` and standard rates; no fine-tuned models/embeddings; unavailable for GPT-6 Astra with EU data residency | | Ultrafast | — | — | limited preview for GPT-5.6 Sol (announced 2026-08-13), up to 14× faster | | Regional processing | `us.`/`eu.`/`ae.` … prefixed domains | +10% uplift | models released on/after 2026-03-05 that are data-residency eligible | ## Text models (Flagship + legacy text) ### Standard — short context (≤272K) | model | input | cached_input | cache_write | output | unit | notes | |---|---|---|---|---|---|---| | `gpt-6-astra` | $10 | $1 | $12.5 | $50 | per 1M tokens | short context rate | | `gpt-5.6-sol` | $4 | $0.4 | $5 | $20 | per 1M tokens | short context rate | | `gpt-5.6-terra` | $2 | $0.2 | $2.5 | $12 | per 1M tokens | short context rate | | `gpt-5.6-luna` | $0.2 | $0.02 | $0.25 | $1.2 | per 1M tokens | short context rate | | `gpt-5.5` | $5 | $0.5 | | $30 | per 1M tokens | <272K context length; short context rate | | `gpt-5.5-pro` | $30 | | | $180 | per 1M tokens | <272K context length; short context rate | | `gpt-5.4` | $2.5 | $0.25 | | $15 | per 1M tokens | <272K context length; short context rate | | `gpt-5.4-mini` | $0.75 | $0.075 | | $4.5 | per 1M tokens | short context rate | | `gpt-5.4-nano` | $0.2 | $0.02 | | $1.25 | per 1M tokens | short context rate | | `gpt-5.4-pro` | $30 | | | $180 | per 1M tokens | <272K context length; short context rate | | `gpt-5.2` | $1.75 | $0.175 | | $14 | per 1M tokens | short context rate | | `gpt-5.2-pro` | $21 | | | $168 | per 1M tokens | short context rate | | `gpt-5.1` | $1.25 | $0.125 | | $10 | per 1M tokens | short context rate | | `gpt-5` | $1.25 | $0.125 | | $10 | per 1M tokens | short context rate | | `gpt-5-mini` | $0.25 | $0.025 | | $2 | per 1M tokens | short context rate | | `gpt-5-nano` | $0.05 | $0.005 | | $0.4 | per 1M tokens | short context rate | | `gpt-5-pro` | $15 | | | $120 | per 1M tokens | short context rate | | `gpt-4.1` | $2 | $0.5 | | $8 | per 1M tokens | short context rate | | `gpt-4.1-mini` | $0.4 | $0.1 | | $1.6 | per 1M tokens | short context rate | | `gpt-4.1-nano` | $0.1 | $0.025 | | $0.4 | per 1M tokens | short context rate | | `gpt-4o` | $2.5 | $1.25 | | $10 | per 1M tokens | short context rate | | `gpt-4o-2024-05-13` | $5 | | | $15 | per 1M tokens | short context rate | | `gpt-4o-mini` | $0.15 | $0.075 | | $0.6 | per 1M tokens | short context rate | | `o1` | $15 | $7.5 | | $60 | per 1M tokens | short context rate | | `o1-pro` | $150 | | | $600 | per 1M tokens | short context rate | | `o3-pro` | $20 | | | $80 | per 1M tokens | short context rate | | `o3` | $2 | $0.5 | | $8 | per 1M tokens | short context rate | | `o4-mini` | $1.1 | $0.275 | | $4.4 | per 1M tokens | short context rate | | `o3-mini` | $1.1 | $0.55 | | $4.4 | per 1M tokens | short context rate | | `gpt-4-turbo-2024-04-09` | $10 | | | $30 | per 1M tokens | short context rate | | `gpt-4-0613` | $30 | | | $60 | per 1M tokens | short context rate | | `gpt-3.5-turbo` | $0.5 | | | $1.5 | per 1M tokens | short context rate | | `gpt-3.5-turbo-0125` | $0.5 | | | $1.5 | per 1M tokens | short context rate | | `gpt-3.5-turbo-1106` | $1 | | | $2 | per 1M tokens | short context rate | | `gpt-3.5-turbo-instruct` | $1.5 | | | $2 | per 1M tokens | short context rate | | `davinci-002` | $2 | | | $2 | per 1M tokens | short context rate | | `babbage-002` | $0.4 | | | $0.4 | per 1M tokens | short context rate | ### Standard — long context (>272K) | model | input | cached_input | cache_write | output | unit | notes | |---|---|---|---|---|---|---| | `gpt-6-astra` | $20 | $2 | $25 | $75 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-sol` | $8 | $0.8 | $10 | $30 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-terra` | $4 | $0.4 | $5 | $18 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-luna` | $0.4 | $0.04 | $0.5 | $1.8 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.5` | $10 | $1 | | $45 | per 1M tokens | <272K context length; long context (>272K input tokens) rate | | `gpt-5.5-pro` | $60 | | | $270 | per 1M tokens | <272K context length; long context (>272K input tokens) rate | | `gpt-5.4` | $5 | $0.5 | | $22.5 | per 1M tokens | <272K context length; long context (>272K input tokens) rate | | `gpt-5.4-pro` | $60 | | | $270 | per 1M tokens | <272K context length; long context (>272K input tokens) rate | ### Batch — short context (≤272K) | model | input | cached_input | cache_write | output | unit | notes | |---|---|---|---|---|---|---| | `gpt-6-astra` | $5 | $0.5 | $6.25 | $25 | per 1M tokens | short context rate | | `gpt-5.6-sol` | $2 | $0.2 | $2.5 | $10 | per 1M tokens | short context rate | | `gpt-5.6-terra` | $1 | $0.1 | $1.25 | $6 | per 1M tokens | short context rate | | `gpt-5.6-luna` | $0.1 | $0.01 | $0.125 | $0.6 | per 1M tokens | short context rate | | `gpt-5.5` | $2.5 | $0.25 | | $15 | per 1M tokens | <272K context length; short context rate | | `gpt-5.5-pro` | $15 | | | $90 | per 1M tokens | <272K context length; short context rate | | `gpt-5.4` | $1.25 | $0.13 | | $7.5 | per 1M tokens | <272K context length; short context rate | | `gpt-5.4-mini` | $0.375 | $0.0375 | | $2.25 | per 1M tokens | short context rate | | `gpt-5.4-nano` | $0.1 | $0.01 | | $0.625 | per 1M tokens | short context rate | | `gpt-5.4-pro` | $15 | | | $90 | per 1M tokens | <272K context length; short context rate | | `gpt-5.2` | $0.875 | $0.0875 | | $7 | per 1M tokens | short context rate | | `gpt-5.2-pro` | $10.5 | | | $84 | per 1M tokens | short context rate | | `gpt-5.1` | $0.625 | $0.0625 | | $5 | per 1M tokens | short context rate | | `gpt-5` | $0.625 | $0.0625 | | $5 | per 1M tokens | short context rate | | `gpt-5-mini` | $0.125 | $0.0125 | | $1 | per 1M tokens | short context rate | | `gpt-5-nano` | $0.025 | $0.0025 | | $0.2 | per 1M tokens | short context rate | | `gpt-5-pro` | $7.5 | | | $60 | per 1M tokens | short context rate | | `gpt-4.1` | $1 | | | $4 | per 1M tokens | short context rate | | `gpt-4.1-mini` | $0.2 | | | $0.8 | per 1M tokens | short context rate | | `gpt-4.1-nano` | $0.05 | | | $0.2 | per 1M tokens | short context rate | | `gpt-4o` | $1.25 | | | $5 | per 1M tokens | short context rate | | `gpt-4o-2024-05-13` | $2.5 | | | $7.5 | per 1M tokens | short context rate | | `gpt-4o-mini` | $0.075 | | | $0.3 | per 1M tokens | short context rate | | `o1` | $7.5 | | | $30 | per 1M tokens | short context rate | | `o1-pro` | $75 | | | $300 | per 1M tokens | short context rate | | `o3-pro` | $10 | | | $40 | per 1M tokens | short context rate | | `o3` | $1 | | | $4 | per 1M tokens | short context rate | | `o4-mini` | $0.55 | | | $2.2 | per 1M tokens | short context rate | | `o3-mini` | $0.55 | | | $2.2 | per 1M tokens | short context rate | | `gpt-4-turbo-2024-04-09` | $5 | | | $15 | per 1M tokens | short context rate | | `gpt-4-0613` | $15 | | | $30 | per 1M tokens | short context rate | | `gpt-3.5-turbo-0125` | $0.25 | | | $0.75 | per 1M tokens | short context rate | | `gpt-3.5-turbo-1106` | $1 | | | $2 | per 1M tokens | short context rate | | `davinci-002` | $1 | | | $1 | per 1M tokens | short context rate | | `babbage-002` | $0.2 | | | $0.2 | per 1M tokens | short context rate | ### Batch — long context (>272K) | model | input | cached_input | cache_write | output | unit | notes | |---|---|---|---|---|---|---| | `gpt-6-astra` | $10 | $1 | $12.5 | $37.5 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-sol` | $4 | $0.4 | $5 | $15 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-terra` | $2 | $0.2 | $2.5 | $9 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-luna` | $0.2 | $0.02 | $0.25 | $0.9 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.5` | $5 | $0.5 | | $22.5 | per 1M tokens | <272K context length; long context (>272K input tokens) rate | | `gpt-5.4` | $2.5 | $0.25 | | $11.25 | per 1M tokens | <272K context length; long context (>272K input tokens) rate | | `gpt-5.4-pro` | $30 | | | $135 | per 1M tokens | <272K context length; long context (>272K input tokens) rate | ### Flex — short context (≤272K) | model | input | cached_input | cache_write | output | unit | notes | |---|---|---|---|---|---|---| | `gpt-6-astra` | $5 | $0.5 | $6.25 | $25 | per 1M tokens | short context rate | | `gpt-5.6-sol` | $2 | $0.2 | $2.5 | $10 | per 1M tokens | short context rate | | `gpt-5.6-terra` | $1 | $0.1 | $1.25 | $6 | per 1M tokens | short context rate | | `gpt-5.6-luna` | $0.1 | $0.01 | $0.125 | $0.6 | per 1M tokens | short context rate | | `gpt-5.5` | $2.5 | $0.25 | | $15 | per 1M tokens | <272K context length; short context rate | | `gpt-5.5-pro` | $15 | | | $90 | per 1M tokens | <272K context length; short context rate | | `gpt-5.4` | $1.25 | $0.13 | | $7.5 | per 1M tokens | <272K context length; short context rate | | `gpt-5.4-mini` | $0.375 | $0.0375 | | $2.25 | per 1M tokens | short context rate | | `gpt-5.4-nano` | $0.1 | $0.01 | | $0.625 | per 1M tokens | short context rate | | `gpt-5.4-pro` | $15 | | | $90 | per 1M tokens | <272K context length; short context rate | | `gpt-5.2` | $0.875 | $0.0875 | | $7 | per 1M tokens | short context rate | | `gpt-5.1` | $0.625 | $0.0625 | | $5 | per 1M tokens | short context rate | | `gpt-5` | $0.625 | $0.0625 | | $5 | per 1M tokens | short context rate | | `gpt-5-mini` | $0.125 | $0.0125 | | $1 | per 1M tokens | short context rate | | `gpt-5-nano` | $0.025 | $0.0025 | | $0.2 | per 1M tokens | short context rate | | `o3` | $1 | $0.25 | | $4 | per 1M tokens | short context rate | | `o4-mini` | $0.55 | $0.138 | | $2.2 | per 1M tokens | short context rate | ### Flex — long context (>272K) | model | input | cached_input | cache_write | output | unit | notes | |---|---|---|---|---|---|---| | `gpt-6-astra` | $10 | $1 | $12.5 | $37.5 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-sol` | $4 | $0.4 | $5 | $15 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-terra` | $2 | $0.2 | $2.5 | $9 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-luna` | $0.2 | $0.02 | $0.25 | $0.9 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.5` | $5 | $0.5 | | $22.5 | per 1M tokens | <272K context length; long context (>272K input tokens) rate | | `gpt-5.4` | $2.5 | $0.25 | | $11.25 | per 1M tokens | <272K context length; long context (>272K input tokens) rate | | `gpt-5.4-pro` | $30 | | | $135 | per 1M tokens | <272K context length; long context (>272K input tokens) rate | ### Fast — short context (≤272K) | model | input | cached_input | cache_write | output | unit | notes | |---|---|---|---|---|---|---| | `gpt-6-astra` | $20 | $2 | $25 | $100 | per 1M tokens | short context rate | | `gpt-5.6-sol` | $8 | $0.8 | $10 | $40 | per 1M tokens | short context rate | | `gpt-5.6-terra` | $4 | $0.4 | $5 | $24 | per 1M tokens | short context rate | | `gpt-5.6-luna` | $0.4 | $0.04 | $0.5 | $2.4 | per 1M tokens | short context rate | | `gpt-5.5` | $12.5 | $1.25 | | $75 | per 1M tokens | <272K context length; short context rate | | `gpt-5.4` | $5 | $0.5 | | $30 | per 1M tokens | <272K context length; short context rate | | `gpt-5.4-mini` | $1.5 | $0.15 | | $9 | per 1M tokens | short context rate | | `gpt-5.2` | $3.5 | $0.35 | | $28 | per 1M tokens | short context rate | | `gpt-5.1` | $2.5 | $0.25 | | $20 | per 1M tokens | short context rate | | `gpt-5` | $2.5 | $0.25 | | $20 | per 1M tokens | short context rate | | `gpt-5-mini` | $0.45 | $0.045 | | $3.6 | per 1M tokens | short context rate | | `gpt-4.1` | $3.5 | $0.875 | | $14 | per 1M tokens | short context rate | | `gpt-4.1-mini` | $0.7 | $0.175 | | $2.8 | per 1M tokens | short context rate | | `gpt-4.1-nano` | $0.2 | $0.05 | | $0.8 | per 1M tokens | short context rate | | `gpt-4o` | $4.25 | $2.125 | | $17 | per 1M tokens | short context rate | | `gpt-4o-2024-05-13` | $8.75 | | | $26.25 | per 1M tokens | short context rate | | `gpt-4o-mini` | $0.25 | $0.125 | | $1 | per 1M tokens | short context rate | | `o3` | $3.5 | $0.875 | | $14 | per 1M tokens | short context rate | | `o4-mini` | $2 | $0.5 | | $8 | per 1M tokens | short context rate | ### Fast — long context (>272K) | model | input | cached_input | cache_write | output | unit | notes | |---|---|---|---|---|---|---| | `gpt-6-astra` | $40 | $4 | $50 | $150 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-sol` | $16 | $1.6 | $20 | $60 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-terra` | $8 | $0.8 | $10 | $36 | per 1M tokens | long context (>272K input tokens) rate | | `gpt-5.6-luna` | $0.8 | $0.08 | $1 | $3.6 | per 1M tokens | long context (>272K input tokens) rate | ### Cyber / Daybreak models (standard) | model | input | cached_input | cache_write | output | unit | notes | |---|---|---|---|---|---|---| | `gpt-5.6-sol` | $4 | $0.4 | $5 | $20 | per 1M tokens | short context rate | | `gpt-5.6-cyber` | $12.5 | $1.25 | $15.625 | $75 | per 1M tokens | short context rate | | `gpt-5.5-cyber` | $12.5 | $1.25 | | $75 | per 1M tokens | short context rate | `gpt-daybreak-blue-latest` → `gpt-5.6-sol`, `gpt-daybreak-red-latest` → `gpt-5.6-cyber` (aliases, priced as the underlying model). ### GPT-Live sessions (per minute, billed per second) | model | session_duration | unit | notes | |---|---|---|---| | `gpt-live-1` | $0.05 | per minute (billed per second) | | ### Realtime & audio models | model | audio_input | audio_cached_input | audio_output | text_input | text_cached_input | text_output | image_input | image_cached_input | unit | notes | |---|---|---|---|---|---|---|---|---|---|---| | `gpt-realtime-2.1 [audio]` | $32 | $0.4 | $64 | | | | | | per 1M tokens | | | `gpt-realtime-2.1 [text]` | | | | $4 | $0.4 | $24 | | | per 1M tokens | | | `gpt-realtime-2.1 [image]` | | | | | | | $5 | $0.5 | per 1M tokens | | | `gpt-realtime-2.1-mini [audio]` | $10 | $0.3 | $20 | | | | | | per 1M tokens | | | `gpt-realtime-2.1-mini [text]` | | | | $0.6 | $0.06 | $2.4 | | | per 1M tokens | | | `gpt-realtime-2.1-mini [image]` | | | | | | | $0.8 | $0.08 | per 1M tokens | | | `gpt-realtime-2 [audio]` | $32 | $0.4 | $64 | | | | | | per 1M tokens | | | `gpt-realtime-2 [text]` | | | | $4 | $0.4 | $24 | | | per 1M tokens | | | `gpt-realtime-2 [image]` | | | | | | | $5 | $0.5 | per 1M tokens | | | `gpt-realtime-1.5 [audio]` | $32 | $0.4 | $64 | | | | | | per 1M tokens | | | `gpt-realtime-1.5 [text]` | | | | $4 | $0.4 | $16 | | | per 1M tokens | | | `gpt-realtime-1.5 [image]` | | | | | | | $5 | $0.5 | per 1M tokens | | | `gpt-realtime-mini [audio]` | $10 | $0.3 | $20 | | | | | | per 1M tokens | | | `gpt-realtime-mini [text]` | | | | $0.6 | $0.06 | $2.4 | | | per 1M tokens | | | `gpt-realtime-mini [image]` | | | | | | | $0.8 | $0.08 | per 1M tokens | | | `gpt-realtime [audio]` | $32 | $0.4 | $64 | | | | | | per 1M tokens | | | `gpt-realtime [text]` | | | | $4 | $0.4 | $16 | | | per 1M tokens | | | `gpt-realtime [image]` | | | | | | | $5 | $0.5 | per 1M tokens | | | `gpt-audio-1.5 [audio]` | $32 | | $64 | | | | | | per 1M tokens | | | `gpt-audio-1.5 [text]` | | | | $2.5 | | $10 | | | per 1M tokens | | | `gpt-audio-mini [audio]` | $10 | | $20 | | | | | | per 1M tokens | | | `gpt-audio-mini [text]` | | | | $0.6 | | $2.4 | | | per 1M tokens | | | `gpt-audio [audio]` | $32 | | $64 | | | | | | per 1M tokens | | | `gpt-audio [text]` | | | | $2.5 | | $10 | | | per 1M tokens | | | `gpt-4o-mini-tts [audio]` | | | $12 | | | | | | per 1M tokens | | | `gpt-4o-mini-tts [text]` | | | | $0.6 | | | | | per 1M tokens | | | `tts-1 [text]` | | | | $15 | | | | | per 1M characters | | | `tts-1-hd [text]` | | | | $30 | | | | | per 1M characters | | ### Image models — standard | model | image_input | image_cached_input | image_output | text_input | text_cached_input | text_output | unit | notes | |---|---|---|---|---|---|---|---|---| | `gpt-image-2.5-sunburst [image]` | $8 | $2 | $30 | | | | per 1M tokens | | | `gpt-image-2.5-sunburst [text]` | | | | $5 | $1.25 | | per 1M tokens | | | `gpt-image-2.5-flare [image]` | $8 | $2 | $30 | | | | per 1M tokens | | | `gpt-image-2.5-flare [text]` | | | | $5 | $1.25 | | per 1M tokens | | | `gpt-image-2 [image]` | $8 | $2 | $30 | | | | per 1M tokens | | | `gpt-image-2 [text]` | | | | $5 | $1.25 | | per 1M tokens | | | `gpt-image-1.5 [image]` | $8 | $2 | $32 | | | | per 1M tokens | | | `gpt-image-1.5 [text]` | | | | $5 | $1.25 | $10 | per 1M tokens | | | `gpt-image-1-mini [image]` | $2.5 | $0.25 | $8 | | | | per 1M tokens | | | `gpt-image-1-mini [text]` | | | | $2 | $0.2 | | per 1M tokens | | | `gpt-image-1 [image]` | $10 | $2.5 | $40 | | | | per 1M tokens | | | `gpt-image-1 [text]` | | | | $5 | $1.25 | | per 1M tokens | | | `chatgpt-image-latest [image]` | $8 | $2 | $32 | | | | per 1M tokens | | | `chatgpt-image-latest [text]` | | | | $5 | $1.25 | $10 | per 1M tokens | | ### Image models — batch | model | image_input | image_cached_input | image_output | text_input | text_cached_input | text_output | unit | notes | |---|---|---|---|---|---|---|---|---| | `gpt-image-2 [image]` | $4 | $1 | $15 | | | | per 1M tokens | | | `gpt-image-2 [text]` | | | | $2.5 | $0.625 | | per 1M tokens | | | `gpt-image-1.5 [image]` | $4 | $1 | $16 | | | | per 1M tokens | | | `gpt-image-1.5 [text]` | | | | $2.5 | $0.63 | $5 | per 1M tokens | | | `gpt-image-1-mini [image]` | $1.25 | $0.13 | $4 | | | | per 1M tokens | | | `gpt-image-1-mini [text]` | | | | $1 | $0.1 | | per 1M tokens | | | `gpt-image-1 [image]` | $5 | $1.25 | $20 | | | | per 1M tokens | | | `gpt-image-1 [text]` | | | | $2.5 | $0.63 | | per 1M tokens | | | `chatgpt-image-latest [image]` | $4 | $1 | $16 | | | | per 1M tokens | | | `chatgpt-image-latest [text]` | | | | $2.5 | $0.63 | $5 | per 1M tokens | | ### Video (Sora 2) — standard, per second | model | video_output | unit | notes | |---|---|---|---| | `sora-2 [720p]` | $0.1 | per second | | | `sora-2-pro [720p]` | $0.3 | per second | | | `sora-2-pro [1024p]` | $0.5 | per second | | | `sora-2-pro [1080p]` | $0.7 | per second | | ### Video (Sora 2) — batch, per second | model | video_output | unit | notes | |---|---|---|---| | `sora-2 [720p]` | $0.05 | per second | | | `sora-2-pro [720p]` | $0.15 | per second | | | `sora-2-pro [1024p]` | $0.25 | per second | | | `sora-2-pro [1080p]` | $0.35 | per second | | ### Transcription / translation models | model | audio_input | text_output | audio_duration | unit | notes | |---|---|---|---|---|---| | `gpt-realtime-translate` | | | $0.034 | per minute | estimated cost per minute of audio | | `gpt-live-transcribe` | | | $0.017 | per minute | estimated cost per minute of audio | | `gpt-realtime-whisper` | | | $0.017 | per minute | estimated cost per minute of audio | | `gpt-transcribe` | | | $0.0045 | per minute | estimated cost per minute of audio | | `gpt-4o-transcribe` | $2.5 | $10 | $0.006 | per minute | estimated cost per minute of audio | | `gpt-4o-mini-transcribe` | $1.25 | $5 | $0.003 | per minute | estimated cost per minute of audio | | `gpt-4o-transcribe-diarize` | $2.5 | $10 | $0.006 | per minute | estimated cost per minute of audio | | `whisper-1` | | | $0.006 | per minute | estimated cost per minute of audio | ### Specialized models — standard | model | input | cached_input | output | unit | notes | |---|---|---|---|---|---| | `chat-latest` | $5 | $0.5 | $30 | per 1M tokens | | | `gpt-5.3-codex` | $1.75 | $0.175 | $14 | per 1M tokens | | | `gpt-rosalind-research` | $5 | $0.5 | $25 | per 1M tokens | | | `gpt-5-search-api` | $1.25 | $0.125 | $10 | per 1M tokens | | | `text-embedding-3-small` | $0.02 | | | per 1M tokens | | | `text-embedding-3-large` | $0.13 | | | per 1M tokens | | | `text-embedding-ada-002` | $0.1 | | | per 1M tokens | | | `omni-moderation-latest` | $0 | | | per 1M tokens | | ### Specialized models — fast | model | input | cached_input | output | unit | notes | |---|---|---|---|---|---| | `gpt-5.3-codex` | $3.5 | $0.35 | $28 | per 1M tokens | | ### Fine-tuning — standard (platform winding down; see deprecations) | model | fine_tuning_training | fine_tuned_input | fine_tuned_cached_input | fine_tuned_output | unit | notes | |---|---|---|---|---|---|---| | `o4-mini-2025-04-16` | $100 | $2 | $0.5 | $8 | per 1M tokens | data sharing | | `gpt-4.1-2025-04-14` | $25 | $3 | $0.75 | $12 | per 1M tokens | | | `gpt-4.1-mini-2025-04-14` | $5 | $0.8 | $0.2 | $3.2 | per 1M tokens | | | `gpt-4.1-nano-2025-04-14` | $1.5 | $0.2 | $0.05 | $0.8 | per 1M tokens | | | `gpt-4o-2024-08-06` | $25 | $3.75 | $1.875 | $15 | per 1M tokens | | | `gpt-4o-mini-2024-07-18` | $3 | $0.3 | $0.15 | $1.2 | per 1M tokens | | | `gpt-3.5-turbo` | $8 | $3 | | $6 | per 1M tokens | legacy | | `davinci-002` | $6 | $12 | | $12 | per 1M tokens | legacy | | `babbage-002` | $0.4 | $1.6 | | $1.6 | per 1M tokens | legacy | ### Fine-tuning — batch inference | model | fine_tuning_training | fine_tuned_input | fine_tuned_cached_input | fine_tuned_output | unit | notes | |---|---|---|---|---|---|---| | `o4-mini-2025-04-16` | $100 | $1 | $0.25 | $4 | per 1M tokens | data sharing | | `gpt-4.1-2025-04-14` | $25 | $1.5 | $0.5 | $6 | per 1M tokens | | | `gpt-4.1-mini-2025-04-14` | $5 | $0.4 | $0.1 | $1.6 | per 1M tokens | | | `gpt-4.1-nano-2025-04-14` | $1.5 | $0.1 | $0.025 | $0.4 | per 1M tokens | | | `gpt-4o-2024-08-06` | $25 | $2.225 | $0.9 | $12.5 | per 1M tokens | | | `gpt-4o-mini-2024-07-18` | $3 | $0.15 | $0.075 | $0.6 | per 1M tokens | | | `gpt-3.5-turbo` | $8 | $1.5 | | $3 | per 1M tokens | legacy | | `davinci-002` | $6 | $6 | | $6 | per 1M tokens | legacy | | `babbage-002` | $0.4 | $0.8 | | $0.9 | per 1M tokens | legacy | ### Built-in tools | tool | details | price | unit | full text | |---|---|---|---|---| | Web search | Web search (all models) | $10 | per 1k calls | $10.00 / 1k calls + Search content tokens billed at model rates. | | Web search | Image Web search (all models) | $10 | per 1k calls | $10.00 / 1k calls + Search content tokens billed at model rates. | | Web search | Web search preview (reasoning models, including `gpt-5`, `o-series`) | $10 | per 1k calls | $10.00 / 1k calls + Search content tokens billed at model rates. | | Web search | Web search preview (non-reasoning models) | $25 | per 1k calls | $25.00 / 1k calls + Search content tokens are free. | | Containers | Hosted Shell and Code Interpreter | $0.03 | per 20-minute session per container (by size) | 1 GB $0.03, 4 GB $0.12, 16 GB $0.48, 64 GB $1.92 per 20-minute session per container. | | File search | Storage | $0.1 | per GB per day | $0.10 / GB per day (1 GB free) | | File search | Tool call | $2.5 | per 1k calls | $2.50 / 1k calls | | Agent Kit | ChatKit file and image upload storage | $0.1 | per GB per day | $0.10 / GB-day after 1 GB free per account per month | ### Prices only on model pages (not in pricing.md) | model | dimension | price | unit | notes | |---|---|---|---|---| | `chatgpt-4o-latest` | input | $5 | per 1M tokens | | | `chatgpt-4o-latest` | output | $15 | per 1M tokens | | | `codex-mini-latest` | input | $1.5 | per 1M tokens | | | `codex-mini-latest` | cached_input | $0.375 | per 1M tokens | | | `codex-mini-latest` | output | $6 | per 1M tokens | | | `computer-use-preview` | input | $3 | per 1M tokens | | | `computer-use-preview` | output | $12 | per 1M tokens | | | `gpt-4-turbo-preview` | input | $10 | per 1M tokens | | | `gpt-4-turbo-preview` | output | $30 | per 1M tokens | | | `gpt-4-turbo` | input | $10 | per 1M tokens | | | `gpt-4-turbo` | output | $30 | per 1M tokens | | | `gpt-4.5-preview` | input | $75 | per 1M tokens | | | `gpt-4.5-preview` | cached_input | $37.5 | per 1M tokens | | | `gpt-4.5-preview` | output | $150 | per 1M tokens | | | `gpt-4` | input | $30 | per 1M tokens | | | `gpt-4` | output | $60 | per 1M tokens | | | `gpt-4o-audio-preview` | input | $2.5 | per 1M tokens | | | `gpt-4o-audio-preview` | output | $10 | per 1M tokens | | | `gpt-4o-audio-preview` | audio_input | $40 | per 1M tokens | | | `gpt-4o-audio-preview` | audio_output | $80 | per 1M tokens | | | `gpt-4o-mini-audio-preview` | input | $0.15 | per 1M tokens | | | `gpt-4o-mini-audio-preview` | output | $0.6 | per 1M tokens | | | `gpt-4o-mini-audio-preview` | audio_input | $10 | per 1M tokens | | | `gpt-4o-mini-audio-preview` | audio_output | $20 | per 1M tokens | | | `gpt-4o-mini-realtime-preview` | input | $0.6 | per 1M tokens | | | `gpt-4o-mini-realtime-preview` | cached_input | $0.3 | per 1M tokens | | | `gpt-4o-mini-realtime-preview` | output | $2.4 | per 1M tokens | | | `gpt-4o-mini-realtime-preview` | audio_input | $10 | per 1M tokens | | | `gpt-4o-mini-realtime-preview` | audio_cached_input | $0.3 | per 1M tokens | | | `gpt-4o-mini-realtime-preview` | audio_output | $20 | per 1M tokens | | | `gpt-4o-mini-search-preview` | input | $0.15 | per 1M tokens | | | `gpt-4o-mini-search-preview` | output | $0.6 | per 1M tokens | | | `gpt-4o-realtime-preview` | input | $5 | per 1M tokens | | | `gpt-4o-realtime-preview` | cached_input | $2.5 | per 1M tokens | | | `gpt-4o-realtime-preview` | output | $20 | per 1M tokens | | | `gpt-4o-realtime-preview` | audio_input | $40 | per 1M tokens | | | `gpt-4o-realtime-preview` | audio_cached_input | $2.5 | per 1M tokens | | | `gpt-4o-realtime-preview` | audio_output | $80 | per 1M tokens | | | `gpt-4o-search-preview` | input | $2.5 | per 1M tokens | | | `gpt-4o-search-preview` | output | $10 | per 1M tokens | | | `gpt-5-chat-latest` | input | $1.25 | per 1M tokens | | | `gpt-5-chat-latest` | cached_input | $0.125 | per 1M tokens | | | `gpt-5-chat-latest` | output | $10 | per 1M tokens | | | `gpt-5-codex` | input | $1.25 | per 1M tokens | | | `gpt-5-codex` | cached_input | $0.125 | per 1M tokens | | | `gpt-5-codex` | output | $10 | per 1M tokens | | | `gpt-5.1-chat-latest` | input | $1.25 | per 1M tokens | | | `gpt-5.1-chat-latest` | cached_input | $0.125 | per 1M tokens | | | `gpt-5.1-chat-latest` | output | $10 | per 1M tokens | | | `gpt-5.1-codex-max` | input | $1.25 | per 1M tokens | | | `gpt-5.1-codex-max` | cached_input | $0.125 | per 1M tokens | | | `gpt-5.1-codex-max` | output | $10 | per 1M tokens | | | `gpt-5.1-codex-mini` | input | $0.25 | per 1M tokens | | | `gpt-5.1-codex-mini` | cached_input | $0.025 | per 1M tokens | | | `gpt-5.1-codex-mini` | output | $2 | per 1M tokens | | | `gpt-5.1-codex` | input | $1.25 | per 1M tokens | | | `gpt-5.1-codex` | cached_input | $0.125 | per 1M tokens | | | `gpt-5.1-codex` | output | $10 | per 1M tokens | | | `gpt-5.2-chat-latest` | input | $1.75 | per 1M tokens | | | `gpt-5.2-chat-latest` | cached_input | $0.175 | per 1M tokens | | | `gpt-5.2-chat-latest` | output | $14 | per 1M tokens | | | `gpt-5.2-codex` | input | $1.75 | per 1M tokens | | | `gpt-5.2-codex` | cached_input | $0.175 | per 1M tokens | | | `gpt-5.2-codex` | output | $14 | per 1M tokens | | | `gpt-5.3-chat-latest` | input | $1.75 | per 1M tokens | | | `gpt-5.3-chat-latest` | cached_input | $0.175 | per 1M tokens | | | `gpt-5.3-chat-latest` | output | $14 | per 1M tokens | | | `o1-mini` | input | $1.1 | per 1M tokens | | | `o1-mini` | cached_input | $0.55 | per 1M tokens | | | `o1-mini` | output | $4.4 | per 1M tokens | | | `o1-preview` | input | $15 | per 1M tokens | | | `o1-preview` | cached_input | $7.5 | per 1M tokens | | | `o1-preview` | output | $60 | per 1M tokens | | | `o3-deep-research` | input | $10 | per 1M tokens | | | `o3-deep-research` | cached_input | $2.5 | per 1M tokens | | | `o3-deep-research` | output | $40 | per 1M tokens | | | `o4-mini-deep-research` | input | $2 | per 1M tokens | | | `o4-mini-deep-research` | cached_input | $0.5 | per 1M tokens | | | `o4-mini-deep-research` | output | $8 | per 1M tokens | | ### Pricing rules (derived records `rule:*`) | rule | dimension | value | unit | notes | |---|---|---|---|---| | rule:batch | discount | -50 | percent vs standard | Batch API: 50% lower cost, 24h completion window; per-model batch tables on the pricing page | | rule:flex | discount | -50 | percent vs standard | Flex processing (beta): tokens priced at Batch API rates; slower, may return 429 resource_unavailable | | rule:fast | premium | 100 | percent vs standard | Fast mode (ex-Priority processing, renamed 2026-07-30): 2x standard token rates for GPT-5.6 Sol / GPT-6 Astra; service_tier 'fast' or 'priority' | | rule:cache_write | cache_write | 125 | percent of uncached input rate | GPT-5.6 and later: cache writes billed at 1.25x uncached input; earlier models: no cache-write charge | | rule:cache_read_gpt-5.6+ | cached_input | 10 | percent of uncached input rate | GPT-5.6 and later: cache reads at 0.1x uncached input rate | | rule:long_context | multiplier | | x standard | Prompts >272K input tokens: 2x input (and cache) rates, 1.5x output rate for the full request (GPT-5.4 / 5.5 / 5.6 / 6 Astra) | | rule:data_residency | uplift | 10 | percent | Regional processing endpoints (us./eu./ae. …api.openai.com): 10% uplift for models released on or after 2026-03-05 that are eligible for data residency | | rule:container_billing | session | | per minute, 5-minute minimum | Since 2026-06-02 eligible container sessions (Code Interpreter / Hosted Shell) are billed per minute with a 5-minute minimum instead of the full 20-minute rate | | rule:web_search_content_tokens_mini | input | | 8,000 input tokens per call | gpt-4o-mini and gpt-4.1-mini with the non-preview web search tool: search content tokens billed as a fixed block of 8,000 input tokens per call | | rule:pro_mode | output | | standard token rates | reasoning.mode: pro (GPT-5.6) aggregates all model work and bills it at the model's standard token rates (more tokens than standard mode) | | gpt-rosalind-research | billing_start | | date | Billing begins 2026-10-05; cache-write pricing does not apply | | gpt-5.6-sol | promotion | | note | Promotional pricing ($4 in / $20 out) available at least through 2026-11-21 | ## Footnotes copied from the pricing page - Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. [OpenAI models in Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock) are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either `service_tier: "priority"` or `service_tier: "fast"` in your API requests. [Learn more about Fast mode](https://developers.openai.com/api/docs/guides/fast-mode). GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026. - Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. - Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. - Fast mode is unavailable for GPT-6 Astra with EU data residency. Use Standard processing for those requests. See [Fast mode compatibility](https://developers.openai.com/api/docs/guides/fast-mode). Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. - the Daybreak program, these aliases will be updated to point to the latest - [GPT-Live 1](https://developers.openai.com/api/docs/models/gpt-live-1) voice sessions are billed per second, - $10.00 / 1k calls + Search content tokens billed at model rates. - Tokens used for built-in tools are billed at the chosen model's per-token rates. GB refers to binary gigabytes (also known as gibibytes), where 1 GB is 2^30 bytes. Web search content tokens are tokens retrieved from the search index and fed to the model alongside your prompt to generate an answer. For gpt-4o-mini and gpt-4.1-mini with the non-preview web search tool, search content tokens are billed as a fixed block of 8,000 input tokens per call. File search tool call pricing applies to the Responses API only. Container pricing includes Hosted Shell and Code Interpreter. Eligible container sessions will be billed by the minute, with a 5-minute minimum per session. Responses API, Chat Completions API, Realtime API, Batch API, and Assistants API are not priced separately. Tokens are billed at the chosen model's input and output rates. - Billing for `gpt-rosalind-research` begins on October 5, 2026. Cache-write pricing does not apply to this model. Access is limited to approved internal research through the [trusted-access program](https://help.openai.com/en/articles/20001193-gpt-rosalind-for-life-sciences-research). All eligible organizations will continue to get access to the latest GPT-Rosalind models as they’re released. Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. - Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our [Your data](https://developers.openai.com/api/docs/guides/your-data) guide for supported regions and processing details. - OpenAI is winding down the fine-tuning platform. The platform is no longer - Tokens used for model grading in reinforcement fine-tuning are billed at that model's per-token rate. Inference discounts are available if you enable data sharing when creating the fine-tune job. Learn more. ## Caveats - Prices are documentation values as of 2026-09-18; the pricing page states promotional pricing for GPT-5.6 Sol through at least 2026-11-21. - Fine-tuning: platform is winding down (no new orgs since 2026-05-07; job creation ends 2027-01-06); inference on fine-tuned models continues until the base model is deprecated. - Realtime/Live sessions, image tokens and video seconds are billed on different units — check `unit` in every record.