SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
36.0 KB

# OpenAI — Pricing (API Atlas)

Status: DOCUMENTED (prices copied from the official pricing page and model pages; not billed-verified beyond the minimal probes). Machine-readable twin: generated/fragments/pricing/openai-pricing.json.

Sources: https://developers.openai.com/api/docs/pricing · https://developers.openai.com/api/docs/models/ · https://developers.openai.com/api/docs/guides/batch · https://developers.openai.com/api/docs/guides/flex-processing · https://developers.openai.com/api/docs/guides/fast-mode · https://developers.openai.com/api/docs/guides/prompt-caching · https://developers.openai.com/api/docs/guides/your-data

Last verified: 2026-09-18

All prices USD. Token prices are per 1M tokens. cached_input = cache read; cache_write (GPT-5.6+/GPT-6 only) = 1.25× uncached input. Long context = prompts >272K input tokens (2× input/cache, 1.5× output for the whole request).

# Service tiers

Tier service_tier Price rule Notes
Standard default / omitted list price
Batch Batch API (/v1/batches, completion_window: 24h) 50% of standard separate, much higher queue limits; 24h turnaround
Flex flex Batch rates (50%) beta, limited models; slower; may return 429 resource_unavailable; raise client timeout (15 min recommended)
Fast (ex-Priority) fast or priority 2× standard for GPT-5.6 Sol / GPT-6 Astra (per-model table) renamed 2026-07-30; up to 2.5× faster; downgraded requests return service_tier: default and standard rates; no fine-tuned models/embeddings; unavailable for GPT-6 Astra with EU data residency
Ultrafast — — limited preview for GPT-5.6 Sol (announced 2026-08-13), up to 14× faster
Regional processing us./eu./ae. … prefixed domains +10% uplift models released on/after 2026-03-05 that are data-residency eligible

# Text models (Flagship + legacy text)

# Standard — short context (≤272K)

model input cached_input cache_write output unit notes
gpt-6-astra $10 $1 $12.5 $50 per 1M tokens short context rate
gpt-5.6-sol $4 $0.4 $5 $20 per 1M tokens short context rate
gpt-5.6-terra $2 $0.2 $2.5 $12 per 1M tokens short context rate
gpt-5.6-luna $0.2 $0.02 $0.25 $1.2 per 1M tokens short context rate
gpt-5.5 $5 $0.5 $30 per 1M tokens <272K context length; short context rate
gpt-5.5-pro $30 $180 per 1M tokens <272K context length; short context rate
gpt-5.4 $2.5 $0.25 $15 per 1M tokens <272K context length; short context rate
gpt-5.4-mini $0.75 $0.075 $4.5 per 1M tokens short context rate
gpt-5.4-nano $0.2 $0.02 $1.25 per 1M tokens short context rate
gpt-5.4-pro $30 $180 per 1M tokens <272K context length; short context rate
gpt-5.2 $1.75 $0.175 $14 per 1M tokens short context rate
gpt-5.2-pro $21 $168 per 1M tokens short context rate
gpt-5.1 $1.25 $0.125 $10 per 1M tokens short context rate
gpt-5 $1.25 $0.125 $10 per 1M tokens short context rate
gpt-5-mini $0.25 $0.025 $2 per 1M tokens short context rate
gpt-5-nano $0.05 $0.005 $0.4 per 1M tokens short context rate
gpt-5-pro $15 $120 per 1M tokens short context rate
gpt-4.1 $2 $0.5 $8 per 1M tokens short context rate
gpt-4.1-mini $0.4 $0.1 $1.6 per 1M tokens short context rate
gpt-4.1-nano $0.1 $0.025 $0.4 per 1M tokens short context rate
gpt-4o $2.5 $1.25 $10 per 1M tokens short context rate
gpt-4o-2024-05-13 $5 $15 per 1M tokens short context rate
gpt-4o-mini $0.15 $0.075 $0.6 per 1M tokens short context rate
o1 $15 $7.5 $60 per 1M tokens short context rate
o1-pro $150 $600 per 1M tokens short context rate
o3-pro $20 $80 per 1M tokens short context rate
o3 $2 $0.5 $8 per 1M tokens short context rate
o4-mini $1.1 $0.275 $4.4 per 1M tokens short context rate
o3-mini $1.1 $0.55 $4.4 per 1M tokens short context rate
gpt-4-turbo-2024-04-09 $10 $30 per 1M tokens short context rate
gpt-4-0613 $30 $60 per 1M tokens short context rate
gpt-3.5-turbo $0.5 $1.5 per 1M tokens short context rate
gpt-3.5-turbo-0125 $0.5 $1.5 per 1M tokens short context rate
gpt-3.5-turbo-1106 $1 $2 per 1M tokens short context rate
gpt-3.5-turbo-instruct $1.5 $2 per 1M tokens short context rate
davinci-002 $2 $2 per 1M tokens short context rate
babbage-002 $0.4 $0.4 per 1M tokens short context rate

# Standard — long context (>272K)

model input cached_input cache_write output unit notes
gpt-6-astra $20 $2 $25 $75 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-sol $8 $0.8 $10 $30 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-terra $4 $0.4 $5 $18 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-luna $0.4 $0.04 $0.5 $1.8 per 1M tokens long context (>272K input tokens) rate
gpt-5.5 $10 $1 $45 per 1M tokens <272K context length; long context (>272K input tokens) rate
gpt-5.5-pro $60 $270 per 1M tokens <272K context length; long context (>272K input tokens) rate
gpt-5.4 $5 $0.5 $22.5 per 1M tokens <272K context length; long context (>272K input tokens) rate
gpt-5.4-pro $60 $270 per 1M tokens <272K context length; long context (>272K input tokens) rate

# Batch — short context (≤272K)

model input cached_input cache_write output unit notes
gpt-6-astra $5 $0.5 $6.25 $25 per 1M tokens short context rate
gpt-5.6-sol $2 $0.2 $2.5 $10 per 1M tokens short context rate
gpt-5.6-terra $1 $0.1 $1.25 $6 per 1M tokens short context rate
gpt-5.6-luna $0.1 $0.01 $0.125 $0.6 per 1M tokens short context rate
gpt-5.5 $2.5 $0.25 $15 per 1M tokens <272K context length; short context rate
gpt-5.5-pro $15 $90 per 1M tokens <272K context length; short context rate
gpt-5.4 $1.25 $0.13 $7.5 per 1M tokens <272K context length; short context rate
gpt-5.4-mini $0.375 $0.0375 $2.25 per 1M tokens short context rate
gpt-5.4-nano $0.1 $0.01 $0.625 per 1M tokens short context rate
gpt-5.4-pro $15 $90 per 1M tokens <272K context length; short context rate
gpt-5.2 $0.875 $0.0875 $7 per 1M tokens short context rate
gpt-5.2-pro $10.5 $84 per 1M tokens short context rate
gpt-5.1 $0.625 $0.0625 $5 per 1M tokens short context rate
gpt-5 $0.625 $0.0625 $5 per 1M tokens short context rate
gpt-5-mini $0.125 $0.0125 $1 per 1M tokens short context rate
gpt-5-nano $0.025 $0.0025 $0.2 per 1M tokens short context rate
gpt-5-pro $7.5 $60 per 1M tokens short context rate
gpt-4.1 $1 $4 per 1M tokens short context rate
gpt-4.1-mini $0.2 $0.8 per 1M tokens short context rate
gpt-4.1-nano $0.05 $0.2 per 1M tokens short context rate
gpt-4o $1.25 $5 per 1M tokens short context rate
gpt-4o-2024-05-13 $2.5 $7.5 per 1M tokens short context rate
gpt-4o-mini $0.075 $0.3 per 1M tokens short context rate
o1 $7.5 $30 per 1M tokens short context rate
o1-pro $75 $300 per 1M tokens short context rate
o3-pro $10 $40 per 1M tokens short context rate
o3 $1 $4 per 1M tokens short context rate
o4-mini $0.55 $2.2 per 1M tokens short context rate
o3-mini $0.55 $2.2 per 1M tokens short context rate
gpt-4-turbo-2024-04-09 $5 $15 per 1M tokens short context rate
gpt-4-0613 $15 $30 per 1M tokens short context rate
gpt-3.5-turbo-0125 $0.25 $0.75 per 1M tokens short context rate
gpt-3.5-turbo-1106 $1 $2 per 1M tokens short context rate
davinci-002 $1 $1 per 1M tokens short context rate
babbage-002 $0.2 $0.2 per 1M tokens short context rate

# Batch — long context (>272K)

model input cached_input cache_write output unit notes
gpt-6-astra $10 $1 $12.5 $37.5 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-sol $4 $0.4 $5 $15 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-terra $2 $0.2 $2.5 $9 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-luna $0.2 $0.02 $0.25 $0.9 per 1M tokens long context (>272K input tokens) rate
gpt-5.5 $5 $0.5 $22.5 per 1M tokens <272K context length; long context (>272K input tokens) rate
gpt-5.4 $2.5 $0.25 $11.25 per 1M tokens <272K context length; long context (>272K input tokens) rate
gpt-5.4-pro $30 $135 per 1M tokens <272K context length; long context (>272K input tokens) rate

# Flex — short context (≤272K)

model input cached_input cache_write output unit notes
gpt-6-astra $5 $0.5 $6.25 $25 per 1M tokens short context rate
gpt-5.6-sol $2 $0.2 $2.5 $10 per 1M tokens short context rate
gpt-5.6-terra $1 $0.1 $1.25 $6 per 1M tokens short context rate
gpt-5.6-luna $0.1 $0.01 $0.125 $0.6 per 1M tokens short context rate
gpt-5.5 $2.5 $0.25 $15 per 1M tokens <272K context length; short context rate
gpt-5.5-pro $15 $90 per 1M tokens <272K context length; short context rate
gpt-5.4 $1.25 $0.13 $7.5 per 1M tokens <272K context length; short context rate
gpt-5.4-mini $0.375 $0.0375 $2.25 per 1M tokens short context rate
gpt-5.4-nano $0.1 $0.01 $0.625 per 1M tokens short context rate
gpt-5.4-pro $15 $90 per 1M tokens <272K context length; short context rate
gpt-5.2 $0.875 $0.0875 $7 per 1M tokens short context rate
gpt-5.1 $0.625 $0.0625 $5 per 1M tokens short context rate
gpt-5 $0.625 $0.0625 $5 per 1M tokens short context rate
gpt-5-mini $0.125 $0.0125 $1 per 1M tokens short context rate
gpt-5-nano $0.025 $0.0025 $0.2 per 1M tokens short context rate
o3 $1 $0.25 $4 per 1M tokens short context rate
o4-mini $0.55 $0.138 $2.2 per 1M tokens short context rate

# Flex — long context (>272K)

model input cached_input cache_write output unit notes
gpt-6-astra $10 $1 $12.5 $37.5 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-sol $4 $0.4 $5 $15 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-terra $2 $0.2 $2.5 $9 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-luna $0.2 $0.02 $0.25 $0.9 per 1M tokens long context (>272K input tokens) rate
gpt-5.5 $5 $0.5 $22.5 per 1M tokens <272K context length; long context (>272K input tokens) rate
gpt-5.4 $2.5 $0.25 $11.25 per 1M tokens <272K context length; long context (>272K input tokens) rate
gpt-5.4-pro $30 $135 per 1M tokens <272K context length; long context (>272K input tokens) rate

# Fast — short context (≤272K)

model input cached_input cache_write output unit notes
gpt-6-astra $20 $2 $25 $100 per 1M tokens short context rate
gpt-5.6-sol $8 $0.8 $10 $40 per 1M tokens short context rate
gpt-5.6-terra $4 $0.4 $5 $24 per 1M tokens short context rate
gpt-5.6-luna $0.4 $0.04 $0.5 $2.4 per 1M tokens short context rate
gpt-5.5 $12.5 $1.25 $75 per 1M tokens <272K context length; short context rate
gpt-5.4 $5 $0.5 $30 per 1M tokens <272K context length; short context rate
gpt-5.4-mini $1.5 $0.15 $9 per 1M tokens short context rate
gpt-5.2 $3.5 $0.35 $28 per 1M tokens short context rate
gpt-5.1 $2.5 $0.25 $20 per 1M tokens short context rate
gpt-5 $2.5 $0.25 $20 per 1M tokens short context rate
gpt-5-mini $0.45 $0.045 $3.6 per 1M tokens short context rate
gpt-4.1 $3.5 $0.875 $14 per 1M tokens short context rate
gpt-4.1-mini $0.7 $0.175 $2.8 per 1M tokens short context rate
gpt-4.1-nano $0.2 $0.05 $0.8 per 1M tokens short context rate
gpt-4o $4.25 $2.125 $17 per 1M tokens short context rate
gpt-4o-2024-05-13 $8.75 $26.25 per 1M tokens short context rate
gpt-4o-mini $0.25 $0.125 $1 per 1M tokens short context rate
o3 $3.5 $0.875 $14 per 1M tokens short context rate
o4-mini $2 $0.5 $8 per 1M tokens short context rate

# Fast — long context (>272K)

model input cached_input cache_write output unit notes
gpt-6-astra $40 $4 $50 $150 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-sol $16 $1.6 $20 $60 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-terra $8 $0.8 $10 $36 per 1M tokens long context (>272K input tokens) rate
gpt-5.6-luna $0.8 $0.08 $1 $3.6 per 1M tokens long context (>272K input tokens) rate

# Cyber / Daybreak models (standard)

model input cached_input cache_write output unit notes
gpt-5.6-sol $4 $0.4 $5 $20 per 1M tokens short context rate
gpt-5.6-cyber $12.5 $1.25 $15.625 $75 per 1M tokens short context rate
gpt-5.5-cyber $12.5 $1.25 $75 per 1M tokens short context rate

gpt-daybreak-blue-latest → gpt-5.6-sol, gpt-daybreak-red-latest → gpt-5.6-cyber (aliases, priced as the underlying model).

# GPT-Live sessions (per minute, billed per second)

model session_duration unit notes
gpt-live-1 $0.05 per minute (billed per second)

# Realtime & audio models

model audio_input audio_cached_input audio_output text_input text_cached_input text_output image_input image_cached_input unit notes
gpt-realtime-2.1 [audio] $32 $0.4 $64 per 1M tokens
gpt-realtime-2.1 [text] $4 $0.4 $24 per 1M tokens
gpt-realtime-2.1 [image] $5 $0.5 per 1M tokens
gpt-realtime-2.1-mini [audio] $10 $0.3 $20 per 1M tokens
gpt-realtime-2.1-mini [text] $0.6 $0.06 $2.4 per 1M tokens
gpt-realtime-2.1-mini [image] $0.8 $0.08 per 1M tokens
gpt-realtime-2 [audio] $32 $0.4 $64 per 1M tokens
gpt-realtime-2 [text] $4 $0.4 $24 per 1M tokens
gpt-realtime-2 [image] $5 $0.5 per 1M tokens
gpt-realtime-1.5 [audio] $32 $0.4 $64 per 1M tokens
gpt-realtime-1.5 [text] $4 $0.4 $16 per 1M tokens
gpt-realtime-1.5 [image] $5 $0.5 per 1M tokens
gpt-realtime-mini [audio] $10 $0.3 $20 per 1M tokens
gpt-realtime-mini [text] $0.6 $0.06 $2.4 per 1M tokens
gpt-realtime-mini [image] $0.8 $0.08 per 1M tokens
gpt-realtime [audio] $32 $0.4 $64 per 1M tokens
gpt-realtime [text] $4 $0.4 $16 per 1M tokens
gpt-realtime [image] $5 $0.5 per 1M tokens
gpt-audio-1.5 [audio] $32 $64 per 1M tokens
gpt-audio-1.5 [text] $2.5 $10 per 1M tokens
gpt-audio-mini [audio] $10 $20 per 1M tokens
gpt-audio-mini [text] $0.6 $2.4 per 1M tokens
gpt-audio [audio] $32 $64 per 1M tokens
gpt-audio [text] $2.5 $10 per 1M tokens
gpt-4o-mini-tts [audio] $12 per 1M tokens
gpt-4o-mini-tts [text] $0.6 per 1M tokens
tts-1 [text] $15 per 1M characters
tts-1-hd [text] $30 per 1M characters

# Image models — standard

model image_input image_cached_input image_output text_input text_cached_input text_output unit notes
gpt-image-2.5-sunburst [image] $8 $2 $30 per 1M tokens
gpt-image-2.5-sunburst [text] $5 $1.25 per 1M tokens
gpt-image-2.5-flare [image] $8 $2 $30 per 1M tokens
gpt-image-2.5-flare [text] $5 $1.25 per 1M tokens
gpt-image-2 [image] $8 $2 $30 per 1M tokens
gpt-image-2 [text] $5 $1.25 per 1M tokens
gpt-image-1.5 [image] $8 $2 $32 per 1M tokens
gpt-image-1.5 [text] $5 $1.25 $10 per 1M tokens
gpt-image-1-mini [image] $2.5 $0.25 $8 per 1M tokens
gpt-image-1-mini [text] $2 $0.2 per 1M tokens
gpt-image-1 [image] $10 $2.5 $40 per 1M tokens
gpt-image-1 [text] $5 $1.25 per 1M tokens
chatgpt-image-latest [image] $8 $2 $32 per 1M tokens
chatgpt-image-latest [text] $5 $1.25 $10 per 1M tokens

# Image models — batch

model image_input image_cached_input image_output text_input text_cached_input text_output unit notes
gpt-image-2 [image] $4 $1 $15 per 1M tokens
gpt-image-2 [text] $2.5 $0.625 per 1M tokens
gpt-image-1.5 [image] $4 $1 $16 per 1M tokens
gpt-image-1.5 [text] $2.5 $0.63 $5 per 1M tokens
gpt-image-1-mini [image] $1.25 $0.13 $4 per 1M tokens
gpt-image-1-mini [text] $1 $0.1 per 1M tokens
gpt-image-1 [image] $5 $1.25 $20 per 1M tokens
gpt-image-1 [text] $2.5 $0.63 per 1M tokens
chatgpt-image-latest [image] $4 $1 $16 per 1M tokens
chatgpt-image-latest [text] $2.5 $0.63 $5 per 1M tokens

# Video (Sora 2) — standard, per second

model video_output unit notes
sora-2 [720p] $0.1 per second
sora-2-pro [720p] $0.3 per second
sora-2-pro [1024p] $0.5 per second
sora-2-pro [1080p] $0.7 per second

# Video (Sora 2) — batch, per second

model video_output unit notes
sora-2 [720p] $0.05 per second
sora-2-pro [720p] $0.15 per second
sora-2-pro [1024p] $0.25 per second
sora-2-pro [1080p] $0.35 per second

# Transcription / translation models

model audio_input text_output audio_duration unit notes
gpt-realtime-translate $0.034 per minute estimated cost per minute of audio
gpt-live-transcribe $0.017 per minute estimated cost per minute of audio
gpt-realtime-whisper $0.017 per minute estimated cost per minute of audio
gpt-transcribe $0.0045 per minute estimated cost per minute of audio
gpt-4o-transcribe $2.5 $10 $0.006 per minute estimated cost per minute of audio
gpt-4o-mini-transcribe $1.25 $5 $0.003 per minute estimated cost per minute of audio
gpt-4o-transcribe-diarize $2.5 $10 $0.006 per minute estimated cost per minute of audio
whisper-1 $0.006 per minute estimated cost per minute of audio

# Specialized models — standard

model input cached_input output unit notes
chat-latest $5 $0.5 $30 per 1M tokens
gpt-5.3-codex $1.75 $0.175 $14 per 1M tokens
gpt-rosalind-research $5 $0.5 $25 per 1M tokens
gpt-5-search-api $1.25 $0.125 $10 per 1M tokens
text-embedding-3-small $0.02 per 1M tokens
text-embedding-3-large $0.13 per 1M tokens
text-embedding-ada-002 $0.1 per 1M tokens
omni-moderation-latest $0 per 1M tokens

# Specialized models — fast

model input cached_input output unit notes
gpt-5.3-codex $3.5 $0.35 $28 per 1M tokens

# Fine-tuning — standard (platform winding down; see deprecations)

model fine_tuning_training fine_tuned_input fine_tuned_cached_input fine_tuned_output unit notes
o4-mini-2025-04-16 $100 $2 $0.5 $8 per 1M tokens data sharing
gpt-4.1-2025-04-14 $25 $3 $0.75 $12 per 1M tokens
gpt-4.1-mini-2025-04-14 $5 $0.8 $0.2 $3.2 per 1M tokens
gpt-4.1-nano-2025-04-14 $1.5 $0.2 $0.05 $0.8 per 1M tokens
gpt-4o-2024-08-06 $25 $3.75 $1.875 $15 per 1M tokens
gpt-4o-mini-2024-07-18 $3 $0.3 $0.15 $1.2 per 1M tokens
gpt-3.5-turbo $8 $3 $6 per 1M tokens legacy
davinci-002 $6 $12 $12 per 1M tokens legacy
babbage-002 $0.4 $1.6 $1.6 per 1M tokens legacy

# Fine-tuning — batch inference

model fine_tuning_training fine_tuned_input fine_tuned_cached_input fine_tuned_output unit notes
o4-mini-2025-04-16 $100 $1 $0.25 $4 per 1M tokens data sharing
gpt-4.1-2025-04-14 $25 $1.5 $0.5 $6 per 1M tokens
gpt-4.1-mini-2025-04-14 $5 $0.4 $0.1 $1.6 per 1M tokens
gpt-4.1-nano-2025-04-14 $1.5 $0.1 $0.025 $0.4 per 1M tokens
gpt-4o-2024-08-06 $25 $2.225 $0.9 $12.5 per 1M tokens
gpt-4o-mini-2024-07-18 $3 $0.15 $0.075 $0.6 per 1M tokens
gpt-3.5-turbo $8 $1.5 $3 per 1M tokens legacy
davinci-002 $6 $6 $6 per 1M tokens legacy
babbage-002 $0.4 $0.8 $0.9 per 1M tokens legacy

# Built-in tools

tool details price unit full text
Web search Web search (all models) $10 per 1k calls $10.00 / 1k calls + Search content tokens billed at model rates.
Web search Image Web search (all models) $10 per 1k calls $10.00 / 1k calls + Search content tokens billed at model rates.
Web search Web search preview (reasoning models, including gpt-5, o-series) $10 per 1k calls $10.00 / 1k calls + Search content tokens billed at model rates.
Web search Web search preview (non-reasoning models) $25 per 1k calls $25.00 / 1k calls + Search content tokens are free.
Containers Hosted Shell and Code Interpreter $0.03 per 20-minute session per container (by size) 1 GB $0.03, 4 GB $0.12, 16 GB $0.48, 64 GB $1.92 per 20-minute session per container.
File search Storage $0.1 per GB per day $0.10 / GB per day (1 GB free)
File search Tool call $2.5 per 1k calls $2.50 / 1k calls
Agent Kit ChatKit file and image upload storage $0.1 per GB per day $0.10 / GB-day after 1 GB free per account per month

# Prices only on model pages (not in pricing.md)

model dimension price unit notes
chatgpt-4o-latest input $5 per 1M tokens
chatgpt-4o-latest output $15 per 1M tokens
codex-mini-latest input $1.5 per 1M tokens
codex-mini-latest cached_input $0.375 per 1M tokens
codex-mini-latest output $6 per 1M tokens
computer-use-preview input $3 per 1M tokens
computer-use-preview output $12 per 1M tokens
gpt-4-turbo-preview input $10 per 1M tokens
gpt-4-turbo-preview output $30 per 1M tokens
gpt-4-turbo input $10 per 1M tokens
gpt-4-turbo output $30 per 1M tokens
gpt-4.5-preview input $75 per 1M tokens
gpt-4.5-preview cached_input $37.5 per 1M tokens
gpt-4.5-preview output $150 per 1M tokens
gpt-4 input $30 per 1M tokens
gpt-4 output $60 per 1M tokens
gpt-4o-audio-preview input $2.5 per 1M tokens
gpt-4o-audio-preview output $10 per 1M tokens
gpt-4o-audio-preview audio_input $40 per 1M tokens
gpt-4o-audio-preview audio_output $80 per 1M tokens
gpt-4o-mini-audio-preview input $0.15 per 1M tokens
gpt-4o-mini-audio-preview output $0.6 per 1M tokens
gpt-4o-mini-audio-preview audio_input $10 per 1M tokens
gpt-4o-mini-audio-preview audio_output $20 per 1M tokens
gpt-4o-mini-realtime-preview input $0.6 per 1M tokens
gpt-4o-mini-realtime-preview cached_input $0.3 per 1M tokens
gpt-4o-mini-realtime-preview output $2.4 per 1M tokens
gpt-4o-mini-realtime-preview audio_input $10 per 1M tokens
gpt-4o-mini-realtime-preview audio_cached_input $0.3 per 1M tokens
gpt-4o-mini-realtime-preview audio_output $20 per 1M tokens
gpt-4o-mini-search-preview input $0.15 per 1M tokens
gpt-4o-mini-search-preview output $0.6 per 1M tokens
gpt-4o-realtime-preview input $5 per 1M tokens
gpt-4o-realtime-preview cached_input $2.5 per 1M tokens
gpt-4o-realtime-preview output $20 per 1M tokens
gpt-4o-realtime-preview audio_input $40 per 1M tokens
gpt-4o-realtime-preview audio_cached_input $2.5 per 1M tokens
gpt-4o-realtime-preview audio_output $80 per 1M tokens
gpt-4o-search-preview input $2.5 per 1M tokens
gpt-4o-search-preview output $10 per 1M tokens
gpt-5-chat-latest input $1.25 per 1M tokens
gpt-5-chat-latest cached_input $0.125 per 1M tokens
gpt-5-chat-latest output $10 per 1M tokens
gpt-5-codex input $1.25 per 1M tokens
gpt-5-codex cached_input $0.125 per 1M tokens
gpt-5-codex output $10 per 1M tokens
gpt-5.1-chat-latest input $1.25 per 1M tokens
gpt-5.1-chat-latest cached_input $0.125 per 1M tokens
gpt-5.1-chat-latest output $10 per 1M tokens
gpt-5.1-codex-max input $1.25 per 1M tokens
gpt-5.1-codex-max cached_input $0.125 per 1M tokens
gpt-5.1-codex-max output $10 per 1M tokens
gpt-5.1-codex-mini input $0.25 per 1M tokens
gpt-5.1-codex-mini cached_input $0.025 per 1M tokens
gpt-5.1-codex-mini output $2 per 1M tokens
gpt-5.1-codex input $1.25 per 1M tokens
gpt-5.1-codex cached_input $0.125 per 1M tokens
gpt-5.1-codex output $10 per 1M tokens
gpt-5.2-chat-latest input $1.75 per 1M tokens
gpt-5.2-chat-latest cached_input $0.175 per 1M tokens
gpt-5.2-chat-latest output $14 per 1M tokens
gpt-5.2-codex input $1.75 per 1M tokens
gpt-5.2-codex cached_input $0.175 per 1M tokens
gpt-5.2-codex output $14 per 1M tokens
gpt-5.3-chat-latest input $1.75 per 1M tokens
gpt-5.3-chat-latest cached_input $0.175 per 1M tokens
gpt-5.3-chat-latest output $14 per 1M tokens
o1-mini input $1.1 per 1M tokens
o1-mini cached_input $0.55 per 1M tokens
o1-mini output $4.4 per 1M tokens
o1-preview input $15 per 1M tokens
o1-preview cached_input $7.5 per 1M tokens
o1-preview output $60 per 1M tokens
o3-deep-research input $10 per 1M tokens
o3-deep-research cached_input $2.5 per 1M tokens
o3-deep-research output $40 per 1M tokens
o4-mini-deep-research input $2 per 1M tokens
o4-mini-deep-research cached_input $0.5 per 1M tokens
o4-mini-deep-research output $8 per 1M tokens

# Pricing rules (derived records rule:*)

rule dimension value unit notes
rule:batch discount -50 percent vs standard Batch API: 50% lower cost, 24h completion window; per-model batch tables on the pricing page
rule:flex discount -50 percent vs standard Flex processing (beta): tokens priced at Batch API rates; slower, may return 429 resource_unavailable
rule:fast premium 100 percent vs standard Fast mode (ex-Priority processing, renamed 2026-07-30): 2x standard token rates for GPT-5.6 Sol / GPT-6 Astra; service_tier 'fast' or 'priority'
rule:cache_write cache_write 125 percent of uncached input rate GPT-5.6 and later: cache writes billed at 1.25x uncached input; earlier models: no cache-write charge
rule:cache_read_gpt-5.6+ cached_input 10 percent of uncached input rate GPT-5.6 and later: cache reads at 0.1x uncached input rate
rule:long_context multiplier x standard Prompts >272K input tokens: 2x input (and cache) rates, 1.5x output rate for the full request (GPT-5.4 / 5.5 / 5.6 / 6 Astra)
rule:data_residency uplift 10 percent Regional processing endpoints (us./eu./ae. …api.openai.com): 10% uplift for models released on or after 2026-03-05 that are eligible for data residency
rule:container_billing session per minute, 5-minute minimum Since 2026-06-02 eligible container sessions (Code Interpreter / Hosted Shell) are billed per minute with a 5-minute minimum instead of the full 20-minute rate
rule:web_search_content_tokens_mini input 8,000 input tokens per call gpt-4o-mini and gpt-4.1-mini with the non-preview web search tool: search content tokens billed as a fixed block of 8,000 input tokens per call
rule:pro_mode output standard token rates reasoning.mode: pro (GPT-5.6) aggregates all model work and bills it at the model's standard token rates (more tokens than standard mode)
gpt-rosalind-research billing_start date Billing begins 2026-10-05; cache-write pricing does not apply
gpt-5.6-sol promotion note Promotional pricing ($4 in / $20 out) available at least through 2026-11-21

# Footnotes copied from the pricing page

  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our Your data guide for supported regions and processing details. OpenAI models in Amazon Bedrock are billed through AWS. Bedrock pricing in commercial regions matches OpenAI direct pricing for equivalent services. Priority processing was renamed Fast mode on July 30, 2026. You can use either service_tier: "priority" or service_tier: "fast" in your API requests. Learn more about Fast mode. GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our Your data guide for supported regions and processing details.
  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our Your data guide for supported regions and processing details.
  • Fast mode is unavailable for GPT-6 Astra with EU data residency. Use Standard processing for those requests. See Fast mode compatibility. Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our Your data guide for supported regions and processing details.
  • the Daybreak program, these aliases will be updated to point to the latest
  • GPT-Live 1 voice sessions are billed per second,
  • $10.00 / 1k calls + Search content tokens billed at model rates.
  • Tokens used for built-in tools are billed at the chosen model's per-token rates. GB refers to binary gigabytes (also known as gibibytes), where 1 GB is 2^30 bytes. Web search content tokens are tokens retrieved from the search index and fed to the model alongside your prompt to generate an answer. For gpt-4o-mini and gpt-4.1-mini with the non-preview web search tool, search content tokens are billed as a fixed block of 8,000 input tokens per call. File search tool call pricing applies to the Responses API only. Container pricing includes Hosted Shell and Code Interpreter. Eligible container sessions will be billed by the minute, with a 5-minute minimum per session. Responses API, Chat Completions API, Realtime API, Batch API, and Assistants API are not priced separately. Tokens are billed at the chosen model's input and output rates.
  • Billing for gpt-rosalind-research begins on October 5, 2026. Cache-write pricing does not apply to this model. Access is limited to approved internal research through the trusted-access program. All eligible organizations will continue to get access to the latest GPT-Rosalind models as they’re released. Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our Your data guide for supported regions and processing details.
  • Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our Your data guide for supported regions and processing details.
  • OpenAI is winding down the fine-tuning platform. The platform is no longer
  • Tokens used for model grading in reinforcement fine-tuning are billed at that model's per-token rate. Inference discounts are available if you enable data sharing when creating the fine-tune job. Learn more.

# Caveats

  • Prices are documentation values as of 2026-09-18; the pricing page states promotional pricing for GPT-5.6 Sol through at least 2026-11-21.
  • Fine-tuning: platform is winding down (no new orgs since 2026-05-07; job creation ends 2027-01-06); inference on fine-tuned models continues until the base model is deprecated.
  • Realtime/Live sessions, image tokens and video seconds are billed on different units — check unit in every record.