OpenAI Audio API — text-to-speech, speech-to-text, translation, custom voices, audio in Chat Completions
Status: LIVE_VERIFIED (/v1/audio/speech incl. SSE, /v1/audio/transcriptions ×5 variants incl. streaming/diarization/logprobs, /v1/audio/translations) · DOCUMENTED/UNVERIFIED for voices & voice consents (list returned 404 for our key — likely access-gated) · LIVE_DISCOVERED for Chat Completions audio (400 shape captured).
Sources: https://developers.openai.com/api/docs/guides/audio · …/text-to-speech · …/speech-to-text · …/transcription · …/custom-voices · …/audio-chat-completions · https://developers.openai.com/api/reference/resources/audio/subresources/speech/methods/create · …/transcriptions/methods/create · …/transcriptions/streaming-events · …/translations/methods/create · …/voices/methods/create · …/voice_consents/methods/{create,list,retrieve,update,delete} · model pages (gpt-4o-mini-tts, tts-1, tts-1-hd, gpt-transcribe, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize, whisper-1, gpt-audio-1.5, gpt-audio, gpt-audio-mini, gpt-4o-(mini-)audio-preview) · openapi-master.yaml
Last verified: 2026-09-18
Twins: generated/fragments/endpoints/openai-realtime-live-audio.json (api_family audio, chat-audio), generated/fragments/parameters/openai-audio.json (53), generated/fragments/streaming-events/openai-audio-transcription.json (5).
1. Workflow chooser (official)
| Need | Use |
|---|---|
| Conversational voice app | GPT-Live (live.md) or Realtime (realtime.md) |
| Transcript of a recorded file | POST /v1/audio/transcriptions — start with gpt-transcribe |
| Live captions (mic/call) | Realtime transcription session — gpt-live-transcribe |
| Speaker labels | gpt-4o-transcribe-diarize + response_format: diarized_json (files only) |
| Word timestamps / SRT / VTT | whisper-1 + timestamp_granularities[] / response_format srt|vtt |
| Translate a recording to English | POST /v1/audio/translations (whisper-1 only) |
| Narration / generated speech | POST /v1/audio/speech — gpt-4o-mini-tts |
| Audio in an existing chat app | Chat Completions with gpt-audio-* and modalities |
2. Model matrix (from model pages, 2026-09-18 — pricing owned by the models agent; cite docs/models/*)
| Model | Task | In → Out | Endpoints (Supported) | Price (model page) | Notes |
|---|---|---|---|---|---|
gpt-4o-mini-tts (-2025-03-20, default -2025-12-15) |
TTS | text → audio | /v1/audio/speech |
$0.60/1M text in, $12/1M audio out | instructions supported, ≤2000 input tokens, SSE streaming, 13 voices, custom voices |
tts-1 |
TTS (speed) | text → audio | /v1/audio/speech |
$15 / 1M characters | 9 voices, no instructions, no SSE |
tts-1-hd |
TTS (quality) | text → audio | /v1/audio/speech |
$30 / 1M characters | idem |
gpt-transcribe |
STT (recommended) | audio(+text) → text | /v1/audio/transcriptions, realtime transcription sessions |
$0.0045 / min | prompt, keywords[], languages[], detected languages in output, streaming |
gpt-4o-transcribe |
STT | audio → text | /v1/audio/transcriptions, /v1/realtime |
$2.5 in / $10 out per 1M tokens | include[]=logprobs, streaming |
gpt-4o-mini-transcribe (-2025-03-20, default -2025-12-15) |
STT | audio → text | idem | $1.25 / $5 per 1M | LIVE: 1.1 s clip = 11 audio input tokens, 4 output |
gpt-4o-transcribe-diarize |
STT + speakers | audio → text | /v1/audio/transcriptions |
$2.5 / $10 per 1M | diarized_json, known_speaker_names[]/references[], no prompt; chunking_strategy for >30 s |
whisper-1 |
STT/translation | audio → text | /v1/audio/transcriptions, /v1/audio/translations |
$0.006 / min | verbose_json, timestamp_granularities, srt/vtt, 98 languages, 224-token prompt, no streaming |
gpt-live-transcribe |
live STT | audio → text | realtime transcription sessions | $0.017 / min | Realtime only (see realtime.md §9) |
gpt-realtime-whisper |
live STT | audio → text | realtime transcription sessions | $0.017 / min | VAD must be null |
gpt-audio-1.5 / gpt-audio (-2025-08-28) |
chat audio | text+audio → text+audio | /v1/chat/completions |
text $2.5/$10, audio $32/$64 per 1M | 128k ctx, 16,384 max out |
gpt-audio-mini (-2025-10-06, default -2025-12-15) |
chat audio | idem | /v1/chat/completions |
text $0.6/$2.4 per 1M | requires audio in input or output |
gpt-4o-audio-preview / gpt-4o-mini-audio-preview |
chat audio (preview) | idem | /v1/chat/completions |
$2.5/$10 text, $40/$80 audio · $0.15/$0.6 text, $10/$20 audio | PREVIEW |
3. Text-to-speech — POST /v1/audio/speech (LIVE_VERIFIED)
| Param | Type / enum | Default | Notes |
|---|---|---|---|
model |
gpt-4o-mini-tts, gpt-4o-mini-tts-2025-12-15, tts-1, tts-1-hd |
— | required |
input |
string ≤ 4096 chars | — | required |
voice |
alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin, cedar or {id: "voice_…"} |
— | required; tts-1/-hd only alloy, ash, coral, echo, fable, onyx, nova, sage, shimmer; best quality marin/cedar; voices optimized for English |
instructions |
string | — | tone/accent/emotion/speed/whisper…; not for tts-1/-hd |
response_format |
mp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE raw) |
mp3 | wav/pcm fastest to first byte |
speed |
0.25–4.0 | 1.0 | |
stream_format |
audio (chunked bytes, default) | sse |
audio | sse emits speech.audio.delta {audio: b64} … speech.audio.done {usage} then data: [DONE]; not for tts-1/-hd |
Observed: {"model":"gpt-4o-mini-tts","input":"OK","voice":"alloy","response_format":"mp3"} → 200, 17,664 bytes (tmp-live/ok.mp3, ≈1.1 s), headers openai-processing-ms: 952, x-ratelimit-limit-requests: 30000. Output is non-deterministic: re-runs of the same request produced 16,128 B (alloy/mp3) and 19,968 B (marin/mp3 + instructions); cedar/wav → 76,844 B. SSE variant (pcm) → 4 × speech.audio.delta (12.8k–25.6k base64 chars each) + speech.audio.done {usage: {input_tokens: 1, output_tokens: 33, total_tokens: 34}} + data: [DONE]. Usage-policy note: disclose to end users that the voice is AI-generated. Languages: follows Whisper's list (Afrikaans … Welsh).
4. Speech-to-text — POST /v1/audio/transcriptions (multipart/form-data, LIVE_VERIFIED)
| Param | Type / enum | Default | Model support |
|---|---|---|---|
file |
binary ≤ 25 MB; flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm | — | all |
model |
gpt-transcribe, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-mini-transcribe-2025-12-15, gpt-4o-transcribe-diarize, whisper-1 |
— | |
language |
ISO-639-1 | — | single-language models (not with languages) |
languages[] |
ISO-639-1 / selected 639-3 / zh-cn … |
— | gpt-transcribe (replaces language) |
keywords[] |
strings (one line, no < > CR LF) |
— | gpt-transcribe |
prompt |
string | — | all but diarize (whisper: 224 tokens) |
response_format |
json (default), text, srt, verbose_json, vtt, diarized_json |
json | gpt-4o-*: json/text (+diarized_json for diarize); whisper: all classic formats |
temperature |
0–1 | 0 | |
include[] |
logprobs |
— | gpt-4o-transcribe / mini, json only |
timestamp_granularities[] |
word, segment |
["segment"] |
whisper-1 + verbose_json |
stream |
boolean | false | all but whisper-1 |
chunking_strategy |
"auto" | {type: server_vad, prefix_padding_ms 300, silence_duration_ms 200, threshold 0.5} |
none (whole file) | gpt-4o-* (required for diarize > 30 s) |
known_speaker_names[] / known_speaker_references[] |
≤4 names + data-URL clips 2–10 s | — | diarize |
Response shapes (all observed on tmp-live/ok.mp3):
json(gpt-4o-mini-transcribe):{"text":"OK.","usage":{"type":"tokens","total_tokens":15,"input_tokens":11,"input_token_details":{"text_tokens":0,"audio_tokens":11},"output_tokens":4}}; withinclude[]=logprobs:+ "logprobs":[{"token":"OK","logprob":-0.0095,"bytes":[79,75]},…].gpt-transcribeadds"languages":[{"code":"en"}](empty array if unsure; documented, not tested).verbose_json(whisper-1, word+segment):{"task":"transcribe","language":"english","duration":1.1,"text":"OK.","segments":[{id, seek, start, end, text, tokens[], temperature, avg_logprob, compression_ratio, no_speech_prob}],"words":[{"word":"OK","start":0.0,"end":0.5}],"usage":{"type":"duration","seconds":2}}(duration usage rounds up to whole seconds).diarized_json(gpt-4o-transcribe-diarize):{"text":"Okay.","task":"transcribe","duration":1.104,"segments":[{"type":"transcript.text.segment","id":"seg_0","speaker":"A","start":0.0,"end":0.35,"text":" Okay."}],"usage":{"type":"tokens","total_tokens":87,"input_tokens":11,"output_tokens":76}}.stream=true(SSE):transcript.text.delta {delta, logprobs?, segment_id?}× n →transcript.text.done {text, languages?, logprobs?, usage}; diarized streams addtranscript.text.segment {id, start, end, text, speaker}. Observed:delta "OK",delta ".",done.
Longer inputs: split ≤25 MB chunks at sentence boundaries (PyDub example in guide). Post-processing with a text model is the documented way around whisper's 224-token prompt.
5. Translation — POST /v1/audio/translations (LIVE_VERIFIED)
file, model (whisper-1 only), prompt, response_format json|text|srt|verbose_json|vtt (default json), temperature 0–1. Output is English only. Observed: {"text":"OK."}; verbose_json returns {duration, language: "english", text, segments[]}.
6. Custom voices & voice consents (DOCUMENTED; access-gated — "limited to eligible customers")
POST /v1/audio/voice_consentsmultipartname,language(BCP 47, e.g.en-US),recording(≤10 MiB; audio/mpeg, wav, x-wav, ogg, aac, flac, webm, mp4) containing exactly one of the 17 consent phrases (de, en, es, fr, hi, id, it, ja, ko, nl, pl, pt, ru, uk, vi, zh — e.g. en: "I am the owner of this voice and I consent to OpenAI using this voice to create a synthetic voice model.") →{id: "cons_…", object: "audio.voice_consent", name, language, created_at}. One consent can back several voices.POST /v1/audio/voicesmultipartname,consent(cons_…),audio_sample(≤30 s, ≥5 s speech / ≥15 tokens, ≤10 MiB; same speaker) →{id: "voice_…", object: "audio.voice", name, created_at}. Max 20 voices per org. Browseraudio/webm;codecs=opusmust be sent asaudio/webm.- Use
voice: {"id": "voice_…"}in/v1/audio/speech, Realtimesession.audio.output.voice, Livesession.audio.output.voice, or Chataudio.voice. - Consents:
GET /v1/audio/voice_consents?after&limit(1–100, default 20)→{object: list, data[], first_id, last_id, has_more};GET|POST(update name)|DELETE /v1/audio/voice_consents/{consent_id}(delete →{id, object, deleted: true}). There is no list/retrieve/delete endpoint for voices in the spec (GET /v1/audio/voices→ 404 "Endpoint not found."); voices appear in the platform Audio → Voices tab. Scopes:api.voices.read/api.voices.write.
LIVE 2026-09-18: GET /v1/audio/voice_consents?limit=5 → 404 Endpoint not found. with our key. Recorded as FAILED_VERIFICATION; the most plausible cause is that the custom-voice program is not enabled for this organization (a 404 does not mean the route does not exist). Creation endpoints were deliberately not called.
7. Audio in Chat Completions (pointer; endpoint owned by the chat agent)
- Output:
{"model":"gpt-audio-mini","modalities":["text","audio"],"audio":{"voice":"alloy","format":"wav|aac|mp3|flac|opus|pcm16"},"messages":[…]}→choices[0].message.audio {id, data (b64), transcript, expires_at}. - Input: content part
{"type":"input_audio","input_audio":{"data":"<base64>","format":"wav|mp3"}}. - LIVE:
gpt-audio-mini,modalities:["text"], text-only prompt → 400{"code":"invalid_value","param":"model","message":"This model requires that either input content or output modality contain audio."}— the audio models cannot be used as pure text models.
8. Live verification log (2026-09-18)
| Call | Status | Cost est. |
|---|---|---|
| POST /v1/audio/speech mp3 "OK" (gpt-4o-mini-tts) | 200, 17,664 B | ≈$0.0006 |
| POST /v1/audio/speech sse+pcm | 200, 4 deltas + done | ≈$0.0006 |
| POST /v1/audio/transcriptions gpt-4o-mini-transcribe json | 200 | <$0.0001 |
| … whisper-1 verbose_json word+segment | 200 | <$0.0001 |
| … gpt-4o-transcribe-diarize diarized_json | 200 | <$0.0001 |
| … gpt-4o-mini-transcribe include[]=logprobs | 200 | <$0.0001 |
| … gpt-4o-mini-transcribe stream=true | 200 SSE | <$0.0001 |
| POST /v1/audio/translations whisper-1 | 200 | <$0.0001 |
| GET /v1/audio/voices (undocumented probe) | 404 Endpoint not found | 0 |
| GET /v1/audio/voice_consents?limit=5 | 404 Endpoint not found | 0 |
| POST /v1/chat/completions gpt-audio-mini modalities=[text] | 400 invalid_value | 0 |
Total audio probe cost ≈ $0.002. Raw sanitized outputs: tmp-live/realtime-audio/*.json, tmp-live/ok.mp3.