# OpenAI Audio API — text-to-speech, speech-to-text, translation, custom voices, audio in Chat Completions **Status**: LIVE_VERIFIED (`/v1/audio/speech` incl. SSE, `/v1/audio/transcriptions` ×5 variants incl. streaming/diarization/logprobs, `/v1/audio/translations`) · DOCUMENTED/UNVERIFIED for voices & voice consents (list returned 404 for our key — likely access-gated) · LIVE_DISCOVERED for Chat Completions audio (400 shape captured). **Sources**: https://developers.openai.com/api/docs/guides/audio · …/text-to-speech · …/speech-to-text · …/transcription · …/custom-voices · …/audio-chat-completions · https://developers.openai.com/api/reference/resources/audio/subresources/speech/methods/create · …/transcriptions/methods/create · …/transcriptions/streaming-events · …/translations/methods/create · …/voices/methods/create · …/voice_consents/methods/{create,list,retrieve,update,delete} · model pages (`gpt-4o-mini-tts`, `tts-1`, `tts-1-hd`, `gpt-transcribe`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe`, `gpt-4o-transcribe-diarize`, `whisper-1`, `gpt-audio-1.5`, `gpt-audio`, `gpt-audio-mini`, `gpt-4o-(mini-)audio-preview`) · `openapi-master.yaml` **Last verified**: 2026-09-18 **Twins**: `generated/fragments/endpoints/openai-realtime-live-audio.json` (api_family `audio`, `chat-audio`), `generated/fragments/parameters/openai-audio.json` (53), `generated/fragments/streaming-events/openai-audio-transcription.json` (5). ## 1. Workflow chooser (official) | Need | Use | |---|---| | Conversational voice app | GPT-Live ([`live.md`](live.md)) or Realtime ([`realtime.md`](realtime.md)) | | Transcript of a recorded file | `POST /v1/audio/transcriptions` — start with `gpt-transcribe` | | Live captions (mic/call) | Realtime transcription session — `gpt-live-transcribe` | | Speaker labels | `gpt-4o-transcribe-diarize` + `response_format: diarized_json` (files only) | | Word timestamps / SRT / VTT | `whisper-1` + `timestamp_granularities[]` / `response_format srt\|vtt` | | Translate a recording to English | `POST /v1/audio/translations` (`whisper-1` only) | | Narration / generated speech | `POST /v1/audio/speech` — `gpt-4o-mini-tts` | | Audio in an existing chat app | Chat Completions with `gpt-audio-*` and `modalities` | ## 2. Model matrix (from model pages, 2026-09-18 — pricing owned by the models agent; cite `docs/models/*`) | Model | Task | In → Out | Endpoints (Supported) | Price (model page) | Notes | |---|---|---|---|---|---| | `gpt-4o-mini-tts` (`-2025-03-20`, default `-2025-12-15`) | TTS | text → audio | `/v1/audio/speech` | $0.60/1M text in, $12/1M audio out | `instructions` supported, ≤2000 input tokens, SSE streaming, 13 voices, custom voices | | `tts-1` | TTS (speed) | text → audio | `/v1/audio/speech` | $15 / 1M characters | 9 voices, no `instructions`, no SSE | | `tts-1-hd` | TTS (quality) | text → audio | `/v1/audio/speech` | $30 / 1M characters | idem | | `gpt-transcribe` | STT (recommended) | audio(+text) → text | `/v1/audio/transcriptions`, realtime transcription sessions | $0.0045 / min | `prompt`, `keywords[]`, `languages[]`, detected `languages` in output, streaming | | `gpt-4o-transcribe` | STT | audio → text | `/v1/audio/transcriptions`, `/v1/realtime` | $2.5 in / $10 out per 1M tokens | `include[]=logprobs`, streaming | | `gpt-4o-mini-transcribe` (`-2025-03-20`, default `-2025-12-15`) | STT | audio → text | idem | $1.25 / $5 per 1M | LIVE: 1.1 s clip = 11 audio input tokens, 4 output | | `gpt-4o-transcribe-diarize` | STT + speakers | audio → text | `/v1/audio/transcriptions` | $2.5 / $10 per 1M | `diarized_json`, `known_speaker_names[]/references[]`, no `prompt`; `chunking_strategy` for >30 s | | `whisper-1` | STT/translation | audio → text | `/v1/audio/transcriptions`, `/v1/audio/translations` | $0.006 / min | `verbose_json`, `timestamp_granularities`, `srt`/`vtt`, 98 languages, 224-token prompt, no streaming | | `gpt-live-transcribe` | live STT | audio → text | realtime transcription sessions | $0.017 / min | Realtime only (see realtime.md §9) | | `gpt-realtime-whisper` | live STT | audio → text | realtime transcription sessions | $0.017 / min | VAD must be null | | `gpt-audio-1.5` / `gpt-audio` (`-2025-08-28`) | chat audio | text+audio → text+audio | `/v1/chat/completions` | text $2.5/$10, audio $32/$64 per 1M | 128k ctx, 16,384 max out | | `gpt-audio-mini` (`-2025-10-06`, default `-2025-12-15`) | chat audio | idem | `/v1/chat/completions` | text $0.6/$2.4 per 1M | requires audio in input or output | | `gpt-4o-audio-preview` / `gpt-4o-mini-audio-preview` | chat audio (preview) | idem | `/v1/chat/completions` | $2.5/$10 text, $40/$80 audio · $0.15/$0.6 text, $10/$20 audio | PREVIEW | ## 3. Text-to-speech — `POST /v1/audio/speech` (LIVE_VERIFIED) | Param | Type / enum | Default | Notes | |---|---|---|---| | `model` | `gpt-4o-mini-tts`, `gpt-4o-mini-tts-2025-12-15`, `tts-1`, `tts-1-hd` | — | required | | `input` | string ≤ 4096 chars | — | required | | `voice` | `alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin, cedar` or `{id: "voice_…"}` | — | required; tts-1/-hd only `alloy, ash, coral, echo, fable, onyx, nova, sage, shimmer`; best quality `marin`/`cedar`; voices optimized for English | | `instructions` | string | — | tone/accent/emotion/speed/whisper…; **not** for tts-1/-hd | | `response_format` | `mp3` (default), `opus`, `aac`, `flac`, `wav`, `pcm` (24 kHz 16-bit LE raw) | mp3 | `wav`/`pcm` fastest to first byte | | `speed` | 0.25–4.0 | 1.0 | | | `stream_format` | `audio` (chunked bytes, default) \| `sse` | audio | `sse` emits `speech.audio.delta {audio: b64}` … `speech.audio.done {usage}` then `data: [DONE]`; not for tts-1/-hd | Observed: `{"model":"gpt-4o-mini-tts","input":"OK","voice":"alloy","response_format":"mp3"}` → 200, **17,664 bytes** (`tmp-live/ok.mp3`, ≈1.1 s), headers `openai-processing-ms: 952`, `x-ratelimit-limit-requests: 30000`. Output is non-deterministic: re-runs of the same request produced 16,128 B (alloy/mp3) and 19,968 B (marin/mp3 + instructions); `cedar`/wav → 76,844 B. SSE variant (`pcm`) → 4 × `speech.audio.delta` (12.8k–25.6k base64 chars each) + `speech.audio.done {usage: {input_tokens: 1, output_tokens: 33, total_tokens: 34}}` + `data: [DONE]`. Usage-policy note: disclose to end users that the voice is AI-generated. Languages: follows Whisper's list (Afrikaans … Welsh). ## 4. Speech-to-text — `POST /v1/audio/transcriptions` (multipart/form-data, LIVE_VERIFIED) | Param | Type / enum | Default | Model support | |---|---|---|---| | `file` | binary ≤ 25 MB; flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm | — | all | | `model` | `gpt-transcribe`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe`, `gpt-4o-mini-transcribe-2025-12-15`, `gpt-4o-transcribe-diarize`, `whisper-1` | — | | | `language` | ISO-639-1 | — | single-language models (not with `languages`) | | `languages[]` | ISO-639-1 / selected 639-3 / `zh-cn` … | — | `gpt-transcribe` (replaces `language`) | | `keywords[]` | strings (one line, no `<` `>` CR LF) | — | `gpt-transcribe` | | `prompt` | string | — | all but diarize (whisper: 224 tokens) | | `response_format` | `json` (default), `text`, `srt`, `verbose_json`, `vtt`, `diarized_json` | json | gpt-4o-*: `json`/`text` (+`diarized_json` for diarize); whisper: all classic formats | | `temperature` | 0–1 | 0 | | | `include[]` | `logprobs` | — | gpt-4o-transcribe / mini, `json` only | | `timestamp_granularities[]` | `word`, `segment` | `["segment"]` | whisper-1 + `verbose_json` | | `stream` | boolean | false | all but whisper-1 | | `chunking_strategy` | `"auto"` \| `{type: server_vad, prefix_padding_ms 300, silence_duration_ms 200, threshold 0.5}` | none (whole file) | gpt-4o-* (required for diarize > 30 s) | | `known_speaker_names[]` / `known_speaker_references[]` | ≤4 names + data-URL clips 2–10 s | — | diarize | Response shapes (all observed on `tmp-live/ok.mp3`): - `json` (gpt-4o-mini-transcribe): `{"text":"OK.","usage":{"type":"tokens","total_tokens":15,"input_tokens":11,"input_token_details":{"text_tokens":0,"audio_tokens":11},"output_tokens":4}}`; with `include[]=logprobs`: `+ "logprobs":[{"token":"OK","logprob":-0.0095,"bytes":[79,75]},…]`. `gpt-transcribe` adds `"languages":[{"code":"en"}]` (empty array if unsure; documented, not tested). - `verbose_json` (whisper-1, word+segment): `{"task":"transcribe","language":"english","duration":1.1,"text":"OK.","segments":[{id, seek, start, end, text, tokens[], temperature, avg_logprob, compression_ratio, no_speech_prob}],"words":[{"word":"OK","start":0.0,"end":0.5}],"usage":{"type":"duration","seconds":2}}` (duration usage rounds up to whole seconds). - `diarized_json` (gpt-4o-transcribe-diarize): `{"text":"Okay.","task":"transcribe","duration":1.104,"segments":[{"type":"transcript.text.segment","id":"seg_0","speaker":"A","start":0.0,"end":0.35,"text":" Okay."}],"usage":{"type":"tokens","total_tokens":87,"input_tokens":11,"output_tokens":76}}`. - `stream=true` (SSE): `transcript.text.delta {delta, logprobs?, segment_id?}` × n → `transcript.text.done {text, languages?, logprobs?, usage}`; diarized streams add `transcript.text.segment {id, start, end, text, speaker}`. Observed: `delta "OK"`, `delta "."`, `done`. Longer inputs: split ≤25 MB chunks at sentence boundaries (PyDub example in guide). Post-processing with a text model is the documented way around whisper's 224-token prompt. ## 5. Translation — `POST /v1/audio/translations` (LIVE_VERIFIED) `file`, `model` (`whisper-1` only), `prompt`, `response_format` `json|text|srt|verbose_json|vtt` (default json), `temperature` 0–1. Output is **English only**. Observed: `{"text":"OK."}`; `verbose_json` returns `{duration, language: "english", text, segments[]}`. ## 6. Custom voices & voice consents (DOCUMENTED; access-gated — "limited to eligible customers") 1. `POST /v1/audio/voice_consents` multipart `name`, `language` (BCP 47, e.g. `en-US`), `recording` (≤10 MiB; audio/mpeg, wav, x-wav, ogg, aac, flac, webm, mp4) containing **exactly** one of the 17 consent phrases (de, en, es, fr, hi, id, it, ja, ko, nl, pl, pt, ru, uk, vi, zh — e.g. en: "I am the owner of this voice and I consent to OpenAI using this voice to create a synthetic voice model.") → `{id: "cons_…", object: "audio.voice_consent", name, language, created_at}`. One consent can back several voices. 2. `POST /v1/audio/voices` multipart `name`, `consent` (`cons_…`), `audio_sample` (≤30 s, ≥5 s speech / ≥15 tokens, ≤10 MiB; same speaker) → `{id: "voice_…", object: "audio.voice", name, created_at}`. Max **20 voices per org**. Browser `audio/webm;codecs=opus` must be sent as `audio/webm`. 3. Use `voice: {"id": "voice_…"}` in `/v1/audio/speech`, Realtime `session.audio.output.voice`, Live `session.audio.output.voice`, or Chat `audio.voice`. 4. Consents: `GET /v1/audio/voice_consents?after&limit(1–100, default 20)` → `{object: list, data[], first_id, last_id, has_more}`; `GET|POST(update name)|DELETE /v1/audio/voice_consents/{consent_id}` (delete → `{id, object, deleted: true}`). There is **no** list/retrieve/delete endpoint for voices in the spec (`GET /v1/audio/voices` → 404 "Endpoint not found."); voices appear in the platform Audio → Voices tab. Scopes: `api.voices.read` / `api.voices.write`. LIVE 2026-09-18: `GET /v1/audio/voice_consents?limit=5` → **404 `Endpoint not found.`** with our key. Recorded as FAILED_VERIFICATION; the most plausible cause is that the custom-voice program is not enabled for this organization (a 404 does not mean the route does not exist). Creation endpoints were deliberately not called. ## 7. Audio in Chat Completions (pointer; endpoint owned by the chat agent) - Output: `{"model":"gpt-audio-mini","modalities":["text","audio"],"audio":{"voice":"alloy","format":"wav|aac|mp3|flac|opus|pcm16"},"messages":[…]}` → `choices[0].message.audio {id, data (b64), transcript, expires_at}`. - Input: content part `{"type":"input_audio","input_audio":{"data":"","format":"wav|mp3"}}`. - LIVE: `gpt-audio-mini`, `modalities:["text"]`, text-only prompt → **400** `{"code":"invalid_value","param":"model","message":"This model requires that either input content or output modality contain audio."}` — the audio models cannot be used as pure text models. ## 8. Live verification log (2026-09-18) | Call | Status | Cost est. | |---|---|---| | POST /v1/audio/speech mp3 "OK" (gpt-4o-mini-tts) | 200, 17,664 B | ≈$0.0006 | | POST /v1/audio/speech sse+pcm | 200, 4 deltas + done | ≈$0.0006 | | POST /v1/audio/transcriptions gpt-4o-mini-transcribe json | 200 | <$0.0001 | | … whisper-1 verbose_json word+segment | 200 | <$0.0001 | | … gpt-4o-transcribe-diarize diarized_json | 200 | <$0.0001 | | … gpt-4o-mini-transcribe include[]=logprobs | 200 | <$0.0001 | | … gpt-4o-mini-transcribe stream=true | 200 SSE | <$0.0001 | | POST /v1/audio/translations whisper-1 | 200 | <$0.0001 | | GET /v1/audio/voices (undocumented probe) | 404 Endpoint not found | 0 | | GET /v1/audio/voice_consents?limit=5 | 404 Endpoint not found | 0 | | POST /v1/chat/completions gpt-audio-mini modalities=[text] | 400 invalid_value | 0 | Total audio probe cost ≈ $0.002. Raw sanitized outputs: `tmp-live/realtime-audio/*.json`, `tmp-live/ok.mp3`.