SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
14 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
13.1 KB

# OpenAI Audio API — text-to-speech, speech-to-text, translation, custom voices, audio in Chat Completions

Status: LIVE_VERIFIED (/v1/audio/speech incl. SSE, /v1/audio/transcriptions ×5 variants incl. streaming/diarization/logprobs, /v1/audio/translations) · DOCUMENTED/UNVERIFIED for voices & voice consents (list returned 404 for our key — likely access-gated) · LIVE_DISCOVERED for Chat Completions audio (400 shape captured). Sources: https://developers.openai.com/api/docs/guides/audio · …/text-to-speech · …/speech-to-text · …/transcription · …/custom-voices · …/audio-chat-completions · https://developers.openai.com/api/reference/resources/audio/subresources/speech/methods/create · …/transcriptions/methods/create · …/transcriptions/streaming-events · …/translations/methods/create · …/voices/methods/create · …/voice_consents/methods/{create,list,retrieve,update,delete} · model pages (gpt-4o-mini-tts, tts-1, tts-1-hd, gpt-transcribe, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize, whisper-1, gpt-audio-1.5, gpt-audio, gpt-audio-mini, gpt-4o-(mini-)audio-preview) · openapi-master.yaml Last verified: 2026-09-18 Twins: generated/fragments/endpoints/openai-realtime-live-audio.json (api_family audio, chat-audio), generated/fragments/parameters/openai-audio.json (53), generated/fragments/streaming-events/openai-audio-transcription.json (5).

# 1. Workflow chooser (official)

Need Use
Conversational voice app GPT-Live (live.md) or Realtime (realtime.md)
Transcript of a recorded file POST /v1/audio/transcriptions — start with gpt-transcribe
Live captions (mic/call) Realtime transcription session — gpt-live-transcribe
Speaker labels gpt-4o-transcribe-diarize + response_format: diarized_json (files only)
Word timestamps / SRT / VTT whisper-1 + timestamp_granularities[] / response_format srt|vtt
Translate a recording to English POST /v1/audio/translations (whisper-1 only)
Narration / generated speech POST /v1/audio/speech — gpt-4o-mini-tts
Audio in an existing chat app Chat Completions with gpt-audio-* and modalities

# 2. Model matrix (from model pages, 2026-09-18 — pricing owned by the models agent; cite docs/models/*)

Model Task In → Out Endpoints (Supported) Price (model page) Notes
gpt-4o-mini-tts (-2025-03-20, default -2025-12-15) TTS text → audio /v1/audio/speech $0.60/1M text in, $12/1M audio out instructions supported, ≤2000 input tokens, SSE streaming, 13 voices, custom voices
tts-1 TTS (speed) text → audio /v1/audio/speech $15 / 1M characters 9 voices, no instructions, no SSE
tts-1-hd TTS (quality) text → audio /v1/audio/speech $30 / 1M characters idem
gpt-transcribe STT (recommended) audio(+text) → text /v1/audio/transcriptions, realtime transcription sessions $0.0045 / min prompt, keywords[], languages[], detected languages in output, streaming
gpt-4o-transcribe STT audio → text /v1/audio/transcriptions, /v1/realtime $2.5 in / $10 out per 1M tokens include[]=logprobs, streaming
gpt-4o-mini-transcribe (-2025-03-20, default -2025-12-15) STT audio → text idem $1.25 / $5 per 1M LIVE: 1.1 s clip = 11 audio input tokens, 4 output
gpt-4o-transcribe-diarize STT + speakers audio → text /v1/audio/transcriptions $2.5 / $10 per 1M diarized_json, known_speaker_names[]/references[], no prompt; chunking_strategy for >30 s
whisper-1 STT/translation audio → text /v1/audio/transcriptions, /v1/audio/translations $0.006 / min verbose_json, timestamp_granularities, srt/vtt, 98 languages, 224-token prompt, no streaming
gpt-live-transcribe live STT audio → text realtime transcription sessions $0.017 / min Realtime only (see realtime.md §9)
gpt-realtime-whisper live STT audio → text realtime transcription sessions $0.017 / min VAD must be null
gpt-audio-1.5 / gpt-audio (-2025-08-28) chat audio text+audio → text+audio /v1/chat/completions text $2.5/$10, audio $32/$64 per 1M 128k ctx, 16,384 max out
gpt-audio-mini (-2025-10-06, default -2025-12-15) chat audio idem /v1/chat/completions text $0.6/$2.4 per 1M requires audio in input or output
gpt-4o-audio-preview / gpt-4o-mini-audio-preview chat audio (preview) idem /v1/chat/completions $2.5/$10 text, $40/$80 audio · $0.15/$0.6 text, $10/$20 audio PREVIEW

# 3. Text-to-speech — POST /v1/audio/speech (LIVE_VERIFIED)

Param Type / enum Default Notes
model gpt-4o-mini-tts, gpt-4o-mini-tts-2025-12-15, tts-1, tts-1-hd — required
input string ≤ 4096 chars — required
voice alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin, cedar or {id: "voice_…"} — required; tts-1/-hd only alloy, ash, coral, echo, fable, onyx, nova, sage, shimmer; best quality marin/cedar; voices optimized for English
instructions string — tone/accent/emotion/speed/whisper…; not for tts-1/-hd
response_format mp3 (default), opus, aac, flac, wav, pcm (24 kHz 16-bit LE raw) mp3 wav/pcm fastest to first byte
speed 0.25–4.0 1.0
stream_format audio (chunked bytes, default) | sse audio sse emits speech.audio.delta {audio: b64} … speech.audio.done {usage} then data: [DONE]; not for tts-1/-hd

Observed: {"model":"gpt-4o-mini-tts","input":"OK","voice":"alloy","response_format":"mp3"} → 200, 17,664 bytes (tmp-live/ok.mp3, ≈1.1 s), headers openai-processing-ms: 952, x-ratelimit-limit-requests: 30000. Output is non-deterministic: re-runs of the same request produced 16,128 B (alloy/mp3) and 19,968 B (marin/mp3 + instructions); cedar/wav → 76,844 B. SSE variant (pcm) → 4 × speech.audio.delta (12.8k–25.6k base64 chars each) + speech.audio.done {usage: {input_tokens: 1, output_tokens: 33, total_tokens: 34}} + data: [DONE]. Usage-policy note: disclose to end users that the voice is AI-generated. Languages: follows Whisper's list (Afrikaans … Welsh).

# 4. Speech-to-text — POST /v1/audio/transcriptions (multipart/form-data, LIVE_VERIFIED)

Param Type / enum Default Model support
file binary ≤ 25 MB; flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm — all
model gpt-transcribe, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-mini-transcribe-2025-12-15, gpt-4o-transcribe-diarize, whisper-1 —
language ISO-639-1 — single-language models (not with languages)
languages[] ISO-639-1 / selected 639-3 / zh-cn … — gpt-transcribe (replaces language)
keywords[] strings (one line, no < > CR LF) — gpt-transcribe
prompt string — all but diarize (whisper: 224 tokens)
response_format json (default), text, srt, verbose_json, vtt, diarized_json json gpt-4o-*: json/text (+diarized_json for diarize); whisper: all classic formats
temperature 0–1 0
include[] logprobs — gpt-4o-transcribe / mini, json only
timestamp_granularities[] word, segment ["segment"] whisper-1 + verbose_json
stream boolean false all but whisper-1
chunking_strategy "auto" | {type: server_vad, prefix_padding_ms 300, silence_duration_ms 200, threshold 0.5} none (whole file) gpt-4o-* (required for diarize > 30 s)
known_speaker_names[] / known_speaker_references[] ≤4 names + data-URL clips 2–10 s — diarize

Response shapes (all observed on tmp-live/ok.mp3):

  • json (gpt-4o-mini-transcribe): {"text":"OK.","usage":{"type":"tokens","total_tokens":15,"input_tokens":11,"input_token_details":{"text_tokens":0,"audio_tokens":11},"output_tokens":4}}; with include[]=logprobs: + "logprobs":[{"token":"OK","logprob":-0.0095,"bytes":[79,75]},…]. gpt-transcribe adds "languages":[{"code":"en"}] (empty array if unsure; documented, not tested).
  • verbose_json (whisper-1, word+segment): {"task":"transcribe","language":"english","duration":1.1,"text":"OK.","segments":[{id, seek, start, end, text, tokens[], temperature, avg_logprob, compression_ratio, no_speech_prob}],"words":[{"word":"OK","start":0.0,"end":0.5}],"usage":{"type":"duration","seconds":2}} (duration usage rounds up to whole seconds).
  • diarized_json (gpt-4o-transcribe-diarize): {"text":"Okay.","task":"transcribe","duration":1.104,"segments":[{"type":"transcript.text.segment","id":"seg_0","speaker":"A","start":0.0,"end":0.35,"text":" Okay."}],"usage":{"type":"tokens","total_tokens":87,"input_tokens":11,"output_tokens":76}}.
  • stream=true (SSE): transcript.text.delta {delta, logprobs?, segment_id?} × n → transcript.text.done {text, languages?, logprobs?, usage}; diarized streams add transcript.text.segment {id, start, end, text, speaker}. Observed: delta "OK", delta ".", done.

Longer inputs: split ≤25 MB chunks at sentence boundaries (PyDub example in guide). Post-processing with a text model is the documented way around whisper's 224-token prompt.

# 5. Translation — POST /v1/audio/translations (LIVE_VERIFIED)

file, model (whisper-1 only), prompt, response_format json|text|srt|verbose_json|vtt (default json), temperature 0–1. Output is English only. Observed: {"text":"OK."}; verbose_json returns {duration, language: "english", text, segments[]}.

# 6. Custom voices & voice consents (DOCUMENTED; access-gated — "limited to eligible customers")

  1. POST /v1/audio/voice_consents multipart name, language (BCP 47, e.g. en-US), recording (≤10 MiB; audio/mpeg, wav, x-wav, ogg, aac, flac, webm, mp4) containing exactly one of the 17 consent phrases (de, en, es, fr, hi, id, it, ja, ko, nl, pl, pt, ru, uk, vi, zh — e.g. en: "I am the owner of this voice and I consent to OpenAI using this voice to create a synthetic voice model.") → {id: "cons_…", object: "audio.voice_consent", name, language, created_at}. One consent can back several voices.
  2. POST /v1/audio/voices multipart name, consent (cons_…), audio_sample (≤30 s, ≥5 s speech / ≥15 tokens, ≤10 MiB; same speaker) → {id: "voice_…", object: "audio.voice", name, created_at}. Max 20 voices per org. Browser audio/webm;codecs=opus must be sent as audio/webm.
  3. Use voice: {"id": "voice_…"} in /v1/audio/speech, Realtime session.audio.output.voice, Live session.audio.output.voice, or Chat audio.voice.
  4. Consents: GET /v1/audio/voice_consents?after&limit(1–100, default 20) → {object: list, data[], first_id, last_id, has_more}; GET|POST(update name)|DELETE /v1/audio/voice_consents/{consent_id} (delete → {id, object, deleted: true}). There is no list/retrieve/delete endpoint for voices in the spec (GET /v1/audio/voices → 404 "Endpoint not found."); voices appear in the platform Audio → Voices tab. Scopes: api.voices.read / api.voices.write.

LIVE 2026-09-18: GET /v1/audio/voice_consents?limit=5 → 404 Endpoint not found. with our key. Recorded as FAILED_VERIFICATION; the most plausible cause is that the custom-voice program is not enabled for this organization (a 404 does not mean the route does not exist). Creation endpoints were deliberately not called.

# 7. Audio in Chat Completions (pointer; endpoint owned by the chat agent)

  • Output: {"model":"gpt-audio-mini","modalities":["text","audio"],"audio":{"voice":"alloy","format":"wav|aac|mp3|flac|opus|pcm16"},"messages":[…]} → choices[0].message.audio {id, data (b64), transcript, expires_at}.
  • Input: content part {"type":"input_audio","input_audio":{"data":"<base64>","format":"wav|mp3"}}.
  • LIVE: gpt-audio-mini, modalities:["text"], text-only prompt → 400 {"code":"invalid_value","param":"model","message":"This model requires that either input content or output modality contain audio."} — the audio models cannot be used as pure text models.

# 8. Live verification log (2026-09-18)

Call Status Cost est.
POST /v1/audio/speech mp3 "OK" (gpt-4o-mini-tts) 200, 17,664 B ≈$0.0006
POST /v1/audio/speech sse+pcm 200, 4 deltas + done ≈$0.0006
POST /v1/audio/transcriptions gpt-4o-mini-transcribe json 200 <$0.0001
… whisper-1 verbose_json word+segment 200 <$0.0001
… gpt-4o-transcribe-diarize diarized_json 200 <$0.0001
… gpt-4o-mini-transcribe include[]=logprobs 200 <$0.0001
… gpt-4o-mini-transcribe stream=true 200 SSE <$0.0001
POST /v1/audio/translations whisper-1 200 <$0.0001
GET /v1/audio/voices (undocumented probe) 404 Endpoint not found 0
GET /v1/audio/voice_consents?limit=5 404 Endpoint not found 0
POST /v1/chat/completions gpt-audio-mini modalities=[text] 400 invalid_value 0

Total audio probe cost ≈ $0.002. Raw sanitized outputs: tmp-live/realtime-audio/*.json, tmp-live/ok.mp3.