Voice agents on OpenAI — architectures (GPT-Live full duplex · Realtime speech-to-speech · chained STT→LLM→TTS)
Status: DOCUMENTED (architecture guidance from the official guides) · component APIs verified as noted in realtime.md, live.md, audio.md.
Sources: https://developers.openai.com/api/docs/guides/voice-agents · …/guides/audio · …/guides/live · …/guides/realtime · …/guides/live-migration · …/guides/voice-latency-cost · …/guides/live-delegation
Last verified: 2026-09-18
1. Three architectures
| Architecture | Best for | How speech meets reasoning | Latency profile | Cost model | Control points | Status |
|---|---|---|---|---|---|---|
GPT-Live (/v1/live/sessions, gpt-live-1) |
full-duplex conversations with a separate backend; keep an existing text agent | Live model listens/speaks continuously and delegates tasks (Responses delegation = OpenAI-run backend; client delegation = your agent) | lowest perceived latency: assistant can talk while backend works; no turn commit | $0.05/min voice + backend tokens/tools | prompts split: conversation style (Live) vs business rules/tools (backend); transcripts via deltas; sideband for guardrails | LIVE_DISCOVERED (route), sessions UNVERIFIED |
Realtime API (/v1/realtime, gpt-realtime-2.1 …) |
speech, reasoning and tool use in one model/session | one speech-to-speech model interprets audio, calls function/MCP tools, answers in speech | low: first-audio fast, VAD turn taking, barge-in | tokens: text + audio (1 tok/100 ms in, 1 tok/50 ms out) + image; input transcription extra | full session state (items), VAD, truncation, out-of-band responses, tools/MCP, sideband via call_id |
LIVE_VERIFIED (text-only turn) |
| Chained pipeline (STT → text agent → TTS) | inspect/transform text between stages; replace components independently | POST /v1/audio/transcriptions → any text model/agent (Responses/Chat) → POST /v1/audio/speech |
highest: three sequential calls; TTS streaming (wav/pcm, stream_format) mitigates |
STT per minute/tokens + LLM tokens + TTS tokens/characters | transcript stored, policy checks before reply, speech only after approved answer | components LIVE_VERIFIED |
Rule from the guide: choose the audio architecture first, then design the agent workflow (tools, handoffs, guardrails, observability) exactly as for text.
2. Speech-to-speech (Realtime) vs chained — when to pick which
- Pick Realtime when interaction must feel conversational (interruptions, natural turn taking, realtime tool use) and the model may hear tone/inflection. Start from the Agents SDK
RealtimeAgent+RealtimeSession(WebRTC in browser via anek_client secret, WebSocket on servers), then drop to raw events (docs/openai/realtime.md) when you need control. - Pick chained when each stage must be visible/replaceable (compliance review of the transcript, RAG before answering, multiple LLM providers, storing text). Downsides: no barge-in without extra engineering, STT errors propagate, latency adds up. Recommended components:
gpt-transcribe(orgpt-live-transcribefor live mic), your text agent,gpt-4o-mini-ttswithinstructionsfor tone. - Pick GPT-Live when you already have a text agent/tool loop and want a voice front-end with full duplex; migrate a Realtime app by splitting the prompt (style →
session.instructions, rules/tools →delegation.responses.instructionsor your client backend), replacinginput_audio_buffer.append→session.input_audio.append,response.output_audio.delta→session.output_audio.delta, transcript events →session.input_transcript.delta/session.output_transcript.delta, and removing manual commits / voice-turn triggers (no end-of-response event exists in Live).
3. Transport choice (shared by Live and Realtime)
| Client | Transport | Notes |
|---|---|---|
| Browser / mobile | WebRTC | media tracks + oai-events data channel; Realtime: ephemeral key or unified /v1/realtime/calls; Live: your server posts the SDP to /v1/live/sessions |
| Server audio pipeline / telephony bridge | WebSocket | base64 PCM16 24 kHz (Live also 16 kHz PCM and G.711 8 kHz) |
| Phone | SIP | Realtime: sip:<proj>@sip.api.openai.com + realtime.call.incoming webhook; Live: trunk → live.transport.incoming webhook + /v1/live/sessions/{id}/accept; partners: Twilio, Telnyx, LiveKit, Daily/Pipecat |
| Backend observer/controller | sideband WebSocket | Realtime ?call_id=; Live /attach |
4. Evaluation & latency (guide summary)
Test task outcomes and conversation: intent preservation, tool arguments, permissions, final state; audible response timing (median + p95), unwanted silence, overlap, yielding to interruptions; recognition across accents/noise/language switches; session reliability (drops, timeouts). Progression: Crawl (synthetic single-turn audio) → Walk (human recordings) → Run (simulated multi-turn caller). Measure backend stages separately (delegation receipt, backend start, first useful result, tool start/end, result submission, audio arrival, playback). Cookbooks: cookbook/examples/audio/voice_agent_evaluation (GPT-Live), cookbook/examples/realtime_eval_guide (Realtime).
5. Cost levers
Realtime: cached input (keep history static; instructions/tools first), truncation.retention_ratio + token_limits.post_instructions, max_output_tokens, mini models, deleting/summarizing old items, VAD filters silence. Live: shorter waits (Fast mode service_tier: priority, parallel tools, speculative lookups from transcript fragments), close idle sessions and resume via input history or fork, pre-load context before the session. Chained: pick per-minute STT (gpt-transcribe $0.0045/min) and stream TTS.
6. Safety & compliance notes
Disclose AI-generated voices (usage policy). Send OpenAI-Safety-Identifier (hashed user id) on connection/creation requests. Keep API keys and tool credentials server-side; browsers get only ephemeral ek_ tokens (Realtime) or talk to your server (Live). Custom voices need recorded consent phrases (see audio.md §6).