# Voice agents on OpenAI — architectures (GPT-Live full duplex · Realtime speech-to-speech · chained STT→LLM→TTS) **Status**: DOCUMENTED (architecture guidance from the official guides) · component APIs verified as noted in [`realtime.md`](realtime.md), [`live.md`](live.md), [`audio.md`](audio.md). **Sources**: https://developers.openai.com/api/docs/guides/voice-agents · …/guides/audio · …/guides/live · …/guides/realtime · …/guides/live-migration · …/guides/voice-latency-cost · …/guides/live-delegation **Last verified**: 2026-09-18 ## 1. Three architectures | Architecture | Best for | How speech meets reasoning | Latency profile | Cost model | Control points | Status | |---|---|---|---|---|---|---| | **GPT-Live** (`/v1/live/sessions`, `gpt-live-1`) | full-duplex conversations with a separate backend; keep an existing text agent | Live model listens/speaks continuously and **delegates** tasks (Responses delegation = OpenAI-run backend; client delegation = your agent) | lowest perceived latency: assistant can talk while backend works; no turn commit | $0.05/min voice + backend tokens/tools | prompts split: conversation style (Live) vs business rules/tools (backend); transcripts via deltas; sideband for guardrails | LIVE_DISCOVERED (route), sessions UNVERIFIED | | **Realtime API** (`/v1/realtime`, `gpt-realtime-2.1` …) | speech, reasoning and tool use in **one** model/session | one speech-to-speech model interprets audio, calls function/MCP tools, answers in speech | low: first-audio fast, VAD turn taking, barge-in | tokens: text + audio (1 tok/100 ms in, 1 tok/50 ms out) + image; input transcription extra | full session state (items), VAD, truncation, out-of-band responses, tools/MCP, sideband via `call_id` | LIVE_VERIFIED (text-only turn) | | **Chained pipeline** (STT → text agent → TTS) | inspect/transform text between stages; replace components independently | `POST /v1/audio/transcriptions` → any text model/agent (Responses/Chat) → `POST /v1/audio/speech` | highest: three sequential calls; TTS streaming (`wav`/`pcm`, `stream_format`) mitigates | STT per minute/tokens + LLM tokens + TTS tokens/characters | transcript stored, policy checks before reply, speech only after approved answer | components LIVE_VERIFIED | Rule from the guide: choose the audio architecture first, then design the agent workflow (tools, handoffs, guardrails, observability) exactly as for text. ## 2. Speech-to-speech (Realtime) vs chained — when to pick which - Pick **Realtime** when interaction must feel conversational (interruptions, natural turn taking, realtime tool use) and the model may hear tone/inflection. Start from the Agents SDK `RealtimeAgent` + `RealtimeSession` (WebRTC in browser via an `ek_` client secret, WebSocket on servers), then drop to raw events (`docs/openai/realtime.md`) when you need control. - Pick **chained** when each stage must be visible/replaceable (compliance review of the transcript, RAG before answering, multiple LLM providers, storing text). Downsides: no barge-in without extra engineering, STT errors propagate, latency adds up. Recommended components: `gpt-transcribe` (or `gpt-live-transcribe` for live mic), your text agent, `gpt-4o-mini-tts` with `instructions` for tone. - Pick **GPT-Live** when you already have a text agent/tool loop and want a voice front-end with full duplex; migrate a Realtime app by splitting the prompt (style → `session.instructions`, rules/tools → `delegation.responses.instructions` or your client backend), replacing `input_audio_buffer.append`→`session.input_audio.append`, `response.output_audio.delta`→`session.output_audio.delta`, transcript events → `session.input_transcript.delta` / `session.output_transcript.delta`, and removing manual commits / voice-turn triggers (no end-of-response event exists in Live). ## 3. Transport choice (shared by Live and Realtime) | Client | Transport | Notes | |---|---|---| | Browser / mobile | **WebRTC** | media tracks + `oai-events` data channel; Realtime: ephemeral key or unified `/v1/realtime/calls`; Live: your server posts the SDP to `/v1/live/sessions` | | Server audio pipeline / telephony bridge | **WebSocket** | base64 PCM16 24 kHz (Live also 16 kHz PCM and G.711 8 kHz) | | Phone | **SIP** | Realtime: `sip:@sip.api.openai.com` + `realtime.call.incoming` webhook; Live: trunk → `live.transport.incoming` webhook + `/v1/live/sessions/{id}/accept`; partners: Twilio, Telnyx, LiveKit, Daily/Pipecat | | Backend observer/controller | **sideband WebSocket** | Realtime `?call_id=`; Live `/attach` | ## 4. Evaluation & latency (guide summary) Test task outcomes **and** conversation: intent preservation, tool arguments, permissions, final state; audible response timing (median + p95), unwanted silence, overlap, yielding to interruptions; recognition across accents/noise/language switches; session reliability (drops, timeouts). Progression: *Crawl* (synthetic single-turn audio) → *Walk* (human recordings) → *Run* (simulated multi-turn caller). Measure backend stages separately (delegation receipt, backend start, first useful result, tool start/end, result submission, audio arrival, playback). Cookbooks: `cookbook/examples/audio/voice_agent_evaluation` (GPT-Live), `cookbook/examples/realtime_eval_guide` (Realtime). ## 5. Cost levers Realtime: cached input (keep history static; instructions/tools first), `truncation.retention_ratio` + `token_limits.post_instructions`, `max_output_tokens`, mini models, deleting/summarizing old items, VAD filters silence. Live: shorter waits (Fast mode `service_tier: priority`, parallel tools, speculative lookups from transcript fragments), close idle sessions and resume via `input` history or fork, pre-load context before the session. Chained: pick per-minute STT (`gpt-transcribe` $0.0045/min) and stream TTS. ## 6. Safety & compliance notes Disclose AI-generated voices (usage policy). Send `OpenAI-Safety-Identifier` (hashed user id) on connection/creation requests. Keep API keys and tool credentials server-side; browsers get only ephemeral `ek_` tokens (Realtime) or talk to your server (Live). Custom voices need recorded consent phrases (see `audio.md` §6).