SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
6.1 KB

# Voice agents on OpenAI — architectures (GPT-Live full duplex · Realtime speech-to-speech · chained STT→LLM→TTS)

Status: DOCUMENTED (architecture guidance from the official guides) · component APIs verified as noted in realtime.md, live.md, audio.md. Sources: https://developers.openai.com/api/docs/guides/voice-agents · …/guides/audio · …/guides/live · …/guides/realtime · …/guides/live-migration · …/guides/voice-latency-cost · …/guides/live-delegation Last verified: 2026-09-18

# 1. Three architectures

Architecture Best for How speech meets reasoning Latency profile Cost model Control points Status
GPT-Live (/v1/live/sessions, gpt-live-1) full-duplex conversations with a separate backend; keep an existing text agent Live model listens/speaks continuously and delegates tasks (Responses delegation = OpenAI-run backend; client delegation = your agent) lowest perceived latency: assistant can talk while backend works; no turn commit $0.05/min voice + backend tokens/tools prompts split: conversation style (Live) vs business rules/tools (backend); transcripts via deltas; sideband for guardrails LIVE_DISCOVERED (route), sessions UNVERIFIED
Realtime API (/v1/realtime, gpt-realtime-2.1 …) speech, reasoning and tool use in one model/session one speech-to-speech model interprets audio, calls function/MCP tools, answers in speech low: first-audio fast, VAD turn taking, barge-in tokens: text + audio (1 tok/100 ms in, 1 tok/50 ms out) + image; input transcription extra full session state (items), VAD, truncation, out-of-band responses, tools/MCP, sideband via call_id LIVE_VERIFIED (text-only turn)
Chained pipeline (STT → text agent → TTS) inspect/transform text between stages; replace components independently POST /v1/audio/transcriptions → any text model/agent (Responses/Chat) → POST /v1/audio/speech highest: three sequential calls; TTS streaming (wav/pcm, stream_format) mitigates STT per minute/tokens + LLM tokens + TTS tokens/characters transcript stored, policy checks before reply, speech only after approved answer components LIVE_VERIFIED

Rule from the guide: choose the audio architecture first, then design the agent workflow (tools, handoffs, guardrails, observability) exactly as for text.

# 2. Speech-to-speech (Realtime) vs chained — when to pick which

  • Pick Realtime when interaction must feel conversational (interruptions, natural turn taking, realtime tool use) and the model may hear tone/inflection. Start from the Agents SDK RealtimeAgent + RealtimeSession (WebRTC in browser via an ek_ client secret, WebSocket on servers), then drop to raw events (docs/openai/realtime.md) when you need control.
  • Pick chained when each stage must be visible/replaceable (compliance review of the transcript, RAG before answering, multiple LLM providers, storing text). Downsides: no barge-in without extra engineering, STT errors propagate, latency adds up. Recommended components: gpt-transcribe (or gpt-live-transcribe for live mic), your text agent, gpt-4o-mini-tts with instructions for tone.
  • Pick GPT-Live when you already have a text agent/tool loop and want a voice front-end with full duplex; migrate a Realtime app by splitting the prompt (style → session.instructions, rules/tools → delegation.responses.instructions or your client backend), replacing input_audio_buffer.append→session.input_audio.append, response.output_audio.delta→session.output_audio.delta, transcript events → session.input_transcript.delta / session.output_transcript.delta, and removing manual commits / voice-turn triggers (no end-of-response event exists in Live).

# 3. Transport choice (shared by Live and Realtime)

Client Transport Notes
Browser / mobile WebRTC media tracks + oai-events data channel; Realtime: ephemeral key or unified /v1/realtime/calls; Live: your server posts the SDP to /v1/live/sessions
Server audio pipeline / telephony bridge WebSocket base64 PCM16 24 kHz (Live also 16 kHz PCM and G.711 8 kHz)
Phone SIP Realtime: sip:<proj>@sip.api.openai.com + realtime.call.incoming webhook; Live: trunk → live.transport.incoming webhook + /v1/live/sessions/{id}/accept; partners: Twilio, Telnyx, LiveKit, Daily/Pipecat
Backend observer/controller sideband WebSocket Realtime ?call_id=; Live /attach

# 4. Evaluation & latency (guide summary)

Test task outcomes and conversation: intent preservation, tool arguments, permissions, final state; audible response timing (median + p95), unwanted silence, overlap, yielding to interruptions; recognition across accents/noise/language switches; session reliability (drops, timeouts). Progression: Crawl (synthetic single-turn audio) → Walk (human recordings) → Run (simulated multi-turn caller). Measure backend stages separately (delegation receipt, backend start, first useful result, tool start/end, result submission, audio arrival, playback). Cookbooks: cookbook/examples/audio/voice_agent_evaluation (GPT-Live), cookbook/examples/realtime_eval_guide (Realtime).

# 5. Cost levers

Realtime: cached input (keep history static; instructions/tools first), truncation.retention_ratio + token_limits.post_instructions, max_output_tokens, mini models, deleting/summarizing old items, VAD filters silence. Live: shorter waits (Fast mode service_tier: priority, parallel tools, speculative lookups from transcript fragments), close idle sessions and resume via input history or fork, pre-load context before the session. Chained: pick per-minute STT (gpt-transcribe $0.0045/min) and stream TTS.

# 6. Safety & compliance notes

Disclose AI-generated voices (usage policy). Send OpenAI-Safety-Identifier (hashed user id) on connection/creation requests. Keep API keys and tool credentials server-side; browsers get only ephemeral ek_ tokens (Realtime) or talk to your server (Live). Custom voices need recorded consent phrases (see audio.md §6).