OpenAI Realtime API (GA) — connections, sessions, events, tools, SIP, transcription, translation
Status: DOCUMENTED + LIVE_VERIFIED (client secrets, translation client secrets, WebSocket text-only turn on gpt-realtime-mini) · LEGACY/FAILED_VERIFICATION for the beta /v1/realtime/sessions endpoints · UNVERIFIED for WebRTC, SIP and sideband (need a browser peer / SIP trunk).
Sources: https://developers.openai.com/api/docs/guides/realtime · …/guides/voice-websockets?api=realtime · …/guides/voice-webrtc?api=realtime · …/guides/voice-sip?api=realtime · …/guides/voice-server-controls?api=realtime · …/guides/realtime-conversations · …/guides/realtime-vad · …/guides/realtime-mcp · …/guides/realtime-transcription · …/guides/realtime-translation · …/guides/voice-latency-cost?api=realtime · https://developers.openai.com/api/reference/resources/realtime/client-events · …/server-events · …/subresources/client_secrets/methods/create · …/subresources/calls/methods/{create,accept,reject,refer,hangup} · https://developers.openai.com/api/reference/realtime-beta/overview · sources/openai/openapi/openapi-master.yaml
Last verified: 2026-09-18
Machine-readable twins: generated/fragments/endpoints/openai-realtime-live-audio.json, generated/fragments/parameters/openai-realtime.json (560 params), generated/fragments/streaming-events/openai-realtime.json (65 events). Full event tables: realtime-events.md.
1. What it is
A stateful, low-latency speech-to-speech (and text/image-in) session with a Realtime model (gpt-realtime-2.1, gpt-realtime-2.1-mini, gpt-realtime-2, gpt-realtime-1.5, gpt-realtime, gpt-realtime-mini, legacy gpt-4o-(mini-)realtime-preview). The server keeps a Session (config), a Conversation (items) and produces Responses (items appended to the conversation). Clients drive it with JSON client events and react to server events. Max session duration: 60 minutes. Three session types share the surface: realtime (voice agent), transcription (input transcription only, no model replies), translation (dedicated /v1/realtime/translations endpoint, continuous interpreter, gpt-realtime-translate).
Sibling product: GPT-Live (/v1/live/sessions, full-duplex, delegation to a backend) — see live.md. Same transports, different handshakes, credentials and event names.
2. Connection methods matrix
| Transport | URL | Auth | Config carried by | Audio path | Events path | Interruption/truncation | Who | Status |
|---|---|---|---|---|---|---|---|---|
| WebSocket (new session) | wss://api.openai.com/v1/realtime?model=<model> |
Authorization: Bearer <sk-… standard key> (server) or Bearer <ek_… client secret>; browsers without header support use subprotocols realtime, openai-insecure-api-key.<ek_…>, optional openai-organization.<org>, openai-project.<proj> |
session.update client event (or the client secret's session) |
base64 in input_audio_buffer.append / response.output_audio.delta (≤15 MB per chunk) |
same socket | client must stop playback + send conversation.item.truncate |
server-to-server; Deno/Workers | LIVE_VERIFIED (text-only turn) |
| WebSocket sideband (existing call) | wss://api.openai.com/v1/realtime?call_id=rtc_… |
standard key | already set by call/accept (model ignored) |
none (media stays on WebRTC/SIP) | same socket | server-managed | app server monitoring a WebRTC/SIP call | UNVERIFIED |
| WebRTC — ephemeral key | browser POST https://api.openai.com/v1/realtime/calls body application/sdp |
Bearer <ek_…> minted by your server via POST /v1/realtime/client_secrets |
client secret's session (overridable via session.update on the data channel) |
media tracks (Opus 48 kHz negotiated) | data channel label oai-events |
server truncates unplayed audio automatically | browsers, mobile | UNVERIFIED |
| WebRTC — unified interface | your server POST /v1/realtime/calls multipart/form-data parts sdp (application/sdp) + session (application/json) |
standard key on your server | session form part |
media tracks | data channel oai-events |
server | browsers via your server (server in the critical path) | UNVERIFIED |
| SIP | trunk → sip:<PROJECT_ID>@sip.api.openai.com;transport=tls (EU: sip-eu.api.openai.com) |
project webhook realtime.call.incoming + standard key for `/v1/realtime/calls/{call_id}/accept |
reject | refer | hangup` | accept body (same fields as a client-secret session, top level) |
SRTP via provider (signaling TLS tcp/5061; media UDP from 13.79.45.80/28, 23.98.140.64/28, 40.67.149.176/28, 40.83.204.240/28) |
sideband WebSocket ?call_id= |
| Translation WebSocket | wss://api.openai.com/v1/realtime/translations?model=gpt-realtime-translate |
standard key or ek_ from POST /v1/realtime/translations/client_secrets |
session.update {audio.output.language} |
session.input_audio_buffer.append / session.output_audio.delta (24 kHz PCM16) |
same socket | n/a (no responses) | server media | client secret LIVE_VERIFIED; socket UNVERIFIED |
| Translation WebRTC | POST https://api.openai.com/v1/realtime/translations/calls (application/sdp) — guide only, not in the OpenAPI spec |
Bearer <ek_…> |
client secret | media tracks | data channel oai-events |
n/a | browsers | UNVERIFIED |
Headers on any connection/creation request: OpenAI-Safety-Identifier: <hashed end-user id> (recommended; when minting a client secret, set it on the /client_secrets request — it is bound to the token). Do not send OpenAI-Beta: realtime=v1 to the GA interface.
2.1 Ephemeral keys (POST /v1/realtime/client_secrets) — LIVE_VERIFIED
Request: {"expires_after": {"anchor": "created_at", "seconds": 10–7200 (default 600)}, "session": {…realtime or transcription session…}}.
Observed response (2026-09-18, session={type:realtime, model:gpt-realtime-mini}): HTTP 200, value: "ek_…", expires_at = now + 600 s, and a fully resolved session object showing the server defaults:
{"type":"realtime","object":"realtime.session","id":"sess_…","model":"gpt-realtime-mini",
"output_modalities":["audio"],
"instructions":"Your knowledge cutoff is 2023-10. You are a helpful, witty, and friendly AI. …",
"tools":[],"tool_choice":"auto","max_output_tokens":"inf","tracing":null,"truncation":"auto","prompt":null,"expires_at":0,
"audio":{"input":{"format":{"type":"audio/pcm","rate":24000},"transcription":null,"noise_reduction":null,
"turn_detection":{"type":"server_vad","threshold":0.5,"prefix_padding_ms":300,"silence_duration_ms":200,
"idle_timeout_ms":null,"create_response":true,"interrupt_response":true}},
"output":{"format":{"type":"audio/pcm","rate":24000},"voice":"alloy","speed":1.0}},
"include":null}A transcription-session secret ({"type":"transcription","audio":{"input":{"transcription":{"model":"gpt-4o-mini-transcribe"}}}}, TTL 60 s) returned object: "realtime.transcription_session" with turn_detection defaulting to server_vad (200 ms silence) — LIVE_VERIFIED. A secret can create multiple sessions until it expires; the session may outlive the secret.
2.2 Legacy beta endpoints
POST /v1/realtime/sessions and POST /v1/realtime/transcription_sessions are still in openapi-master.yaml and in the model pages' endpoint tables, but on 2026-09-18 both returned 404 Invalid URL even with OpenAI-Beta: realtime=v1 → status LEGACY, BETA, FAILED_VERIFICATION. Migrate: remove the beta header, use /v1/realtime/client_secrets, use /v1/realtime/calls for WebRTC, set session.type, move output audio config under session.audio.output, and adopt the GA event names (response.output_text.delta, response.output_audio.delta, response.output_audio_transcript.delta; conversation.item.added/done instead of conversation.item.created).
3. Session object (shared by client secret, session.update, /calls session part, /calls/{id}/accept)
Flattened with types/enums/defaults in generated/fragments/parameters/openai-realtime.json. Key fields (realtime type):
| Field | Type / enum | Default | Notes |
|---|---|---|---|
type |
realtime | transcription |
— | required discriminator |
model |
realtime model id | — | cannot be changed by session.update |
output_modalities |
["audio"] | ["text"] |
["audio"] |
audio implies a transcript too; text-only verified live |
instructions |
string | server default prompt (see above) | pass "" to clear |
audio.input.format |
{type: audio/pcm, rate: 24000} | {type: audio/pcmu} | {type: audio/pcma} |
pcm 24 kHz | G.711 8 kHz for telephony |
audio.input.transcription |
{model, language, languages[], keywords[], prompt, delay} or null |
null (off) |
models: whisper-1, gpt-transcribe, gpt-live-transcribe, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize, gpt-realtime-whisper; delay ∈ minimal/low/medium/high/xhigh; billed separately |
audio.input.noise_reduction |
{type: near_field | far_field} or null |
null |
filters the input buffer before VAD/model |
audio.input.turn_detection |
{type: server_vad, threshold, prefix_padding_ms, silence_duration_ms, idle_timeout_ms, create_response, interrupt_response} | {type: semantic_vad, eagerness: low/medium/high/auto, create_response, interrupt_response} | null |
server_vad 0.5 / 300 / 200 ms, create+interrupt true | see §5 |
audio.output.format |
same union as input | pcm 24 kHz | |
audio.output.voice |
alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar or {id: "voice_…"} (custom) |
alloy |
locked after the first audio output; marin/cedar recommended |
audio.output.speed |
0.25–1.5 | 1.0 | |
tools[] |
{type: function, name, description, parameters} | {type: mcp, server_label, server_url | connector_id | tunnel_id, authorization, headers, allowed_tools, require_approval, server_description, defer_loading} |
[] |
§6 |
tool_choice |
auto | none | required | {type: function, name} | {type: mcp, server_label, name} |
auto |
|
parallel_tool_calls |
boolean | — | reasoning models only (gpt-realtime-2…) |
reasoning.effort |
minimal/low/medium/high/xhigh | low | reasoning-capable models |
max_output_tokens |
1–4096 | "inf" |
"inf" |
per response, incl. tool calls |
truncation |
auto | disabled | {type: retention_ratio, retention_ratio 0–1, token_limits: {post_instructions}} |
auto |
cost control; see cost guide |
tracing |
"auto" | {workflow_name, group_id, metadata} | null |
null |
Traces dashboard |
prompt |
{id, version, variables} |
null |
reusable prompt template |
include |
["item.input_audio_transcription.logprobs"] |
null |
session.update merges fields (only present fields change; clear with "", [], null); server answers with session.updated carrying the full effective config. voice and model are immutable (voice only until first audio).
4. Session lifecycle & core event flow
Observed live (WebSocket, gpt-realtime-mini, output_modalities: ["text"], 3.6 s total, usage 13 input / 3 output text tokens):
→ session.update → conversation.item.create(user "Reply with OK.") → response.create({output_modalities:[text], max_output_tokens:16})
← session.created
← session.updated
← conversation.item.added (user item, status completed)
← conversation.item.done
← response.created (status in_progress)
← response.output_item.added (assistant message, content [])
← conversation.item.added (assistant item, in_progress)
← response.content_part.added (part {type:text, text:""})
← response.output_text.delta ("OK", plus an `obfuscation` padding field)
← response.output_text.done
← response.content_part.done
← conversation.item.done
← response.output_item.done
← response.done (status completed, usage{input_tokens, output_tokens, *_token_details incl. cached_tokens})Audio turn (documented): the same skeleton with response.output_audio.delta / response.output_audio.done and response.output_audio_transcript.delta / .done instead of the text events; rate_limits.updated is documented at response start (not observed in our text run). With VAD on, the server itself creates items and responses (input_audio_buffer.speech_started → .speech_stopped → .committed → conversation.item.added → response.created …). Every client event may carry an event_id which is echoed in error.event_id when it fails.
Full-audio input alternatives: conversation.item.create with content: [{type: input_audio, audio: <b64>}] (whole message), or streaming via input_audio_buffer.append + commit (VAD off). conversation.item.retrieve returns an item with its audio; conversation.item.delete removes it; conversation.item.truncate cuts assistant audio at audio_end_ms.
Responses outside the default conversation: response.create {response: {conversation: "none", metadata, input: [...items or {type:item_reference,id}], instructions, tools, output_modalities, …}} → results arrive with the same response.* events but are not appended to the conversation (input: [] = no context).
Errors: server event error {error: {type, code, message, param, event_id}}; e.g. sending an unknown event type yields invalid_request_error with the offending event_id. HTTP-side errors observed: POST /v1/realtime/calls/rtc_bogus/hangup → 404 {type: invalid_request_error, code: call_id_not_found, param: ""}.
5. Voice activity detection, interruptions, push-to-talk
| Mode | Fields | Behaviour |
|---|---|---|
server_vad (default) |
threshold 0–1 (0.5), prefix_padding_ms (300), silence_duration_ms (200; guide example 500), idle_timeout_ms, create_response (true), interrupt_response (true) |
silence-based chunking; events input_audio_buffer.speech_started / speech_stopped / committed; auto response + auto interrupt |
semantic_vad |
eagerness low/medium/high/auto (=medium), create_response, interrupt_response |
classifier estimates end-of-utterance; longer wait after "ummm…" |
null |
— | manual: input_audio_buffer.append → commit → response.create; input_audio_buffer.clear before a new turn |
VAD on, create_response:false, interrupt_response:false |
keep turn detection but decide when to respond (moderation/RAG) |
create_response/interrupt_response are conversation-only; in transcription sessions VAD only chunks audio. gpt-realtime-whisper requires turn_detection: null.
Interruptions/truncation: with VAD the server cancels the in-progress response on new speech (response.cancelled status in response.done, output_audio_buffer.* events on WebRTC/SIP). WebRTC/SIP: server knows playback position and truncates automatically. WebSocket: client must stop playback, measure played ms and send conversation.item.truncate {item_id, content_index, audio_end_ms} → conversation.item.truncated (audio cut; transcript of the unplayed part removed, not re-aligned). Push-to-talk: WS → turn_detection:null, response.cancel, truncate, append, commit, response.create; WebRTC/SIP → also input_audio_buffer.clear on press and output_audio_buffer.clear to drop unplayed audio.
6. Tools and MCP in Realtime
- Function tools (
session.toolsor per-turnresponse.tools): model emitsresponse.function_call_arguments.delta/.done;response.done.output[i]hastype: function_call, name, arguments (JSON string), call_id. Reply withconversation.item.create {item: {type: function_call_output, call_id, output: "<json string>"}}thenresponse.create. - MCP tools (executed by the Realtime API itself):
{type: mcp, server_label, server_url | connector_id (deprecated for models after 2026-09-01; e.g. connector_googlecalendar, gmail, dropbox, googledrive, microsoftteams, outlookcalendar, outlookemail, sharepoint) | tunnel_id (Secure MCP Tunnel), authorization, headers, allowed_tools {tool_names, read_only}, require_approval always|never|{always,never}, server_description, defer_loading}. Flow:mcp_list_tools.in_progress→.completed(or.failed;conversation.item.donewithitem.type: mcp_list_toolslists imported tools) → on callresponse.mcp_call_arguments.delta/.done→ optionalmcp_approval_requestitem (answer viaconversation.item.create {item:{type: mcp_approval_response, approval_request_id, approve, reason}}) →response.mcp_call.in_progress→response.output_item.done(item.type: mcp_call) orresponse.mcp_call.failed.response.donecan precede MCP completion; send anotherresponse.createafterwards — follow-ups are not automatic.server_labelalone can re-reference a definition within the session. Validation failures: duplicateserver_label, bothserver_urlandconnector_id,authorization+headers.Authorization, invalid connector id. - Remote MCP servers see only what the model sends in the call, not the whole conversation; narrow
allowed_tools, require approval for side effects.
7. Sideband / server-side controls
WebRTC: the POST /v1/realtime/calls response has a Location: /v1/realtime/calls/rtc_xxx header → your server opens wss://api.openai.com/v1/realtime?call_id=rtc_xxx with the standard key and receives/sends the same events (tool execution, session.update, monitoring) while the browser keeps the media and the ephemeral key. SIP: the webhook gives call_id; after accept, open the same URL. The sideband lives for the life of the call.
8. SIP calls (Realtime)
- Create a project webhook for
realtime.call.incoming(platform settings → Webhooks). Payload:{object: event, id: evt_…, type: realtime.call.incoming, created_at, data: {call_id, sip_headers: [{name, value}]}}; headerswebhook-id,webhook-timestamp,webhook-signature(verify withclient.webhooks.unwrap). - Point the trunk at
sip:<proj_…>@sip.api.openai.com;transport=tls(sip-eu.for EU residency). POST /v1/realtime/calls/{call_id}/acceptwith the session config at top level (type,model,instructions,audio.output.voice,tools…) → 200 once ringing; or…/reject {status_code}(default 603 Decline; 486 busy).- Attach
wss://api.openai.com/v1/realtime?call_id=…, typically sendresponse.create {response:{instructions:"greet…"}}first. …/refer {target_uri: "tel:+1…" | "sip:agent@example.com"}(SIP REFER) and…/hangup(also ends WebRTC calls). Outbound calls are not created via the API.
9. Transcription-only sessions (type: transcription)
Config: session.audio.input.{format, transcription{model, prompt, keywords[], languages[] | language, delay}, turn_detection}; no responses are generated. Recommended model gpt-live-transcribe (low latency, delay tuning, no timestamps/speaker labels/confidence), gpt-transcribe for committed-turn transcription with detected languages (WebSocket only), legacy gpt-4o-(mini-)transcribe, gpt-realtime-whisper (VAD must be null, $0.017/min). Events: input_audio_buffer.committed, conversation.item.added, conversation.item.input_audio_transcription.delta (item_id, content_index, delta), .segment (diarization-style segments where supported), .completed (transcript, usage, languages for gpt-transcribe, optional logprobs with include), .failed. Completion ordering across turns is not guaranteed — match on item_id. Language codes: ISO 639-1, selected ISO 639-3 (eng, yue, cmn), zh-cn/zh-tw/zh-hk. Keywords: one line each, no <, >, CR, LF. Speaker diarization is not available in Realtime sessions (use /v1/audio/transcriptions with gpt-4o-transcribe-diarize). Transcription sessions are billed by audio duration.
10. Translation sessions (gpt-realtime-translate) — client secret LIVE_VERIFIED
Dedicated endpoint (/v1/realtime/translations), no conversation/response lifecycle, no response.create; stream audio continuously including silence. Client events (3): session.update {session:{audio:{output:{language}, input:{transcription, noise_reduction}}}}, session.input_audio_buffer.append {audio}, session.close (flushes and ends with session.closed; only supported for translation sessions). Server events (6): session.created, session.updated, session.closed, session.input_transcript.delta, session.output_transcript.delta, session.output_audio.delta, plus error. Observed client-secret session: {"type":"translation","model":"gpt-realtime-translate","audio":{"input":{"noise_reduction":null,"transcription":null},"output":{"language":"es"}}} (default target language es). One session per (source track × target language). Pricing: $0.034/min (model page).
11. Images, voices, formats
Image input: conversation.item.create user message with a content part of type input_image (gpt-realtime-2, gpt-realtime and later support images). Audio formats: PCM16 24 kHz mono little-endian (default), G.711 μ-law/A-law 8 kHz. Realtime voices: alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar (+ custom {id}); the TTS endpoint additionally has fable, onyx, nova.
12. Costs (pointers — pricing fragment owned by the models agent)
Per-response billing on input+output tokens across text/audio/image; user audio = 1 token / 100 ms, assistant audio = 1 token / 50 ms; whole conversation is re-sent each turn (cached input discounts apply best-effort — keep history static, put instructions/tools first). Input transcription billed separately at the transcription model's rate. Transcription and translation sessions are billed by duration. Model pages (2026-09-18): gpt-realtime-2.1 text $4/$0.4 cached/$24 out per 1M, audio $32/$0.4/$64; gpt-realtime-mini text $0.6/$0.06/$2.4; gpt-realtime-2.1-mini audio $10/$0.3/$20; gpt-4o-realtime-preview audio $40/$2.5/$80. Control cost with max_output_tokens, truncation.retention_ratio + token_limits.post_instructions, conversation.item.delete, mini models. Read usage in response.done.response.usage and conversation.item.input_audio_transcription.completed.usage.
13. Live verification log (2026-09-18)
| Call | Result |
|---|---|
POST /v1/realtime/client_secrets (realtime, gpt-realtime-mini) |
200, ek_…, expiry +600 s |
POST /v1/realtime/client_secrets (transcription, 60 s) |
200, realtime.transcription_session |
POST /v1/realtime/translations/client_secrets |
200, type: translation, language es |
POST /v1/realtime/sessions (+beta header) |
404 Invalid URL |
POST /v1/realtime/transcription_sessions (+beta header) |
404 Invalid URL |
POST /v1/realtime/calls/rtc_bogus/hangup |
404 call_id_not_found |
WS wss://api.openai.com/v1/realtime?model=gpt-realtime-mini |
101; 3 client events → 14 server events → response.done (13 in / 3 out tokens, ≈$0.00001) |
Estimated cost of the whole realtime probe set: < $0.005.
14. Uncertainties
- WebRTC (
/v1/realtime/calls), SIP and?call_id=sideband not exercised (no peer/trunk).Locationheader behaviour taken from the guide. /v1/realtime/translations/callsexists only in the translation guide, not in the OpenAPI spec.- Server-events reference still lists
conversation.createdandconversation.item.created(pre-GA names) at the end of the page; GA sockets emittedconversation.item.added/done. - Model pages still label the transcription route
v1/realtime/transcription_sessionsalthough the REST endpoint 404s; transcription sessions are created via client secrets /session.update.