# OpenAI Realtime API (GA) — connections, sessions, events, tools, SIP, transcription, translation **Status**: DOCUMENTED + LIVE_VERIFIED (client secrets, translation client secrets, WebSocket text-only turn on `gpt-realtime-mini`) · LEGACY/FAILED_VERIFICATION for the beta `/v1/realtime/sessions` endpoints · UNVERIFIED for WebRTC, SIP and sideband (need a browser peer / SIP trunk). **Sources**: https://developers.openai.com/api/docs/guides/realtime · …/guides/voice-websockets?api=realtime · …/guides/voice-webrtc?api=realtime · …/guides/voice-sip?api=realtime · …/guides/voice-server-controls?api=realtime · …/guides/realtime-conversations · …/guides/realtime-vad · …/guides/realtime-mcp · …/guides/realtime-transcription · …/guides/realtime-translation · …/guides/voice-latency-cost?api=realtime · https://developers.openai.com/api/reference/resources/realtime/client-events · …/server-events · …/subresources/client_secrets/methods/create · …/subresources/calls/methods/{create,accept,reject,refer,hangup} · https://developers.openai.com/api/reference/realtime-beta/overview · `sources/openai/openapi/openapi-master.yaml` **Last verified**: 2026-09-18 **Machine-readable twins**: `generated/fragments/endpoints/openai-realtime-live-audio.json`, `generated/fragments/parameters/openai-realtime.json` (560 params), `generated/fragments/streaming-events/openai-realtime.json` (65 events). Full event tables: [`realtime-events.md`](realtime-events.md). ## 1. What it is A stateful, low-latency speech-to-speech (and text/image-in) session with a Realtime model (`gpt-realtime-2.1`, `gpt-realtime-2.1-mini`, `gpt-realtime-2`, `gpt-realtime-1.5`, `gpt-realtime`, `gpt-realtime-mini`, legacy `gpt-4o-(mini-)realtime-preview`). The server keeps a **Session** (config), a **Conversation** (items) and produces **Responses** (items appended to the conversation). Clients drive it with JSON **client events** and react to **server events**. Max session duration: **60 minutes**. Three session *types* share the surface: `realtime` (voice agent), `transcription` (input transcription only, no model replies), `translation` (dedicated `/v1/realtime/translations` endpoint, continuous interpreter, `gpt-realtime-translate`). Sibling product: **GPT-Live** (`/v1/live/sessions`, full-duplex, delegation to a backend) — see [`live.md`](live.md). Same transports, **different handshakes, credentials and event names**. ## 2. Connection methods matrix | Transport | URL | Auth | Config carried by | Audio path | Events path | Interruption/truncation | Who | Status | |---|---|---|---|---|---|---|---|---| | **WebSocket (new session)** | `wss://api.openai.com/v1/realtime?model=` | `Authorization: Bearer ` (server) **or** `Bearer `; browsers without header support use subprotocols `realtime`, `openai-insecure-api-key.`, optional `openai-organization.`, `openai-project.` | `session.update` client event (or the client secret's session) | base64 in `input_audio_buffer.append` / `response.output_audio.delta` (≤15 MB per chunk) | same socket | **client** must stop playback + send `conversation.item.truncate` | server-to-server; Deno/Workers | LIVE_VERIFIED (text-only turn) | | **WebSocket sideband (existing call)** | `wss://api.openai.com/v1/realtime?call_id=rtc_…` | standard key | already set by call/accept (`model` ignored) | none (media stays on WebRTC/SIP) | same socket | server-managed | app server monitoring a WebRTC/SIP call | UNVERIFIED | | **WebRTC — ephemeral key** | browser `POST https://api.openai.com/v1/realtime/calls` body `application/sdp` | `Bearer ` minted by your server via `POST /v1/realtime/client_secrets` | client secret's `session` (overridable via `session.update` on the data channel) | media tracks (Opus 48 kHz negotiated) | data channel label **`oai-events`** | **server** truncates unplayed audio automatically | browsers, mobile | UNVERIFIED | | **WebRTC — unified interface** | your server `POST /v1/realtime/calls` `multipart/form-data` parts `sdp` (`application/sdp`) + `session` (`application/json`) | standard key on your server | `session` form part | media tracks | data channel `oai-events` | server | browsers via your server (server in the critical path) | UNVERIFIED | | **SIP** | trunk → `sip:@sip.api.openai.com;transport=tls` (EU: `sip-eu.api.openai.com`) | project webhook `realtime.call.incoming` + standard key for `/v1/realtime/calls/{call_id}/accept|reject|refer|hangup` | `accept` body (same fields as a client-secret `session`, top level) | SRTP via provider (signaling TLS tcp/5061; media UDP from `13.79.45.80/28`, `23.98.140.64/28`, `40.67.149.176/28`, `40.83.204.240/28`) | sideband WebSocket `?call_id=` | server | phone calls | UNVERIFIED | | **Translation WebSocket** | `wss://api.openai.com/v1/realtime/translations?model=gpt-realtime-translate` | standard key or ek_ from `POST /v1/realtime/translations/client_secrets` | `session.update {audio.output.language}` | `session.input_audio_buffer.append` / `session.output_audio.delta` (24 kHz PCM16) | same socket | n/a (no responses) | server media | client secret LIVE_VERIFIED; socket UNVERIFIED | | **Translation WebRTC** | `POST https://api.openai.com/v1/realtime/translations/calls` (`application/sdp`) — guide only, **not in the OpenAPI spec** | `Bearer ` | client secret | media tracks | data channel `oai-events` | n/a | browsers | UNVERIFIED | Headers on any connection/creation request: `OpenAI-Safety-Identifier: ` (recommended; when minting a client secret, set it on the `/client_secrets` request — it is bound to the token). **Do not** send `OpenAI-Beta: realtime=v1` to the GA interface. ### 2.1 Ephemeral keys (`POST /v1/realtime/client_secrets`) — LIVE_VERIFIED Request: `{"expires_after": {"anchor": "created_at", "seconds": 10–7200 (default 600)}, "session": {…realtime or transcription session…}}`. Observed response (2026-09-18, `session={type:realtime, model:gpt-realtime-mini}`): HTTP 200, `value: "ek_…"`, `expires_at = now + 600 s`, and a fully resolved `session` object showing the **server defaults**: ```json {"type":"realtime","object":"realtime.session","id":"sess_…","model":"gpt-realtime-mini", "output_modalities":["audio"], "instructions":"Your knowledge cutoff is 2023-10. You are a helpful, witty, and friendly AI. …", "tools":[],"tool_choice":"auto","max_output_tokens":"inf","tracing":null,"truncation":"auto","prompt":null,"expires_at":0, "audio":{"input":{"format":{"type":"audio/pcm","rate":24000},"transcription":null,"noise_reduction":null, "turn_detection":{"type":"server_vad","threshold":0.5,"prefix_padding_ms":300,"silence_duration_ms":200, "idle_timeout_ms":null,"create_response":true,"interrupt_response":true}}, "output":{"format":{"type":"audio/pcm","rate":24000},"voice":"alloy","speed":1.0}}, "include":null} ``` A transcription-session secret (`{"type":"transcription","audio":{"input":{"transcription":{"model":"gpt-4o-mini-transcribe"}}}}`, TTL 60 s) returned `object: "realtime.transcription_session"` with `turn_detection` defaulting to `server_vad` (200 ms silence) — LIVE_VERIFIED. A secret can create **multiple** sessions until it expires; the session may outlive the secret. ### 2.2 Legacy beta endpoints `POST /v1/realtime/sessions` and `POST /v1/realtime/transcription_sessions` are still in `openapi-master.yaml` and in the model pages' endpoint tables, but on 2026-09-18 both returned **404 `Invalid URL`** even with `OpenAI-Beta: realtime=v1` → status `LEGACY, BETA, FAILED_VERIFICATION`. Migrate: remove the beta header, use `/v1/realtime/client_secrets`, use `/v1/realtime/calls` for WebRTC, set `session.type`, move output audio config under `session.audio.output`, and adopt the GA event names (`response.output_text.delta`, `response.output_audio.delta`, `response.output_audio_transcript.delta`; `conversation.item.added/done` instead of `conversation.item.created`). ## 3. Session object (shared by client secret, `session.update`, `/calls` `session` part, `/calls/{id}/accept`) Flattened with types/enums/defaults in `generated/fragments/parameters/openai-realtime.json`. Key fields (realtime type): | Field | Type / enum | Default | Notes | |---|---|---|---| | `type` | `realtime` \| `transcription` | — | required discriminator | | `model` | realtime model id | — | cannot be changed by `session.update` | | `output_modalities` | `["audio"]` \| `["text"]` | `["audio"]` | audio implies a transcript too; text-only verified live | | `instructions` | string | server default prompt (see above) | pass `""` to clear | | `audio.input.format` | `{type: audio/pcm, rate: 24000}` \| `{type: audio/pcmu}` \| `{type: audio/pcma}` | pcm 24 kHz | G.711 8 kHz for telephony | | `audio.input.transcription` | `{model, language, languages[], keywords[], prompt, delay}` or `null` | `null` (off) | models: `whisper-1`, `gpt-transcribe`, `gpt-live-transcribe`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe`, `gpt-4o-transcribe-diarize`, `gpt-realtime-whisper`; `delay` ∈ minimal/low/medium/high/xhigh; billed separately | | `audio.input.noise_reduction` | `{type: near_field \| far_field}` or `null` | `null` | filters the input buffer before VAD/model | | `audio.input.turn_detection` | `{type: server_vad, threshold, prefix_padding_ms, silence_duration_ms, idle_timeout_ms, create_response, interrupt_response}` \| `{type: semantic_vad, eagerness: low/medium/high/auto, create_response, interrupt_response}` \| `null` | server_vad 0.5 / 300 / 200 ms, create+interrupt true | see §5 | | `audio.output.format` | same union as input | pcm 24 kHz | | | `audio.output.voice` | `alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar` or `{id: "voice_…"}` (custom) | `alloy` | locked after the first audio output; `marin`/`cedar` recommended | | `audio.output.speed` | 0.25–1.5 | 1.0 | | | `tools[]` | `{type: function, name, description, parameters}` \| `{type: mcp, server_label, server_url \| connector_id \| tunnel_id, authorization, headers, allowed_tools, require_approval, server_description, defer_loading}` | `[]` | §6 | | `tool_choice` | `auto` \| `none` \| `required` \| `{type: function, name}` \| `{type: mcp, server_label, name}` | `auto` | | | `parallel_tool_calls` | boolean | — | reasoning models only (`gpt-realtime-2`…) | | `reasoning.effort` | minimal/low/medium/high/xhigh | low | reasoning-capable models | | `max_output_tokens` | 1–4096 \| `"inf"` | `"inf"` | per response, incl. tool calls | | `truncation` | `auto` \| `disabled` \| `{type: retention_ratio, retention_ratio 0–1, token_limits: {post_instructions}}` | `auto` | cost control; see cost guide | | `tracing` | `"auto"` \| `{workflow_name, group_id, metadata}` \| `null` | `null` | Traces dashboard | | `prompt` | `{id, version, variables}` | `null` | reusable prompt template | | `include` | `["item.input_audio_transcription.logprobs"]` | `null` | | `session.update` merges fields (only present fields change; clear with `""`, `[]`, `null`); server answers with `session.updated` carrying the full effective config. `voice` and `model` are immutable (voice only until first audio). ## 4. Session lifecycle & core event flow Observed live (WebSocket, `gpt-realtime-mini`, `output_modalities: ["text"]`, 3.6 s total, usage 13 input / 3 output text tokens): ``` → session.update → conversation.item.create(user "Reply with OK.") → response.create({output_modalities:[text], max_output_tokens:16}) ← session.created ← session.updated ← conversation.item.added (user item, status completed) ← conversation.item.done ← response.created (status in_progress) ← response.output_item.added (assistant message, content []) ← conversation.item.added (assistant item, in_progress) ← response.content_part.added (part {type:text, text:""}) ← response.output_text.delta ("OK", plus an `obfuscation` padding field) ← response.output_text.done ← response.content_part.done ← conversation.item.done ← response.output_item.done ← response.done (status completed, usage{input_tokens, output_tokens, *_token_details incl. cached_tokens}) ``` Audio turn (documented): the same skeleton with `response.output_audio.delta` / `response.output_audio.done` and `response.output_audio_transcript.delta` / `.done` instead of the text events; `rate_limits.updated` is documented at response start (not observed in our text run). With VAD on, the server itself creates items and responses (`input_audio_buffer.speech_started` → `.speech_stopped` → `.committed` → `conversation.item.added` → `response.created` …). Every client event may carry an `event_id` which is echoed in `error.event_id` when it fails. Full-audio input alternatives: `conversation.item.create` with `content: [{type: input_audio, audio: }]` (whole message), or streaming via `input_audio_buffer.append` + `commit` (VAD off). `conversation.item.retrieve` returns an item with its audio; `conversation.item.delete` removes it; `conversation.item.truncate` cuts assistant audio at `audio_end_ms`. Responses outside the default conversation: `response.create {response: {conversation: "none", metadata, input: [...items or {type:item_reference,id}], instructions, tools, output_modalities, …}}` → results arrive with the same `response.*` events but are **not** appended to the conversation (`input: []` = no context). Errors: server event `error {error: {type, code, message, param, event_id}}`; e.g. sending an unknown event type yields `invalid_request_error` with the offending `event_id`. HTTP-side errors observed: `POST /v1/realtime/calls/rtc_bogus/hangup` → 404 `{type: invalid_request_error, code: call_id_not_found, param: ""}`. ## 5. Voice activity detection, interruptions, push-to-talk | Mode | Fields | Behaviour | |---|---|---| | `server_vad` (default) | `threshold` 0–1 (0.5), `prefix_padding_ms` (300), `silence_duration_ms` (200; guide example 500), `idle_timeout_ms`, `create_response` (true), `interrupt_response` (true) | silence-based chunking; events `input_audio_buffer.speech_started` / `speech_stopped` / `committed`; auto response + auto interrupt | | `semantic_vad` | `eagerness` low/medium/high/auto (=medium), `create_response`, `interrupt_response` | classifier estimates end-of-utterance; longer wait after "ummm…" | | `null` | — | manual: `input_audio_buffer.append` → `commit` → `response.create`; `input_audio_buffer.clear` before a new turn | | VAD on, `create_response:false`, `interrupt_response:false` | | keep turn detection but decide when to respond (moderation/RAG) | `create_response`/`interrupt_response` are conversation-only; in transcription sessions VAD only chunks audio. `gpt-realtime-whisper` requires `turn_detection: null`. **Interruptions/truncation**: with VAD the server cancels the in-progress response on new speech (`response.cancelled` status in `response.done`, `output_audio_buffer.*` events on WebRTC/SIP). WebRTC/SIP: server knows playback position and truncates automatically. WebSocket: client must stop playback, measure played ms and send `conversation.item.truncate {item_id, content_index, audio_end_ms}` → `conversation.item.truncated` (audio cut; transcript of the unplayed part removed, not re-aligned). Push-to-talk: WS → `turn_detection:null`, `response.cancel`, `truncate`, `append`, `commit`, `response.create`; WebRTC/SIP → also `input_audio_buffer.clear` on press and `output_audio_buffer.clear` to drop unplayed audio. ## 6. Tools and MCP in Realtime - **Function tools** (`session.tools` or per-turn `response.tools`): model emits `response.function_call_arguments.delta/.done`; `response.done.output[i]` has `type: function_call, name, arguments (JSON string), call_id`. Reply with `conversation.item.create {item: {type: function_call_output, call_id, output: ""}}` then `response.create`. - **MCP tools** (executed by the Realtime API itself): `{type: mcp, server_label, server_url | connector_id (deprecated for models after 2026-09-01; e.g. connector_googlecalendar, gmail, dropbox, googledrive, microsoftteams, outlookcalendar, outlookemail, sharepoint) | tunnel_id (Secure MCP Tunnel), authorization, headers, allowed_tools {tool_names, read_only}, require_approval always|never|{always,never}, server_description, defer_loading}`. Flow: `mcp_list_tools.in_progress` → `.completed` (or `.failed`; `conversation.item.done` with `item.type: mcp_list_tools` lists imported tools) → on call `response.mcp_call_arguments.delta/.done` → optional `mcp_approval_request` item (answer via `conversation.item.create {item:{type: mcp_approval_response, approval_request_id, approve, reason}}`) → `response.mcp_call.in_progress` → `response.output_item.done` (`item.type: mcp_call`) or `response.mcp_call.failed`. `response.done` can precede MCP completion; **send another `response.create`** afterwards — follow-ups are not automatic. `server_label` alone can re-reference a definition within the session. Validation failures: duplicate `server_label`, both `server_url` and `connector_id`, `authorization` + `headers.Authorization`, invalid connector id. - Remote MCP servers see only what the model sends in the call, not the whole conversation; narrow `allowed_tools`, require approval for side effects. ## 7. Sideband / server-side controls WebRTC: the `POST /v1/realtime/calls` response has a `Location: /v1/realtime/calls/rtc_xxx` header → your server opens `wss://api.openai.com/v1/realtime?call_id=rtc_xxx` with the standard key and receives/sends the same events (tool execution, `session.update`, monitoring) while the browser keeps the media and the ephemeral key. SIP: the webhook gives `call_id`; after `accept`, open the same URL. The sideband lives for the life of the call. ## 8. SIP calls (Realtime) 1. Create a project webhook for `realtime.call.incoming` (platform settings → Webhooks). Payload: `{object: event, id: evt_…, type: realtime.call.incoming, created_at, data: {call_id, sip_headers: [{name, value}]}}`; headers `webhook-id`, `webhook-timestamp`, `webhook-signature` (verify with `client.webhooks.unwrap`). 2. Point the trunk at `sip:@sip.api.openai.com;transport=tls` (`sip-eu.` for EU residency). 3. `POST /v1/realtime/calls/{call_id}/accept` with the session config at top level (`type`, `model`, `instructions`, `audio.output.voice`, `tools`…) → 200 once ringing; or `…/reject {status_code}` (default 603 Decline; 486 busy). 4. Attach `wss://api.openai.com/v1/realtime?call_id=…`, typically send `response.create {response:{instructions:"greet…"}}` first. 5. `…/refer {target_uri: "tel:+1…" | "sip:agent@example.com"}` (SIP REFER) and `…/hangup` (also ends WebRTC calls). Outbound calls are not created via the API. ## 9. Transcription-only sessions (`type: transcription`) Config: `session.audio.input.{format, transcription{model, prompt, keywords[], languages[] | language, delay}, turn_detection}`; no responses are generated. Recommended model `gpt-live-transcribe` (low latency, `delay` tuning, no timestamps/speaker labels/confidence), `gpt-transcribe` for committed-turn transcription with detected `languages` (WebSocket only), legacy `gpt-4o-(mini-)transcribe`, `gpt-realtime-whisper` (VAD must be null, $0.017/min). Events: `input_audio_buffer.committed`, `conversation.item.added`, `conversation.item.input_audio_transcription.delta` (`item_id`, `content_index`, `delta`), `.segment` (diarization-style segments where supported), `.completed` (`transcript`, `usage`, `languages` for gpt-transcribe, optional `logprobs` with `include`), `.failed`. Completion ordering across turns is not guaranteed — match on `item_id`. Language codes: ISO 639-1, selected ISO 639-3 (`eng`, `yue`, `cmn`), `zh-cn/zh-tw/zh-hk`. Keywords: one line each, no `<`, `>`, CR, LF. Speaker diarization is **not** available in Realtime sessions (use `/v1/audio/transcriptions` with `gpt-4o-transcribe-diarize`). Transcription sessions are billed by audio duration. ## 10. Translation sessions (`gpt-realtime-translate`) — client secret LIVE_VERIFIED Dedicated endpoint (`/v1/realtime/translations`), no conversation/response lifecycle, no `response.create`; stream audio continuously including silence. Client events (3): `session.update {session:{audio:{output:{language}, input:{transcription, noise_reduction}}}}`, `session.input_audio_buffer.append {audio}`, `session.close` (flushes and ends with `session.closed`; **only supported for translation sessions**). Server events (6): `session.created`, `session.updated`, `session.closed`, `session.input_transcript.delta`, `session.output_transcript.delta`, `session.output_audio.delta`, plus `error`. Observed client-secret session: `{"type":"translation","model":"gpt-realtime-translate","audio":{"input":{"noise_reduction":null,"transcription":null},"output":{"language":"es"}}}` (default target language `es`). One session per (source track × target language). Pricing: $0.034/min (model page). ## 11. Images, voices, formats Image input: `conversation.item.create` user message with a content part of type `input_image` (`gpt-realtime-2`, `gpt-realtime` and later support images). Audio formats: PCM16 24 kHz mono little-endian (default), G.711 μ-law/A-law 8 kHz. Realtime voices: `alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar` (+ custom `{id}`); the TTS endpoint additionally has `fable, onyx, nova`. ## 12. Costs (pointers — pricing fragment owned by the models agent) Per-response billing on input+output tokens across text/audio/image; **user audio = 1 token / 100 ms, assistant audio = 1 token / 50 ms**; whole conversation is re-sent each turn (cached input discounts apply best-effort — keep history static, put instructions/tools first). Input transcription billed separately at the transcription model's rate. Transcription and translation sessions are billed by duration. Model pages (2026-09-18): `gpt-realtime-2.1` text $4/$0.4 cached/$24 out per 1M, audio $32/$0.4/$64; `gpt-realtime-mini` text $0.6/$0.06/$2.4; `gpt-realtime-2.1-mini` audio $10/$0.3/$20; `gpt-4o-realtime-preview` audio $40/$2.5/$80. Control cost with `max_output_tokens`, `truncation.retention_ratio` + `token_limits.post_instructions`, `conversation.item.delete`, mini models. Read usage in `response.done.response.usage` and `conversation.item.input_audio_transcription.completed.usage`. ## 13. Live verification log (2026-09-18) | Call | Result | |---|---| | `POST /v1/realtime/client_secrets` (realtime, gpt-realtime-mini) | 200, `ek_…`, expiry +600 s | | `POST /v1/realtime/client_secrets` (transcription, 60 s) | 200, `realtime.transcription_session` | | `POST /v1/realtime/translations/client_secrets` | 200, `type: translation`, language `es` | | `POST /v1/realtime/sessions` (+beta header) | 404 Invalid URL | | `POST /v1/realtime/transcription_sessions` (+beta header) | 404 Invalid URL | | `POST /v1/realtime/calls/rtc_bogus/hangup` | 404 `call_id_not_found` | | `WS wss://api.openai.com/v1/realtime?model=gpt-realtime-mini` | 101; 3 client events → 14 server events → `response.done` (13 in / 3 out tokens, ≈$0.00001) | Estimated cost of the whole realtime probe set: < $0.005. ## 14. Uncertainties - WebRTC (`/v1/realtime/calls`), SIP and `?call_id=` sideband not exercised (no peer/trunk). `Location` header behaviour taken from the guide. - `/v1/realtime/translations/calls` exists only in the translation guide, not in the OpenAPI spec. - Server-events reference still lists `conversation.created` and `conversation.item.created` (pre-GA names) at the end of the page; GA sockets emitted `conversation.item.added/done`. - Model pages still label the transcription route `v1/realtime/transcription_sessions` although the REST endpoint 404s; transcription sessions are created via client secrets / `session.update`.