# OpenAI Realtime / Live / Audio — full event reference **Status**: DOCUMENTED (all events, parsed from the official reference pages) · LIVE_VERIFIED where marked ✅ (observed 2026-09-18 with our key). Machine-readable twin: `generated/fragments/streaming-events/openai-{realtime,live,audio-transcription}.json` (each record carries the full field tree to depth 2 + the official example). **Sources**: https://developers.openai.com/api/reference/resources/realtime/client-events · …/realtime/server-events · …/realtime/translation-client-events · …/realtime/translation-server-events · https://developers.openai.com/api/reference/resources/live/primary-websocket · …/live/fork-websocket · …/live/sideband-websocket · https://developers.openai.com/api/reference/resources/audio/subresources/transcriptions/streaming-events · …/audio/subresources/speech/methods/create **Last verified**: 2026-09-18 Legend: `field?` = optional; `a{b, c}` = object with listed sub-fields; ✅ = observed live. ## 1. Realtime API (`wss://api.openai.com/v1/realtime`) — 11 client events, 45 server events ### 1.1 Observed server-event sequence (text-only turn, gpt-realtime-mini, 2026-09-18) Client sent: `session.update` → `conversation.item.create` → `response.create`. Server sent, in order: ``` 1. session.created 2. session.updated 3. conversation.item.added 4. conversation.item.done 5. response.created 6. response.output_item.added 7. conversation.item.added 8. response.content_part.added 9. response.output_text.delta 10. response.output_text.done 11. response.content_part.done 12. conversation.item.done 13. response.output_item.done 14. response.done ``` Note: `conversation.item.added` is emitted twice (once for the user item, once for the assistant item, the latter right after `response.output_item.added`). `rate_limits.updated` was NOT emitted in this run (it is documented as emitted at the start of a response). No `response.output_audio*` events because `output_modalities` was `["text"]`. ### 1.2 Client events (client → server) | Event | Live | Description | Payload fields | |---|---|---|---| | `session.update` | ✅ | Send this event to update the session’s configuration. The client may send this event at any time to update any field except for `voice` and `model`. `voice` can be updated only if there have been no other audio outputs yet. When the server receives a `session.update`, it will respond with a `session.updated` event showing the full, effective configuration. Only the fields that are present in the | `session{RealtimeSessionCreateRequest | RealtimeTranscriptionSessionCreateRequest}`, `type`, `event_id`? | | `input_audio_buffer.append` | | Send this event to append audio bytes to the input audio buffer. The audio buffer is temporary storage you can write to and later commit. A "commit" will create a new user message item in the conversation history from the buffer content and clear the buffer. Input audio transcription (if enabled) will be generated when the buffer is committed. If VAD is enabled the audio buffer is used to detect s | `audio`, `type`, `event_id`? | | `input_audio_buffer.commit` | | Send this event to commit the user input audio buffer, which will create a new user message item in the conversation. This event will produce an error if the input audio buffer is empty. When in Server VAD mode, the client does not need to send this event, the server will commit the audio buffer automatically. Committing the input audio buffer will trigger input audio transcription (if enable | `type`, `event_id`? | | `input_audio_buffer.clear` | | Send this event to clear the audio bytes in the buffer. The server will respond with an `input_audio_buffer.cleared` event. | `type`, `event_id`? | | `conversation.item.create` | ✅ | Add a new Item to the Conversation's context, including messages, function calls, and function call responses. This event can be used both to populate a "history" of the conversation and to add new items mid-stream, but has the current limitation that it cannot populate assistant audio messages. If successful, the server will emit a `conversation.item.added` event and, when the item is finalized, | `item{RealtimeConversationItemSystemMessage | RealtimeConversationItemUserMessage | RealtimeConversationItemAssistantMessage | RealtimeConversationItemFunctionCall | RealtimeConversationItemFunctionCallOutput | RealtimeMcpApprovalResponse | RealtimeMcpListTools | RealtimeMcpToolCall | RealtimeMcpApprovalRequest}`, `type`, `event_id`?, `previous_item_id`? | | `conversation.item.retrieve` | | Send this event when you want to retrieve the server's representation of a specific item in the conversation history. This is useful, for example, to inspect user audio after noise cancellation and VAD. The server will respond with a `conversation.item.retrieved` event, unless the item does not exist in the conversation history, in which case the server will respond with an error. | `item_id`, `type`, `event_id`? | | `conversation.item.truncate` | | Send this event to truncate a previous assistant message’s audio. The server will produce audio faster than realtime, so this event is useful when the user interrupts to truncate audio that has already been sent to the client but not yet played. This will synchronize the server's understanding of the audio with the client's playback. Truncating audio will delete the server-side text transcript to | `audio_end_ms`, `content_index`, `item_id`, `type`, `event_id`? | | `conversation.item.delete` | | Send this event when you want to remove any item from the conversation history. The server will respond with a `conversation.item.deleted` event, unless the item does not exist in the conversation history, in which case the server will respond with an error. | `item_id`, `type`, `event_id`? | | `response.create` | ✅ | This event instructs the server to create a Response, which means triggering model inference. When in Server VAD mode, the server will create Responses automatically. A Response will include at least one Item, and may have two, in which case the second will be a function call. These Items will be appended to the conversation history by default. The server will respond with a `response.created` eve | `type`, `event_id`?, `response{audio, conversation, input, instructions, max_output_tokens, metadata, output_modalities, parallel_tool_calls…}`? | | `response.cancel` | | Send this event to cancel an in-progress response. The server will respond with a `response.done` event with a status of `response.status=cancelled`. If there is no response to cancel, the server will respond with an error. It's safe to call `response.cancel` even if no response is in progress, an error will be returned the session will remain unaffected. | `type`, `event_id`?, `response_id`? | | `output_audio_buffer.clear` | | **WebRTC/SIP Only:** Emit to cut off the current audio response. This will trigger the server to stop generating audio and emit a `output_audio_buffer.cleared` event. This event should be preceded by a `response.cancel` client event to stop the generation of the current response. [Learn more](https://developers.openai.com/api/docs/guides/realtime-conversations#client-and-server-events-for-audio-in | `type`, `event_id`? | ### 1.3 Server events (server → client) | Event | Live | Description | Payload fields | |---|---|---|---| | `session.created` | ✅ | Returned when a Session is created. Emitted automatically when a new connection is established as the first server event. This event will contain the default Session configuration. | `event_id`, `session{RealtimeSessionCreateResponse | RealtimeTranscriptionSessionCreateResponse}`, `type` | | `session.updated` | ✅ | Returned when a session is updated with a `session.update` event, unless there is an error. | `event_id`, `session{RealtimeSessionCreateResponse | RealtimeTranscriptionSessionCreateResponse}`, `type` | | `conversation.item.added` | ✅ | Sent by the server when an Item is added to the default Conversation. This can happen in several cases: - When the client sends a `conversation.item.create` event. - When the input audio buffer is committed. In this case the item will be a user message containing the audio from the buffer. - When the model is generating a Response. In this case the `conversation.item.added` event will be sent when | `event_id`, `item{RealtimeConversationItemSystemMessage | RealtimeConversationItemUserMessage | RealtimeConversationItemAssistantMessage | RealtimeConversationItemFunctionCall | RealtimeConversationItemFunctionCallOutput | RealtimeMcpApprovalResponse | RealtimeMcpListTools | RealtimeMcpToolCall | RealtimeMcpApprovalRequest}`, `type`, `previous_item_id`? | | `conversation.item.done` | ✅ | Returned when a conversation item is finalized. The event will include the full content of the Item except for audio data, which can be retrieved separately with a `conversation.item.retrieve` event if needed. | `event_id`, `item{RealtimeConversationItemSystemMessage | RealtimeConversationItemUserMessage | RealtimeConversationItemAssistantMessage | RealtimeConversationItemFunctionCall | RealtimeConversationItemFunctionCallOutput | RealtimeMcpApprovalResponse | RealtimeMcpListTools | RealtimeMcpToolCall | RealtimeMcpApprovalRequest}`, `type`, `previous_item_id`? | | `conversation.item.retrieved` | | Returned when a conversation item is retrieved with `conversation.item.retrieve`. This is provided as a way to fetch the server's representation of an item, for example to get access to the post-processed audio data after noise cancellation and VAD. It includes the full content of the Item, including audio data. | `event_id`, `item{RealtimeConversationItemSystemMessage | RealtimeConversationItemUserMessage | RealtimeConversationItemAssistantMessage | RealtimeConversationItemFunctionCall | RealtimeConversationItemFunctionCallOutput | RealtimeMcpApprovalResponse | RealtimeMcpListTools | RealtimeMcpToolCall | RealtimeMcpApprovalRequest}`, `type` | | `conversation.item.input_audio_transcription.completed` | | This event is the output of audio transcription for user audio written to the user audio buffer. Transcription begins when the input audio buffer is committed by the client or server (when VAD is enabled). Transcription runs asynchronously with Response creation, so this event may come before or after the Response events. Realtime API models accept audio natively, and thus input transcription is a | `content_index`, `event_id`, `item_id`, `transcript`, `type`, `usage{Tokens | Duration}`, `languages{code}`?, `logprobs{token, bytes, logprob}`? | | `conversation.item.input_audio_transcription.delta` | | Returned when the text value of an input audio transcription content part is updated with incremental transcription results. | `event_id`, `item_id`, `type`, `content_index`?, `delta`?, `logprobs{token, bytes, logprob}`? | | `conversation.item.input_audio_transcription.segment` | | Returned when an input audio transcription segment is identified for an item. | `id`, `content_index`, `end`, `event_id`, `item_id`, `speaker`, `start`, `text`, `type` | | `conversation.item.input_audio_transcription.failed` | | Returned when input audio transcription is configured, and a transcription request for a user message failed. These events are separate from other `error` events so that the client can identify the related Item. | `content_index`, `error{code, message, param, type}`, `event_id`, `item_id`, `type` | | `conversation.item.truncated` | | Returned when an earlier assistant audio message item is truncated by the client with a `conversation.item.truncate` event. This event is used to synchronize the server's understanding of the audio with the client's playback. This action will truncate the audio and remove the server-side text transcript to ensure there is no text in the context that hasn't been heard by the user. | `audio_end_ms`, `content_index`, `event_id`, `item_id`, `type` | | `conversation.item.deleted` | | Returned when an item in the conversation is deleted by the client with a `conversation.item.delete` event. This event is used to synchronize the server's understanding of the conversation history with the client's view. | `event_id`, `item_id`, `type` | | `input_audio_buffer.committed` | | Returned when an input audio buffer is committed, either by the client or automatically in server VAD mode. The `item_id` property is the ID of the user message item that will be created, thus a `conversation.item.created` event will also be sent to the client. | `event_id`, `item_id`, `type`, `previous_item_id`? | | `input_audio_buffer.dtmf_event_received` | | **SIP Only:** Returned when an DTMF event is received. A DTMF event is a message that represents a telephone keypad press (0–9, *, #, A–D). The `event` property is the keypad that the user press. The `received_at` is the UTC Unix Timestamp that the server received the event. | `event`, `received_at`, `type` | | `input_audio_buffer.cleared` | | Returned when the input audio buffer is cleared by the client with a `input_audio_buffer.clear` event. | `event_id`, `type` | | `input_audio_buffer.speech_started` | | Sent by the server when in `server_vad` mode to indicate that speech has been detected in the audio buffer. This can happen any time audio is added to the buffer (unless speech is already detected). The client may want to use this event to interrupt audio playback or provide visual feedback to the user. The client should expect to receive a `input_audio_buffer.speech_stopped` event when speech sto | `audio_start_ms`, `event_id`, `item_id`, `type` | | `input_audio_buffer.speech_stopped` | | Returned in `server_vad` mode when the server detects the end of speech in the audio buffer. The server will also send an `conversation.item.created` event with the user message item that is created from the audio buffer. | `audio_end_ms`, `event_id`, `item_id`, `type` | | `input_audio_buffer.timeout_triggered` | | Returned when the Server VAD timeout is triggered for the input audio buffer. This is configured with `idle_timeout_ms` in the `turn_detection` settings of the session, and it indicates that there hasn't been any speech detected for the configured duration. The `audio_start_ms` and `audio_end_ms` fields indicate the segment of audio after the last model response up to the triggering time, as an of | `audio_end_ms`, `audio_start_ms`, `event_id`, `item_id`, `type` | | `output_audio_buffer.started` | | **WebRTC/SIP Only:** Emitted when the server begins streaming audio to the client. This event is emitted after an audio content part has been added (`response.content_part.added`) to the response. [Learn more](https://developers.openai.com/api/docs/guides/realtime-conversations#client-and-server-events-for-audio-in-webrtc). | `event_id`, `response_id`, `type` | | `output_audio_buffer.stopped` | | **WebRTC/SIP Only:** Emitted when the output audio buffer has been completely drained on the server, and no more audio is forthcoming. This event is emitted after the full response data has been sent to the client (`response.done`). [Learn more](https://developers.openai.com/api/docs/guides/realtime-conversations#client-and-server-events-for-audio-in-webrtc). | `event_id`, `response_id`, `type` | | `output_audio_buffer.cleared` | | **WebRTC/SIP Only:** Emitted when the output audio buffer is cleared. This happens either in VAD mode when the user has interrupted (`input_audio_buffer.speech_started`), or when the client has emitted the `output_audio_buffer.clear` event to manually cut off the current audio response. [Learn more](https://developers.openai.com/api/docs/guides/realtime-conversations#client-and-server-events-for-a | `event_id`, `response_id`, `type` | | `response.created` | ✅ | Returned when a new Response is created. The first event of response creation, where the response is in an initial state of `in_progress`. | `event_id`, `response{id, audio, conversation_id, max_output_tokens, metadata, object, output, output_modalities…}`, `type` | | `response.done` | ✅ | Returned when a Response is done streaming. Always emitted, no matter the final state. The Response object included in the `response.done` event will include all output Items in the Response but will omit the raw audio data. Clients should check the `status` field of the Response to determine if it was successful (`completed`) or if there was another outcome: `cancelled`, `failed`, or `incomplete` | `event_id`, `response{id, audio, conversation_id, max_output_tokens, metadata, object, output, output_modalities…}`, `type` | | `response.output_item.added` | ✅ | Returned when a new Item is created during Response generation. | `event_id`, `item{RealtimeConversationItemSystemMessage | RealtimeConversationItemUserMessage | RealtimeConversationItemAssistantMessage | RealtimeConversationItemFunctionCall | RealtimeConversationItemFunctionCallOutput | RealtimeMcpApprovalResponse | RealtimeMcpListTools | RealtimeMcpToolCall | RealtimeMcpApprovalRequest}`, `output_index`, `response_id`, `type` | | `response.output_item.done` | ✅ | Returned when an Item is done streaming. Also emitted when a Response is interrupted, incomplete, or cancelled. | `event_id`, `item{RealtimeConversationItemSystemMessage | RealtimeConversationItemUserMessage | RealtimeConversationItemAssistantMessage | RealtimeConversationItemFunctionCall | RealtimeConversationItemFunctionCallOutput | RealtimeMcpApprovalResponse | RealtimeMcpListTools | RealtimeMcpToolCall | RealtimeMcpApprovalRequest}`, `output_index`, `response_id`, `type` | | `response.content_part.added` | ✅ | Returned when a new content part is added to an assistant message item during response generation. | `content_index`, `event_id`, `item_id`, `output_index`, `part{audio, text, transcript, type}`, `response_id`, `type` | | `response.content_part.done` | ✅ | Returned when a content part is done streaming in an assistant message item. Also emitted when a Response is interrupted, incomplete, or cancelled. | `content_index`, `event_id`, `item_id`, `output_index`, `part{audio, text, transcript, type}`, `response_id`, `type` | | `response.output_text.delta` | ✅ | Returned when the text value of an "output_text" content part is updated. | `content_index`, `delta`, `event_id`, `item_id`, `output_index`, `response_id`, `type` | | `response.output_text.done` | ✅ | Returned when the text value of an "output_text" content part is done streaming. Also emitted when a Response is interrupted, incomplete, or cancelled. | `content_index`, `event_id`, `item_id`, `output_index`, `response_id`, `text`, `type` | | `response.output_audio_transcript.delta` | | Returned when the model-generated transcription of audio output is updated. | `content_index`, `delta`, `event_id`, `item_id`, `output_index`, `response_id`, `type` | | `response.output_audio_transcript.done` | | Returned when the model-generated transcription of audio output is done streaming. Also emitted when a Response is interrupted, incomplete, or cancelled. | `content_index`, `event_id`, `item_id`, `output_index`, `response_id`, `transcript`, `type` | | `response.output_audio.delta` | | Returned when the model-generated audio is updated. | `content_index`, `delta`, `event_id`, `item_id`, `output_index`, `response_id`, `type` | | `response.output_audio.done` | | Returned when the model-generated audio is done. Also emitted when a Response is interrupted, incomplete, or cancelled. | `content_index`, `event_id`, `item_id`, `output_index`, `response_id`, `type` | | `response.function_call_arguments.delta` | | Returned when the model-generated function call arguments are updated. | `call_id`, `delta`, `event_id`, `item_id`, `output_index`, `response_id`, `type` | | `response.function_call_arguments.done` | | Returned when the model-generated function call arguments are done streaming. Also emitted when a Response is interrupted, incomplete, or cancelled. | `arguments`, `call_id`, `event_id`, `item_id`, `name`, `output_index`, `response_id`, `type` | | `response.mcp_call_arguments.delta` | | Returned when MCP tool call arguments are updated during response generation. | `delta`, `event_id`, `item_id`, `output_index`, `response_id`, `type`, `obfuscation`? | | `response.mcp_call_arguments.done` | | Returned when MCP tool call arguments are finalized during response generation. | `arguments`, `event_id`, `item_id`, `output_index`, `response_id`, `type` | | `response.mcp_call.in_progress` | | Returned when an MCP tool call has started and is in progress. | `event_id`, `item_id`, `output_index`, `type` | | `response.mcp_call.completed` | | Returned when an MCP tool call has completed successfully. | `event_id`, `item_id`, `output_index`, `type` | | `response.mcp_call.failed` | | Returned when an MCP tool call has failed. | `event_id`, `item_id`, `output_index`, `type` | | `mcp_list_tools.in_progress` | | Returned when listing MCP tools is in progress for an item. | `event_id`, `item_id`, `type` | | `mcp_list_tools.completed` | | Returned when listing MCP tools has completed for an item. | `event_id`, `item_id`, `type` | | `mcp_list_tools.failed` | | Returned when listing MCP tools has failed for an item. | `event_id`, `item_id`, `type` | | `rate_limits.updated` | | Emitted at the beginning of a Response to indicate the updated rate limits. When a Response is created some tokens will be "reserved" for the output tokens, the rate limits shown here reflect that reservation, which is then adjusted accordingly once the Response is completed. | `event_id`, `rate_limits{limit, name, remaining, reset_seconds}`, `type` | | `conversation.created` | | Returned when a conversation is created. Emitted right after session creation. | `conversation{id, object}`, `event_id`, `type` | | `conversation.item.created` | | Returned when a conversation item is created. There are several scenarios that produce this event: - The server is generating a Response, which if successful will produce either one or two Items, which will be of type `message` (role `assistant`) or type `function_call`. - The input audio buffer has been committed, either by the client or the server (in `server_vad` mode). The server will take the | `event_id`, `item{RealtimeConversationItemSystemMessage | RealtimeConversationItemUserMessage | RealtimeConversationItemAssistantMessage | RealtimeConversationItemFunctionCall | RealtimeConversationItemFunctionCallOutput | RealtimeMcpApprovalResponse | RealtimeMcpListTools | RealtimeMcpToolCall | RealtimeMcpApprovalRequest}`, `type`, `previous_item_id`? | ### 1.4 Translation session client events (`/v1/realtime/translations`) | Event | Live | Description | Payload fields | |---|---|---|---| | `session.update` | | Send this event to update the translation session configuration. Translation sessions support updates to `audio.output.language`, `audio.input.transcription`, and `audio.input.noise_reduction`. | `session{audio}`, `type`, `event_id`? | | `session.input_audio_buffer.append` | | Send this event to append audio bytes to the translation session input audio buffer. WebSocket translation sessions accept base64-encoded 24 kHz PCM16 mono little-endian raw audio bytes. Unsupported websocket audio formats return a validation error because lower-quality audio materially degrades translation quality. Translation consumes 200 ms engine frames. For best realtime behavior, append audi | `audio`, `type`, `event_id`? | | `session.close` | | Gracefully close the realtime translation session. The server flushes pending input audio and emits any remaining translated output before closing the session. | `type`, `event_id`? | ### 1.5 Translation session server events | Event | Live | Description | Payload fields | |---|---|---|---| | `session.created` | | Returned when a translation session is created. Emitted automatically when a new connection is established as the first server event. This event contains the default translation session configuration. | `event_id`, `session{id, audio, expires_at, model, type}`, `type` | | `session.updated` | | Returned when a translation session is updated with a `session.update` event, unless there is an error. | `event_id`, `session{id, audio, expires_at, model, type}`, `type` | | `session.closed` | | Returned when a realtime translation session is closed. | `event_id`, `type` | | `session.input_transcript.delta` | | Returned when optional source-language transcript text is available. This event is emitted only when `audio.input.transcription` is configured. Transcript deltas are append-only text fragments. Clients should not insert unconditional spaces between deltas. | `delta`, `event_id`, `type`, `elapsed_ms`? | | `session.output_transcript.delta` | | Returned when translated transcript text is available. Transcript deltas are append-only text fragments. Clients should not insert unconditional spaces between deltas. | `delta`, `event_id`, `type`, `elapsed_ms`? | | `session.output_audio.delta` | | Returned when translated output audio is available. The `delta` contains a PCM16 audio chunk whose length can vary. Clients should decode and queue the complete delta instead of assuming a fixed byte or sample count. | `delta`, `event_id`, `type`, `channels`?, `elapsed_ms`?, `format`?, `sample_rate`? | ### 1.6 Realtime event families at a glance | Family | Client events | Server events | |---|---|---| | `session.*` | `session.update` | `session.created`, `session.updated` | | `input_audio_buffer.*` | `input_audio_buffer.append`, `input_audio_buffer.commit`, `input_audio_buffer.clear` | `input_audio_buffer.committed`, `input_audio_buffer.dtmf_event_received`, `input_audio_buffer.cleared`, `input_audio_buffer.speech_started`, `input_audio_buffer.speech_stopped`, `input_audio_buffer.timeout_triggered` | | `conversation.*` | `conversation.item.create`, `conversation.item.retrieve`, `conversation.item.truncate`, `conversation.item.delete` | `conversation.item.added`, `conversation.item.done`, `conversation.item.retrieved`, `conversation.item.input_audio_transcription.completed`, `conversation.item.input_audio_transcription.delta`, `conversation.item.input_audio_transcription.segment`, `conversation.item.input_audio_transcription.failed`, `conversation.item.truncated`, `conversation.item.deleted`, `conversation.created`, `conversation.item.created` | | `response.*` | `response.create`, `response.cancel` | `response.created`, `response.done`, `response.output_item.added`, `response.output_item.done`, `response.content_part.added`, `response.content_part.done`, `response.output_text.delta`, `response.output_text.done`, `response.output_audio_transcript.delta`, `response.output_audio_transcript.done`, `response.output_audio.delta`, `response.output_audio.done`, `response.function_call_arguments.delta`, `response.function_call_arguments.done`, `response.mcp_call_arguments.delta`, `response.mcp_call_arguments.done`, `response.mcp_call.in_progress`, `response.mcp_call.completed`, `response.mcp_call.failed` | | `output_audio_buffer.*` | `output_audio_buffer.clear` | `output_audio_buffer.started`, `output_audio_buffer.stopped`, `output_audio_buffer.cleared` | | `mcp_list_tools.*` | — | `mcp_list_tools.in_progress`, `mcp_list_tools.completed`, `mcp_list_tools.failed` | | `rate_limits.*` | — | `rate_limits.updated` | ## 2. Live API (GPT-Live) WebSocket events — 11 client, 20 server (DOCUMENTED only; no Live session was opened) Connections: **primary** `wss://api.openai.com/v1/live/sessions` · **fork** `wss://api.openai.com/v1/live/sessions/{session_id}/fork` · **sideband** `wss://api.openai.com/v1/live/sessions/{session_id}/attach`. WebRTC sessions receive the same server events (minus audio) on the data channel `oai-events`. The `connections` column shows which reference pages list the event. ### 2.1 Client events | Event | Connections | Description | Payload fields | |---|---|---|---| | `session.start` | primary, fork | Start a Live session on a primary WebSocket. Send this event before other commands and wait for `session.started`. | `session{model, audio, client, delegation, input, instructions, store}`, `type`, `event_id`? | | `session.update` | primary, fork, sideband | Update the delegation settings of an active Live session. The server acknowledges accepted changes with `session.updated`. | `session{delegation}`, `type`, `event_id`? | | `session.input_audio.append` | primary, fork | Send audio to a Live session over its primary WebSocket. WebRTC and SIP sessions send audio over their media transport. | `audio`, `type`, `event_id`? | | `session.input_audio.mute` | primary, fork, sideband | Mute audio input to the Live model without closing the session. The server acknowledges with `session.input_audio.muted`. | `type`, `event_id`? | | `session.input_audio.unmute` | primary, fork, sideband | Resume audio input to a Live model after muting it. The server acknowledges with `session.input_audio.unmuted`. | `type`, `event_id`? | | `session.instructions.append` | primary, fork, sideband | Append instructions to the Live conversation while it is running, optionally associating them with an existing client delegation. | `content`, `delegation_id`, `type`, `event_id`? | | `session.thinking.append` | primary, fork, sideband | Provide silent reasoning or progress context to the Live model, optionally for an existing client delegation. | `content`, `delegation_id`, `type`, `event_id`? | | `session.commentary.append` | primary, fork, sideband | Provide context the Live model can communicate to the user, optionally for an existing client delegation. | `content`, `delegation_id`, `type`, `event_id`? | | `response.item.create` | primary, fork, sideband | Add an input item to the Live session’s Responses backend. Requires Responses delegation; use `response.create` to request a response. | `item{EasyInputMessage | Message | ResponseOutputMessage | FileSearchCall | ComputerCall | ComputerCallOutput | WebSearchCall | FunctionCall | FunctionCallOutput | ToolSearchCall | ToolSearchOutput | AdditionalTools | ConfigurationUpdate | Reasoning | Compaction | ImageGenerationCall | CodeInterpreterCall | LocalShellCall | LocalShellCallOutput | ShellCall | ShellCallOutput | ApplyPatchCall | ApplyPatchCallOutput | McpListTools | McpApprovalRequest | McpApprovalResponse | McpCall | CustomToolCallOutput | CustomToolCall | CompactionTrigger | ItemReference | Program | ProgramOutput}`, `type`, `event_id`? | | `response.create` | primary, fork, sideband | Request a response from the Live session’s Responses backend, or continue a delegated response waiting for tool results. Requires Responses delegation. | `type`, `event_id`? | | `session.close` | primary, fork, sideband | Request that the Live session close. The terminal `session.closed` event contains the close reason and final usage. | `type`, `event_id`? | ### 2.2 Server events | Event | Connections | Description | Payload fields | |---|---|---|---| | `session.started` | primary, fork, sideband | Returned when a Live session has started. Contains the resolved session configuration, including server defaults. | `event_id`, `session{id, expires_at, model, status, audio, client, delegation, input…}`, `type`, `client_event_id`? | | `session.updated` | primary, fork, sideband | Returned when a Live session update is accepted. Contains the resolved session configuration after the update. | `event_id`, `session{id, expires_at, model, status, audio, client, delegation, input…}`, `type`, `client_event_id`? | | `session.input_audio.muted` | primary, fork, sideband | Returned when a session.input_audio.mute command is accepted. Input audio is no longer sent to the model; sideband audio reflection continues. | `event_id`, `type`, `client_event_id`? | | `session.input_audio.unmuted` | primary, fork, sideband | Returned when a session.input_audio.unmute command is accepted. Input audio is sent to the model again. | `event_id`, `type`, `client_event_id`? | | `session.instructions.appended` | primary, fork, sideband | Returned when a session.instructions.append command is accepted into the Live session timeline. Acknowledges the appended instructions without guaranteeing that the model has acted on them. | `end_ms`, `event_id`, `start_ms`, `type`, `client_event_id`? | | `session.thinking.appended` | primary, fork, sideband | Returned when a session.thinking.append command is accepted into the Live session timeline. Acknowledges the added reasoning context without guaranteeing any spoken output. | `end_ms`, `event_id`, `start_ms`, `type`, `client_event_id`? | | `session.commentary.appended` | primary, fork, sideband | Returned when a session.commentary.append command is accepted into the Live session timeline. Acknowledges the added commentary without guaranteeing exact wording or completed audio playback. | `end_ms`, `event_id`, `start_ms`, `type`, `client_event_id`? | | `session.output_audio.delta` | primary, fork | An audio chunk generated by the Live model. Decode and play primary WebSocket chunks in delivery order using the configured session audio format. Sideband connections receive reflected output audio with timestamps. | `delta`, `type`, `end_ms`?, `start_ms`? | | `session.input_transcript.delta` | primary, fork, sideband | A transcript fragment for user input audio in the Live session. Accumulate fragments in delivery order; these events do not define complete turns or include a transcript-done event. | `delta`, `end_ms`, `event_id`, `start_ms`, `type`, `client_event_id`? | | `session.output_transcript.delta` | primary, fork, sideband | A transcript fragment for assistant output audio in the Live session. Accumulate fragments in delivery order; these events do not define complete turns or include a transcript-done event. | `delta`, `end_ms`, `event_id`, `start_ms`, `type`, `client_event_id`? | | `session.delegation.created` | primary, fork, sideband | Returned when the Live model delegates work to your application or a Responses backend. Contains delegation metadata and the position on the session timeline where the work was delegated. | `delegation{id, target, type, response_id}`, `event_id`, `offset_ms`, `type`, `client_event_id`? | | `response.event` | primary, fork, sideband | A streaming Responses API event from a backend delegated to by the Live session. Use the outer delegation_id to associate the nested stream with its Live delegation. | `event`, `event_id`, `type`, `client_event_id`?, `delegation_id`? | | `session.usage.updated` | primary, fork, sideband | Reports cumulative Live audio usage and, when available, the most recent context-window usage. Delegated Responses token usage is reported separately in response.event events. | `event_id`, `type`, `usage{seconds}`, `client_event_id`?, `context_window{usage_ratio}`? | | `session.closed` | primary, fork, sideband | Returned after the Live session finishes finalizing, with the close reason, final session snapshot, and cumulative audio usage. A connection closing without this event does not confirm successful finalization. | `event_id`, `reason{"close_requested"}`, `session{id, expires_at, model, status, audio, client, delegation, input…}`, `type`, `usage{seconds}`, `client_event_id`? | | `session.input_audio.append` | primary, fork, sideband | Input audio received from the primary transport and reflected to a Live sideband connection before model-input muting. | `audio`, `type` | | `transport.dtmf.received` | primary, fork, sideband | A SIP DTMF keypress received from the caller. Delivered only to sideband observers. | `event`, `event_id`, `type` | | `transport.dtmf.send` | primary, fork, sideband | A SIP DTMF keypress successfully sent by the hosted tool. Delivered only to sideband observers; this is not a client command. | `event`, `event_id`, `type` | | `transport.ringing` | primary, fork, sideband | The outbound SIP provider leg is ringing or providing early media. Delivered only to sideband observers. | `event_id`, `session_id`, `type` | | `transport.answered` | primary, fork, sideband | The outbound SIP provider leg answered and media is established. Delivered only to sideband observers. | `event_id`, `session_id`, `type` | | `transport.failed` | primary, fork, sideband | An asynchronous outbound SIP setup failure. Delivered only to sideband observers. | `error{code, message, type, param}`, `event_id`, `session_id`, `type` | ## 3. Audio REST streaming events (SSE) | Endpoint | Event | Live | Description | Payload fields | |---|---|---|---|---| | `POST /v1/audio/transcriptions (stream=true, SSE)` | `transcript.text.segment` | ✅ | Emitted when a diarized transcription returns a completed segment with speaker information. Only emitted when you [create a transcription](https://developers.openai.com/api/reference/resources/audio/subresources/transcriptions/methods/create) with `stream` set to `true` and `response_format` set to | `id`, `end`, `speaker`, `start`, `text`, `type` | | `POST /v1/audio/transcriptions (stream=true, SSE)` | `transcript.text.delta` | ✅ | Emitted when there is an additional text delta. This is also the first event emitted when the transcription starts. Only emitted when you [create a transcription](https://developers.openai.com/api/reference/resources/audio/subresources/transcriptions/methods/create) with the `Stream` parameter set t | `delta`, `type`, `logprobs{token, bytes, logprob}`?, `segment_id`? | | `POST /v1/audio/transcriptions (stream=true, SSE)` | `transcript.text.done` | ✅ | Emitted when the transcription is complete. Contains the complete transcription text. Only emitted when you [create a transcription](https://developers.openai.com/api/reference/resources/audio/subresources/transcriptions/methods/create) with the `Stream` parameter set to `true`. | `text`, `type`, `languages{code}`?, `logprobs{token, bytes, logprob}`?, `usage{input_tokens, output_tokens, total_tokens, type, input_token_details}`? | | `POST /v1/audio/speech (stream_format=sse, SSE)` | `speech.audio.delta` | ✅ | Emitted for each chunk of audio data generated during speech synthesis (stream_format=sse). | `type`, `audio` | | `POST /v1/audio/speech (stream_format=sse, SSE)` | `speech.audio.done` | ✅ | Emitted when the speech synthesis is complete and all audio has been streamed. | `type`, `usage` | Observed 2026-09-18: `POST /v1/audio/transcriptions` `stream=true` (gpt-4o-mini-transcribe, 1.1 s clip) → `transcript.text.delta` ×2 → `transcript.text.done` (with `usage.type: tokens`). `POST /v1/audio/speech` `stream_format: sse` (gpt-4o-mini-tts, pcm) → `speech.audio.delta` ×4 → `speech.audio.done` → `data: [DONE]`.