SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
28.7 KB

# xAI Voice APIs — Speech to Speech (realtime), Text to Speech, Speech to Text, Custom Voices, SIP

Status: WS wss://api.x.ai/v1/realtime DOCUMENTED + LIVE_VERIFIED (text-only turn) · POST /v1/realtime/client_secrets LIVE_VERIFIED · POST /v1/tts LIVE_VERIFIED · GET /v1/tts/voices[/{id}] LIVE_VERIFIED · POST /v1/stt LIVE_VERIFIED · GET /v1/custom-voices LIVE_VERIFIED · WS /v1/tts, WS /v1/stt, POST /v2/phone-numbers, /v1/realtime/calls/{call_id}/refer|hangup, custom-voice create/get/patch/delete/audio DOCUMENTED (create = ACCOUNT_RESTRICTED, Enterprise) · models: grok-voice-latest → grok-voice-think-fast-2.0 (GA 2026-07), grok-voice-think-fast-1.0 LEGACY (2026-04), grok-voice-transcribe-2.0 (2026-09, default) / -1.0. Sources: Voice reference · Voice overview · Speech to Speech · SIP · Ephemeral tokens · Text to Speech · Speech to Text · Custom Voices · model cards developers/models/{speech-to-speech,text-to-speech,speech-to-text} · WebSocket schemas voice-realtime.ws.json, tts-streaming.ws.json, stt-streaming.ws.json (local copies sources/xai/openapi/*.ws.json) · Pricing · Rate limits. Last verified: 2026-09-18 (raw tmp-live/xai-media/voice-ws-session.json, tts-*.json, stt-ok.json, realtime-client-secrets*.json). Machine-readable: generated/fragments/endpoints/xai-media-voice-skills.json (api_family voice), parameters/xai-voice.json (116 params), streaming-events/xai-voice.json (61 events, api = realtime | tts | stt), objects/xai-media-objects.json, status-lifecycles/xai-lifecycles.json.

# Surface

Endpoint Mode Model(s) Price Status
WS wss://api.x.ai/v1/realtime Speech to Speech, bidirectional JSON (+ optional binary audio frames) grok-voice-latest = grok-voice-think-fast-2.0 $0.08 / min audio ($4.80 / h, billable_audio_seconds rounded up) + $0.004 per text conversation.item.create (function_call_output free) LIVE_VERIFIED
POST /v1/realtime/client_secrets ephemeral token for browsers/mobile — free LIVE_VERIFIED
POST /v2/phone-numbers, POST /v1/realtime/calls/{call_id}/refer, …/hangup SIP telephony (BYO trunk → sip.voice.x.ai), webhook realtime.call.incoming, transfer, hang-up S2S session bound by ?call_id= audio minutes DOCUMENTED
POST /v1/tts unary TTS → audio bytes or JSON envelope roster of 28 voices $15 / 1M chars LIVE_VERIFIED
WS wss://api.x.ai/v1/tts streaming TTS (text.delta → audio.delta) same same DOCUMENTED
GET /v1/tts/voices, /v1/tts/voices/{voice_id} built-in voice roster — free LIVE_VERIFIED
POST /v1/stt batch transcription (multipart) grok-voice-transcribe-2.0 (default) / -1.0 $0.10 / h LIVE_VERIFIED
WS wss://api.x.ai/v1/stt streaming transcription (binary audio in) same $0.20 / h DOCUMENTED
POST/GET/PATCH/DELETE /v1/custom-voices[/{voice_id}[/audio]] voice cloning from ≤ 120 s clip custom voice_id (8 chars) usable everywhere free up to 30 voices/team (console); API create Enterprise-only; US except Illinois GET list LIVE_VERIFIED

Rate limits (documented per credit tier T0–T4): S2S concurrent sessions 10/20/50/100/200 (max session 120 min); TTS 50–500 RPS and 100–500 concurrent streams; STT 10–40 RPS, 100–500 concurrent. Region us-east-1. Audio "processed in real time and never stored or used for training".

# Speech to Speech (wss://api.x.ai/v1/realtime)

Connect: wss://api.x.ai/v1/realtime?model=grok-voice-latest[&reasoning.effort=high|none][&conversation_id=<id>] or ?call_id=<sip call>. Auth: Authorization: Bearer <API key> (server) or ephemeral secret either as Bearer or as WebSocket subprotocol xai-client-secret.<token> (browser; verified live — the server negotiates that subprotocol back). Client secret: POST /v1/realtime/client_secrets {expires_after:{seconds ≤ 3600, default 600}, session?:{model, reasoning:{effort}}} → {value, expires_at} (value 111 chars, prefix xai-realtime).

Session (session.update) — key fields (full list in parameters/xai-voice.json): instructions, voice (built-in id or custom id; server default reported as xai_ara, docs say eve), reasoning.effort (high default | none), turn_detection ({type:"server_vad", threshold 0.1–0.9 (0.85), silence_duration_ms, prefix_padding_ms (333), idle_timeout_ms} or null for manual turns), audio.input|output.format {type: audio/pcm|audio/pcmu|audio/pcma|audio/opus, rate 8000…48000 (24000)}, audio.input|output.transport json|binary, audio.input.transcription {language_hint, keyterms[], model:"grok-transcribe"}, audio.output.speed 0.7–1.5, resumption.enabled, replace {phrase: spoken}, tools[] (file_search, web_search, x_search, mcp, function). session.updated echoes an effective config that also contains undocumented keys (tool_choice:"auto", enable_noise_suppression, enable_phonetic_spelling, keep_context, temperature:-1, max_response_output_tokens:"inf", input_audio_format:"not specified").

Turn flow (manual, text): conversation.item.create {item:{type:"message", role:"user", content:[{type:"input_text", text}]}} → response.create [{response:{instructions, metadata}}]. Audio: input_audio_buffer.append {audio: base64} (or binary frames) → with server_vad the server commits and responds automatically (speech_started/speech_stopped/committed); without VAD send input_audio_buffer.commit then response.create; input_audio_buffer.clear discards. Tools: response.function_call_arguments.done {call_id, name, arguments} → conversation.item.create {item:{type:"function_call_output", call_id, output}} (one per call, all before) → response.create. Server-side tools (web_search, x_search, file_search, mcp) run without client action (mcp_list_tools.*, response.mcp_call* events). xAI extensions: force_message item (verbatim TTS line, interruptible), resumption (replay by conversation_id, 30 min expiry, replayed:true items), replace, binary transport, idle_timeout_ms → input_audio_buffer.timeout_triggered, SIP input_audio_buffer.dtmf_event_received.

OpenAI Realtime compatibility: event names match GA OpenAI names (response.output_audio.delta, response.output_text.delta alias of response.text.delta); renamed conversation.item.input_audio_transcription.updated (cumulative, needs grok-transcribe); not supported: conversation.item.retrieve, output_audio_buffer.*, conversation.item.done, rate_limits.updated, transcription .failed/.segment.

# Live probe (2026-09-18) — text in, audio out

session.update (turn_detection null, effort none, voice eve, pcm 16 kHz in/out) → conversation.item.create "Reply with OK." → response.create. 18 server events in 1.23 s:

text
session.created → conversation.created → ping (undocumented) → session.updated → conversation.item.added (user)
→ response.created → response.output_item.added → conversation.item.added (assistant) → response.content_part.added {part:{type:"audio",transcript:""}}
→ response.output_audio.delta ×3 (audio_duration_ms 550/92/68, first carries latency "0.49") → response.output_audio_transcript.delta "OK"
→ response.output_audio_transcript.done → response.content_part.done → response.output_audio.done → response.output_item.done → response.done

response.done carries a top-level usage (not in the schema, which only shows response.usage = {}): {input_tokens:4, input_token_details:{text_tokens:4, audio_tokens:0, grok_tokens:0}, output_tokens:37, output_token_details:{text_tokens:1, audio_tokens:36, grok_tokens:0}, total_tokens:41, output_audio_seconds:0.71, billable_audio_seconds:1} → cost ≈ 1 s × $0.08/60 + $0.004 ≈ $0.005. Every event carries previous_item_id; audio deltas add rid, ts, audio_duration_ms, latency; response.status_details is the string "unimplemented". 22.7 KB of PCM16 (0.71 s @ 16 kHz) decoded. A second connection authenticated only with the ephemeral secret subprotocol received session.created + conversation.created.

# Text to Speech

POST /v1/tts body: text (≤ 60,000 chars; tags [pause] [long-pause] [laugh] [chuckle] [giggle] [cry] [sigh] [breath] …, wrappers <whisper> <soft> <loud> <slow> <fast> <emphasis> <singing> …), voice_id (default eve), language required (en, zh, pt-BR, … or auto; missing → 422 text/plain missing field \language`), output_format {codec mp3|wav|pcm|mulaw|alaw, sample_rate 8000–48000 (24000), bit_rate 32k–192k (128k, mp3)}, speed 0.7–1.5, optimize_streaming_latency 0|1(|2), text_normalization, with_timestamps, replace(≤ 200 entries; respelling or/IPA/). Response: raw audio (Content-Type: audio/wavobserved, 25 KB for "OK." at 16 kHz) or, withwith_timestamps, JSON {audio (base64), content_type, duration, audio_timestamps:{graph_chars[], graph_times:[[start,end],…]}}— livegraph_timesis an array of pairs (guide), not{start,end}objects (reference page). Streaming:wss://api.x.ai/v1/tts?language=en&voice=eve&codec=mp3…withtext.delta* → text.done→audio.delta* → audio.done` (multi-utterance, no length limit).

Voices (GET /v1/tts/voices, live 28): altair, ara, atlas, aurora, carina, castor, celeste, cosmo, eve, helios, helix, iris, kepler, leo, liora, lumen, luna, lux, naksh, orion, perseus, rex, rigel, sal, sirius, ursa, zagan, zenith — each {voice_id, name, language:"multilingual", gender} (gender and multilingual undocumented; docs example lists 26 with language:"en"). Custom voices are listed only by GET /v1/custom-voices (live {voices:[], total_count:0, cap:30} — total_count/cap undocumented).

# Speech to Text

POST /v1/stt multipart: file (≤ 500 MB, last field; wav/mp3/ogg/opus/flac/aac/mp4/m4a/mkv auto-detected; raw pcm/mulaw/alaw need audio_format + sample_rate) or url; language (+ format=true for inverse text normalization), multichannel/channels (2–8), diarize, keyterm (repeatable ≤ 100 × 50 chars), filler_words, vad_threshold (0.5), model. Response {text, language, duration, words:[{text,start,end,confidence?,speaker?}], channels?}; live: the 0.79 s TTS clip → {"text":"Okay.","language":"en","duration":0.79,"words":[{"text":"Okay.","start":0.122,"end":0.466}]} (no confidence). Streaming wss://api.x.ai/v1/stt?sample_rate=16000&encoding=pcm&interim_results=true&endpointing=400&smart_turn=0.7…: wait for transcript.created, send binary frames (100 ms chunks; opus = one packet per frame), optional {"type":"finalize"}, then {"type":"audio.done"} → transcript.done and close. transcript.partial flags: is_final=false interim → is_final=true,speech_final=false chunk final → speech_final=true utterance final (end_of_turn_confidence with Smart Turn).

# SIP telephony (documented)

  1. POST /v2/phone-numbers {origin:"byo_trunk", name, phone_number:"+1…", webhook:{url, auth_url?, auth_token?} | agent_id, sip_auth:{auth_username, auth_password} | {allowed_addresses:[CIDR]}} → number record with sip_host: sip.voice.x.ai and a one-time webhook.dispatch_signing_secret. Point the carrier (Twilio Elastic SIP trunk, etc.) at sip:{number}@sip.voice.x.ai;transport=tls.
  2. Incoming call → signed realtime.call.incoming webhook (Standard Webhooks headers webhook-id, webhook-timestamp, webhook-signature) with data.call_id.
  3. Open wss://api.x.ai/v1/realtime?call_id=<id> with the API key (ephemeral secrets not allowed), session.update, response.create to greet. Audio is bridged (G.711 μ-law/A-law native). DTMF digits are buffered and flushed to the model on #, 2.5 s idle or speech; input_audio_buffer.dtmf_event_received is the audit trail.
  4. POST /v1/realtime/calls/{call_id}/refer {target_uri:"tel:+…"|"sip:…"} (blocking) and POST …/hangup.

# Errors observed

case HTTP body
TTS without language 422 text/plain Failed to deserialize the JSON body into the target type: missing field \language` at line 1 column 34`
realtime invalid tool config / bad event WS error event {type:"error", error:{type: invalid_request_error|invalid_event|internal_error|timeout|max_duration, code, message, param?, event_id?}} — session stays open for most errors

# Examples / tests

examples/xai/voice/: tts.sh (LIVE_VERIFIED, ~$0.00005), voices.sh (free), stt.py (LIVE_VERIFIED on the TTS clip), realtime_text_turn.py (websockets; LIVE_VERIFIED, ~$0.005), realtime_text_turn.ts (ws; UNVERIFIED), client_secret.sh (free). Tests: tests/xai/test_voice.py — voice roster, client secret and the TTS 422 are free and always run; the S2S text turn is gated by RUN_REALTIME_TESTS, TTS/STT by RUN_AUDIO_TESTS.

# Event reference

Generated from generated/fragments/streaming-events/xai-voice.json (official *.ws.json schemas + live observations). Fields column = top-level schema properties.

# Speech to Speech — wss://api.x.ai/v1/realtime — client→server (9)

event status description fields
session.update DOCUMENTED + LIVE_VERIFIED Update session configuration such as system prompt, voice, audio format, turn detection, and tools. type, session {model, instructions, reasoning, voice, turn_detection, resumption, audio, tools, replace}
input_audio_buffer.append DOCUMENTED Append chunks of base64-encoded audio data to the input buffer. The server does not send back a corresponding message. type, audio
input_audio_buffer.commit DOCUMENTED Commit the audio buffer as a user message. Only available when turn_detection type is null. Confirmed by input_audio_buffer.committed from the server. type
conversation.item.create DOCUMENTED + LIVE_VERIFIED Create a new conversation item. Can be a user text message, an assistant text message for history seeding, a function call for seeding tool-use history, or a function call output. type, previous_item_id, item
input_audio_buffer.clear DOCUMENTED Clear the input audio buffer. Use this to discard any pending audio data without committing it. type
conversation.item.delete DOCUMENTED Delete a conversation item by ID. The server confirms deletion with a conversation.item.deleted event. type, item_id
conversation.item.truncate DOCUMENTED Truncate a previous assistant audio message item. Removes audio and transcript content after the specified duration, keeping only the content up to that point. The server confirms with a conversation.item.truncated event. type, item_id, content_index, audio_end_ms
response.create DOCUMENTED + LIVE_VERIFIED Request the server to create a new assistant response. This is handled automatically when using server-side VAD. type, response {modalities, instructions, metadata}
response.cancel DOCUMENTED Cancel an in-progress response. In VAD mode, interruptions are automatic — use this for manual cancel in non-VAD mode. type, response_id

# Speech to Speech — wss://api.x.ai/v1/realtime — server→client (39)

event status description fields
session.created DOCUMENTED + LIVE_VERIFIED Sent automatically on WebSocket connection. Contains the session configuration. event_id, type, session {id, object, model, instructions, reasoning, voice, modalities, turn_detection, tools, replace}
conversation.created DOCUMENTED + LIVE_VERIFIED The first message on connection. Notifies the client that a conversation session has been created. event_id, type, conversation {id, object}
session.updated DOCUMENTED + LIVE_VERIFIED Acknowledges the client's session.update message that the session has been configured. event_id, type, session {id, object, model, instructions, voice, modalities, turn_detection, tools, replace}
input_audio_buffer.speech_started DOCUMENTED Notifies that the server's VAD detected the start of speech. Only available with server_vad turn detection. event_id, type, item_id, audio_start_ms
input_audio_buffer.speech_stopped DOCUMENTED Notifies that the server's VAD detected the end of speech. Only available with server_vad turn detection. event_id, type, item_id, audio_end_ms
input_audio_buffer.committed DOCUMENTED Input audio buffer has been committed as a user message. event_id, type, previous_item_id, item_id
input_audio_buffer.timeout_triggered DOCUMENTED The turn_detection.idle_timeout_ms idle timer fired: no user speech was detected for the configured duration after the assistant finished responding. The server commits a silent user turn and generates a proactive check-in. event_id, type, item_id, previous_item_id, audio_start_ms, audio_end_ms
input_audio_buffer.cleared DOCUMENTED Confirms the input audio buffer has been cleared. event_id, type
conversation.item.deleted DOCUMENTED Confirms a conversation item has been deleted. event_id, type, item_id
conversation.item.added DOCUMENTED + LIVE_VERIFIED A new user or assistant message has been added to the conversation history. event_id, type, previous_item_id, item
conversation.item.truncated DOCUMENTED Confirms that a conversation item has been truncated. Sent in response to a conversation.item.truncate client event. event_id, type, item_id, content_index, audio_end_ms, transcript
conversation.item.input_audio_transcription.completed DOCUMENTED Audio transcription for the user's input has been completed. event_id, type, item_id, transcript
conversation.item.input_audio_transcription.updated DOCUMENTED Streaming transcription update for the user's audio input. Emitted as the user speaks, providing the cumulative transcript so far before the final completed event. Note that this is the cumulative transcript which may have corrections to previous updated transcripts — this is different from a transcript delta. Only emitted when audio.input.transcription.model is set to grok-transcribe in the session configuration. Useful for displaying live captions. event_id, type, item_id, transcript
input_audio_buffer.dtmf_event_received DOCUMENTED A DTMF tone (phone keypress) was detected on a SIP session. SIP only — not emitted on direct WebSocket connections. Digits are buffered server-side and flushed as a text message to the model on # key, 2.5s idle, or when the user begins speaking. event_id, type, event, received_at
response.created DOCUMENTED + LIVE_VERIFIED A new assistant response turn is in progress. Audio deltas from this turn share the same response_id. event_id, type, response {id, object, status, output, metadata}
response.output_item.added DOCUMENTED + LIVE_VERIFIED A new assistant response item is added to the message history. event_id, type, response_id, output_index, item
response.output_item.done DOCUMENTED + LIVE_VERIFIED An output item is complete. event_id, type, response_id, output_index, item
response.content_part.added DOCUMENTED + LIVE_VERIFIED A content part starts within an output item. event_id, type, response_id, item_id, output_index, content_index, part {type, transcript}
response.content_part.done DOCUMENTED + LIVE_VERIFIED A content part finishes. event_id, type, response_id, item_id, output_index, content_index, part {type, transcript}
response.output_audio_transcript.delta DOCUMENTED + LIVE_VERIFIED Streaming text transcript delta of the assistant's audio response. event_id, type, response_id, item_id, output_index, content_index, delta
response.output_audio_transcript.done DOCUMENTED + LIVE_VERIFIED The audio transcript for this assistant turn has finished generating. event_id, type, response_id, item_id, output_index, content_index, transcript
response.output_audio.delta DOCUMENTED + LIVE_VERIFIED Streaming base64-encoded audio delta of the assistant's response. event_id, type, response_id, item_id, output_index, content_index, delta
response.output_audio.done DOCUMENTED + LIVE_VERIFIED Audio generation for this assistant turn has finished. event_id, type, response_id, item_id, output_index, content_index
response.text.delta DOCUMENTED Text-mode output delta (when using text modality). type ∈ response.text.delta, event_id, response_id, item_id, output_index, content_index, delta
response.output_text.delta DOCUMENTED Text-mode output delta using the OpenAI GA event name. Functionally identical to response.text.delta. Clients should handle both event names for maximum compatibility. type ∈ response.output_text.delta, event_id, response_id, item_id, output_index, content_index, delta
response.function_call_arguments.delta DOCUMENTED Streaming function call arguments. event_id, type, response_id, item_id, output_index, call_id, delta
response.function_call_arguments.done DOCUMENTED A function call has been triggered with complete arguments. Your code should execute the function and return results via conversation.item.create with type function_call_output. event_id, type, response_id, item_id, output_index, call_id, name, arguments
mcp_list_tools.in_progress DOCUMENTED MCP tool discovery has started. event_id, type, item_id
mcp_list_tools.completed DOCUMENTED MCP tool discovery succeeded. event_id, type, item_id
mcp_list_tools.failed DOCUMENTED MCP tool discovery failed. event_id, type, item_id, error {type, message}
response.mcp_call_arguments.delta DOCUMENTED MCP call arguments streaming. event_id, type, response_id, item_id, call_id, delta
response.mcp_call_arguments.done DOCUMENTED MCP call arguments finalized. event_id, type, response_id, item_id, call_id, name, arguments
response.mcp_call.in_progress DOCUMENTED MCP server HTTP call starting. event_id, type, item_id, output_index
response.mcp_call.completed DOCUMENTED MCP tool execution succeeded. event_id, type, item_id, output_index
response.mcp_call.failed DOCUMENTED MCP tool execution failed. event_id, type, item_id, output_index, error {type, message}
response.done DOCUMENTED + LIVE_VERIFIED The assistant's response is completed. Sent after all audio and transcript deltas. Ready for the client to add a new conversation item. event_id, type, response {id, object, status, usage, metadata}
error DOCUMENTED Sent when an error occurs. Contains error code and message. Most errors are recoverable and the session stays open. event_id, type, error {type, code, message, param, event_id}
ping LIVE_DISCOVERED Keep-alive / clock event sent right after conversation.created (and presumably periodically). Not in the official ws.json schema. type, event_id, timestamp, previous_item_id
response.audio.delta DOCUMENTED + UNVERIFIED Legacy alias of response.output_audio.delta mentioned in the Speech-to-Speech guide (audio transport table); not emitted in the live probe (only response.output_audio.delta was).

# Speech to Speech — wss://api.x.ai/v1/realtime — server→client (webhook) (1)

event status description fields
realtime.call.incoming DOCUMENTED Signed webhook (Standard Webhooks v1: webhook-id, webhook-timestamp, webhook-signature headers, HMAC-SHA256 with dispatch_signing_secret) POSTed to the phone number's webhook URL when a SIP call arrives; data.call_id is then used as ?call_id= on wss://api.x.ai/v1/realtime. type, data

# Streaming TTS — wss://api.x.ai/v1/tts — client→server (2)

event status description fields
text.delta DOCUMENTED Send a chunk of text to be synthesized. Text is processed incrementally — audio generation begins as soon as enough text is buffered. Individual deltas are capped at 60,000 characters. type, delta
text.done DOCUMENTED Signal that all text for this utterance has been sent. The server will finish generating audio and send audio.done. After receiving audio.done, you can start a new utterance with another text.delta. type

# Streaming TTS — wss://api.x.ai/v1/tts — server→client (3)

event status description fields
audio.delta DOCUMENTED A chunk of base64-encoded audio data. Decode and append to your audio buffer or pipe directly to playback. The format matches the codec and sample_rate specified in the query parameters. When the connection was opened with with_timestamps=true, the event also carries audio_timestamps and audio_duration for the characters that fall inside this chunk. type, delta, audio_timestamps {graph_chars, graph_times}, audio_duration
audio.done DOCUMENTED Audio generation for this utterance is complete. The connection remains open for multi-utterance — send another text.delta to start a new synthesis, or close the connection. type, trace_id
error DOCUMENTED An error occurred during synthesis. The connection may be closed after this message. type, message

# Streaming STT — wss://api.x.ai/v1/stt — client→server (3)

event status description fields
Binary frame (audio) DOCUMENTED Send raw audio as binary WebSocket frames in the encoding specified by the encoding query parameter. Audio should be streamed in real-time-paced chunks (e.g. 100 ms at a time). No base64 encoding — send raw bytes directly. With encoding=opus, each binary frame must contain exactly one raw Opus packet — never concatenate packets or split one across frames. An undecodable frame sends an error event and closes the session.
finalize DOCUMENTED Force the current utterance to finalize as speech_final immediately, without waiting for VAD endpointing or Smart Turn. The session stays open so you can continue streaming audio. Accepts finalize or Finalize as the type value. When multichannel=true, optional channel (0-based) limits the finalize to that channel; omit channel to finalize every channel. type ∈ finalize|Finalize, channel
audio.done DOCUMENTED Signal that all audio has been sent. The server flushes any remaining buffered audio, emits final transcript events, and sends a transcript.done event. The connection closes after transcript.done. type

# Streaming STT — wss://api.x.ai/v1/stt — server→client (4)

event status description fields
transcript.created DOCUMENTED Sent immediately after the WebSocket connection is established and the server is ready to receive audio. Wait for this event before sending audio — the server needs to initialize its ASR backend. type, id
transcript.partial DOCUMENTED A transcript result for a portion of the audio stream. Two boolean fields convey state: interim (is_final=false) means text may still change, chunk final (is_final=true, speech_final=false) means the chunk is locked, and utterance final (is_final=true, speech_final=true) means the speaker stopped talking. type, text, words, is_final, speech_final, start, duration, channel_index, end_of_turn_confidence
transcript.done DOCUMENTED Final transcript after audio.done. duration always present. One per channel when multichannel=true. Connection closes after this event. type, text, words, duration, channel_index
error DOCUMENTED An error occurred during the session. Most errors (pipeline failures, stream timeouts, undecodable audio frames) close the connection. Only client message parse errors keep the connection open. type, message