xAI Voice APIs — Speech to Speech (realtime), Text to Speech, Speech to Text, Custom Voices, SIP
Status: WS wss://api.x.ai/v1/realtime DOCUMENTED + LIVE_VERIFIED (text-only turn) · POST /v1/realtime/client_secrets LIVE_VERIFIED · POST /v1/tts LIVE_VERIFIED · GET /v1/tts/voices[/{id}] LIVE_VERIFIED · POST /v1/stt LIVE_VERIFIED · GET /v1/custom-voices LIVE_VERIFIED · WS /v1/tts, WS /v1/stt, POST /v2/phone-numbers, /v1/realtime/calls/{call_id}/refer|hangup, custom-voice create/get/patch/delete/audio DOCUMENTED (create = ACCOUNT_RESTRICTED, Enterprise) · models: grok-voice-latest → grok-voice-think-fast-2.0 (GA 2026-07), grok-voice-think-fast-1.0 LEGACY (2026-04), grok-voice-transcribe-2.0 (2026-09, default) / -1.0.
Sources: Voice reference · Voice overview · Speech to Speech · SIP · Ephemeral tokens · Text to Speech · Speech to Text · Custom Voices · model cards developers/models/{speech-to-speech,text-to-speech,speech-to-text} · WebSocket schemas voice-realtime.ws.json, tts-streaming.ws.json, stt-streaming.ws.json (local copies sources/xai/openapi/*.ws.json) · Pricing · Rate limits.
Last verified: 2026-09-18 (raw tmp-live/xai-media/voice-ws-session.json, tts-*.json, stt-ok.json, realtime-client-secrets*.json).
Machine-readable: generated/fragments/endpoints/xai-media-voice-skills.json (api_family voice), parameters/xai-voice.json (116 params), streaming-events/xai-voice.json (61 events, api = realtime | tts | stt), objects/xai-media-objects.json, status-lifecycles/xai-lifecycles.json.
Surface
| Endpoint | Mode | Model(s) | Price | Status |
|---|---|---|---|---|
WS wss://api.x.ai/v1/realtime |
Speech to Speech, bidirectional JSON (+ optional binary audio frames) | grok-voice-latest = grok-voice-think-fast-2.0 |
$0.08 / min audio ($4.80 / h, billable_audio_seconds rounded up) + $0.004 per text conversation.item.create (function_call_output free) |
LIVE_VERIFIED |
POST /v1/realtime/client_secrets |
ephemeral token for browsers/mobile | — | free | LIVE_VERIFIED |
POST /v2/phone-numbers, POST /v1/realtime/calls/{call_id}/refer, …/hangup |
SIP telephony (BYO trunk → sip.voice.x.ai), webhook realtime.call.incoming, transfer, hang-up |
S2S session bound by ?call_id= |
audio minutes | DOCUMENTED |
POST /v1/tts |
unary TTS → audio bytes or JSON envelope | roster of 28 voices | $15 / 1M chars | LIVE_VERIFIED |
WS wss://api.x.ai/v1/tts |
streaming TTS (text.delta → audio.delta) | same | same | DOCUMENTED |
GET /v1/tts/voices, /v1/tts/voices/{voice_id} |
built-in voice roster | — | free | LIVE_VERIFIED |
POST /v1/stt |
batch transcription (multipart) | grok-voice-transcribe-2.0 (default) / -1.0 |
$0.10 / h | LIVE_VERIFIED |
WS wss://api.x.ai/v1/stt |
streaming transcription (binary audio in) | same | $0.20 / h | DOCUMENTED |
POST/GET/PATCH/DELETE /v1/custom-voices[/{voice_id}[/audio]] |
voice cloning from ≤ 120 s clip | custom voice_id (8 chars) usable everywhere |
free up to 30 voices/team (console); API create Enterprise-only; US except Illinois | GET list LIVE_VERIFIED |
Rate limits (documented per credit tier T0–T4): S2S concurrent sessions 10/20/50/100/200 (max session 120 min); TTS 50–500 RPS and 100–500 concurrent streams; STT 10–40 RPS, 100–500 concurrent. Region us-east-1. Audio "processed in real time and never stored or used for training".
Speech to Speech (wss://api.x.ai/v1/realtime)
Connect: wss://api.x.ai/v1/realtime?model=grok-voice-latest[&reasoning.effort=high|none][&conversation_id=<id>] or ?call_id=<sip call>. Auth: Authorization: Bearer <API key> (server) or ephemeral secret either as Bearer or as WebSocket subprotocol xai-client-secret.<token> (browser; verified live — the server negotiates that subprotocol back). Client secret: POST /v1/realtime/client_secrets {expires_after:{seconds ≤ 3600, default 600}, session?:{model, reasoning:{effort}}} → {value, expires_at} (value 111 chars, prefix xai-realtime).
Session (session.update) — key fields (full list in parameters/xai-voice.json): instructions, voice (built-in id or custom id; server default reported as xai_ara, docs say eve), reasoning.effort (high default | none), turn_detection ({type:"server_vad", threshold 0.1–0.9 (0.85), silence_duration_ms, prefix_padding_ms (333), idle_timeout_ms} or null for manual turns), audio.input|output.format {type: audio/pcm|audio/pcmu|audio/pcma|audio/opus, rate 8000…48000 (24000)}, audio.input|output.transport json|binary, audio.input.transcription {language_hint, keyterms[], model:"grok-transcribe"}, audio.output.speed 0.7–1.5, resumption.enabled, replace {phrase: spoken}, tools[] (file_search, web_search, x_search, mcp, function). session.updated echoes an effective config that also contains undocumented keys (tool_choice:"auto", enable_noise_suppression, enable_phonetic_spelling, keep_context, temperature:-1, max_response_output_tokens:"inf", input_audio_format:"not specified").
Turn flow (manual, text): conversation.item.create {item:{type:"message", role:"user", content:[{type:"input_text", text}]}} → response.create [{response:{instructions, metadata}}]. Audio: input_audio_buffer.append {audio: base64} (or binary frames) → with server_vad the server commits and responds automatically (speech_started/speech_stopped/committed); without VAD send input_audio_buffer.commit then response.create; input_audio_buffer.clear discards. Tools: response.function_call_arguments.done {call_id, name, arguments} → conversation.item.create {item:{type:"function_call_output", call_id, output}} (one per call, all before) → response.create. Server-side tools (web_search, x_search, file_search, mcp) run without client action (mcp_list_tools.*, response.mcp_call* events). xAI extensions: force_message item (verbatim TTS line, interruptible), resumption (replay by conversation_id, 30 min expiry, replayed:true items), replace, binary transport, idle_timeout_ms → input_audio_buffer.timeout_triggered, SIP input_audio_buffer.dtmf_event_received.
OpenAI Realtime compatibility: event names match GA OpenAI names (response.output_audio.delta, response.output_text.delta alias of response.text.delta); renamed conversation.item.input_audio_transcription.updated (cumulative, needs grok-transcribe); not supported: conversation.item.retrieve, output_audio_buffer.*, conversation.item.done, rate_limits.updated, transcription .failed/.segment.
Live probe (2026-09-18) — text in, audio out
session.update (turn_detection null, effort none, voice eve, pcm 16 kHz in/out) → conversation.item.create "Reply with OK." → response.create. 18 server events in 1.23 s:
session.created → conversation.created → ping (undocumented) → session.updated → conversation.item.added (user)
→ response.created → response.output_item.added → conversation.item.added (assistant) → response.content_part.added {part:{type:"audio",transcript:""}}
→ response.output_audio.delta ×3 (audio_duration_ms 550/92/68, first carries latency "0.49") → response.output_audio_transcript.delta "OK"
→ response.output_audio_transcript.done → response.content_part.done → response.output_audio.done → response.output_item.done → response.doneresponse.done carries a top-level usage (not in the schema, which only shows response.usage = {}): {input_tokens:4, input_token_details:{text_tokens:4, audio_tokens:0, grok_tokens:0}, output_tokens:37, output_token_details:{text_tokens:1, audio_tokens:36, grok_tokens:0}, total_tokens:41, output_audio_seconds:0.71, billable_audio_seconds:1} → cost ≈ 1 s × $0.08/60 + $0.004 ≈ $0.005. Every event carries previous_item_id; audio deltas add rid, ts, audio_duration_ms, latency; response.status_details is the string "unimplemented". 22.7 KB of PCM16 (0.71 s @ 16 kHz) decoded. A second connection authenticated only with the ephemeral secret subprotocol received session.created + conversation.created.
Text to Speech
POST /v1/tts body: text (≤ 60,000 chars; tags [pause] [long-pause] [laugh] [chuckle] [giggle] [cry] [sigh] [breath] …, wrappers <whisper> <soft> <loud> <slow> <fast> <emphasis> <singing> …), voice_id (default eve), language required (en, zh, pt-BR, … or auto; missing → 422 text/plain missing field \language`), output_format {codec mp3|wav|pcm|mulaw|alaw, sample_rate 8000–48000 (24000), bit_rate 32k–192k (128k, mp3)}, speed 0.7–1.5, optimize_streaming_latency 0|1(|2), text_normalization, with_timestamps, replace(≤ 200 entries; respelling or/IPA/). Response: raw audio (Content-Type: audio/wavobserved, 25 KB for "OK." at 16 kHz) or, withwith_timestamps, JSON {audio (base64), content_type, duration, audio_timestamps:{graph_chars[], graph_times:[[start,end],…]}}— livegraph_timesis an array of pairs (guide), not{start,end}objects (reference page). Streaming:wss://api.x.ai/v1/tts?language=en&voice=eve&codec=mp3…withtext.delta* → text.done→audio.delta* → audio.done` (multi-utterance, no length limit).
Voices (GET /v1/tts/voices, live 28): altair, ara, atlas, aurora, carina, castor, celeste, cosmo, eve, helios, helix, iris, kepler, leo, liora, lumen, luna, lux, naksh, orion, perseus, rex, rigel, sal, sirius, ursa, zagan, zenith — each {voice_id, name, language:"multilingual", gender} (gender and multilingual undocumented; docs example lists 26 with language:"en"). Custom voices are listed only by GET /v1/custom-voices (live {voices:[], total_count:0, cap:30} — total_count/cap undocumented).
Speech to Text
POST /v1/stt multipart: file (≤ 500 MB, last field; wav/mp3/ogg/opus/flac/aac/mp4/m4a/mkv auto-detected; raw pcm/mulaw/alaw need audio_format + sample_rate) or url; language (+ format=true for inverse text normalization), multichannel/channels (2–8), diarize, keyterm (repeatable ≤ 100 × 50 chars), filler_words, vad_threshold (0.5), model. Response {text, language, duration, words:[{text,start,end,confidence?,speaker?}], channels?}; live: the 0.79 s TTS clip → {"text":"Okay.","language":"en","duration":0.79,"words":[{"text":"Okay.","start":0.122,"end":0.466}]} (no confidence). Streaming wss://api.x.ai/v1/stt?sample_rate=16000&encoding=pcm&interim_results=true&endpointing=400&smart_turn=0.7…: wait for transcript.created, send binary frames (100 ms chunks; opus = one packet per frame), optional {"type":"finalize"}, then {"type":"audio.done"} → transcript.done and close. transcript.partial flags: is_final=false interim → is_final=true,speech_final=false chunk final → speech_final=true utterance final (end_of_turn_confidence with Smart Turn).
SIP telephony (documented)
POST /v2/phone-numbers {origin:"byo_trunk", name, phone_number:"+1…", webhook:{url, auth_url?, auth_token?} | agent_id, sip_auth:{auth_username, auth_password} | {allowed_addresses:[CIDR]}}→ number record withsip_host: sip.voice.x.aiand a one-timewebhook.dispatch_signing_secret. Point the carrier (Twilio Elastic SIP trunk, etc.) atsip:{number}@sip.voice.x.ai;transport=tls.- Incoming call → signed
realtime.call.incomingwebhook (Standard Webhooks headerswebhook-id,webhook-timestamp,webhook-signature) withdata.call_id. - Open
wss://api.x.ai/v1/realtime?call_id=<id>with the API key (ephemeral secrets not allowed),session.update,response.createto greet. Audio is bridged (G.711 μ-law/A-law native). DTMF digits are buffered and flushed to the model on#, 2.5 s idle or speech;input_audio_buffer.dtmf_event_receivedis the audit trail. POST /v1/realtime/calls/{call_id}/refer {target_uri:"tel:+…"|"sip:…"}(blocking) andPOST …/hangup.
Errors observed
| case | HTTP | body |
|---|---|---|
TTS without language |
422 text/plain |
Failed to deserialize the JSON body into the target type: missing field \language` at line 1 column 34` |
| realtime invalid tool config / bad event | WS error event |
{type:"error", error:{type: invalid_request_error|invalid_event|internal_error|timeout|max_duration, code, message, param?, event_id?}} — session stays open for most errors |
Examples / tests
examples/xai/voice/: tts.sh (LIVE_VERIFIED, ~$0.00005), voices.sh (free), stt.py (LIVE_VERIFIED on the TTS clip), realtime_text_turn.py (websockets; LIVE_VERIFIED, ~$0.005), realtime_text_turn.ts (ws; UNVERIFIED), client_secret.sh (free). Tests: tests/xai/test_voice.py — voice roster, client secret and the TTS 422 are free and always run; the S2S text turn is gated by RUN_REALTIME_TESTS, TTS/STT by RUN_AUDIO_TESTS.
Event reference
Generated from generated/fragments/streaming-events/xai-voice.json (official *.ws.json schemas + live observations). Fields column = top-level schema properties.
Speech to Speech — wss://api.x.ai/v1/realtime — client→server (9)
| event | status | description | fields |
|---|---|---|---|
session.update |
DOCUMENTED + LIVE_VERIFIED | Update session configuration such as system prompt, voice, audio format, turn detection, and tools. | type, session {model, instructions, reasoning, voice, turn_detection, resumption, audio, tools, replace} |
input_audio_buffer.append |
DOCUMENTED | Append chunks of base64-encoded audio data to the input buffer. The server does not send back a corresponding message. | type, audio |
input_audio_buffer.commit |
DOCUMENTED | Commit the audio buffer as a user message. Only available when turn_detection type is null. Confirmed by input_audio_buffer.committed from the server. |
type |
conversation.item.create |
DOCUMENTED + LIVE_VERIFIED | Create a new conversation item. Can be a user text message, an assistant text message for history seeding, a function call for seeding tool-use history, or a function call output. | type, previous_item_id, item |
input_audio_buffer.clear |
DOCUMENTED | Clear the input audio buffer. Use this to discard any pending audio data without committing it. | type |
conversation.item.delete |
DOCUMENTED | Delete a conversation item by ID. The server confirms deletion with a conversation.item.deleted event. |
type, item_id |
conversation.item.truncate |
DOCUMENTED | Truncate a previous assistant audio message item. Removes audio and transcript content after the specified duration, keeping only the content up to that point. The server confirms with a conversation.item.truncated event. |
type, item_id, content_index, audio_end_ms |
response.create |
DOCUMENTED + LIVE_VERIFIED | Request the server to create a new assistant response. This is handled automatically when using server-side VAD. | type, response {modalities, instructions, metadata} |
response.cancel |
DOCUMENTED | Cancel an in-progress response. In VAD mode, interruptions are automatic — use this for manual cancel in non-VAD mode. | type, response_id |
Speech to Speech — wss://api.x.ai/v1/realtime — server→client (39)
| event | status | description | fields |
|---|---|---|---|
session.created |
DOCUMENTED + LIVE_VERIFIED | Sent automatically on WebSocket connection. Contains the session configuration. | event_id, type, session {id, object, model, instructions, reasoning, voice, modalities, turn_detection, tools, replace} |
conversation.created |
DOCUMENTED + LIVE_VERIFIED | The first message on connection. Notifies the client that a conversation session has been created. | event_id, type, conversation {id, object} |
session.updated |
DOCUMENTED + LIVE_VERIFIED | Acknowledges the client's session.update message that the session has been configured. | event_id, type, session {id, object, model, instructions, voice, modalities, turn_detection, tools, replace} |
input_audio_buffer.speech_started |
DOCUMENTED | Notifies that the server's VAD detected the start of speech. Only available with server_vad turn detection. | event_id, type, item_id, audio_start_ms |
input_audio_buffer.speech_stopped |
DOCUMENTED | Notifies that the server's VAD detected the end of speech. Only available with server_vad turn detection. | event_id, type, item_id, audio_end_ms |
input_audio_buffer.committed |
DOCUMENTED | Input audio buffer has been committed as a user message. | event_id, type, previous_item_id, item_id |
input_audio_buffer.timeout_triggered |
DOCUMENTED | The turn_detection.idle_timeout_ms idle timer fired: no user speech was detected for the configured duration after the assistant finished responding. The server commits a silent user turn and generates a proactive check-in. |
event_id, type, item_id, previous_item_id, audio_start_ms, audio_end_ms |
input_audio_buffer.cleared |
DOCUMENTED | Confirms the input audio buffer has been cleared. | event_id, type |
conversation.item.deleted |
DOCUMENTED | Confirms a conversation item has been deleted. | event_id, type, item_id |
conversation.item.added |
DOCUMENTED + LIVE_VERIFIED | A new user or assistant message has been added to the conversation history. | event_id, type, previous_item_id, item |
conversation.item.truncated |
DOCUMENTED | Confirms that a conversation item has been truncated. Sent in response to a conversation.item.truncate client event. |
event_id, type, item_id, content_index, audio_end_ms, transcript |
conversation.item.input_audio_transcription.completed |
DOCUMENTED | Audio transcription for the user's input has been completed. | event_id, type, item_id, transcript |
conversation.item.input_audio_transcription.updated |
DOCUMENTED | Streaming transcription update for the user's audio input. Emitted as the user speaks, providing the cumulative transcript so far before the final completed event. Note that this is the cumulative transcript which may have corrections to previous updated transcripts — this is different from a transcript delta. Only emitted when audio.input.transcription.model is set to grok-transcribe in the session configuration. Useful for displaying live captions. |
event_id, type, item_id, transcript |
input_audio_buffer.dtmf_event_received |
DOCUMENTED | A DTMF tone (phone keypress) was detected on a SIP session. SIP only — not emitted on direct WebSocket connections. Digits are buffered server-side and flushed as a text message to the model on # key, 2.5s idle, or when the user begins speaking. |
event_id, type, event, received_at |
response.created |
DOCUMENTED + LIVE_VERIFIED | A new assistant response turn is in progress. Audio deltas from this turn share the same response_id. | event_id, type, response {id, object, status, output, metadata} |
response.output_item.added |
DOCUMENTED + LIVE_VERIFIED | A new assistant response item is added to the message history. | event_id, type, response_id, output_index, item |
response.output_item.done |
DOCUMENTED + LIVE_VERIFIED | An output item is complete. | event_id, type, response_id, output_index, item |
response.content_part.added |
DOCUMENTED + LIVE_VERIFIED | A content part starts within an output item. | event_id, type, response_id, item_id, output_index, content_index, part {type, transcript} |
response.content_part.done |
DOCUMENTED + LIVE_VERIFIED | A content part finishes. | event_id, type, response_id, item_id, output_index, content_index, part {type, transcript} |
response.output_audio_transcript.delta |
DOCUMENTED + LIVE_VERIFIED | Streaming text transcript delta of the assistant's audio response. | event_id, type, response_id, item_id, output_index, content_index, delta |
response.output_audio_transcript.done |
DOCUMENTED + LIVE_VERIFIED | The audio transcript for this assistant turn has finished generating. | event_id, type, response_id, item_id, output_index, content_index, transcript |
response.output_audio.delta |
DOCUMENTED + LIVE_VERIFIED | Streaming base64-encoded audio delta of the assistant's response. | event_id, type, response_id, item_id, output_index, content_index, delta |
response.output_audio.done |
DOCUMENTED + LIVE_VERIFIED | Audio generation for this assistant turn has finished. | event_id, type, response_id, item_id, output_index, content_index |
response.text.delta |
DOCUMENTED | Text-mode output delta (when using text modality). | type ∈ response.text.delta, event_id, response_id, item_id, output_index, content_index, delta |
response.output_text.delta |
DOCUMENTED | Text-mode output delta using the OpenAI GA event name. Functionally identical to response.text.delta. Clients should handle both event names for maximum compatibility. |
type ∈ response.output_text.delta, event_id, response_id, item_id, output_index, content_index, delta |
response.function_call_arguments.delta |
DOCUMENTED | Streaming function call arguments. | event_id, type, response_id, item_id, output_index, call_id, delta |
response.function_call_arguments.done |
DOCUMENTED | A function call has been triggered with complete arguments. Your code should execute the function and return results via conversation.item.create with type function_call_output. |
event_id, type, response_id, item_id, output_index, call_id, name, arguments |
mcp_list_tools.in_progress |
DOCUMENTED | MCP tool discovery has started. | event_id, type, item_id |
mcp_list_tools.completed |
DOCUMENTED | MCP tool discovery succeeded. | event_id, type, item_id |
mcp_list_tools.failed |
DOCUMENTED | MCP tool discovery failed. | event_id, type, item_id, error {type, message} |
response.mcp_call_arguments.delta |
DOCUMENTED | MCP call arguments streaming. | event_id, type, response_id, item_id, call_id, delta |
response.mcp_call_arguments.done |
DOCUMENTED | MCP call arguments finalized. | event_id, type, response_id, item_id, call_id, name, arguments |
response.mcp_call.in_progress |
DOCUMENTED | MCP server HTTP call starting. | event_id, type, item_id, output_index |
response.mcp_call.completed |
DOCUMENTED | MCP tool execution succeeded. | event_id, type, item_id, output_index |
response.mcp_call.failed |
DOCUMENTED | MCP tool execution failed. | event_id, type, item_id, output_index, error {type, message} |
response.done |
DOCUMENTED + LIVE_VERIFIED | The assistant's response is completed. Sent after all audio and transcript deltas. Ready for the client to add a new conversation item. | event_id, type, response {id, object, status, usage, metadata} |
error |
DOCUMENTED | Sent when an error occurs. Contains error code and message. Most errors are recoverable and the session stays open. | event_id, type, error {type, code, message, param, event_id} |
ping |
LIVE_DISCOVERED | Keep-alive / clock event sent right after conversation.created (and presumably periodically). Not in the official ws.json schema. | type, event_id, timestamp, previous_item_id |
response.audio.delta |
DOCUMENTED + UNVERIFIED | Legacy alias of response.output_audio.delta mentioned in the Speech-to-Speech guide (audio transport table); not emitted in the live probe (only response.output_audio.delta was). |
Speech to Speech — wss://api.x.ai/v1/realtime — server→client (webhook) (1)
| event | status | description | fields |
|---|---|---|---|
realtime.call.incoming |
DOCUMENTED | Signed webhook (Standard Webhooks v1: webhook-id, webhook-timestamp, webhook-signature headers, HMAC-SHA256 with dispatch_signing_secret) POSTed to the phone number's webhook URL when a SIP call arrives; data.call_id is then used as ?call_id= on wss://api.x.ai/v1/realtime. |
type, data |
Streaming TTS — wss://api.x.ai/v1/tts — client→server (2)
| event | status | description | fields |
|---|---|---|---|
text.delta |
DOCUMENTED | Send a chunk of text to be synthesized. Text is processed incrementally — audio generation begins as soon as enough text is buffered. Individual deltas are capped at 60,000 characters. | type, delta |
text.done |
DOCUMENTED | Signal that all text for this utterance has been sent. The server will finish generating audio and send audio.done. After receiving audio.done, you can start a new utterance with another text.delta. |
type |
Streaming TTS — wss://api.x.ai/v1/tts — server→client (3)
| event | status | description | fields |
|---|---|---|---|
audio.delta |
DOCUMENTED | A chunk of base64-encoded audio data. Decode and append to your audio buffer or pipe directly to playback. The format matches the codec and sample_rate specified in the query parameters. When the connection was opened with with_timestamps=true, the event also carries audio_timestamps and audio_duration for the characters that fall inside this chunk. |
type, delta, audio_timestamps {graph_chars, graph_times}, audio_duration |
audio.done |
DOCUMENTED | Audio generation for this utterance is complete. The connection remains open for multi-utterance — send another text.delta to start a new synthesis, or close the connection. |
type, trace_id |
error |
DOCUMENTED | An error occurred during synthesis. The connection may be closed after this message. | type, message |
Streaming STT — wss://api.x.ai/v1/stt — client→server (3)
| event | status | description | fields |
|---|---|---|---|
Binary frame (audio) |
DOCUMENTED | Send raw audio as binary WebSocket frames in the encoding specified by the encoding query parameter. Audio should be streamed in real-time-paced chunks (e.g. 100 ms at a time). No base64 encoding — send raw bytes directly. With encoding=opus, each binary frame must contain exactly one raw Opus packet — never concatenate packets or split one across frames. An undecodable frame sends an error event and closes the session. |
|
finalize |
DOCUMENTED | Force the current utterance to finalize as speech_final immediately, without waiting for VAD endpointing or Smart Turn. The session stays open so you can continue streaming audio. Accepts finalize or Finalize as the type value. When multichannel=true, optional channel (0-based) limits the finalize to that channel; omit channel to finalize every channel. |
type ∈ finalize|Finalize, channel |
audio.done |
DOCUMENTED | Signal that all audio has been sent. The server flushes any remaining buffered audio, emits final transcript events, and sends a transcript.done event. The connection closes after transcript.done. |
type |
Streaming STT — wss://api.x.ai/v1/stt — server→client (4)
| event | status | description | fields |
|---|---|---|---|
transcript.created |
DOCUMENTED | Sent immediately after the WebSocket connection is established and the server is ready to receive audio. Wait for this event before sending audio — the server needs to initialize its ASR backend. | type, id |
transcript.partial |
DOCUMENTED | A transcript result for a portion of the audio stream. Two boolean fields convey state: interim (is_final=false) means text may still change, chunk final (is_final=true, speech_final=false) means the chunk is locked, and utterance final (is_final=true, speech_final=true) means the speaker stopped talking. |
type, text, words, is_final, speech_final, start, duration, channel_index, end_of_turn_confidence |
transcript.done |
DOCUMENTED | Final transcript after audio.done. duration always present. One per channel when multichannel=true. Connection closes after this event. |
type, text, words, duration, channel_index |
error |
DOCUMENTED | An error occurred during the session. Most errors (pipeline failures, stream timeouts, undecodable audio frames) close the connection. Only client message parse errors keep the connection open. | type, message |