Gemini Live API — full message/event reference (client and server)
Status: DOCUMENTED + LIVE_VERIFIED (2026-09-18 run — see "Live verification" section at the end) · PREVIEW (Live API)
Sources: https://ai.google.dev/api/live · https://ai.google.dev/gemini-api/docs/live-api/get-started-websocket · …/live-api/capabilities · …/live-api/session-management · …/live-api/tools · …/live-api/thinking · …/live-api/live-transcribe · …/live-api/live-translate · discovery v1beta rev. 20260918 · SDK types python-genai-types.py (LiveClient*, LiveServer*)
Last verified: 2026-09-18
Machine-readable twin: generated/fragments/streaming-events/gemini-live.json (32 records). Overview, models, auth, pricing: live-api.md.
Conventions: field names are the wire JSON (camelCase). SDK equivalents are snake_case in Python (server_content.turn_complete) and camelCase in Node. "Constructed" examples follow the documented shape with illustrative values; the others are copied from official docs.
1. Message envelopes and ordering rules
| Direction | Envelope | Rule |
|---|---|---|
| client→server | BidiGenerateContentClientMessage |
exactly one of setup, clientContent, realtimeInput, toolResponse |
| server→client | BidiGenerateContentServerMessage |
optional usageMetadata + exactly one of setupComplete, serverContent, toolCall, toolCallCancellation, goAway, sessionResumptionUpdate (union flattened to top level) |
Ordering (documented):
setupmust be the first client message; wait forsetupCompletebefore sending anything else.- Realtime streams (audio, video, text) are concurrent; no ordering guarantee across streams; mixing
clientContentandrealtimeInputis best-effort. - Transcriptions (
inputTranscription,interimInputTranscription,outputTranscription) are sent independently ofmodelTurn; the lastoutputTranscriptionof a turn precedesgenerationComplete/interrupted, which precedeturnComplete. - Normal turn:
modelTurn* →generationComplete→ (playback delay) →turnComplete(+interactionStatus). Interrupted turn:interrupted→turnComplete(nogenerationComplete). interactionStatusalways accompaniesturnComplete; on extended-thinking,IN_PROGRESSmeans more output / tool calls follow.goAwayandsessionResumptionUpdatecan arrive at any time;usageMetadataaccompanies other messages periodically.
Typical text-only turn (order hints used in the JSON twin): setup 0 → setupComplete 1 → clientContent 2 → serverContent.modelTurn 3 → serverContent.generationComplete 4 → serverContent.turnComplete 5 → usageMetadata 6.
2. Client messages
2.1 setup — BidiGenerateContentSetup
| Field | Type | Req. | Enum / notes |
|---|---|---|---|
model |
string | yes | models/{model} |
generationConfig |
GenerationConfig subset | no | see live-api.md §4; unsupported: responseLogprobs, responseMimeType, logprobs, responseSchema, responseJsonSchema, stopSequence, skipResponseCache, routingConfig, audioTimestamp |
generationConfig.responseModalities[] |
enum[] | no | TEXT, AUDIO |
generationConfig.speechConfig |
SpeechConfig | no | voiceConfig.prebuiltVoiceConfig.voiceName, languageCode, multiSpeakerVoiceConfig (UNVERIFIED in Live) |
generationConfig.mediaResolution |
enum | no | MEDIA_RESOLUTION_LOW | MEDIUM | HIGH |
generationConfig.thinkingConfig |
ThinkingConfig | no | thinkingLevel (MINIMAL/LOW/MEDIUM/HIGH), thinkingBudget, includeThoughts |
generationConfig.translationConfig |
TranslationConfig | no | targetLanguageCode, echoTargetLanguage |
generationConfig.enableAffectiveDialog |
bool | no | 2.5 only |
systemInstruction |
Content | no | text parts only |
tools[] |
Tool[] | no | functionDeclarations[]{name, description, parameters, behavior BLOCKING|NON_BLOCKING}, googleSearch {} |
realtimeInputConfig |
RealtimeInputConfig | no | automaticActivityDetection{disabled, startOfSpeechSensitivity, endOfSpeechSensitivity, prefixPaddingMs, silenceDurationMs}, activityHandling, turnCoverage |
sessionResumption |
SessionResumptionConfig | no | handle (transparent SDK-only) |
contextWindowCompression |
ContextWindowCompressionConfig | no | triggerTokens, slidingWindow.targetTokens |
inputAudioTranscription / outputAudioTranscription |
AudioTranscriptionConfig | no | languageCodes[], customVocabulary[], mode VERBATIM|SMART, wordTimestamp, diarization |
proactivity |
ProactivityConfig | no | proactiveAudio |
historyConfig |
HistoryConfig | no | initialHistoryInClientContent |
labels |
map | no | discovery-only, UNVERIFIED |
{"setup": {"model": "models/gemini-3.8-live-extended-thinking",
"generationConfig": {"responseModalities": ["AUDIO"],
"speechConfig": {"voiceConfig": {"prebuiltVoiceConfig": {"voiceName": "Puck"}}},
"thinkingConfig": {"thinkingLevel": "LOW"}},
"tools": [{"functionDeclarations": [{"name": "searchFlights", "description": "Searches for flights between cities.",
"behavior": "NON_BLOCKING", "parameters": {"type": "OBJECT", "properties": {"destination": {"type": "STRING"}}, "required": ["destination"]}}]}],
"inputAudioTranscription": {}, "outputAudioTranscription": {}, "sessionResumption": {},
"contextWindowCompression": {"triggerTokens": "25600", "slidingWindow": {"targetTokens": "12800"}}}}Sent: once, immediately after the socket opens. Then wait for setupComplete.
2.2 clientContent — BidiGenerateContentClientContent
| Field | Type | Req. | Notes |
|---|---|---|---|
turns[] |
Content[] | no | {role: "user"|"model", parts: [{text}, {inlineData}, …]}; appended unconditionally to history; explicit roles supported all session long (3.1+) |
turnComplete |
bool | no | true → generate now (unconditionally interrupts active generation); absent/false → server waits for more messages |
{"clientContent": {"turns": [{"role": "user", "parts": [{"text": "Hello world!"}]}], "turnComplete": true}}Sent: to inject text turns or restore history (also before realtime starts when historyConfig.initialHistoryInClientContent is true). A clientContent message interrupts any current generation.
2.3 realtimeInput — BidiGenerateContentRealtimeInput
| Field | Type | Req. | Notes |
|---|---|---|---|
audio |
Blob {mimeType, data} |
one of | audio/pcm;rate=16000, raw 16-bit LE PCM mono, base64 |
video |
Blob | one of | image/jpeg / image/png frame, ≤ 1 FPS |
text |
string | one of | realtime text stream |
activityStart |
{} |
one of | manual VAD only (automaticActivityDetection.disabled: true) |
activityEnd |
{} |
one of | manual VAD only; server finalizes immediately |
audioStreamEnd |
bool | one of | automatic VAD only; mic off / hybrid VAD end-of-speech; flushes cached audio |
mediaResolution |
enum | no | per-message override |
mediaChunks[] |
Blob[] | deprecated | only first chunk used |
{"realtimeInput": {"audio": {"data": "UklGRiQAAABXQVZF...", "mimeType": "audio/pcm;rate=16000"}}}
{"realtimeInput": {"video": {"data": "<base64 JPEG>", "mimeType": "image/jpeg"}}}
{"realtimeInput": {"text": "Hello, how are you?"}}
{"realtimeInput": {"activityStart": {}}}
{"realtimeInput": {"activityEnd": {}}}
{"realtimeInput": {"audioStreamEnd": true}}Sent: continuously; does not interrupt generation by itself (VAD does). End of turn is derived from activity. Always treated as user input (cannot populate history).
2.4 toolResponse — BidiGenerateContentToolResponse
| Field | Type | Req. | Notes |
|---|---|---|---|
functionResponses[] |
FunctionResponse[] | yes | one per toolCall.functionCalls[] |
functionResponses[].id |
string | yes | = FunctionCall.id |
functionResponses[].name |
string | yes | = FunctionCall.name |
functionResponses[].response |
object | yes | any JSON (output, result, error …) |
functionResponses[].scheduling |
enum | no | INTERRUPT | WHEN_IDLE (default) | SILENT — NON_BLOCKING calls only; not on 3.8-extended-thinking |
functionResponses[].willContinue |
bool | no | NON_BLOCKING only; generator semantics |
functionResponses[].parts[] |
FunctionResponsePart[] | no | multimodal results |
{"toolResponse": {"functionResponses": [{"id": "call_123", "name": "searchFlights",
"response": {"output": {"flight": "DL 145", "price": "$145"}, "scheduling": "INTERRUPT"}}]}}Sent: after executing a toolCall. BLOCKING calls: the model waits; NON_BLOCKING: conversation continues meanwhile.
3. Server messages
3.1 setupComplete — BidiGenerateContentSetupComplete
No fields (reference). SDK types add sessionId, voiceConsentSignature (UNVERIFIED).
{"setupComplete": {}}Sent: once, in response to setup.
3.2 serverContent — BidiGenerateContentServerContent
| Field | Type | Enum / notes | When |
|---|---|---|---|
modelTurn |
Content | parts[]: inlineData {mimeType: "audio/pcm;rate=24000", data}, text, thought parts, executableCode… — several parts per message possible |
while generating |
generationComplete |
bool | absent in interrupted turns | model done generating |
turnComplete |
bool | turn over (after playback if realtime playback assumed) | |
interactionStatus |
enum IN_PROGRESS | IDLE (| REQUIRES_ACTION deprecated) |
always with turnComplete; also seen with toolCall on extended thinking |
with turnComplete |
interrupted |
bool | stop and flush playback | barge-in / clientContent |
groundingMetadata |
GroundingMetadata | webSearchQueries, groundingChunks, groundingSupports, searchEntryPoint |
Google Search used |
inputTranscription |
Transcription {text, languageCode} |
finalized user speech | inputAudioTranscription set |
interimInputTranscription |
Transcription | partial hypotheses, frequent | transcribe-live |
outputTranscription |
Transcription | model speech transcript | outputAudioTranscription set |
urlContextMetadata |
UrlContextMetadata {urlMetadata[]} |
tool listed unsupported on Live | — |
waitingForInput |
bool | model expects the user to continue | — |
speechState |
enum (deprecated) | use VoiceActivity | — |
turnCompleteReason |
enum (SDK-only) | MALFORMED_FUNCTION_CALL, RESPONSE_REJECTED, NEED_MORE_INPUT, PROHIBITED_INPUT_CONTENT, … UNVERIFIED |
— |
{"serverContent": {"modelTurn": {"parts": [{"inlineData": {"mimeType": "audio/pcm;rate=24000", "data": "..."}}]}}}
{"serverContent": {"inputTranscription": {"text": "Hello, how are you?", "languageCode": "en"}}}
{"serverContent": {"outputTranscription": {"text": "I'm doing well, thanks!"}}}
{"serverContent": {"generationComplete": true}}
{"serverContent": {"turnComplete": true, "interactionStatus": "IDLE"}}
{"serverContent": {"interrupted": true}}Extended-thinking filler (docs):
{"serverContent": {"modelTurn": {"parts": [{"inlineData": {"mimeType": "audio/pcm;rate=24000", "data": "..."}}]},
"turnComplete": true, "interactionStatus": "IN_PROGRESS"}} 3.3 toolCall — BidiGenerateContentToolCall
| Field | Type | Notes |
|---|---|---|
functionCalls[] |
FunctionCall[] | {id, name, args}; parallel calls possible |
{"toolCall": {"functionCalls": [{"id": "call_123", "name": "searchFlights", "args": {"destination": "Seattle"}}]}}Sent: when the model decides to call functions; on extended thinking it may be tagged "interactionStatus": "IN_PROGRESS" (docs example places it at top level next to toolCall).
3.4 toolCallCancellation — BidiGenerateContentToolCallCancellation
| Field | Type | Notes |
|---|---|---|
ids[] |
string[] | calls that must not run / should be undone |
{"toolCallCancellation": {"ids": ["call_123"]}}Sent: only when the client interrupts a server turn (barge-in) while calls are pending. Constructed example.
3.5 usageMetadata — UsageMetadata (companion field)
| Field | Type | Notes |
|---|---|---|
promptTokenCount |
int32 | includes cached content |
cachedContentTokenCount |
int32 | |
responseTokenCount |
int32 | all candidates |
toolUsePromptTokenCount |
int32 | |
thoughtsTokenCount |
int32 | thinking models |
totalTokenCount |
int32 | prompt + response (+ thoughts + tool-use) |
promptTokensDetails[] |
ModalityTokenCount[] | {modality: TEXT|IMAGE|VIDEO|AUDIO|DOCUMENT, tokenCount} |
cacheTokensDetails[] |
ModalityTokenCount[] | |
responseTokensDetails[] |
ModalityTokenCount[] | |
toolUsePromptTokensDetails[] |
ModalityTokenCount[] |
{"usageMetadata": {"promptTokenCount": 1250, "responseTokenCount": 420, "totalTokenCount": 1670,
"promptTokensDetails": [{"modality": "AUDIO", "tokenCount": 1200}, {"modality": "TEXT", "tokenCount": 50}],
"responseTokensDetails": [{"modality": "AUDIO", "tokenCount": 420}]}}(Constructed values.) Sent: alongside other server messages, "periodically" (SDK sample loops on message.usage_metadata). The modality breakdown is how audio vs text vs image/video billing is split; audio ≈ 25 tokens/s. Since the whole context is re-billed each turn, promptTokenCount grows with the session unless contextWindowCompression caps it.
3.6 goAway — GoAway
| Field | Type | Notes |
|---|---|---|
timeLeft |
Duration string (e.g. "30s") |
time before the connection is terminated as ABORTED; never below a model-specific minimum (not published) |
{"goAway": {"timeLeft": "30s"}}(Constructed value.) Sent: before the server resets the connection (lifetime ≈ 10 min). Action: finish/queue work, then reconnect with sessionResumption.handle.
3.7 sessionResumptionUpdate — SessionResumptionUpdate
| Field | Type | Notes |
|---|---|---|
newHandle |
string | resumable state handle; empty if resumable is false |
resumable |
bool | false while generating / executing function calls |
lastConsumedClientMessageIndex |
int64 (SDK-only, transparent mode) |
index of last client message included in the state — UNVERIFIED |
{"sessionResumptionUpdate": {"newHandle": "<opaque handle>", "resumable": true}}Sent: periodically, only if setup.sessionResumption was present. Keep the latest handle with resumable: true; it stays valid 2 h after the last session termination.
3.8 SDK-only server fields (UNVERIFIED)
voiceActivity {voiceActivityType: ACTIVITY_START|ACTIVITY_END, audioOffset} and voiceActivityDetectionSignal {vadSignalType: VAD_SIGNAL_TYPE_SOS|EOS} ("allowlisted only", tied to setup.explicitVadSignal). Present in LiveServerMessage SDK types, absent from the public reference.
4. usageMetadata modality breakdown — reading it
| Question | Where |
|---|---|
| How much audio did I send this turn (incl. re-billed history)? | promptTokensDetails[modality=AUDIO].tokenCount |
| Video/image frames cost | `promptTokensDetails[modality=IMAGE |
| Spoken output cost | responseTokensDetails[modality=AUDIO] |
| Transcripts / TEXT modality | responseTokensDetails[modality=TEXT] (billed at text output rate) |
| Reasoning | thoughtsTokenCount (included in output price) |
| Tool results fed back | toolUsePromptTokenCount / toolUsePromptTokensDetails[] |
5. goAway / sessionResumptionUpdate timing
t=0 setup → setupComplete ; sessionResumptionUpdate{newHandle=h1, resumable=true}
… periodic sessionResumptionUpdate (resumable=false while the model speaks / runs tools)
t≈10 min goAway{timeLeft} → client drains playback, opens a new WebSocket
t≈10 min setup{…, sessionResumption:{handle:h_latest}} → setupComplete → continue (config may change, model may not)- Session limits without compression: 15 min audio-only / 2 min audio+video → enable
contextWindowCompressionfor unlimited sessions. - Handles expire 2 h after the last session termination.
- Ephemeral tokens: reconnecting with a handle does not consume a
uses; it must happen beforeexpireTime.
Live verification (2026-09-18)
Observed 2026-09-18 on gemini-2.5-flash-native-audio-latest (AUDIO turn, tmp-live/gemini-tools/j1_*.json, j4_*.json, j5_ws_toolcall.json):
→ setup {model, generationConfig.responseModalities:[AUDIO], outputAudioTranscription:{}, sessionResumption:{}}
← setupComplete {}
← sessionResumptionUpdate {newHandle, resumable:true}
→ clientContent {turns:[{role:user, parts:[{text}]}], turnComplete:true}
← serverContent {modelTurn:{parts:[{text:"…", thought:true}]}}
← toolCall {functionCalls:[{name, args, id:"function-call-…"}]} (only when a NON_BLOCKING declaration is set)
← serverContent {outputTranscription:{text:"OK"}}
← serverContent {modelTurn:{parts:[{inlineData:{mimeType:"audio/pcm;rate=24000", data}}]}} × N
← serverContent {generationComplete:true}
← serverContent {turnComplete:true} + usageMetadata {promptTokenCount, responseTokenCount, totalTokenCount, promptTokensDetails, responseTokensDetails:[{modality:AUDIO}], thoughtsTokenCount}
← sessionResumptionUpdate {newHandle, resumable:true}Rejections: responseModalities:[TEXT] → close 1007 on every current live model; setup.toolConfig → close 1007; ephemeral token on the non-Constrained method → close 1008. Not observed: interrupted, goAway, toolCallCancellation, inputTranscription, groundingMetadata, voiceActivity*.