Python 88.3%
TypeScript 7.6%
Shell 4.1%
1# Gemini Live API — full message/event reference (client and server)23**Status:** DOCUMENTED + LIVE_VERIFIED (2026-09-18 run — see "Live verification" section at the end) · PREVIEW (Live API)4**Sources:** https://ai.google.dev/api/live · https://ai.google.dev/gemini-api/docs/live-api/get-started-websocket · …/live-api/capabilities · …/live-api/session-management · …/live-api/tools · …/live-api/thinking · …/live-api/live-transcribe · …/live-api/live-translate · discovery `v1beta` rev. 20260918 · SDK types `python-genai-types.py` (`LiveClient*`, `LiveServer*`)5**Last verified:** 2026-09-186**Machine-readable twin:** `generated/fragments/streaming-events/gemini-live.json` (32 records). Overview, models, auth, pricing: [`live-api.md`](live-api.md).78Conventions: field names are the wire JSON (camelCase). SDK equivalents are snake_case in Python (`server_content.turn_complete`) and camelCase in Node. "Constructed" examples follow the documented shape with illustrative values; the others are copied from official docs.910## 1. Message envelopes and ordering rules1112| Direction | Envelope | Rule |13|---|---|---|14| client→server | `BidiGenerateContentClientMessage` | exactly **one** of `setup`, `clientContent`, `realtimeInput`, `toolResponse` |15| server→client | `BidiGenerateContentServerMessage` | optional `usageMetadata` + exactly **one** of `setupComplete`, `serverContent`, `toolCall`, `toolCallCancellation`, `goAway`, `sessionResumptionUpdate` (union flattened to top level) |1617Ordering (documented):181. `setup` must be the first client message; wait for `setupComplete` before sending anything else.192. Realtime streams (audio, video, text) are concurrent; **no ordering guarantee across streams**; mixing `clientContent` and `realtimeInput` is best-effort.203. Transcriptions (`inputTranscription`, `interimInputTranscription`, `outputTranscription`) are sent independently of `modelTurn`; the last `outputTranscription` of a turn precedes `generationComplete` / `interrupted`, which precede `turnComplete`.214. Normal turn: `modelTurn`* → `generationComplete` → (playback delay) → `turnComplete` (+ `interactionStatus`). Interrupted turn: `interrupted` → `turnComplete` (no `generationComplete`).225. `interactionStatus` always accompanies `turnComplete`; on extended-thinking, `IN_PROGRESS` means more output / tool calls follow.236. `goAway` and `sessionResumptionUpdate` can arrive at any time; `usageMetadata` accompanies other messages periodically.2425Typical text-only turn (order hints used in the JSON twin): `setup` 0 → `setupComplete` 1 → `clientContent` 2 → `serverContent.modelTurn` 3 → `serverContent.generationComplete` 4 → `serverContent.turnComplete` 5 → `usageMetadata` 6.2627## 2. Client messages2829### 2.1 `setup` — `BidiGenerateContentSetup`3031| Field | Type | Req. | Enum / notes |32|---|---|---|---|33| `model` | string | yes | `models/{model}` |34| `generationConfig` | GenerationConfig subset | no | see [`live-api.md` §4](live-api.md); unsupported: responseLogprobs, responseMimeType, logprobs, responseSchema, responseJsonSchema, stopSequence, skipResponseCache, routingConfig, audioTimestamp |35| `generationConfig.responseModalities[]` | enum[] | no | `TEXT`, `AUDIO` |36| `generationConfig.speechConfig` | SpeechConfig | no | `voiceConfig.prebuiltVoiceConfig.voiceName`, `languageCode`, `multiSpeakerVoiceConfig` (UNVERIFIED in Live) |37| `generationConfig.mediaResolution` | enum | no | `MEDIA_RESOLUTION_LOW` \| `MEDIUM` \| `HIGH` |38| `generationConfig.thinkingConfig` | ThinkingConfig | no | `thinkingLevel` (`MINIMAL`/`LOW`/`MEDIUM`/`HIGH`), `thinkingBudget`, `includeThoughts` |39| `generationConfig.translationConfig` | TranslationConfig | no | `targetLanguageCode`, `echoTargetLanguage` |40| `generationConfig.enableAffectiveDialog` | bool | no | 2.5 only |41| `systemInstruction` | Content | no | text parts only |42| `tools[]` | Tool[] | no | `functionDeclarations[]{name, description, parameters, behavior BLOCKING\|NON_BLOCKING}`, `googleSearch {}` |43| `realtimeInputConfig` | RealtimeInputConfig | no | `automaticActivityDetection{disabled, startOfSpeechSensitivity, endOfSpeechSensitivity, prefixPaddingMs, silenceDurationMs}`, `activityHandling`, `turnCoverage` |44| `sessionResumption` | SessionResumptionConfig | no | `handle` (`transparent` SDK-only) |45| `contextWindowCompression` | ContextWindowCompressionConfig | no | `triggerTokens`, `slidingWindow.targetTokens` |46| `inputAudioTranscription` / `outputAudioTranscription` | AudioTranscriptionConfig | no | `languageCodes[]`, `customVocabulary[]`, `mode VERBATIM\|SMART`, `wordTimestamp`, `diarization` |47| `proactivity` | ProactivityConfig | no | `proactiveAudio` |48| `historyConfig` | HistoryConfig | no | `initialHistoryInClientContent` |49| `labels` | map | no | discovery-only, UNVERIFIED |5051```json52{"setup": {"model": "models/gemini-3.8-live-extended-thinking",53 "generationConfig": {"responseModalities": ["AUDIO"],54 "speechConfig": {"voiceConfig": {"prebuiltVoiceConfig": {"voiceName": "Puck"}}},55 "thinkingConfig": {"thinkingLevel": "LOW"}},56 "tools": [{"functionDeclarations": [{"name": "searchFlights", "description": "Searches for flights between cities.",57 "behavior": "NON_BLOCKING", "parameters": {"type": "OBJECT", "properties": {"destination": {"type": "STRING"}}, "required": ["destination"]}}]}],58 "inputAudioTranscription": {}, "outputAudioTranscription": {}, "sessionResumption": {},59 "contextWindowCompression": {"triggerTokens": "25600", "slidingWindow": {"targetTokens": "12800"}}}}60```61Sent: once, immediately after the socket opens. Then wait for `setupComplete`.6263### 2.2 `clientContent` — `BidiGenerateContentClientContent`6465| Field | Type | Req. | Notes |66|---|---|---|---|67| `turns[]` | Content[] | no | `{role: "user"\|"model", parts: [{text}, {inlineData}, …]}`; appended unconditionally to history; explicit roles supported all session long (3.1+) |68| `turnComplete` | bool | no | `true` → generate now (unconditionally interrupts active generation); absent/false → server waits for more messages |6970```json71{"clientContent": {"turns": [{"role": "user", "parts": [{"text": "Hello world!"}]}], "turnComplete": true}}72```73Sent: to inject text turns or restore history (also before realtime starts when `historyConfig.initialHistoryInClientContent` is true). A `clientContent` message interrupts any current generation.7475### 2.3 `realtimeInput` — `BidiGenerateContentRealtimeInput`7677| Field | Type | Req. | Notes |78|---|---|---|---|79| `audio` | Blob `{mimeType, data}` | one of | `audio/pcm;rate=16000`, raw 16-bit LE PCM mono, base64 |80| `video` | Blob | one of | `image/jpeg` / `image/png` frame, ≤ 1 FPS |81| `text` | string | one of | realtime text stream |82| `activityStart` | `{}` | one of | manual VAD only (`automaticActivityDetection.disabled: true`) |83| `activityEnd` | `{}` | one of | manual VAD only; server finalizes immediately |84| `audioStreamEnd` | bool | one of | automatic VAD only; mic off / hybrid VAD end-of-speech; flushes cached audio |85| `mediaResolution` | enum | no | per-message override |86| `mediaChunks[]` | Blob[] | deprecated | only first chunk used |8788```json89{"realtimeInput": {"audio": {"data": "UklGRiQAAABXQVZF...", "mimeType": "audio/pcm;rate=16000"}}}90{"realtimeInput": {"video": {"data": "<base64 JPEG>", "mimeType": "image/jpeg"}}}91{"realtimeInput": {"text": "Hello, how are you?"}}92{"realtimeInput": {"activityStart": {}}}93{"realtimeInput": {"activityEnd": {}}}94{"realtimeInput": {"audioStreamEnd": true}}95```96Sent: continuously; does not interrupt generation by itself (VAD does). End of turn is derived from activity. Always treated as user input (cannot populate history).9798### 2.4 `toolResponse` — `BidiGenerateContentToolResponse`99100| Field | Type | Req. | Notes |101|---|---|---|---|102| `functionResponses[]` | FunctionResponse[] | yes | one per `toolCall.functionCalls[]` |103| `functionResponses[].id` | string | yes | = `FunctionCall.id` |104| `functionResponses[].name` | string | yes | = `FunctionCall.name` |105| `functionResponses[].response` | object | yes | any JSON (`output`, `result`, `error` …) |106| `functionResponses[].scheduling` | enum | no | `INTERRUPT` \| `WHEN_IDLE` (default) \| `SILENT` — NON_BLOCKING calls only; not on 3.8-extended-thinking |107| `functionResponses[].willContinue` | bool | no | NON_BLOCKING only; generator semantics |108| `functionResponses[].parts[]` | FunctionResponsePart[] | no | multimodal results |109110```json111{"toolResponse": {"functionResponses": [{"id": "call_123", "name": "searchFlights",112 "response": {"output": {"flight": "DL 145", "price": "$145"}, "scheduling": "INTERRUPT"}}]}}113```114Sent: after executing a `toolCall`. BLOCKING calls: the model waits; NON_BLOCKING: conversation continues meanwhile.115116## 3. Server messages117118### 3.1 `setupComplete` — `BidiGenerateContentSetupComplete`119120No fields (reference). SDK types add `sessionId`, `voiceConsentSignature` (UNVERIFIED).121```json122{"setupComplete": {}}123```124Sent: once, in response to `setup`.125126### 3.2 `serverContent` — `BidiGenerateContentServerContent`127128| Field | Type | Enum / notes | When |129|---|---|---|---|130| `modelTurn` | Content | `parts[]`: `inlineData {mimeType: "audio/pcm;rate=24000", data}`, `text`, `thought` parts, `executableCode`… — several parts per message possible | while generating |131| `generationComplete` | bool | absent in interrupted turns | model done generating |132| `turnComplete` | bool | | turn over (after playback if realtime playback assumed) |133| `interactionStatus` | enum `IN_PROGRESS` \| `IDLE` (\| `REQUIRES_ACTION` deprecated) | always with `turnComplete`; also seen with `toolCall` on extended thinking | with turnComplete |134| `interrupted` | bool | stop and flush playback | barge-in / clientContent |135| `groundingMetadata` | GroundingMetadata | `webSearchQueries`, `groundingChunks`, `groundingSupports`, `searchEntryPoint` | Google Search used |136| `inputTranscription` | Transcription `{text, languageCode}` | finalized user speech | inputAudioTranscription set |137| `interimInputTranscription` | Transcription | partial hypotheses, frequent | transcribe-live |138| `outputTranscription` | Transcription | model speech transcript | outputAudioTranscription set |139| `urlContextMetadata` | UrlContextMetadata `{urlMetadata[]}` | tool listed unsupported on Live | — |140| `waitingForInput` | bool | model expects the user to continue | — |141| `speechState` | enum (deprecated) | use VoiceActivity | — |142| `turnCompleteReason` | enum (SDK-only) | `MALFORMED_FUNCTION_CALL`, `RESPONSE_REJECTED`, `NEED_MORE_INPUT`, `PROHIBITED_INPUT_CONTENT`, … UNVERIFIED | — |143144```json145{"serverContent": {"modelTurn": {"parts": [{"inlineData": {"mimeType": "audio/pcm;rate=24000", "data": "..."}}]}}}146{"serverContent": {"inputTranscription": {"text": "Hello, how are you?", "languageCode": "en"}}}147{"serverContent": {"outputTranscription": {"text": "I'm doing well, thanks!"}}}148{"serverContent": {"generationComplete": true}}149{"serverContent": {"turnComplete": true, "interactionStatus": "IDLE"}}150{"serverContent": {"interrupted": true}}151```152Extended-thinking filler (docs):153```json154{"serverContent": {"modelTurn": {"parts": [{"inlineData": {"mimeType": "audio/pcm;rate=24000", "data": "..."}}]},155 "turnComplete": true, "interactionStatus": "IN_PROGRESS"}}156```157158### 3.3 `toolCall` — `BidiGenerateContentToolCall`159160| Field | Type | Notes |161|---|---|---|162| `functionCalls[]` | FunctionCall[] | `{id, name, args}`; parallel calls possible |163164```json165{"toolCall": {"functionCalls": [{"id": "call_123", "name": "searchFlights", "args": {"destination": "Seattle"}}]}}166```167Sent: when the model decides to call functions; on extended thinking it may be tagged `"interactionStatus": "IN_PROGRESS"` (docs example places it at top level next to `toolCall`).168169### 3.4 `toolCallCancellation` — `BidiGenerateContentToolCallCancellation`170171| Field | Type | Notes |172|---|---|---|173| `ids[]` | string[] | calls that must not run / should be undone |174175```json176{"toolCallCancellation": {"ids": ["call_123"]}}177```178Sent: only when the client interrupts a server turn (barge-in) while calls are pending. Constructed example.179180### 3.5 `usageMetadata` — `UsageMetadata` (companion field)181182| Field | Type | Notes |183|---|---|---|184| `promptTokenCount` | int32 | includes cached content |185| `cachedContentTokenCount` | int32 | |186| `responseTokenCount` | int32 | all candidates |187| `toolUsePromptTokenCount` | int32 | |188| `thoughtsTokenCount` | int32 | thinking models |189| `totalTokenCount` | int32 | prompt + response (+ thoughts + tool-use) |190| `promptTokensDetails[]` | ModalityTokenCount[] | `{modality: TEXT\|IMAGE\|VIDEO\|AUDIO\|DOCUMENT, tokenCount}` |191| `cacheTokensDetails[]` | ModalityTokenCount[] | |192| `responseTokensDetails[]` | ModalityTokenCount[] | |193| `toolUsePromptTokensDetails[]` | ModalityTokenCount[] | |194195```json196{"usageMetadata": {"promptTokenCount": 1250, "responseTokenCount": 420, "totalTokenCount": 1670,197 "promptTokensDetails": [{"modality": "AUDIO", "tokenCount": 1200}, {"modality": "TEXT", "tokenCount": 50}],198 "responseTokensDetails": [{"modality": "AUDIO", "tokenCount": 420}]}}199```200(Constructed values.) Sent: alongside other server messages, "periodically" (SDK sample loops on `message.usage_metadata`). The **modality breakdown** is how audio vs text vs image/video billing is split; audio ≈ 25 tokens/s. Since the whole context is re-billed each turn, `promptTokenCount` grows with the session unless `contextWindowCompression` caps it.201202### 3.6 `goAway` — `GoAway`203204| Field | Type | Notes |205|---|---|---|206| `timeLeft` | Duration string (e.g. `"30s"`) | time before the connection is terminated as ABORTED; never below a model-specific minimum (not published) |207208```json209{"goAway": {"timeLeft": "30s"}}210```211(Constructed value.) Sent: before the server resets the connection (lifetime ≈ 10 min). Action: finish/queue work, then reconnect with `sessionResumption.handle`.212213### 3.7 `sessionResumptionUpdate` — `SessionResumptionUpdate`214215| Field | Type | Notes |216|---|---|---|217| `newHandle` | string | resumable state handle; empty if `resumable` is false |218| `resumable` | bool | false while generating / executing function calls |219| `lastConsumedClientMessageIndex` | int64 (SDK-only, `transparent` mode) | index of last client message included in the state — UNVERIFIED |220221```json222{"sessionResumptionUpdate": {"newHandle": "<opaque handle>", "resumable": true}}223```224Sent: periodically, only if `setup.sessionResumption` was present. Keep the latest handle with `resumable: true`; it stays valid 2 h after the last session termination.225226### 3.8 SDK-only server fields (UNVERIFIED)227228`voiceActivity {voiceActivityType: ACTIVITY_START|ACTIVITY_END, audioOffset}` and `voiceActivityDetectionSignal {vadSignalType: VAD_SIGNAL_TYPE_SOS|EOS}` ("allowlisted only", tied to `setup.explicitVadSignal`). Present in `LiveServerMessage` SDK types, absent from the public reference.229230## 4. `usageMetadata` modality breakdown — reading it231232| Question | Where |233|---|---|234| How much audio did I send this turn (incl. re-billed history)? | `promptTokensDetails[modality=AUDIO].tokenCount` |235| Video/image frames cost | `promptTokensDetails[modality=IMAGE|VIDEO]` (depends on `mediaResolution`, `turnCoverage`) |236| Spoken output cost | `responseTokensDetails[modality=AUDIO]` |237| Transcripts / TEXT modality | `responseTokensDetails[modality=TEXT]` (billed at text output rate) |238| Reasoning | `thoughtsTokenCount` (included in output price) |239| Tool results fed back | `toolUsePromptTokenCount` / `toolUsePromptTokensDetails[]` |240241## 5. `goAway` / `sessionResumptionUpdate` timing242243```244t=0 setup → setupComplete ; sessionResumptionUpdate{newHandle=h1, resumable=true}245… periodic sessionResumptionUpdate (resumable=false while the model speaks / runs tools)246t≈10 min goAway{timeLeft} → client drains playback, opens a new WebSocket247t≈10 min setup{…, sessionResumption:{handle:h_latest}} → setupComplete → continue (config may change, model may not)248```249- Session limits without compression: 15 min audio-only / 2 min audio+video → enable `contextWindowCompression` for unlimited sessions.250- Handles expire **2 h** after the last session termination.251- Ephemeral tokens: reconnecting with a handle does not consume a `uses`; it must happen before `expireTime`.252253## Live verification (2026-09-18)254255Observed 2026-09-18 on `gemini-2.5-flash-native-audio-latest` (AUDIO turn, `tmp-live/gemini-tools/j1_*.json`, `j4_*.json`, `j5_ws_toolcall.json`):256257```258→ setup {model, generationConfig.responseModalities:[AUDIO], outputAudioTranscription:{}, sessionResumption:{}}259← setupComplete {}260← sessionResumptionUpdate {newHandle, resumable:true}261→ clientContent {turns:[{role:user, parts:[{text}]}], turnComplete:true}262← serverContent {modelTurn:{parts:[{text:"…", thought:true}]}}263← toolCall {functionCalls:[{name, args, id:"function-call-…"}]} (only when a NON_BLOCKING declaration is set)264← serverContent {outputTranscription:{text:"OK"}}265← serverContent {modelTurn:{parts:[{inlineData:{mimeType:"audio/pcm;rate=24000", data}}]}} × N266← serverContent {generationComplete:true}267← serverContent {turnComplete:true} + usageMetadata {promptTokenCount, responseTokenCount, totalTokenCount, promptTokensDetails, responseTokensDetails:[{modality:AUDIO}], thoughtsTokenCount}268← sessionResumptionUpdate {newHandle, resumable:true}269```270Rejections: `responseModalities:[TEXT]` → close 1007 on every current live model; `setup.toolConfig` → close 1007; ephemeral token on the non-Constrained method → close 1008. Not observed: `interrupted`, `goAway`, `toolCallCancellation`, `inputTranscription`, `groundingMetadata`, `voiceActivity*`.271