SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
17.4 KB

# Gemini Live API — full message/event reference (client and server)

Status: DOCUMENTED + LIVE_VERIFIED (2026-09-18 run — see "Live verification" section at the end) · PREVIEW (Live API) Sources: https://ai.google.dev/api/live · https://ai.google.dev/gemini-api/docs/live-api/get-started-websocket · …/live-api/capabilities · …/live-api/session-management · …/live-api/tools · …/live-api/thinking · …/live-api/live-transcribe · …/live-api/live-translate · discovery v1beta rev. 20260918 · SDK types python-genai-types.py (LiveClient*, LiveServer*) Last verified: 2026-09-18 Machine-readable twin: generated/fragments/streaming-events/gemini-live.json (32 records). Overview, models, auth, pricing: live-api.md.

Conventions: field names are the wire JSON (camelCase). SDK equivalents are snake_case in Python (server_content.turn_complete) and camelCase in Node. "Constructed" examples follow the documented shape with illustrative values; the others are copied from official docs.

# 1. Message envelopes and ordering rules

Direction Envelope Rule
client→server BidiGenerateContentClientMessage exactly one of setup, clientContent, realtimeInput, toolResponse
server→client BidiGenerateContentServerMessage optional usageMetadata + exactly one of setupComplete, serverContent, toolCall, toolCallCancellation, goAway, sessionResumptionUpdate (union flattened to top level)

Ordering (documented):

  1. setup must be the first client message; wait for setupComplete before sending anything else.
  2. Realtime streams (audio, video, text) are concurrent; no ordering guarantee across streams; mixing clientContent and realtimeInput is best-effort.
  3. Transcriptions (inputTranscription, interimInputTranscription, outputTranscription) are sent independently of modelTurn; the last outputTranscription of a turn precedes generationComplete / interrupted, which precede turnComplete.
  4. Normal turn: modelTurn* → generationComplete → (playback delay) → turnComplete (+ interactionStatus). Interrupted turn: interrupted → turnComplete (no generationComplete).
  5. interactionStatus always accompanies turnComplete; on extended-thinking, IN_PROGRESS means more output / tool calls follow.
  6. goAway and sessionResumptionUpdate can arrive at any time; usageMetadata accompanies other messages periodically.

Typical text-only turn (order hints used in the JSON twin): setup 0 → setupComplete 1 → clientContent 2 → serverContent.modelTurn 3 → serverContent.generationComplete 4 → serverContent.turnComplete 5 → usageMetadata 6.

# 2. Client messages

# 2.1 setup — BidiGenerateContentSetup

Field Type Req. Enum / notes
model string yes models/{model}
generationConfig GenerationConfig subset no see live-api.md §4; unsupported: responseLogprobs, responseMimeType, logprobs, responseSchema, responseJsonSchema, stopSequence, skipResponseCache, routingConfig, audioTimestamp
generationConfig.responseModalities[] enum[] no TEXT, AUDIO
generationConfig.speechConfig SpeechConfig no voiceConfig.prebuiltVoiceConfig.voiceName, languageCode, multiSpeakerVoiceConfig (UNVERIFIED in Live)
generationConfig.mediaResolution enum no MEDIA_RESOLUTION_LOW | MEDIUM | HIGH
generationConfig.thinkingConfig ThinkingConfig no thinkingLevel (MINIMAL/LOW/MEDIUM/HIGH), thinkingBudget, includeThoughts
generationConfig.translationConfig TranslationConfig no targetLanguageCode, echoTargetLanguage
generationConfig.enableAffectiveDialog bool no 2.5 only
systemInstruction Content no text parts only
tools[] Tool[] no functionDeclarations[]{name, description, parameters, behavior BLOCKING|NON_BLOCKING}, googleSearch {}
realtimeInputConfig RealtimeInputConfig no automaticActivityDetection{disabled, startOfSpeechSensitivity, endOfSpeechSensitivity, prefixPaddingMs, silenceDurationMs}, activityHandling, turnCoverage
sessionResumption SessionResumptionConfig no handle (transparent SDK-only)
contextWindowCompression ContextWindowCompressionConfig no triggerTokens, slidingWindow.targetTokens
inputAudioTranscription / outputAudioTranscription AudioTranscriptionConfig no languageCodes[], customVocabulary[], mode VERBATIM|SMART, wordTimestamp, diarization
proactivity ProactivityConfig no proactiveAudio
historyConfig HistoryConfig no initialHistoryInClientContent
labels map no discovery-only, UNVERIFIED
json
{"setup": {"model": "models/gemini-3.8-live-extended-thinking",
  "generationConfig": {"responseModalities": ["AUDIO"],
    "speechConfig": {"voiceConfig": {"prebuiltVoiceConfig": {"voiceName": "Puck"}}},
    "thinkingConfig": {"thinkingLevel": "LOW"}},
  "tools": [{"functionDeclarations": [{"name": "searchFlights", "description": "Searches for flights between cities.",
    "behavior": "NON_BLOCKING", "parameters": {"type": "OBJECT", "properties": {"destination": {"type": "STRING"}}, "required": ["destination"]}}]}],
  "inputAudioTranscription": {}, "outputAudioTranscription": {}, "sessionResumption": {},
  "contextWindowCompression": {"triggerTokens": "25600", "slidingWindow": {"targetTokens": "12800"}}}}

Sent: once, immediately after the socket opens. Then wait for setupComplete.

# 2.2 clientContent — BidiGenerateContentClientContent

Field Type Req. Notes
turns[] Content[] no {role: "user"|"model", parts: [{text}, {inlineData}, …]}; appended unconditionally to history; explicit roles supported all session long (3.1+)
turnComplete bool no true → generate now (unconditionally interrupts active generation); absent/false → server waits for more messages
json
{"clientContent": {"turns": [{"role": "user", "parts": [{"text": "Hello world!"}]}], "turnComplete": true}}

Sent: to inject text turns or restore history (also before realtime starts when historyConfig.initialHistoryInClientContent is true). A clientContent message interrupts any current generation.

# 2.3 realtimeInput — BidiGenerateContentRealtimeInput

Field Type Req. Notes
audio Blob {mimeType, data} one of audio/pcm;rate=16000, raw 16-bit LE PCM mono, base64
video Blob one of image/jpeg / image/png frame, ≤ 1 FPS
text string one of realtime text stream
activityStart {} one of manual VAD only (automaticActivityDetection.disabled: true)
activityEnd {} one of manual VAD only; server finalizes immediately
audioStreamEnd bool one of automatic VAD only; mic off / hybrid VAD end-of-speech; flushes cached audio
mediaResolution enum no per-message override
mediaChunks[] Blob[] deprecated only first chunk used
json
{"realtimeInput": {"audio": {"data": "UklGRiQAAABXQVZF...", "mimeType": "audio/pcm;rate=16000"}}}
{"realtimeInput": {"video": {"data": "<base64 JPEG>", "mimeType": "image/jpeg"}}}
{"realtimeInput": {"text": "Hello, how are you?"}}
{"realtimeInput": {"activityStart": {}}}
{"realtimeInput": {"activityEnd": {}}}
{"realtimeInput": {"audioStreamEnd": true}}

Sent: continuously; does not interrupt generation by itself (VAD does). End of turn is derived from activity. Always treated as user input (cannot populate history).

# 2.4 toolResponse — BidiGenerateContentToolResponse

Field Type Req. Notes
functionResponses[] FunctionResponse[] yes one per toolCall.functionCalls[]
functionResponses[].id string yes = FunctionCall.id
functionResponses[].name string yes = FunctionCall.name
functionResponses[].response object yes any JSON (output, result, error …)
functionResponses[].scheduling enum no INTERRUPT | WHEN_IDLE (default) | SILENT — NON_BLOCKING calls only; not on 3.8-extended-thinking
functionResponses[].willContinue bool no NON_BLOCKING only; generator semantics
functionResponses[].parts[] FunctionResponsePart[] no multimodal results
json
{"toolResponse": {"functionResponses": [{"id": "call_123", "name": "searchFlights",
  "response": {"output": {"flight": "DL 145", "price": "$145"}, "scheduling": "INTERRUPT"}}]}}

Sent: after executing a toolCall. BLOCKING calls: the model waits; NON_BLOCKING: conversation continues meanwhile.

# 3. Server messages

# 3.1 setupComplete — BidiGenerateContentSetupComplete

No fields (reference). SDK types add sessionId, voiceConsentSignature (UNVERIFIED).

json
{"setupComplete": {}}

Sent: once, in response to setup.

# 3.2 serverContent — BidiGenerateContentServerContent

Field Type Enum / notes When
modelTurn Content parts[]: inlineData {mimeType: "audio/pcm;rate=24000", data}, text, thought parts, executableCode… — several parts per message possible while generating
generationComplete bool absent in interrupted turns model done generating
turnComplete bool turn over (after playback if realtime playback assumed)
interactionStatus enum IN_PROGRESS | IDLE (| REQUIRES_ACTION deprecated) always with turnComplete; also seen with toolCall on extended thinking with turnComplete
interrupted bool stop and flush playback barge-in / clientContent
groundingMetadata GroundingMetadata webSearchQueries, groundingChunks, groundingSupports, searchEntryPoint Google Search used
inputTranscription Transcription {text, languageCode} finalized user speech inputAudioTranscription set
interimInputTranscription Transcription partial hypotheses, frequent transcribe-live
outputTranscription Transcription model speech transcript outputAudioTranscription set
urlContextMetadata UrlContextMetadata {urlMetadata[]} tool listed unsupported on Live —
waitingForInput bool model expects the user to continue —
speechState enum (deprecated) use VoiceActivity —
turnCompleteReason enum (SDK-only) MALFORMED_FUNCTION_CALL, RESPONSE_REJECTED, NEED_MORE_INPUT, PROHIBITED_INPUT_CONTENT, … UNVERIFIED —
json
{"serverContent": {"modelTurn": {"parts": [{"inlineData": {"mimeType": "audio/pcm;rate=24000", "data": "..."}}]}}}
{"serverContent": {"inputTranscription": {"text": "Hello, how are you?", "languageCode": "en"}}}
{"serverContent": {"outputTranscription": {"text": "I'm doing well, thanks!"}}}
{"serverContent": {"generationComplete": true}}
{"serverContent": {"turnComplete": true, "interactionStatus": "IDLE"}}
{"serverContent": {"interrupted": true}}

Extended-thinking filler (docs):

json
{"serverContent": {"modelTurn": {"parts": [{"inlineData": {"mimeType": "audio/pcm;rate=24000", "data": "..."}}]},
  "turnComplete": true, "interactionStatus": "IN_PROGRESS"}}

# 3.3 toolCall — BidiGenerateContentToolCall

Field Type Notes
functionCalls[] FunctionCall[] {id, name, args}; parallel calls possible
json
{"toolCall": {"functionCalls": [{"id": "call_123", "name": "searchFlights", "args": {"destination": "Seattle"}}]}}

Sent: when the model decides to call functions; on extended thinking it may be tagged "interactionStatus": "IN_PROGRESS" (docs example places it at top level next to toolCall).

# 3.4 toolCallCancellation — BidiGenerateContentToolCallCancellation

Field Type Notes
ids[] string[] calls that must not run / should be undone
json
{"toolCallCancellation": {"ids": ["call_123"]}}

Sent: only when the client interrupts a server turn (barge-in) while calls are pending. Constructed example.

# 3.5 usageMetadata — UsageMetadata (companion field)

Field Type Notes
promptTokenCount int32 includes cached content
cachedContentTokenCount int32
responseTokenCount int32 all candidates
toolUsePromptTokenCount int32
thoughtsTokenCount int32 thinking models
totalTokenCount int32 prompt + response (+ thoughts + tool-use)
promptTokensDetails[] ModalityTokenCount[] {modality: TEXT|IMAGE|VIDEO|AUDIO|DOCUMENT, tokenCount}
cacheTokensDetails[] ModalityTokenCount[]
responseTokensDetails[] ModalityTokenCount[]
toolUsePromptTokensDetails[] ModalityTokenCount[]
json
{"usageMetadata": {"promptTokenCount": 1250, "responseTokenCount": 420, "totalTokenCount": 1670,
  "promptTokensDetails": [{"modality": "AUDIO", "tokenCount": 1200}, {"modality": "TEXT", "tokenCount": 50}],
  "responseTokensDetails": [{"modality": "AUDIO", "tokenCount": 420}]}}

(Constructed values.) Sent: alongside other server messages, "periodically" (SDK sample loops on message.usage_metadata). The modality breakdown is how audio vs text vs image/video billing is split; audio ≈ 25 tokens/s. Since the whole context is re-billed each turn, promptTokenCount grows with the session unless contextWindowCompression caps it.

# 3.6 goAway — GoAway

Field Type Notes
timeLeft Duration string (e.g. "30s") time before the connection is terminated as ABORTED; never below a model-specific minimum (not published)
json
{"goAway": {"timeLeft": "30s"}}

(Constructed value.) Sent: before the server resets the connection (lifetime ≈ 10 min). Action: finish/queue work, then reconnect with sessionResumption.handle.

# 3.7 sessionResumptionUpdate — SessionResumptionUpdate

Field Type Notes
newHandle string resumable state handle; empty if resumable is false
resumable bool false while generating / executing function calls
lastConsumedClientMessageIndex int64 (SDK-only, transparent mode) index of last client message included in the state — UNVERIFIED
json
{"sessionResumptionUpdate": {"newHandle": "<opaque handle>", "resumable": true}}

Sent: periodically, only if setup.sessionResumption was present. Keep the latest handle with resumable: true; it stays valid 2 h after the last session termination.

# 3.8 SDK-only server fields (UNVERIFIED)

voiceActivity {voiceActivityType: ACTIVITY_START|ACTIVITY_END, audioOffset} and voiceActivityDetectionSignal {vadSignalType: VAD_SIGNAL_TYPE_SOS|EOS} ("allowlisted only", tied to setup.explicitVadSignal). Present in LiveServerMessage SDK types, absent from the public reference.

# 4. usageMetadata modality breakdown — reading it

Question Where
How much audio did I send this turn (incl. re-billed history)? promptTokensDetails[modality=AUDIO].tokenCount
Video/image frames cost `promptTokensDetails[modality=IMAGE
Spoken output cost responseTokensDetails[modality=AUDIO]
Transcripts / TEXT modality responseTokensDetails[modality=TEXT] (billed at text output rate)
Reasoning thoughtsTokenCount (included in output price)
Tool results fed back toolUsePromptTokenCount / toolUsePromptTokensDetails[]

# 5. goAway / sessionResumptionUpdate timing

text
t=0        setup → setupComplete ; sessionResumptionUpdate{newHandle=h1, resumable=true}
…          periodic sessionResumptionUpdate (resumable=false while the model speaks / runs tools)
t≈10 min   goAway{timeLeft} → client drains playback, opens a new WebSocket
t≈10 min   setup{…, sessionResumption:{handle:h_latest}} → setupComplete → continue (config may change, model may not)
  • Session limits without compression: 15 min audio-only / 2 min audio+video → enable contextWindowCompression for unlimited sessions.
  • Handles expire 2 h after the last session termination.
  • Ephemeral tokens: reconnecting with a handle does not consume a uses; it must happen before expireTime.

# Live verification (2026-09-18)

Observed 2026-09-18 on gemini-2.5-flash-native-audio-latest (AUDIO turn, tmp-live/gemini-tools/j1_*.json, j4_*.json, j5_ws_toolcall.json):

text
→ setup {model, generationConfig.responseModalities:[AUDIO], outputAudioTranscription:{}, sessionResumption:{}}
← setupComplete {}
← sessionResumptionUpdate {newHandle, resumable:true}
→ clientContent {turns:[{role:user, parts:[{text}]}], turnComplete:true}
← serverContent {modelTurn:{parts:[{text:"…", thought:true}]}}
← toolCall {functionCalls:[{name, args, id:"function-call-…"}]}            (only when a NON_BLOCKING declaration is set)
← serverContent {outputTranscription:{text:"OK"}}
← serverContent {modelTurn:{parts:[{inlineData:{mimeType:"audio/pcm;rate=24000", data}}]}}   × N
← serverContent {generationComplete:true}
← serverContent {turnComplete:true} + usageMetadata {promptTokenCount, responseTokenCount, totalTokenCount, promptTokensDetails, responseTokensDetails:[{modality:AUDIO}], thoughtsTokenCount}
← sessionResumptionUpdate {newHandle, resumable:true}

Rejections: responseModalities:[TEXT] → close 1007 on every current live model; setup.toolConfig → close 1007; ephemeral token on the non-Constrained method → close 1008. Not observed: interrupted, goAway, toolCallCancellation, inputTranscription, groundingMetadata, voiceActivity*.