Gemini Live API — bidirectional WebSocket (BidiGenerateContent) + ephemeral tokens
Status: DOCUMENTED + LIVE_VERIFIED (2026-09-18 run — see "Live verification" section at the end) · PREVIEW (Google labels the Live API and ephemeral tokens "Preview"; gemini-3.8-live / gemini-3.8-live-extended-thinking are "Stable"/GA since 2026-09-15)
Sources: https://ai.google.dev/api/live (WebSocket reference) · https://ai.google.dev/gemini-api/docs/live-api · …/live-api/get-started-sdk · …/live-api/get-started-websocket · …/live-api/capabilities · …/live-api/session-management · …/live-api/tools · …/live-api/thinking · …/live-api/ephemeral-tokens · …/live-api/live-transcribe · …/live-api/live-translate · …/live-api/best-practices · https://ai.google.dev/gemini-api/docs/models/{gemini-3.8-live, gemini-3.8-live-extended-thinking, gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-preview-12-2025, gemini-3.5-live-translate-preview, gemini-3.5-transcribe} · https://ai.google.dev/gemini-api/docs/pricing · https://ai.google.dev/gemini-api/docs/changelog · discovery document v1beta rev. 20260918 (sources/gemini/discovery-v1beta.json) · SDK types sources/gemini/openapi/python-genai-types.py · live model list sources/gemini/models-api-raw.json
Last verified: 2026-09-18
Machine-readable twins: generated/fragments/streaming-events/gemini-live.json (32 events), generated/fragments/parameters/gemini-live.json (115 params), tmp/gemini-parts/live-endpoints.json (4 endpoints), tmp/gemini-parts/live-objects.json (39 objects). Full per-message field tables: live-events.md.
1. Overview and architecture
The Live API is a stateful WebSocket (WSS) session with a Gemini model: continuous audio / video / text in, native audio (or text, for the transcribe model) out, with barge-in, server-side VAD, function calling, Google Search grounding, transcription, session resumption and context compression. All traffic is JSON text frames; media are base64 Blobs.
Two integration approaches (docs): server-to-server (your backend holds the API key and proxies the stream) and client-to-server (browser/mobile connects directly — lower latency; use ephemeral tokens, never an API key). Partner stacks: Pipecat, LiveKit, Fishjam, ADK, Vision Agents, Voximplant.
sequenceDiagram
participant C as Client
participant S as Gemini Live (WSS v1beta)
C->>S: {"setup": {model, generationConfig, tools, realtimeInputConfig, ...}}
S-->>C: {"setupComplete": {}}
loop conversation
C->>S: {"realtimeInput": {audio|video|text|activityStart|activityEnd|audioStreamEnd}}
C->>S: {"clientContent": {turns[], turnComplete}}
S-->>C: {"serverContent": {inputTranscription | interimInputTranscription}}
S-->>C: {"serverContent": {modelTurn: {parts:[inlineData audio/pcm;rate=24000]}}}
S-->>C: {"serverContent": {outputTranscription}}
S-->>C: {"toolCall": {functionCalls[{id,name,args}]}}
C->>S: {"toolResponse": {functionResponses[{id,name,response,scheduling}]}}
S-->>C: {"serverContent": {interrupted: true}} / {"toolCallCancellation": {ids[]}}
S-->>C: {"serverContent": {generationComplete: true}}
S-->>C: {"serverContent": {turnComplete: true, interactionStatus: IDLE}, "usageMetadata": {...}}
S-->>C: {"sessionResumptionUpdate": {newHandle, resumable}}
end
S-->>C: {"goAway": {timeLeft: "30s"}}
C->>S: reconnect: {"setup": {..., sessionResumption: {handle}}}
Wire rules (reference): a client message has exactly one of setup | clientContent | realtimeInput | toolResponse; a server message may carry usageMetadata and otherwise exactly one of setupComplete | serverContent | toolCall | toolCallCancellation | goAway | sessionResumptionUpdate (the proto messageType union is flattened to the top level). Setup is sent once; configuration cannot change mid-connection (only on resume, model excluded).
2. Models (bidiGenerateContent on this key)
All current Live models are native audio (audio-to-audio) models. The former half-cascade models (gemini-2.0-flash-live-001, gemini-live-2.5-flash-preview) were shut down 2025-12-09 (changelog); no current page mentions "half-cascade".
| Model id | Type / role | Input → output | Thinking | Tools | Async FC | Proactive / affective | Context / output tokens | Session limits | Status (docs) |
|---|---|---|---|---|---|---|---|---|---|
gemini-3.8-live |
native audio, default voice agent | text, image, audio, video → audio (+ transcript) | interleaved; thinkingLevel not supported (omit) |
function calling, Google Search | default NON_BLOCKING; BLOCKING allowed; scheduling supported |
proactive audio always on (false → error); affective dialog removed |
131,072 in / 65,536 out | 15 min audio / 2 min audio+video w/o compression; 128k context | Stable (GA 2026-09-15), API "Preview" |
gemini-3.8-live-extended-thinking |
native audio, background reasoning | same | thinkingLevel LOW/MEDIUM/HIGH (no MINIMAL); includeThoughts |
function calling (async only), Google Search | NON_BLOCKING only (BLOCKING = hard error); no scheduling |
proactive always on; affective removed | 131,072 / 65,536 | same; watch interactionStatus |
Stable (GA 2026-09-15) |
gemini-3.1-flash-live-preview |
native audio, legacy preview | text, image, audio, video → audio | thinkingLevel MINIMAL (default)/LOW/MEDIUM/HIGH |
function calling (sync only), Google Search | not supported | neither supported | 131,072 / 65,536 | same | Preview, "legacy — migrate to 3.8" |
gemini-2.5-flash-native-audio-preview-12-2025 (+ -preview-09-2025, -native-audio-latest) |
native audio (2.5) | audio, video, text → audio and text | thinkingBudget (thinking: true) |
function calling (sync + async), Google Search | supported | both opt-in (v1beta) | 131,072 / 8,192 | 15 min / 2 min; 128k | Preview (2.5 family "no longer available to new users" on this key per CLAUDE.md — verify) |
gemini-3.5-live-translate-preview |
speech-to-speech translation, 70+ languages | audio only → translated audio + transcripts | none | none | — | — | 16,384 in / 32,768 out (models API) | UNVERIFIED (not stated) | Preview (June 2026) |
gemini-3.5-transcribe-live |
streaming speech-to-text | 16-bit PCM audio → TEXT (inputTranscription / interimInputTranscription) |
none | none | — | — | 131,072 / 65,536 | 10 min per session | GA-ish (Aug 2026), version 3.5-transcribe-live-08-2026 |
gemini-robotics-er-2-streaming-preview |
text streaming for robots (out of scope) | audio, video → text | — | — | — | — | 131,072 / 65,536 | — | Preview |
Native-audio limitation: only the AUDIO response modality; get text via outputAudioTranscription. Any Gemini TTS voice can be used; languages are chosen automatically (99 listed).
3. Connection and authentication
| Item | Value |
|---|---|
| URL (API key) | wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent?key=YOUR_API_KEY |
| URL (ephemeral token) | wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContentConstrained?access_token=<token> — or header Authorization: Token <token> (reference page) |
| Version | v1beta only (reference note; ephemeral tokens "only with v1beta"). One snippet in the thinking guide shows a v1alpha URL — treat as legacy/UNVERIFIED. |
| First frame | {"setup": …}; wait for {"setupComplete": {}} before anything else |
| Python SDK | async with client.aio.live.connect(model="gemini-3.8-live", config=types.LiveConnectConfig(...)) as session: then session.send_realtime_input(...), session.send_client_content(...), session.send_tool_response(...), async for msg in session.receive() |
| Node SDK | const session = await ai.live.connect({model, config, callbacks:{onopen,onmessage,onerror,onclose}}); session.sendRealtimeInput({...}), sendClientContent, sendToolResponse, session.close() |
| Ephemeral token via SDK | new GoogleGenAI({ apiKey: token.name }) / genai.Client(api_key=token.name, http_options={"api_version":"v1beta"}) — the SDK switches to the constrained endpoint for you |
Security note (project rule): keys only in .env; the docs' ?key= query form is server-side only. The parent agent does live calls.
4. setup configuration — every field
Wire names (camelCase) = REST JSON; SDK LiveConnectConfig uses snake_case (Python) / camelCase (Node) and flattens generationConfig fields (e.g. response_modalities, speech_config, thinking_config, media_resolution, enable_affective_dialog, translation_config).
| Field | Type | Required | Default | Notes / models |
|---|---|---|---|---|
setup.model |
string | yes | — | models/gemini-3.8-live |
setup.generationConfig |
GenerationConfig subset | no | — | Not supported: responseLogprobs, responseMimeType, logprobs, responseSchema, responseJsonSchema, stopSequence, skipResponseCache, routingConfig, audioTimestamp |
…responseModalities[] |
TEXT | AUDIO |
no | SDK: AUDIO | Native audio models: AUDIO only; transcribe-live: TEXT |
…speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName |
string | no | model default | 30 TTS voices (Zephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirrhoe, Autonoe, Enceladus, Iapetus, Umbriel, Algieba, Despina, Erinome, Algenib, Rasalgethi, Laomedeia, Achernar, Alnilam, Schedar, Gacrux, Pulcherrima, Achird, Zubenelgenubi, Vindemiatrix, Sadachbia, Sadaltager, Sulafat) |
…speechConfig.languageCode |
BCP-47 | no | auto | 30 codes listed in discovery; native audio models ignore/unsupported (steer via system instructions) — UNVERIFIED effect |
…speechConfig.multiSpeakerVoiceConfig |
object | no | — | TTS feature; Live support undocumented (UNVERIFIED) |
…mediaResolution |
MEDIA_RESOLUTION_LOW (64 tok) | MEDIUM (256) | HIGH (zoomed 256) |
no | unspecified | Visual frames only; audio tokenization fixed |
…thinkingConfig.thinkingLevel |
MINIMAL | LOW | MEDIUM | HIGH |
no | 3.1: MINIMAL | 3.1: all four; 3.8-extended-thinking: LOW/MEDIUM/HIGH; 3.8-live: unsupported |
…thinkingConfig.thinkingBudget |
int | no | — | Gemini 2.5 native audio (legacy) |
…thinkingConfig.includeThoughts |
bool | no | false | thought summaries as parts |
…enableAffectiveDialog |
bool | no | false | Gemini 2.5 only (v1beta); 3.1 unsupported; 3.8 removed |
…translationConfig.targetLanguageCode / .echoTargetLanguage |
string / bool | translate: yes / no | "en" / false | gemini-3.5-live-translate-preview only |
…temperature, topP, topK, maxOutputTokens, candidateCount, presencePenalty, frequencyPenalty, seed |
number/int | no | — | listed in the reference example |
setup.systemInstruction |
Content (text parts) | no | — | each part = paragraph; specify the language here |
setup.tools[] |
Tool[] | no | — | functionDeclarations[] (with Live-only behavior), googleSearch: {}; codeExecution / urlContext / googleMaps not supported |
setup.realtimeInputConfig.automaticActivityDetection.disabled |
bool | no | false | true = manual VAD (activityStart/End) |
….startOfSpeechSensitivity |
START_SENSITIVITY_HIGH | LOW |
no | HIGH | |
….endOfSpeechSensitivity |
END_SENSITIVITY_HIGH | LOW |
no | HIGH | |
….prefixPaddingMs |
int32 | no | — | look-back before speech; 0 clips onsets |
….silenceDurationMs |
int32 | no | ≈800 (server) | recommended 500–800 |
setup.realtimeInputConfig.activityHandling |
START_OF_ACTIVITY_INTERRUPTS | NO_INTERRUPTION |
no | interrupts | barge-in |
setup.realtimeInputConfig.turnCoverage |
TURN_INCLUDES_ONLY_ACTIVITY | TURN_INCLUDES_ALL_INPUT | TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO |
no | 2.5: ONLY_ACTIVITY; 3.1+: AUDIO_ACTIVITY_AND_ALL_VIDEO | all video frames billed on 3.1+ |
setup.inputAudioTranscription |
AudioTranscriptionConfig | no | off | {} enables; languageCodes[] (empty = auto), customVocabulary[] (≤1000), mode VERBATIM|SMART; wordTimestamp/diarization not supported over Live |
setup.outputAudioTranscription |
AudioTranscriptionConfig | no | off | {} enables transcript of model audio (billed as text output) |
setup.sessionResumption.handle |
string | no | new session | {} = new resumable session; handle valid 2 h; (transparent SDK-only) |
setup.contextWindowCompression.triggerTokens |
int64 | no | 80 % of context | |
setup.contextWindowCompression.slidingWindow.targetTokens |
int64 | no | triggerTokens/2 | |
setup.proactivity.proactiveAudio |
bool | no | 2.5: false | 3.8: always on (false → error); 3.1: unsupported |
setup.historyConfig.initialHistoryInClientContent |
bool | no | false | ingest clientContent history before realtime starts |
setup.labels |
map | no | — | discovery-only (safety_identifier) — UNVERIFIED |
setup.explicitVadSignal, setup.safetySettings[], setup.avatarConfig |
— | no | — | SDK types only — UNVERIFIED |
5. Client messages (summary)
| Key | Type | Purpose | SDK |
|---|---|---|---|
setup |
BidiGenerateContentSetup | first message, session config | live.connect(model, config) |
clientContent |
{turns[]: Content, turnComplete} |
append history / text turn; interrupts generation; turnComplete:true starts generation |
send_client_content / sendClientContent |
realtimeInput |
{audio | video | text | activityStart | activityEnd | audioStreamEnd | mediaResolution | mediaChunks(deprecated)} |
continuous streams; turn end from VAD | send_realtime_input / sendRealtimeInput |
toolResponse |
{functionResponses[]: {id, name, response, scheduling, willContinue, parts}} |
answer a toolCall |
send_tool_response / sendToolResponse |
6. Server messages (summary)
| Key | Content | When |
|---|---|---|
setupComplete |
{} |
once, after setup |
serverContent |
modelTurn (parts: inlineData audio/pcm;rate=24000, text, thoughts), generationComplete, turnComplete (+ interactionStatus), interrupted, groundingMetadata, inputTranscription, interimInputTranscription, outputTranscription, urlContextMetadata, waitingForInput, speechState (deprecated) |
throughout a turn; several fields/parts per message possible |
toolCall |
functionCalls[] {id, name, args} |
when the model wants a function; may arrive during audio (async) |
toolCallCancellation |
ids[] |
after a barge-in discards pending calls |
usageMetadata |
token counts + per-modality details | alongside other messages, periodically |
goAway |
timeLeft (Duration) |
before the ~10-min connection reset |
sessionResumptionUpdate |
newHandle, resumable |
periodically when sessionResumption is set |
voiceActivity, voiceActivityDetectionSignal |
SDK types only ("allowlisted") | UNVERIFIED |
Full field tables and examples: live-events.md.
7. Audio, video and text formats
| Direction | Format |
|---|---|
| Input audio | raw 16-bit PCM, little-endian, mono, 16 kHz native; Blob.mimeType = "audio/pcm;rate=16000" (other rates are resampled if declared); base64 in realtimeInput.audio.data; chunks 20–40 ms recommended (best practices), 20–100 ms acceptable, 100 ms in the transcribe guide; resample 44.1/48 kHz mics to 16 kHz client-side |
| Output audio | raw 16-bit PCM 24 kHz little-endian; serverContent.modelTurn.parts[].inlineData with mimeType: "audio/pcm;rate=24000" (write WAV with 1 channel, 2 bytes, 24000 Hz); generated faster than realtime — buffer and play |
| Video | individual frames as image/jpeg (or image/png), ≤ 1 frame/s, realtimeInput.video; mediaResolution controls tokens per frame |
| Text | realtimeInput.text (stream) or clientContent.turns[].parts[].text |
| Token rate | audio ≈ 25 tokens per second (both directions); transcribe output ≈ 175 text tokens/min |
8. VAD, activity handling, turn coverage, interruptions
- Automatic VAD (default) on the continuous audio stream. Tune
startOfSpeechSensitivity,endOfSpeechSensitivity,prefixPaddingMs,silenceDurationMs(500–800 ms recommended; 100–200 ms fragments utterances; 2000+ ms adds latency). When the mic pauses > 1 s sendrealtimeInput.audioStreamEnd: trueto flush; resume by sending audio. - Hybrid VAD: keep server VAD (robust start detection with prefix padding) and let a client VAD send
audioStreamEndas an immediate end-of-speech → minimal finalization latency; server VAD is the fallback. - Manual VAD:
automaticActivityDetection.disabled: true; sendactivityStart→ audio →activityEnd. No pre-speech buffer and no silence tolerance server-side — use ≥ 500 ms end-of-speech threshold client-side.audioStreamEndis not used. - Interruptions: with
START_OF_ACTIVITY_INTERRUPTS(default) user speech cancels generation; only content already sent stays in history; the server sendsserverContent.interrupted: true(stop playback, flush queue) andtoolCallCancellationfor discarded calls.NO_INTERRUPTIONdisables barge-in.clientContent.turnComplete: truealso interrupts unconditionally (3.1+). - Turn coverage:
TURN_INCLUDES_ONLY_ACTIVITY(speech only, 2.5 default),TURN_INCLUDES_ALL_INPUT(incl. silence),TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO(3.1+ default; every video frame counts — send frames only when needed).
9. Transcription
setup.inputAudioTranscription: {}→serverContent.inputTranscription {text, languageCode}(finalized) and, ongemini-3.5-transcribe-live,serverContent.interimInputTranscription(fast partial hypotheses).setup.outputAudioTranscription: {}→serverContent.outputTranscriptionfor the model's speech (language inferred). The last output transcription of a turn precedesgenerationComplete/interrupted; ordering vs. audio parts is approximate.- Transcribe-live specifics:
languageCodes[](empty = auto, code-switching),customVocabulary[](≤ 1,000, best ≤ 100),mode: "SMART"(filler removal, formatting; no word annotations). Not supported over Live: word timestamps, diarization. Sessions ≤ 10 min. Response modality["TEXT"]. - Billing: transcript text tokens are charged at text output rate on top of audio tokens.
10. Session management
| Limit | Documented value |
|---|---|
| Connection lifetime | ≈ 10 minutes; the server sends goAway.timeLeft first, then terminates as ABORTED |
| Session (audio-only) | 15 minutes without context window compression |
| Session (audio + video) | 2 minutes without compression |
With contextWindowCompression |
unlimited duration |
| Context window | 128k tokens native audio models, 32k other Live models |
| Resumption handle validity | 2 hours after the last session termination (changelog 2025-04-09 mentioned 24 h server-side state — current docs say 2 h) |
| Transcribe-live session | 10 minutes |
| Concurrent sessions | not documented on the rate-limits page (only "Concurrent batch requests: 100" for Batch) |
- Context window compression:
contextWindowCompression: {triggerTokens, slidingWindow: {targetTokens}}— sliding window discards from the start (always at a USER turn; system instructions kept). Defaults: trigger = 80 % of context, target = trigger/2. Best-practice example: trigger 25,000 / target 8,000 also caps compounding cost. - Session resumption: set
sessionResumption: {}; store the latestsessionResumptionUpdate.newHandlewhenresumable: true(false while generating / running functions); on reconnect sendsessionResumption: {handle}. Configuration exceptmodelmay change on resume. Ephemeral-token flows need this to reconnect every ~10 min withinexpireTime(same token even withuses: 1). - GoAway: use
timeLeftto wrap up or reconnect. generationComplete: model finished generating (UI hook).
11. Tools in Live
| Tool | 3.8 Live | 3.8 Live Extended Thinking | 3.1 Flash Live | 2.5 native audio | Translate / Transcribe |
|---|---|---|---|---|---|
| Function calling | yes — default NON_BLOCKING, BLOCKING allowed, scheduling supported |
async NON_BLOCKING only (BLOCKING = hard error), no scheduling |
synchronous only (model waits for toolResponse) |
sync + async | no |
Google Search (googleSearch: {}) |
yes | yes | yes | yes | no |
| Code execution, URL context, Google Maps | not supported | not supported | not supported | not supported | no |
- Declare functions in
setup.tools[].functionDeclarations[]; add"behavior": "NON_BLOCKING"for async. The server sendstoolCall.functionCalls[](always withid); reply withtoolResponse.functionResponses[](id,name,response). No automatic tool handling in Live. - Async
FunctionResponse.scheduling:INTERRUPT(speak result now),WHEN_IDLE(default, after current output),SILENT(context only).willContinue: trueturns a call into a generator (more responses later); end withwillContinue: false(+SILENTto avoid triggering speech). - Parallel calls:
toolCall.functionCalls[]can contain several calls; on extended thinking the model speaks fillers while calls run (interactionStatus: IN_PROGRESS). - Barge-in during a tool turn →
toolCallCancellation.ids[](undo side effects if possible). - Grounding results in
serverContent.groundingMetadata; 3.x: 5,000 free searches/month then $14 / 1,000.
12. Thinking in Live
gemini-3.1-flash-live-preview:thinkingConfig.thinkingLevelMINIMAL (default) / LOW / MEDIUM / HIGH (replaces 2.5thinkingBudget).gemini-3.8-live: interleaved reasoning, fixed latency profile; omitthinkingConfig.gemini-3.8-live-extended-thinking: background reasoning,thinkingLevelLOW/MEDIUM/HIGH,includeThoughtsfor summaries. Lifecycle changes: the model speaks conversational fillers withturnComplete: true+interactionStatus: "IN_PROGRESS", emits asynctoolCalls (also tagged IN_PROGRESS), then the final answer withinteractionStatus: "IDLE".turnCompletealone no longer means idle — keep listening until IDLE.
13. Proactive audio and affective dialog
proactivity.proactiveAudio: truelets the model ignore irrelevant speech / not answer when no request was made. Gemini 2.5: opt-in (v1beta). Gemini 3.8 (both): permanently enabled,falsereturns an error; billing: input tokens charged the whole time the API listens, output only when it answers. Gemini 3.1: unsupported (bills only while you stream).generationConfig.enableAffectiveDialog: true(adapt tone to the user): Gemini 2.5 only; unsupported on 3.1; removed on 3.8.
14. Voices and languages
- Voices: any Gemini TTS voice via
speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName(30 names, see §4). ThegenerateContentTTS voice set differs slightly. - Languages: 99 supported (BCP-47 list in the capabilities guide, e.g.
en,fr,zh-Hans,pt-BR); native audio models switch languages naturally and do not supportspeechConfig.languageCode; constrain via system instructions.
15. Live translate and Live transcribe
- Translate (
gemini-3.5-live-translate-preview): audio-only input,generationConfig.translationConfig {targetLanguageCode (default "en"), echoTargetLanguage (default false)},responseModalities: ["AUDIO"], optionalinputAudioTranscription/outputAudioTranscription(transcripts carrylanguageCode). No tools, no system instructions, continuous (non-turn) processing; 70+ languages. Limitations: voice replication may drift; language detection struggles with accents/similar languages; background audio may leak with echo on. With ephemeral tokens locktranslationConfigserver-side, or omit it and setlock_additional_fields: []to let the client choose. - Transcribe (
gemini-3.5-transcribe-live):responseModalities: ["TEXT"],inputAudioTranscription {languageCodes, customVocabulary, mode}, interim + final transcripts, auto/hybrid/manual VAD, 10-min sessions, 85+ languages, no diarization / word timestamps over Live. Pricing ≈ $0.009/min blended.
16. Ephemeral tokens (POST /v1beta/auth_tokens)
Flow: client authenticates with your backend → backend calls POST https://generativelanguage.googleapis.com/v1beta/auth_tokens (header x-goog-api-key) → response AuthToken.name is the token → client opens the WebSocket with it (BidiGenerateContentConstrained?access_token=<token> or Authorization: Token <token>; SDK: apiKey: token.name). Live API only, v1beta only, Preview.
Field (REST body = AuthToken) |
Type | Default / constraint |
|---|---|---|
uses |
int32 | 1; 0 = unlimited; resuming a session does not count |
expireTime |
RFC 3339 | now + 30 min; < 20 h; after it, session messages are rejected |
newSessionExpireTime |
RFC 3339 | now + 60 s; < 20 h; after it, new sessions are rejected |
fieldMask |
FieldMask | empty + no setup → client's setup used; empty + setup → token's setup only; non-empty → listed fields overwrite the client's |
bidiGenerateContentSetup |
BidiGenerateContentSetup | locked config (discovery / reference name) |
liveConnectConstraints {model, config} |
object | shape used by the official curl examples and SDKs (live_connect_constraints); not in the discovery schema — REST acceptance UNVERIFIED |
name |
string | output only: the token |
SDK-only: lock_additional_fields[], http_options |
— | extra fields to lock (→ fieldMask); [] unlocks e.g. translationConfig |
Best practices: short expireTime, re-provision on expiry, secure your backend auth, do not use ephemeral tokens for backend-to-Gemini paths. Ephemeral tokens pair with sessionResumption for reconnects every ~10 min.
17. Pricing (paid tier, USD per 1M tokens; free tier: free)
| Model | Text in | Audio in | Image/video in | Text out | Audio out | Notes |
|---|---|---|---|---|---|---|
gemini-3.8-live, gemini-3.8-live-extended-thinking, gemini-3.1-flash-live-preview |
$0.75 | $3.00 (≈ $0.005/min) | $1.00 (≈ $0.002/min) | $4.50 (incl. thinking) | $12.00 (≈ $0.018/min) | Search: 5,000 free/month shared across 3.x, then $14/1,000 |
gemini-2.5-flash-native-audio-preview-12-2025 |
$0.50 | $3.00 | $3.00 | $2.00 | $12.00 | |
gemini-3.5-live-translate-preview |
— | $3.50 (≈ $0.0053/min) | — | — | $21.00 (≈ $0.0315/min) | ≈ $0.0368/min effective, 25 tokens/s |
gemini-3.5-transcribe-live |
— | $3.50 (≈ $0.005/min) | — | $21.00 (≈ $0.004/min) | — | ≈ $0.009/min blended; no Search |
Billing model: per turn, all tokens in the context window are re-billed (compounding); audio kept as audio tokens (≈ 25 tok/s); transcription adds text-output tokens; cap growth with contextWindowCompression; proactive audio bills input while listening.
18. Rate limits and concurrency
The rate-limits page documents RPM/TPM/RPD per model tier but no Live-specific concurrent-session limit; goAway.timeLeft minimum is "specified together with the rate limits for the model" but not published. Observed headers are not applicable (WebSocket). Anything else here is UNVERIFIED.
19. Best practices (docs)
Clear, structured system instructions (persona, task, constraints, tool guidance, language); precise tool definitions; send 20–40 ms audio chunks, don't buffer ~1 s; handle interrupted by flushing playback; resample to 16 kHz; enable compression + resumption; handle goAway and generationComplete; process all parts of each serverContent; on 3.1+ only send video frames when needed.
20. Snippets
Raw WebSocket, text-only turn (Python websockets), key read from the environment (never hard-coded):
import asyncio, json, os, websockets
URL = ("wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage."
"v1beta.GenerativeService.BidiGenerateContent?key=" + os.environ["GEMINI_API_KEY"])
async def main():
async with websockets.connect(URL) as ws:
await ws.send(json.dumps({"setup": {
"model": "models/gemini-3.8-live",
"generationConfig": {"responseModalities": ["AUDIO"]},
"outputAudioTranscription": {}}}))
assert "setupComplete" in json.loads(await ws.recv())
await ws.send(json.dumps({"clientContent": {
"turns": [{"role": "user", "parts": [{"text": "Reply with OK."}]}],
"turnComplete": True}}))
async for raw in ws:
msg = json.loads(raw)
sc = msg.get("serverContent", {})
if "outputTranscription" in sc:
print(sc["outputTranscription"]["text"])
if sc.get("turnComplete"):
print("usage:", msg.get("usageMetadata"))
break
asyncio.run(main())Python SDK:
from google import genai
from google.genai import types
client = genai.Client() # GEMINI_API_KEY from env
config = types.LiveConnectConfig(response_modalities=["AUDIO"], output_audio_transcription={},
session_resumption=types.SessionResumptionConfig())
async with client.aio.live.connect(model="gemini-3.8-live", config=config) as session:
await session.send_client_content(turns={"role": "user", "parts": [{"text": "Reply with OK."}]}, turn_complete=True)
async for msg in session.receive():
if msg.server_content and msg.server_content.output_transcription:
print(msg.server_content.output_transcription.text)
if msg.server_content and msg.server_content.turn_complete:
breakNode (@google/genai):
import { GoogleGenAI, Modality } from "@google/genai";
const ai = new GoogleGenAI({}); // GEMINI_API_KEY from env
const session = await ai.live.connect({
model: "gemini-3.8-live",
config: { responseModalities: [Modality.AUDIO], outputAudioTranscription: {} },
callbacks: { onmessage: (m) => { if (m.serverContent?.outputTranscription) console.log(m.serverContent.outputTranscription.text); } },
});
session.sendClientContent({ turns: [{ role: "user", parts: [{ text: "Reply with OK." }] }], turnComplete: true });Ephemeral token (curl, server side):
curl -X POST "https://generativelanguage.googleapis.com/v1beta/auth_tokens" \
-H "x-goog-api-key: ${GEMINI_API_KEY}" -H "Content-Type: application/json" \
-d '{"uses": 1, "expireTime": "2026-09-18T20:30:00Z",
"liveConnectConstraints": {"model": "models/gemini-3.8-live",
"config": {"sessionResumption": {}, "responseModalities": ["AUDIO"]}}}'Live verification (2026-09-18)
Run of 2026-09-18 (tmp-live/gemini-tools/j*.json, _summary.json; examples examples/gemini/live/* executed; tests/gemini/test_live.py cheap test passes, realtime tests gated by RUN_REALTIME_TESTS).
| Probe | Result |
|---|---|
responseModalities:["TEXT"] on gemini-2.5-flash-native-audio-latest, gemini-3.1-flash-live-preview, gemini-3.8-live |
all rejected: WebSocket close 1007 The requested combination of response modalities (TEXT) is not supported by the model (3.8-live first sent setupComplete + sessionResumptionUpdate, then closed) |
responseModalities:["AUDIO"] + outputAudioTranscription:{} on gemini-2.5-flash-native-audio-latest, clientContent "Reply with OK." |
setupComplete → serverContent.modelTurn parts [{text, thought:true}] → serverContent.outputTranscription{text:"OK"} → serverContent.modelTurn inlineData{mimeType:"audio/pcm;rate=24000"} × 0–6 → serverContent.generationComplete → serverContent.turnComplete with usageMetadata in the same message ({promptTokenCount:373, responseTokenCount:4–21, responseTokensDetails:[{modality:"AUDIO"}], thoughtsTokenCount}) |
sessionResumption:{} + contextWindowCompression:{slidingWindow:{}} |
sessionResumptionUpdate{newHandle:<uuid>, resumable:true} right after setupComplete and again after turnComplete |
setup.toolConfig |
close 1007 Unknown name "toolConfig" at 'setup' — tool mode cannot be forced in Live |
tools:[{functionDeclarations:[{…, behavior:"NON_BLOCKING"}]}] |
toolCall{functionCalls:[{name, args, id:"function-call-11967657623789764093"}]} after the thought part; audio continued (non-blocking) |
SDK client.aio.live.connect(model, config={"response_modalities":["AUDIO"], "output_audio_transcription":{}}) |
sequence server_content ×3 → server_content+usage_metadata, transcript "OK." |
@google/genai ai.live.connect({callbacks}) |
setupComplete → modelTurn.thought → outputTranscription → generationComplete → turnComplete |
POST /v1alpha/auth_tokens and /v1beta/auth_tokens {uses:1, expireTime, newSessionExpireTime} |
both 200 {name:"auth_tokens/…"}; with bidiGenerateContentSetup{model, generationConfig} lock → 200 |
ephemeral token on …v1alpha.GenerativeService.BidiGenerateContent (query or header) |
close 1008 Method doesn't allow unregistered callers |
ephemeral token on …v1alpha.GenerativeService.BidiGenerateContentConstrained |
works with Authorization: Token <name> header and with ?access_token=<name>; locked token works too |
Not exercised: realtime audio/video input, VAD tuning, interruptions, goAway timing (sessions lasted < 10 s), Live translate / transcribe models, half-cascade models (none remain).