SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
33.0 KB

# Gemini Live API — bidirectional WebSocket (BidiGenerateContent) + ephemeral tokens

Status: DOCUMENTED + LIVE_VERIFIED (2026-09-18 run — see "Live verification" section at the end) · PREVIEW (Google labels the Live API and ephemeral tokens "Preview"; gemini-3.8-live / gemini-3.8-live-extended-thinking are "Stable"/GA since 2026-09-15) Sources: https://ai.google.dev/api/live (WebSocket reference) · https://ai.google.dev/gemini-api/docs/live-api · …/live-api/get-started-sdk · …/live-api/get-started-websocket · …/live-api/capabilities · …/live-api/session-management · …/live-api/tools · …/live-api/thinking · …/live-api/ephemeral-tokens · …/live-api/live-transcribe · …/live-api/live-translate · …/live-api/best-practices · https://ai.google.dev/gemini-api/docs/models/{gemini-3.8-live, gemini-3.8-live-extended-thinking, gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-preview-12-2025, gemini-3.5-live-translate-preview, gemini-3.5-transcribe} · https://ai.google.dev/gemini-api/docs/pricing · https://ai.google.dev/gemini-api/docs/changelog · discovery document v1beta rev. 20260918 (sources/gemini/discovery-v1beta.json) · SDK types sources/gemini/openapi/python-genai-types.py · live model list sources/gemini/models-api-raw.json Last verified: 2026-09-18 Machine-readable twins: generated/fragments/streaming-events/gemini-live.json (32 events), generated/fragments/parameters/gemini-live.json (115 params), tmp/gemini-parts/live-endpoints.json (4 endpoints), tmp/gemini-parts/live-objects.json (39 objects). Full per-message field tables: live-events.md.

# 1. Overview and architecture

The Live API is a stateful WebSocket (WSS) session with a Gemini model: continuous audio / video / text in, native audio (or text, for the transcribe model) out, with barge-in, server-side VAD, function calling, Google Search grounding, transcription, session resumption and context compression. All traffic is JSON text frames; media are base64 Blobs.

Two integration approaches (docs): server-to-server (your backend holds the API key and proxies the stream) and client-to-server (browser/mobile connects directly — lower latency; use ephemeral tokens, never an API key). Partner stacks: Pipecat, LiveKit, Fishjam, ADK, Vision Agents, Voximplant.

sequenceDiagram
    participant C as Client
    participant S as Gemini Live (WSS v1beta)
    C->>S: {"setup": {model, generationConfig, tools, realtimeInputConfig, ...}}
    S-->>C: {"setupComplete": {}}
    loop conversation
        C->>S: {"realtimeInput": {audio|video|text|activityStart|activityEnd|audioStreamEnd}}
        C->>S: {"clientContent": {turns[], turnComplete}}
        S-->>C: {"serverContent": {inputTranscription | interimInputTranscription}}
        S-->>C: {"serverContent": {modelTurn: {parts:[inlineData audio/pcm;rate=24000]}}}
        S-->>C: {"serverContent": {outputTranscription}}
        S-->>C: {"toolCall": {functionCalls[{id,name,args}]}}
        C->>S: {"toolResponse": {functionResponses[{id,name,response,scheduling}]}}
        S-->>C: {"serverContent": {interrupted: true}} / {"toolCallCancellation": {ids[]}}
        S-->>C: {"serverContent": {generationComplete: true}}
        S-->>C: {"serverContent": {turnComplete: true, interactionStatus: IDLE}, "usageMetadata": {...}}
        S-->>C: {"sessionResumptionUpdate": {newHandle, resumable}}
    end
    S-->>C: {"goAway": {timeLeft: "30s"}}
    C->>S: reconnect: {"setup": {..., sessionResumption: {handle}}}

Wire rules (reference): a client message has exactly one of setup | clientContent | realtimeInput | toolResponse; a server message may carry usageMetadata and otherwise exactly one of setupComplete | serverContent | toolCall | toolCallCancellation | goAway | sessionResumptionUpdate (the proto messageType union is flattened to the top level). Setup is sent once; configuration cannot change mid-connection (only on resume, model excluded).

# 2. Models (bidiGenerateContent on this key)

All current Live models are native audio (audio-to-audio) models. The former half-cascade models (gemini-2.0-flash-live-001, gemini-live-2.5-flash-preview) were shut down 2025-12-09 (changelog); no current page mentions "half-cascade".

Model id Type / role Input → output Thinking Tools Async FC Proactive / affective Context / output tokens Session limits Status (docs)
gemini-3.8-live native audio, default voice agent text, image, audio, video → audio (+ transcript) interleaved; thinkingLevel not supported (omit) function calling, Google Search default NON_BLOCKING; BLOCKING allowed; scheduling supported proactive audio always on (false → error); affective dialog removed 131,072 in / 65,536 out 15 min audio / 2 min audio+video w/o compression; 128k context Stable (GA 2026-09-15), API "Preview"
gemini-3.8-live-extended-thinking native audio, background reasoning same thinkingLevel LOW/MEDIUM/HIGH (no MINIMAL); includeThoughts function calling (async only), Google Search NON_BLOCKING only (BLOCKING = hard error); no scheduling proactive always on; affective removed 131,072 / 65,536 same; watch interactionStatus Stable (GA 2026-09-15)
gemini-3.1-flash-live-preview native audio, legacy preview text, image, audio, video → audio thinkingLevel MINIMAL (default)/LOW/MEDIUM/HIGH function calling (sync only), Google Search not supported neither supported 131,072 / 65,536 same Preview, "legacy — migrate to 3.8"
gemini-2.5-flash-native-audio-preview-12-2025 (+ -preview-09-2025, -native-audio-latest) native audio (2.5) audio, video, text → audio and text thinkingBudget (thinking: true) function calling (sync + async), Google Search supported both opt-in (v1beta) 131,072 / 8,192 15 min / 2 min; 128k Preview (2.5 family "no longer available to new users" on this key per CLAUDE.md — verify)
gemini-3.5-live-translate-preview speech-to-speech translation, 70+ languages audio only → translated audio + transcripts none none — — 16,384 in / 32,768 out (models API) UNVERIFIED (not stated) Preview (June 2026)
gemini-3.5-transcribe-live streaming speech-to-text 16-bit PCM audio → TEXT (inputTranscription / interimInputTranscription) none none — — 131,072 / 65,536 10 min per session GA-ish (Aug 2026), version 3.5-transcribe-live-08-2026
gemini-robotics-er-2-streaming-preview text streaming for robots (out of scope) audio, video → text — — — — 131,072 / 65,536 — Preview

Native-audio limitation: only the AUDIO response modality; get text via outputAudioTranscription. Any Gemini TTS voice can be used; languages are chosen automatically (99 listed).

# 3. Connection and authentication

Item Value
URL (API key) wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent?key=YOUR_API_KEY
URL (ephemeral token) wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContentConstrained?access_token=<token> — or header Authorization: Token <token> (reference page)
Version v1beta only (reference note; ephemeral tokens "only with v1beta"). One snippet in the thinking guide shows a v1alpha URL — treat as legacy/UNVERIFIED.
First frame {"setup": …}; wait for {"setupComplete": {}} before anything else
Python SDK async with client.aio.live.connect(model="gemini-3.8-live", config=types.LiveConnectConfig(...)) as session: then session.send_realtime_input(...), session.send_client_content(...), session.send_tool_response(...), async for msg in session.receive()
Node SDK const session = await ai.live.connect({model, config, callbacks:{onopen,onmessage,onerror,onclose}}); session.sendRealtimeInput({...}), sendClientContent, sendToolResponse, session.close()
Ephemeral token via SDK new GoogleGenAI({ apiKey: token.name }) / genai.Client(api_key=token.name, http_options={"api_version":"v1beta"}) — the SDK switches to the constrained endpoint for you

Security note (project rule): keys only in .env; the docs' ?key= query form is server-side only. The parent agent does live calls.

# 4. setup configuration — every field

Wire names (camelCase) = REST JSON; SDK LiveConnectConfig uses snake_case (Python) / camelCase (Node) and flattens generationConfig fields (e.g. response_modalities, speech_config, thinking_config, media_resolution, enable_affective_dialog, translation_config).

Field Type Required Default Notes / models
setup.model string yes — models/gemini-3.8-live
setup.generationConfig GenerationConfig subset no — Not supported: responseLogprobs, responseMimeType, logprobs, responseSchema, responseJsonSchema, stopSequence, skipResponseCache, routingConfig, audioTimestamp
…responseModalities[] TEXT | AUDIO no SDK: AUDIO Native audio models: AUDIO only; transcribe-live: TEXT
…speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName string no model default 30 TTS voices (Zephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirrhoe, Autonoe, Enceladus, Iapetus, Umbriel, Algieba, Despina, Erinome, Algenib, Rasalgethi, Laomedeia, Achernar, Alnilam, Schedar, Gacrux, Pulcherrima, Achird, Zubenelgenubi, Vindemiatrix, Sadachbia, Sadaltager, Sulafat)
…speechConfig.languageCode BCP-47 no auto 30 codes listed in discovery; native audio models ignore/unsupported (steer via system instructions) — UNVERIFIED effect
…speechConfig.multiSpeakerVoiceConfig object no — TTS feature; Live support undocumented (UNVERIFIED)
…mediaResolution MEDIA_RESOLUTION_LOW (64 tok) | MEDIUM (256) | HIGH (zoomed 256) no unspecified Visual frames only; audio tokenization fixed
…thinkingConfig.thinkingLevel MINIMAL | LOW | MEDIUM | HIGH no 3.1: MINIMAL 3.1: all four; 3.8-extended-thinking: LOW/MEDIUM/HIGH; 3.8-live: unsupported
…thinkingConfig.thinkingBudget int no — Gemini 2.5 native audio (legacy)
…thinkingConfig.includeThoughts bool no false thought summaries as parts
…enableAffectiveDialog bool no false Gemini 2.5 only (v1beta); 3.1 unsupported; 3.8 removed
…translationConfig.targetLanguageCode / .echoTargetLanguage string / bool translate: yes / no "en" / false gemini-3.5-live-translate-preview only
…temperature, topP, topK, maxOutputTokens, candidateCount, presencePenalty, frequencyPenalty, seed number/int no — listed in the reference example
setup.systemInstruction Content (text parts) no — each part = paragraph; specify the language here
setup.tools[] Tool[] no — functionDeclarations[] (with Live-only behavior), googleSearch: {}; codeExecution / urlContext / googleMaps not supported
setup.realtimeInputConfig.automaticActivityDetection.disabled bool no false true = manual VAD (activityStart/End)
….startOfSpeechSensitivity START_SENSITIVITY_HIGH | LOW no HIGH
….endOfSpeechSensitivity END_SENSITIVITY_HIGH | LOW no HIGH
….prefixPaddingMs int32 no — look-back before speech; 0 clips onsets
….silenceDurationMs int32 no ≈800 (server) recommended 500–800
setup.realtimeInputConfig.activityHandling START_OF_ACTIVITY_INTERRUPTS | NO_INTERRUPTION no interrupts barge-in
setup.realtimeInputConfig.turnCoverage TURN_INCLUDES_ONLY_ACTIVITY | TURN_INCLUDES_ALL_INPUT | TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO no 2.5: ONLY_ACTIVITY; 3.1+: AUDIO_ACTIVITY_AND_ALL_VIDEO all video frames billed on 3.1+
setup.inputAudioTranscription AudioTranscriptionConfig no off {} enables; languageCodes[] (empty = auto), customVocabulary[] (≤1000), mode VERBATIM|SMART; wordTimestamp/diarization not supported over Live
setup.outputAudioTranscription AudioTranscriptionConfig no off {} enables transcript of model audio (billed as text output)
setup.sessionResumption.handle string no new session {} = new resumable session; handle valid 2 h; (transparent SDK-only)
setup.contextWindowCompression.triggerTokens int64 no 80 % of context
setup.contextWindowCompression.slidingWindow.targetTokens int64 no triggerTokens/2
setup.proactivity.proactiveAudio bool no 2.5: false 3.8: always on (false → error); 3.1: unsupported
setup.historyConfig.initialHistoryInClientContent bool no false ingest clientContent history before realtime starts
setup.labels map no — discovery-only (safety_identifier) — UNVERIFIED
setup.explicitVadSignal, setup.safetySettings[], setup.avatarConfig — no — SDK types only — UNVERIFIED

# 5. Client messages (summary)

Key Type Purpose SDK
setup BidiGenerateContentSetup first message, session config live.connect(model, config)
clientContent {turns[]: Content, turnComplete} append history / text turn; interrupts generation; turnComplete:true starts generation send_client_content / sendClientContent
realtimeInput {audio | video | text | activityStart | activityEnd | audioStreamEnd | mediaResolution | mediaChunks(deprecated)} continuous streams; turn end from VAD send_realtime_input / sendRealtimeInput
toolResponse {functionResponses[]: {id, name, response, scheduling, willContinue, parts}} answer a toolCall send_tool_response / sendToolResponse

# 6. Server messages (summary)

Key Content When
setupComplete {} once, after setup
serverContent modelTurn (parts: inlineData audio/pcm;rate=24000, text, thoughts), generationComplete, turnComplete (+ interactionStatus), interrupted, groundingMetadata, inputTranscription, interimInputTranscription, outputTranscription, urlContextMetadata, waitingForInput, speechState (deprecated) throughout a turn; several fields/parts per message possible
toolCall functionCalls[] {id, name, args} when the model wants a function; may arrive during audio (async)
toolCallCancellation ids[] after a barge-in discards pending calls
usageMetadata token counts + per-modality details alongside other messages, periodically
goAway timeLeft (Duration) before the ~10-min connection reset
sessionResumptionUpdate newHandle, resumable periodically when sessionResumption is set
voiceActivity, voiceActivityDetectionSignal SDK types only ("allowlisted") UNVERIFIED

Full field tables and examples: live-events.md.

# 7. Audio, video and text formats

Direction Format
Input audio raw 16-bit PCM, little-endian, mono, 16 kHz native; Blob.mimeType = "audio/pcm;rate=16000" (other rates are resampled if declared); base64 in realtimeInput.audio.data; chunks 20–40 ms recommended (best practices), 20–100 ms acceptable, 100 ms in the transcribe guide; resample 44.1/48 kHz mics to 16 kHz client-side
Output audio raw 16-bit PCM 24 kHz little-endian; serverContent.modelTurn.parts[].inlineData with mimeType: "audio/pcm;rate=24000" (write WAV with 1 channel, 2 bytes, 24000 Hz); generated faster than realtime — buffer and play
Video individual frames as image/jpeg (or image/png), ≤ 1 frame/s, realtimeInput.video; mediaResolution controls tokens per frame
Text realtimeInput.text (stream) or clientContent.turns[].parts[].text
Token rate audio ≈ 25 tokens per second (both directions); transcribe output ≈ 175 text tokens/min

# 8. VAD, activity handling, turn coverage, interruptions

  • Automatic VAD (default) on the continuous audio stream. Tune startOfSpeechSensitivity, endOfSpeechSensitivity, prefixPaddingMs, silenceDurationMs (500–800 ms recommended; 100–200 ms fragments utterances; 2000+ ms adds latency). When the mic pauses > 1 s send realtimeInput.audioStreamEnd: true to flush; resume by sending audio.
  • Hybrid VAD: keep server VAD (robust start detection with prefix padding) and let a client VAD send audioStreamEnd as an immediate end-of-speech → minimal finalization latency; server VAD is the fallback.
  • Manual VAD: automaticActivityDetection.disabled: true; send activityStart → audio → activityEnd. No pre-speech buffer and no silence tolerance server-side — use ≥ 500 ms end-of-speech threshold client-side. audioStreamEnd is not used.
  • Interruptions: with START_OF_ACTIVITY_INTERRUPTS (default) user speech cancels generation; only content already sent stays in history; the server sends serverContent.interrupted: true (stop playback, flush queue) and toolCallCancellation for discarded calls. NO_INTERRUPTION disables barge-in. clientContent.turnComplete: true also interrupts unconditionally (3.1+).
  • Turn coverage: TURN_INCLUDES_ONLY_ACTIVITY (speech only, 2.5 default), TURN_INCLUDES_ALL_INPUT (incl. silence), TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO (3.1+ default; every video frame counts — send frames only when needed).

# 9. Transcription

  • setup.inputAudioTranscription: {} → serverContent.inputTranscription {text, languageCode} (finalized) and, on gemini-3.5-transcribe-live, serverContent.interimInputTranscription (fast partial hypotheses).
  • setup.outputAudioTranscription: {} → serverContent.outputTranscription for the model's speech (language inferred). The last output transcription of a turn precedes generationComplete/interrupted; ordering vs. audio parts is approximate.
  • Transcribe-live specifics: languageCodes[] (empty = auto, code-switching), customVocabulary[] (≤ 1,000, best ≤ 100), mode: "SMART" (filler removal, formatting; no word annotations). Not supported over Live: word timestamps, diarization. Sessions ≤ 10 min. Response modality ["TEXT"].
  • Billing: transcript text tokens are charged at text output rate on top of audio tokens.

# 10. Session management

Limit Documented value
Connection lifetime ≈ 10 minutes; the server sends goAway.timeLeft first, then terminates as ABORTED
Session (audio-only) 15 minutes without context window compression
Session (audio + video) 2 minutes without compression
With contextWindowCompression unlimited duration
Context window 128k tokens native audio models, 32k other Live models
Resumption handle validity 2 hours after the last session termination (changelog 2025-04-09 mentioned 24 h server-side state — current docs say 2 h)
Transcribe-live session 10 minutes
Concurrent sessions not documented on the rate-limits page (only "Concurrent batch requests: 100" for Batch)
  • Context window compression: contextWindowCompression: {triggerTokens, slidingWindow: {targetTokens}} — sliding window discards from the start (always at a USER turn; system instructions kept). Defaults: trigger = 80 % of context, target = trigger/2. Best-practice example: trigger 25,000 / target 8,000 also caps compounding cost.
  • Session resumption: set sessionResumption: {}; store the latest sessionResumptionUpdate.newHandle when resumable: true (false while generating / running functions); on reconnect send sessionResumption: {handle}. Configuration except model may change on resume. Ephemeral-token flows need this to reconnect every ~10 min within expireTime (same token even with uses: 1).
  • GoAway: use timeLeft to wrap up or reconnect. generationComplete: model finished generating (UI hook).

# 11. Tools in Live

Tool 3.8 Live 3.8 Live Extended Thinking 3.1 Flash Live 2.5 native audio Translate / Transcribe
Function calling yes — default NON_BLOCKING, BLOCKING allowed, scheduling supported async NON_BLOCKING only (BLOCKING = hard error), no scheduling synchronous only (model waits for toolResponse) sync + async no
Google Search (googleSearch: {}) yes yes yes yes no
Code execution, URL context, Google Maps not supported not supported not supported not supported no
  • Declare functions in setup.tools[].functionDeclarations[]; add "behavior": "NON_BLOCKING" for async. The server sends toolCall.functionCalls[] (always with id); reply with toolResponse.functionResponses[] (id, name, response). No automatic tool handling in Live.
  • Async FunctionResponse.scheduling: INTERRUPT (speak result now), WHEN_IDLE (default, after current output), SILENT (context only). willContinue: true turns a call into a generator (more responses later); end with willContinue: false (+ SILENT to avoid triggering speech).
  • Parallel calls: toolCall.functionCalls[] can contain several calls; on extended thinking the model speaks fillers while calls run (interactionStatus: IN_PROGRESS).
  • Barge-in during a tool turn → toolCallCancellation.ids[] (undo side effects if possible).
  • Grounding results in serverContent.groundingMetadata; 3.x: 5,000 free searches/month then $14 / 1,000.

# 12. Thinking in Live

  • gemini-3.1-flash-live-preview: thinkingConfig.thinkingLevel MINIMAL (default) / LOW / MEDIUM / HIGH (replaces 2.5 thinkingBudget).
  • gemini-3.8-live: interleaved reasoning, fixed latency profile; omit thinkingConfig.
  • gemini-3.8-live-extended-thinking: background reasoning, thinkingLevel LOW/MEDIUM/HIGH, includeThoughts for summaries. Lifecycle changes: the model speaks conversational fillers with turnComplete: true + interactionStatus: "IN_PROGRESS", emits async toolCalls (also tagged IN_PROGRESS), then the final answer with interactionStatus: "IDLE". turnComplete alone no longer means idle — keep listening until IDLE.

# 13. Proactive audio and affective dialog

  • proactivity.proactiveAudio: true lets the model ignore irrelevant speech / not answer when no request was made. Gemini 2.5: opt-in (v1beta). Gemini 3.8 (both): permanently enabled, false returns an error; billing: input tokens charged the whole time the API listens, output only when it answers. Gemini 3.1: unsupported (bills only while you stream).
  • generationConfig.enableAffectiveDialog: true (adapt tone to the user): Gemini 2.5 only; unsupported on 3.1; removed on 3.8.

# 14. Voices and languages

  • Voices: any Gemini TTS voice via speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName (30 names, see §4). The generateContent TTS voice set differs slightly.
  • Languages: 99 supported (BCP-47 list in the capabilities guide, e.g. en, fr, zh-Hans, pt-BR); native audio models switch languages naturally and do not support speechConfig.languageCode; constrain via system instructions.

# 15. Live translate and Live transcribe

  • Translate (gemini-3.5-live-translate-preview): audio-only input, generationConfig.translationConfig {targetLanguageCode (default "en"), echoTargetLanguage (default false)}, responseModalities: ["AUDIO"], optional inputAudioTranscription / outputAudioTranscription (transcripts carry languageCode). No tools, no system instructions, continuous (non-turn) processing; 70+ languages. Limitations: voice replication may drift; language detection struggles with accents/similar languages; background audio may leak with echo on. With ephemeral tokens lock translationConfig server-side, or omit it and set lock_additional_fields: [] to let the client choose.
  • Transcribe (gemini-3.5-transcribe-live): responseModalities: ["TEXT"], inputAudioTranscription {languageCodes, customVocabulary, mode}, interim + final transcripts, auto/hybrid/manual VAD, 10-min sessions, 85+ languages, no diarization / word timestamps over Live. Pricing ≈ $0.009/min blended.

# 16. Ephemeral tokens (POST /v1beta/auth_tokens)

Flow: client authenticates with your backend → backend calls POST https://generativelanguage.googleapis.com/v1beta/auth_tokens (header x-goog-api-key) → response AuthToken.name is the token → client opens the WebSocket with it (BidiGenerateContentConstrained?access_token=<token> or Authorization: Token <token>; SDK: apiKey: token.name). Live API only, v1beta only, Preview.

Field (REST body = AuthToken) Type Default / constraint
uses int32 1; 0 = unlimited; resuming a session does not count
expireTime RFC 3339 now + 30 min; < 20 h; after it, session messages are rejected
newSessionExpireTime RFC 3339 now + 60 s; < 20 h; after it, new sessions are rejected
fieldMask FieldMask empty + no setup → client's setup used; empty + setup → token's setup only; non-empty → listed fields overwrite the client's
bidiGenerateContentSetup BidiGenerateContentSetup locked config (discovery / reference name)
liveConnectConstraints {model, config} object shape used by the official curl examples and SDKs (live_connect_constraints); not in the discovery schema — REST acceptance UNVERIFIED
name string output only: the token
SDK-only: lock_additional_fields[], http_options — extra fields to lock (→ fieldMask); [] unlocks e.g. translationConfig

Best practices: short expireTime, re-provision on expiry, secure your backend auth, do not use ephemeral tokens for backend-to-Gemini paths. Ephemeral tokens pair with sessionResumption for reconnects every ~10 min.

# 17. Pricing (paid tier, USD per 1M tokens; free tier: free)

Model Text in Audio in Image/video in Text out Audio out Notes
gemini-3.8-live, gemini-3.8-live-extended-thinking, gemini-3.1-flash-live-preview $0.75 $3.00 (≈ $0.005/min) $1.00 (≈ $0.002/min) $4.50 (incl. thinking) $12.00 (≈ $0.018/min) Search: 5,000 free/month shared across 3.x, then $14/1,000
gemini-2.5-flash-native-audio-preview-12-2025 $0.50 $3.00 $3.00 $2.00 $12.00
gemini-3.5-live-translate-preview — $3.50 (≈ $0.0053/min) — — $21.00 (≈ $0.0315/min) ≈ $0.0368/min effective, 25 tokens/s
gemini-3.5-transcribe-live — $3.50 (≈ $0.005/min) — $21.00 (≈ $0.004/min) — ≈ $0.009/min blended; no Search

Billing model: per turn, all tokens in the context window are re-billed (compounding); audio kept as audio tokens (≈ 25 tok/s); transcription adds text-output tokens; cap growth with contextWindowCompression; proactive audio bills input while listening.

# 18. Rate limits and concurrency

The rate-limits page documents RPM/TPM/RPD per model tier but no Live-specific concurrent-session limit; goAway.timeLeft minimum is "specified together with the rate limits for the model" but not published. Observed headers are not applicable (WebSocket). Anything else here is UNVERIFIED.

# 19. Best practices (docs)

Clear, structured system instructions (persona, task, constraints, tool guidance, language); precise tool definitions; send 20–40 ms audio chunks, don't buffer ~1 s; handle interrupted by flushing playback; resample to 16 kHz; enable compression + resumption; handle goAway and generationComplete; process all parts of each serverContent; on 3.1+ only send video frames when needed.

# 20. Snippets

Raw WebSocket, text-only turn (Python websockets), key read from the environment (never hard-coded):

python
import asyncio, json, os, websockets

URL = ("wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage."
       "v1beta.GenerativeService.BidiGenerateContent?key=" + os.environ["GEMINI_API_KEY"])

async def main():
    async with websockets.connect(URL) as ws:
        await ws.send(json.dumps({"setup": {
            "model": "models/gemini-3.8-live",
            "generationConfig": {"responseModalities": ["AUDIO"]},
            "outputAudioTranscription": {}}}))
        assert "setupComplete" in json.loads(await ws.recv())
        await ws.send(json.dumps({"clientContent": {
            "turns": [{"role": "user", "parts": [{"text": "Reply with OK."}]}],
            "turnComplete": True}}))
        async for raw in ws:
            msg = json.loads(raw)
            sc = msg.get("serverContent", {})
            if "outputTranscription" in sc:
                print(sc["outputTranscription"]["text"])
            if sc.get("turnComplete"):
                print("usage:", msg.get("usageMetadata"))
                break

asyncio.run(main())

Python SDK:

python
from google import genai
from google.genai import types
client = genai.Client()  # GEMINI_API_KEY from env
config = types.LiveConnectConfig(response_modalities=["AUDIO"], output_audio_transcription={},
                                 session_resumption=types.SessionResumptionConfig())
async with client.aio.live.connect(model="gemini-3.8-live", config=config) as session:
    await session.send_client_content(turns={"role": "user", "parts": [{"text": "Reply with OK."}]}, turn_complete=True)
    async for msg in session.receive():
        if msg.server_content and msg.server_content.output_transcription:
            print(msg.server_content.output_transcription.text)
        if msg.server_content and msg.server_content.turn_complete:
            break

Node (@google/genai):

ts
import { GoogleGenAI, Modality } from "@google/genai";
const ai = new GoogleGenAI({});  // GEMINI_API_KEY from env
const session = await ai.live.connect({
  model: "gemini-3.8-live",
  config: { responseModalities: [Modality.AUDIO], outputAudioTranscription: {} },
  callbacks: { onmessage: (m) => { if (m.serverContent?.outputTranscription) console.log(m.serverContent.outputTranscription.text); } },
});
session.sendClientContent({ turns: [{ role: "user", parts: [{ text: "Reply with OK." }] }], turnComplete: true });

Ephemeral token (curl, server side):

bash
curl -X POST "https://generativelanguage.googleapis.com/v1beta/auth_tokens" \
  -H "x-goog-api-key: ${GEMINI_API_KEY}" -H "Content-Type: application/json" \
  -d '{"uses": 1, "expireTime": "2026-09-18T20:30:00Z",
       "liveConnectConstraints": {"model": "models/gemini-3.8-live",
         "config": {"sessionResumption": {}, "responseModalities": ["AUDIO"]}}}'

# Live verification (2026-09-18)

Run of 2026-09-18 (tmp-live/gemini-tools/j*.json, _summary.json; examples examples/gemini/live/* executed; tests/gemini/test_live.py cheap test passes, realtime tests gated by RUN_REALTIME_TESTS).

Probe Result
responseModalities:["TEXT"] on gemini-2.5-flash-native-audio-latest, gemini-3.1-flash-live-preview, gemini-3.8-live all rejected: WebSocket close 1007 The requested combination of response modalities (TEXT) is not supported by the model (3.8-live first sent setupComplete + sessionResumptionUpdate, then closed)
responseModalities:["AUDIO"] + outputAudioTranscription:{} on gemini-2.5-flash-native-audio-latest, clientContent "Reply with OK." setupComplete → serverContent.modelTurn parts [{text, thought:true}] → serverContent.outputTranscription{text:"OK"} → serverContent.modelTurn inlineData{mimeType:"audio/pcm;rate=24000"} × 0–6 → serverContent.generationComplete → serverContent.turnComplete with usageMetadata in the same message ({promptTokenCount:373, responseTokenCount:4–21, responseTokensDetails:[{modality:"AUDIO"}], thoughtsTokenCount})
sessionResumption:{} + contextWindowCompression:{slidingWindow:{}} sessionResumptionUpdate{newHandle:<uuid>, resumable:true} right after setupComplete and again after turnComplete
setup.toolConfig close 1007 Unknown name "toolConfig" at 'setup' — tool mode cannot be forced in Live
tools:[{functionDeclarations:[{…, behavior:"NON_BLOCKING"}]}] toolCall{functionCalls:[{name, args, id:"function-call-11967657623789764093"}]} after the thought part; audio continued (non-blocking)
SDK client.aio.live.connect(model, config={"response_modalities":["AUDIO"], "output_audio_transcription":{}}) sequence server_content ×3 → server_content+usage_metadata, transcript "OK."
@google/genai ai.live.connect({callbacks}) setupComplete → modelTurn.thought → outputTranscription → generationComplete → turnComplete
POST /v1alpha/auth_tokens and /v1beta/auth_tokens {uses:1, expireTime, newSessionExpireTime} both 200 {name:"auth_tokens/…"}; with bidiGenerateContentSetup{model, generationConfig} lock → 200
ephemeral token on …v1alpha.GenerativeService.BidiGenerateContent (query or header) close 1008 Method doesn't allow unregistered callers
ephemeral token on …v1alpha.GenerativeService.BidiGenerateContentConstrained works with Authorization: Token <name> header and with ?access_token=<name>; locked token works too

Not exercised: realtime audio/video input, VAD tuning, interruptions, goAway timing (sessions lasted < 10 s), Live translate / transcribe models, half-cascade models (none remain).