# Gemini Live API — bidirectional WebSocket (BidiGenerateContent) + ephemeral tokens **Status:** DOCUMENTED + LIVE_VERIFIED (2026-09-18 run — see "Live verification" section at the end) · PREVIEW (Google labels the Live API and ephemeral tokens "Preview"; `gemini-3.8-live` / `gemini-3.8-live-extended-thinking` are "Stable"/GA since 2026-09-15) **Sources:** https://ai.google.dev/api/live (WebSocket reference) · https://ai.google.dev/gemini-api/docs/live-api · …/live-api/get-started-sdk · …/live-api/get-started-websocket · …/live-api/capabilities · …/live-api/session-management · …/live-api/tools · …/live-api/thinking · …/live-api/ephemeral-tokens · …/live-api/live-transcribe · …/live-api/live-translate · …/live-api/best-practices · https://ai.google.dev/gemini-api/docs/models/{gemini-3.8-live, gemini-3.8-live-extended-thinking, gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-preview-12-2025, gemini-3.5-live-translate-preview, gemini-3.5-transcribe} · https://ai.google.dev/gemini-api/docs/pricing · https://ai.google.dev/gemini-api/docs/changelog · discovery document `v1beta` rev. 20260918 (`sources/gemini/discovery-v1beta.json`) · SDK types `sources/gemini/openapi/python-genai-types.py` · live model list `sources/gemini/models-api-raw.json` **Last verified:** 2026-09-18 **Machine-readable twins:** `generated/fragments/streaming-events/gemini-live.json` (32 events), `generated/fragments/parameters/gemini-live.json` (115 params), `tmp/gemini-parts/live-endpoints.json` (4 endpoints), `tmp/gemini-parts/live-objects.json` (39 objects). Full per-message field tables: [`live-events.md`](live-events.md). ## 1. Overview and architecture The Live API is a **stateful WebSocket (WSS) session** with a Gemini model: continuous audio / video / text in, **native audio** (or text, for the transcribe model) out, with barge-in, server-side VAD, function calling, Google Search grounding, transcription, session resumption and context compression. All traffic is JSON text frames; media are base64 `Blob`s. Two integration approaches (docs): **server-to-server** (your backend holds the API key and proxies the stream) and **client-to-server** (browser/mobile connects directly — lower latency; use **ephemeral tokens**, never an API key). Partner stacks: Pipecat, LiveKit, Fishjam, ADK, Vision Agents, Voximplant. ```mermaid sequenceDiagram participant C as Client participant S as Gemini Live (WSS v1beta) C->>S: {"setup": {model, generationConfig, tools, realtimeInputConfig, ...}} S-->>C: {"setupComplete": {}} loop conversation C->>S: {"realtimeInput": {audio|video|text|activityStart|activityEnd|audioStreamEnd}} C->>S: {"clientContent": {turns[], turnComplete}} S-->>C: {"serverContent": {inputTranscription | interimInputTranscription}} S-->>C: {"serverContent": {modelTurn: {parts:[inlineData audio/pcm;rate=24000]}}} S-->>C: {"serverContent": {outputTranscription}} S-->>C: {"toolCall": {functionCalls[{id,name,args}]}} C->>S: {"toolResponse": {functionResponses[{id,name,response,scheduling}]}} S-->>C: {"serverContent": {interrupted: true}} / {"toolCallCancellation": {ids[]}} S-->>C: {"serverContent": {generationComplete: true}} S-->>C: {"serverContent": {turnComplete: true, interactionStatus: IDLE}, "usageMetadata": {...}} S-->>C: {"sessionResumptionUpdate": {newHandle, resumable}} end S-->>C: {"goAway": {timeLeft: "30s"}} C->>S: reconnect: {"setup": {..., sessionResumption: {handle}}} ``` Wire rules (reference): a client message has **exactly one** of `setup | clientContent | realtimeInput | toolResponse`; a server message may carry `usageMetadata` and otherwise **exactly one** of `setupComplete | serverContent | toolCall | toolCallCancellation | goAway | sessionResumptionUpdate` (the proto `messageType` union is flattened to the top level). Setup is sent once; configuration cannot change mid-connection (only on resume, model excluded). ## 2. Models (bidiGenerateContent on this key) All current Live models are **native audio** (audio-to-audio) models. The former half-cascade models (`gemini-2.0-flash-live-001`, `gemini-live-2.5-flash-preview`) were **shut down 2025-12-09** (changelog); no current page mentions "half-cascade". | Model id | Type / role | Input → output | Thinking | Tools | Async FC | Proactive / affective | Context / output tokens | Session limits | Status (docs) | |---|---|---|---|---|---|---|---|---|---| | `gemini-3.8-live` | native audio, default voice agent | text, image, audio, video → audio (+ transcript) | interleaved; `thinkingLevel` **not supported** (omit) | function calling, Google Search | default `NON_BLOCKING`; `BLOCKING` allowed; scheduling supported | proactive audio **always on** (`false` → error); affective dialog **removed** | 131,072 in / 65,536 out | 15 min audio / 2 min audio+video w/o compression; 128k context | Stable (GA 2026-09-15), API "Preview" | | `gemini-3.8-live-extended-thinking` | native audio, background reasoning | same | `thinkingLevel` LOW/MEDIUM/HIGH (no MINIMAL); `includeThoughts` | function calling (**async only**), Google Search | `NON_BLOCKING` only (BLOCKING = hard error); no scheduling | proactive always on; affective removed | 131,072 / 65,536 | same; watch `interactionStatus` | Stable (GA 2026-09-15) | | `gemini-3.1-flash-live-preview` | native audio, legacy preview | text, image, audio, video → audio | `thinkingLevel` MINIMAL (default)/LOW/MEDIUM/HIGH | function calling (sync only), Google Search | not supported | neither supported | 131,072 / 65,536 | same | Preview, "legacy — migrate to 3.8" | | `gemini-2.5-flash-native-audio-preview-12-2025` (+ `-preview-09-2025`, `-native-audio-latest`) | native audio (2.5) | audio, video, text → audio and text | `thinkingBudget` (thinking: true) | function calling (sync + async), Google Search | supported | both opt-in (v1beta) | 131,072 / 8,192 | 15 min / 2 min; 128k | Preview (2.5 family "no longer available to new users" on this key per CLAUDE.md — verify) | | `gemini-3.5-live-translate-preview` | speech-to-speech translation, 70+ languages | audio only → translated audio + transcripts | none | none | — | — | 16,384 in / 32,768 out (models API) | UNVERIFIED (not stated) | Preview (June 2026) | | `gemini-3.5-transcribe-live` | streaming speech-to-text | 16-bit PCM audio → `TEXT` (inputTranscription / interimInputTranscription) | none | none | — | — | 131,072 / 65,536 | **10 min per session** | GA-ish (Aug 2026), version `3.5-transcribe-live-08-2026` | | `gemini-robotics-er-2-streaming-preview` | text streaming for robots (out of scope) | audio, video → text | — | — | — | — | 131,072 / 65,536 | — | Preview | Native-audio limitation: only the `AUDIO` response modality; get text via `outputAudioTranscription`. Any Gemini TTS voice can be used; languages are chosen automatically (99 listed). ## 3. Connection and authentication | Item | Value | |---|---| | URL (API key) | `wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent?key=YOUR_API_KEY` | | URL (ephemeral token) | `wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContentConstrained?access_token=` — or header `Authorization: Token ` (reference page) | | Version | **v1beta** only (reference note; ephemeral tokens "only with v1beta"). One snippet in the thinking guide shows a `v1alpha` URL — treat as legacy/UNVERIFIED. | | First frame | `{"setup": …}`; wait for `{"setupComplete": {}}` before anything else | | Python SDK | `async with client.aio.live.connect(model="gemini-3.8-live", config=types.LiveConnectConfig(...)) as session:` then `session.send_realtime_input(...)`, `session.send_client_content(...)`, `session.send_tool_response(...)`, `async for msg in session.receive()` | | Node SDK | `const session = await ai.live.connect({model, config, callbacks:{onopen,onmessage,onerror,onclose}})`; `session.sendRealtimeInput({...})`, `sendClientContent`, `sendToolResponse`, `session.close()` | | Ephemeral token via SDK | `new GoogleGenAI({ apiKey: token.name })` / `genai.Client(api_key=token.name, http_options={"api_version":"v1beta"})` — the SDK switches to the constrained endpoint for you | Security note (project rule): keys only in `.env`; the docs' `?key=` query form is server-side only. The parent agent does live calls. ## 4. `setup` configuration — every field Wire names (camelCase) = REST JSON; SDK `LiveConnectConfig` uses snake_case (Python) / camelCase (Node) and flattens `generationConfig` fields (e.g. `response_modalities`, `speech_config`, `thinking_config`, `media_resolution`, `enable_affective_dialog`, `translation_config`). | Field | Type | Required | Default | Notes / models | |---|---|---|---|---| | `setup.model` | string | yes | — | `models/gemini-3.8-live` | | `setup.generationConfig` | GenerationConfig subset | no | — | Not supported: responseLogprobs, responseMimeType, logprobs, responseSchema, responseJsonSchema, stopSequence, skipResponseCache, routingConfig, audioTimestamp | | `…responseModalities[]` | `TEXT` \| `AUDIO` | no | SDK: AUDIO | Native audio models: AUDIO only; transcribe-live: TEXT | | `…speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName` | string | no | model default | 30 TTS voices (Zephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirrhoe, Autonoe, Enceladus, Iapetus, Umbriel, Algieba, Despina, Erinome, Algenib, Rasalgethi, Laomedeia, Achernar, Alnilam, Schedar, Gacrux, Pulcherrima, Achird, Zubenelgenubi, Vindemiatrix, Sadachbia, Sadaltager, Sulafat) | | `…speechConfig.languageCode` | BCP-47 | no | auto | 30 codes listed in discovery; **native audio models ignore/unsupported** (steer via system instructions) — UNVERIFIED effect | | `…speechConfig.multiSpeakerVoiceConfig` | object | no | — | TTS feature; Live support undocumented (UNVERIFIED) | | `…mediaResolution` | `MEDIA_RESOLUTION_LOW` (64 tok) \| `MEDIUM` (256) \| `HIGH` (zoomed 256) | no | unspecified | Visual frames only; audio tokenization fixed | | `…thinkingConfig.thinkingLevel` | `MINIMAL` \| `LOW` \| `MEDIUM` \| `HIGH` | no | 3.1: MINIMAL | 3.1: all four; 3.8-extended-thinking: LOW/MEDIUM/HIGH; **3.8-live: unsupported** | | `…thinkingConfig.thinkingBudget` | int | no | — | Gemini 2.5 native audio (legacy) | | `…thinkingConfig.includeThoughts` | bool | no | false | thought summaries as parts | | `…enableAffectiveDialog` | bool | no | false | Gemini 2.5 only (v1beta); 3.1 unsupported; 3.8 removed | | `…translationConfig.targetLanguageCode` / `.echoTargetLanguage` | string / bool | translate: yes / no | "en" / false | `gemini-3.5-live-translate-preview` only | | `…temperature, topP, topK, maxOutputTokens, candidateCount, presencePenalty, frequencyPenalty, seed` | number/int | no | — | listed in the reference example | | `setup.systemInstruction` | Content (text parts) | no | — | each part = paragraph; specify the language here | | `setup.tools[]` | Tool[] | no | — | `functionDeclarations[]` (with Live-only `behavior`), `googleSearch: {}`; codeExecution / urlContext / googleMaps **not supported** | | `setup.realtimeInputConfig.automaticActivityDetection.disabled` | bool | no | false | true = manual VAD (activityStart/End) | | `….startOfSpeechSensitivity` | `START_SENSITIVITY_HIGH` \| `LOW` | no | HIGH | | | `….endOfSpeechSensitivity` | `END_SENSITIVITY_HIGH` \| `LOW` | no | HIGH | | | `….prefixPaddingMs` | int32 | no | — | look-back before speech; 0 clips onsets | | `….silenceDurationMs` | int32 | no | ≈800 (server) | recommended 500–800 | | `setup.realtimeInputConfig.activityHandling` | `START_OF_ACTIVITY_INTERRUPTS` \| `NO_INTERRUPTION` | no | interrupts | barge-in | | `setup.realtimeInputConfig.turnCoverage` | `TURN_INCLUDES_ONLY_ACTIVITY` \| `TURN_INCLUDES_ALL_INPUT` \| `TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO` | no | 2.5: ONLY_ACTIVITY; 3.1+: AUDIO_ACTIVITY_AND_ALL_VIDEO | all video frames billed on 3.1+ | | `setup.inputAudioTranscription` | AudioTranscriptionConfig | no | off | `{}` enables; `languageCodes[]` (empty = auto), `customVocabulary[]` (≤1000), `mode` VERBATIM\|SMART; `wordTimestamp`/`diarization` not supported over Live | | `setup.outputAudioTranscription` | AudioTranscriptionConfig | no | off | `{}` enables transcript of model audio (billed as text output) | | `setup.sessionResumption.handle` | string | no | new session | `{}` = new resumable session; handle valid 2 h; (`transparent` SDK-only) | | `setup.contextWindowCompression.triggerTokens` | int64 | no | 80 % of context | | | `setup.contextWindowCompression.slidingWindow.targetTokens` | int64 | no | triggerTokens/2 | | | `setup.proactivity.proactiveAudio` | bool | no | 2.5: false | 3.8: always on (false → error); 3.1: unsupported | | `setup.historyConfig.initialHistoryInClientContent` | bool | no | false | ingest clientContent history before realtime starts | | `setup.labels` | map | no | — | discovery-only (`safety_identifier`) — UNVERIFIED | | `setup.explicitVadSignal`, `setup.safetySettings[]`, `setup.avatarConfig` | — | no | — | SDK types only — UNVERIFIED | ## 5. Client messages (summary) | Key | Type | Purpose | SDK | |---|---|---|---| | `setup` | BidiGenerateContentSetup | first message, session config | `live.connect(model, config)` | | `clientContent` | `{turns[]: Content, turnComplete}` | append history / text turn; interrupts generation; `turnComplete:true` starts generation | `send_client_content` / `sendClientContent` | | `realtimeInput` | `{audio \| video \| text \| activityStart \| activityEnd \| audioStreamEnd \| mediaResolution \| mediaChunks(deprecated)}` | continuous streams; turn end from VAD | `send_realtime_input` / `sendRealtimeInput` | | `toolResponse` | `{functionResponses[]: {id, name, response, scheduling, willContinue, parts}}` | answer a `toolCall` | `send_tool_response` / `sendToolResponse` | ## 6. Server messages (summary) | Key | Content | When | |---|---|---| | `setupComplete` | `{}` | once, after setup | | `serverContent` | `modelTurn` (parts: `inlineData` audio/pcm;rate=24000, text, thoughts), `generationComplete`, `turnComplete` (+ `interactionStatus`), `interrupted`, `groundingMetadata`, `inputTranscription`, `interimInputTranscription`, `outputTranscription`, `urlContextMetadata`, `waitingForInput`, `speechState` (deprecated) | throughout a turn; several fields/parts per message possible | | `toolCall` | `functionCalls[] {id, name, args}` | when the model wants a function; may arrive during audio (async) | | `toolCallCancellation` | `ids[]` | after a barge-in discards pending calls | | `usageMetadata` | token counts + per-modality details | alongside other messages, periodically | | `goAway` | `timeLeft` (Duration) | before the ~10-min connection reset | | `sessionResumptionUpdate` | `newHandle`, `resumable` | periodically when `sessionResumption` is set | | `voiceActivity`, `voiceActivityDetectionSignal` | SDK types only ("allowlisted") | UNVERIFIED | Full field tables and examples: [`live-events.md`](live-events.md). ## 7. Audio, video and text formats | Direction | Format | |---|---| | Input audio | raw **16-bit PCM, little-endian, mono, 16 kHz native**; `Blob.mimeType = "audio/pcm;rate=16000"` (other rates are resampled if declared); base64 in `realtimeInput.audio.data`; chunks 20–40 ms recommended (best practices), 20–100 ms acceptable, 100 ms in the transcribe guide; resample 44.1/48 kHz mics to 16 kHz client-side | | Output audio | raw **16-bit PCM 24 kHz** little-endian; `serverContent.modelTurn.parts[].inlineData` with `mimeType: "audio/pcm;rate=24000"` (write WAV with 1 channel, 2 bytes, 24000 Hz); generated faster than realtime — buffer and play | | Video | individual frames as `image/jpeg` (or `image/png`), **≤ 1 frame/s**, `realtimeInput.video`; `mediaResolution` controls tokens per frame | | Text | `realtimeInput.text` (stream) or `clientContent.turns[].parts[].text` | | Token rate | audio ≈ **25 tokens per second** (both directions); transcribe output ≈ 175 text tokens/min | ## 8. VAD, activity handling, turn coverage, interruptions - **Automatic VAD** (default) on the continuous audio stream. Tune `startOfSpeechSensitivity`, `endOfSpeechSensitivity`, `prefixPaddingMs`, `silenceDurationMs` (500–800 ms recommended; 100–200 ms fragments utterances; 2000+ ms adds latency). When the mic pauses > 1 s send `realtimeInput.audioStreamEnd: true` to flush; resume by sending audio. - **Hybrid VAD**: keep server VAD (robust start detection with prefix padding) and let a client VAD send `audioStreamEnd` as an immediate end-of-speech → minimal finalization latency; server VAD is the fallback. - **Manual VAD**: `automaticActivityDetection.disabled: true`; send `activityStart` → audio → `activityEnd`. No pre-speech buffer and no silence tolerance server-side — use ≥ 500 ms end-of-speech threshold client-side. `audioStreamEnd` is not used. - **Interruptions**: with `START_OF_ACTIVITY_INTERRUPTS` (default) user speech cancels generation; only content already sent stays in history; the server sends `serverContent.interrupted: true` (stop playback, flush queue) and `toolCallCancellation` for discarded calls. `NO_INTERRUPTION` disables barge-in. `clientContent.turnComplete: true` also interrupts unconditionally (3.1+). - **Turn coverage**: `TURN_INCLUDES_ONLY_ACTIVITY` (speech only, 2.5 default), `TURN_INCLUDES_ALL_INPUT` (incl. silence), `TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO` (3.1+ default; every video frame counts — send frames only when needed). ## 9. Transcription - `setup.inputAudioTranscription: {}` → `serverContent.inputTranscription {text, languageCode}` (finalized) and, on `gemini-3.5-transcribe-live`, `serverContent.interimInputTranscription` (fast partial hypotheses). - `setup.outputAudioTranscription: {}` → `serverContent.outputTranscription` for the model's speech (language inferred). The last output transcription of a turn precedes `generationComplete`/`interrupted`; ordering vs. audio parts is approximate. - Transcribe-live specifics: `languageCodes[]` (empty = auto, code-switching), `customVocabulary[]` (≤ 1,000, best ≤ 100), `mode: "SMART"` (filler removal, formatting; no word annotations). Not supported over Live: word timestamps, diarization. Sessions ≤ 10 min. Response modality `["TEXT"]`. - Billing: transcript text tokens are charged at text output rate on top of audio tokens. ## 10. Session management | Limit | Documented value | |---|---| | Connection lifetime | ≈ **10 minutes**; the server sends `goAway.timeLeft` first, then terminates as ABORTED | | Session (audio-only) | **15 minutes** without context window compression | | Session (audio + video) | **2 minutes** without compression | | With `contextWindowCompression` | unlimited duration | | Context window | **128k** tokens native audio models, 32k other Live models | | Resumption handle validity | **2 hours** after the last session termination (changelog 2025-04-09 mentioned 24 h server-side state — current docs say 2 h) | | Transcribe-live session | 10 minutes | | Concurrent sessions | not documented on the rate-limits page (only "Concurrent batch requests: 100" for Batch) | - **Context window compression**: `contextWindowCompression: {triggerTokens, slidingWindow: {targetTokens}}` — sliding window discards from the start (always at a USER turn; system instructions kept). Defaults: trigger = 80 % of context, target = trigger/2. Best-practice example: trigger 25,000 / target 8,000 also caps compounding cost. - **Session resumption**: set `sessionResumption: {}`; store the latest `sessionResumptionUpdate.newHandle` when `resumable: true` (false while generating / running functions); on reconnect send `sessionResumption: {handle}`. Configuration except `model` may change on resume. Ephemeral-token flows need this to reconnect every ~10 min within `expireTime` (same token even with `uses: 1`). - **GoAway**: use `timeLeft` to wrap up or reconnect. **generationComplete**: model finished generating (UI hook). ## 11. Tools in Live | Tool | 3.8 Live | 3.8 Live Extended Thinking | 3.1 Flash Live | 2.5 native audio | Translate / Transcribe | |---|---|---|---|---|---| | Function calling | yes — default `NON_BLOCKING`, `BLOCKING` allowed, scheduling supported | async `NON_BLOCKING` only (BLOCKING = hard error), no scheduling | synchronous only (model waits for `toolResponse`) | sync + async | no | | Google Search (`googleSearch: {}`) | yes | yes | yes | yes | no | | Code execution, URL context, Google Maps | not supported | not supported | not supported | not supported | no | - Declare functions in `setup.tools[].functionDeclarations[]`; add `"behavior": "NON_BLOCKING"` for async. The server sends `toolCall.functionCalls[]` (always with `id`); reply with `toolResponse.functionResponses[]` (`id`, `name`, `response`). **No automatic tool handling** in Live. - Async `FunctionResponse.scheduling`: `INTERRUPT` (speak result now), `WHEN_IDLE` (default, after current output), `SILENT` (context only). `willContinue: true` turns a call into a generator (more responses later); end with `willContinue: false` (+ `SILENT` to avoid triggering speech). - Parallel calls: `toolCall.functionCalls[]` can contain several calls; on extended thinking the model speaks fillers while calls run (`interactionStatus: IN_PROGRESS`). - Barge-in during a tool turn → `toolCallCancellation.ids[]` (undo side effects if possible). - Grounding results in `serverContent.groundingMetadata`; 3.x: 5,000 free searches/month then $14 / 1,000. ## 12. Thinking in Live - `gemini-3.1-flash-live-preview`: `thinkingConfig.thinkingLevel` MINIMAL (default) / LOW / MEDIUM / HIGH (replaces 2.5 `thinkingBudget`). - `gemini-3.8-live`: interleaved reasoning, fixed latency profile; **omit `thinkingConfig`**. - `gemini-3.8-live-extended-thinking`: background reasoning, `thinkingLevel` LOW/MEDIUM/HIGH, `includeThoughts` for summaries. Lifecycle changes: the model speaks **conversational fillers** with `turnComplete: true` + `interactionStatus: "IN_PROGRESS"`, emits async `toolCall`s (also tagged IN_PROGRESS), then the final answer with `interactionStatus: "IDLE"`. `turnComplete` alone no longer means idle — keep listening until IDLE. ## 13. Proactive audio and affective dialog - `proactivity.proactiveAudio: true` lets the model ignore irrelevant speech / not answer when no request was made. Gemini 2.5: opt-in (`v1beta`). Gemini 3.8 (both): **permanently enabled**, `false` returns an error; billing: input tokens charged the whole time the API listens, output only when it answers. Gemini 3.1: unsupported (bills only while you stream). - `generationConfig.enableAffectiveDialog: true` (adapt tone to the user): Gemini 2.5 only; unsupported on 3.1; **removed** on 3.8. ## 14. Voices and languages - Voices: any Gemini TTS voice via `speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName` (30 names, see §4). The `generateContent` TTS voice set differs slightly. - Languages: 99 supported (BCP-47 list in the capabilities guide, e.g. `en`, `fr`, `zh-Hans`, `pt-BR`); native audio models switch languages naturally and **do not support** `speechConfig.languageCode`; constrain via system instructions. ## 15. Live translate and Live transcribe - **Translate** (`gemini-3.5-live-translate-preview`): audio-only input, `generationConfig.translationConfig {targetLanguageCode (default "en"), echoTargetLanguage (default false)}`, `responseModalities: ["AUDIO"]`, optional `inputAudioTranscription` / `outputAudioTranscription` (transcripts carry `languageCode`). No tools, no system instructions, continuous (non-turn) processing; 70+ languages. Limitations: voice replication may drift; language detection struggles with accents/similar languages; background audio may leak with echo on. With ephemeral tokens lock `translationConfig` server-side, or omit it and set `lock_additional_fields: []` to let the client choose. - **Transcribe** (`gemini-3.5-transcribe-live`): `responseModalities: ["TEXT"]`, `inputAudioTranscription {languageCodes, customVocabulary, mode}`, interim + final transcripts, auto/hybrid/manual VAD, 10-min sessions, 85+ languages, no diarization / word timestamps over Live. Pricing ≈ $0.009/min blended. ## 16. Ephemeral tokens (`POST /v1beta/auth_tokens`) Flow: client authenticates with **your** backend → backend calls `POST https://generativelanguage.googleapis.com/v1beta/auth_tokens` (header `x-goog-api-key`) → response `AuthToken.name` is the token → client opens the WebSocket with it (`BidiGenerateContentConstrained?access_token=` or `Authorization: Token `; SDK: `apiKey: token.name`). Live API only, **v1beta only**, Preview. | Field (REST body = `AuthToken`) | Type | Default / constraint | |---|---|---| | `uses` | int32 | 1; 0 = unlimited; resuming a session does not count | | `expireTime` | RFC 3339 | now + 30 min; < 20 h; after it, session messages are rejected | | `newSessionExpireTime` | RFC 3339 | now + 60 s; < 20 h; after it, new sessions are rejected | | `fieldMask` | FieldMask | empty + no setup → client's setup used; empty + setup → token's setup only; non-empty → listed fields overwrite the client's | | `bidiGenerateContentSetup` | BidiGenerateContentSetup | locked config (discovery / reference name) | | `liveConnectConstraints {model, config}` | object | shape used by the **official curl examples** and SDKs (`live_connect_constraints`); not in the discovery schema — REST acceptance UNVERIFIED | | `name` | string | output only: the token | | SDK-only: `lock_additional_fields[]`, `http_options` | — | extra fields to lock (→ fieldMask); `[]` unlocks e.g. `translationConfig` | Best practices: short `expireTime`, re-provision on expiry, secure your backend auth, do not use ephemeral tokens for backend-to-Gemini paths. Ephemeral tokens pair with `sessionResumption` for reconnects every ~10 min. ## 17. Pricing (paid tier, USD per 1M tokens; free tier: free) | Model | Text in | Audio in | Image/video in | Text out | Audio out | Notes | |---|---|---|---|---|---|---| | `gemini-3.8-live`, `gemini-3.8-live-extended-thinking`, `gemini-3.1-flash-live-preview` | $0.75 | $3.00 (≈ $0.005/min) | $1.00 (≈ $0.002/min) | $4.50 (incl. thinking) | $12.00 (≈ $0.018/min) | Search: 5,000 free/month shared across 3.x, then $14/1,000 | | `gemini-2.5-flash-native-audio-preview-12-2025` | $0.50 | $3.00 | $3.00 | $2.00 | $12.00 | | | `gemini-3.5-live-translate-preview` | — | $3.50 (≈ $0.0053/min) | — | — | $21.00 (≈ $0.0315/min) | ≈ $0.0368/min effective, 25 tokens/s | | `gemini-3.5-transcribe-live` | — | $3.50 (≈ $0.005/min) | — | $21.00 (≈ $0.004/min) | — | ≈ $0.009/min blended; no Search | Billing model: per turn, **all tokens in the context window** are re-billed (compounding); audio kept as audio tokens (≈ 25 tok/s); transcription adds text-output tokens; cap growth with `contextWindowCompression`; proactive audio bills input while listening. ## 18. Rate limits and concurrency The rate-limits page documents RPM/TPM/RPD per model tier but **no Live-specific concurrent-session limit**; `goAway.timeLeft` minimum is "specified together with the rate limits for the model" but not published. Observed headers are not applicable (WebSocket). Anything else here is UNVERIFIED. ## 19. Best practices (docs) Clear, structured system instructions (persona, task, constraints, tool guidance, language); precise tool definitions; send 20–40 ms audio chunks, don't buffer ~1 s; handle `interrupted` by flushing playback; resample to 16 kHz; enable compression + resumption; handle `goAway` and `generationComplete`; process **all parts** of each `serverContent`; on 3.1+ only send video frames when needed. ## 20. Snippets Raw WebSocket, text-only turn (Python `websockets`), key read from the environment (never hard-coded): ```python import asyncio, json, os, websockets URL = ("wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage." "v1beta.GenerativeService.BidiGenerateContent?key=" + os.environ["GEMINI_API_KEY"]) async def main(): async with websockets.connect(URL) as ws: await ws.send(json.dumps({"setup": { "model": "models/gemini-3.8-live", "generationConfig": {"responseModalities": ["AUDIO"]}, "outputAudioTranscription": {}}})) assert "setupComplete" in json.loads(await ws.recv()) await ws.send(json.dumps({"clientContent": { "turns": [{"role": "user", "parts": [{"text": "Reply with OK."}]}], "turnComplete": True}})) async for raw in ws: msg = json.loads(raw) sc = msg.get("serverContent", {}) if "outputTranscription" in sc: print(sc["outputTranscription"]["text"]) if sc.get("turnComplete"): print("usage:", msg.get("usageMetadata")) break asyncio.run(main()) ``` Python SDK: ```python from google import genai from google.genai import types client = genai.Client() # GEMINI_API_KEY from env config = types.LiveConnectConfig(response_modalities=["AUDIO"], output_audio_transcription={}, session_resumption=types.SessionResumptionConfig()) async with client.aio.live.connect(model="gemini-3.8-live", config=config) as session: await session.send_client_content(turns={"role": "user", "parts": [{"text": "Reply with OK."}]}, turn_complete=True) async for msg in session.receive(): if msg.server_content and msg.server_content.output_transcription: print(msg.server_content.output_transcription.text) if msg.server_content and msg.server_content.turn_complete: break ``` Node (`@google/genai`): ```ts import { GoogleGenAI, Modality } from "@google/genai"; const ai = new GoogleGenAI({}); // GEMINI_API_KEY from env const session = await ai.live.connect({ model: "gemini-3.8-live", config: { responseModalities: [Modality.AUDIO], outputAudioTranscription: {} }, callbacks: { onmessage: (m) => { if (m.serverContent?.outputTranscription) console.log(m.serverContent.outputTranscription.text); } }, }); session.sendClientContent({ turns: [{ role: "user", parts: [{ text: "Reply with OK." }] }], turnComplete: true }); ``` Ephemeral token (curl, server side): ```bash curl -X POST "https://generativelanguage.googleapis.com/v1beta/auth_tokens" \ -H "x-goog-api-key: ${GEMINI_API_KEY}" -H "Content-Type: application/json" \ -d '{"uses": 1, "expireTime": "2026-09-18T20:30:00Z", "liveConnectConstraints": {"model": "models/gemini-3.8-live", "config": {"sessionResumption": {}, "responseModalities": ["AUDIO"]}}}' ``` ## Live verification (2026-09-18) Run of 2026-09-18 (`tmp-live/gemini-tools/j*.json`, `_summary.json`; examples `examples/gemini/live/*` executed; `tests/gemini/test_live.py` cheap test passes, realtime tests gated by `RUN_REALTIME_TESTS`). | Probe | Result | |---|---| | `responseModalities:["TEXT"]` on `gemini-2.5-flash-native-audio-latest`, `gemini-3.1-flash-live-preview`, `gemini-3.8-live` | **all rejected**: WebSocket close **1007** `The requested combination of response modalities (TEXT) is not supported by the model` (3.8-live first sent `setupComplete` + `sessionResumptionUpdate`, then closed) | | `responseModalities:["AUDIO"]` + `outputAudioTranscription:{}` on `gemini-2.5-flash-native-audio-latest`, `clientContent` "Reply with OK." | `setupComplete` → `serverContent.modelTurn` parts `[{text, thought:true}]` → `serverContent.outputTranscription{text:"OK"}` → `serverContent.modelTurn` `inlineData{mimeType:"audio/pcm;rate=24000"}` × 0–6 → `serverContent.generationComplete` → `serverContent.turnComplete` **with `usageMetadata` in the same message** (`{promptTokenCount:373, responseTokenCount:4–21, responseTokensDetails:[{modality:"AUDIO"}], thoughtsTokenCount}`) | | `sessionResumption:{}` + `contextWindowCompression:{slidingWindow:{}}` | `sessionResumptionUpdate{newHandle:, resumable:true}` right after `setupComplete` and again after `turnComplete` | | `setup.toolConfig` | close **1007** `Unknown name "toolConfig" at 'setup'` — tool mode cannot be forced in Live | | `tools:[{functionDeclarations:[{…, behavior:"NON_BLOCKING"}]}]` | `toolCall{functionCalls:[{name, args, id:"function-call-11967657623789764093"}]}` after the thought part; audio continued (non-blocking) | | SDK `client.aio.live.connect(model, config={"response_modalities":["AUDIO"], "output_audio_transcription":{}})` | sequence `server_content ×3 → server_content+usage_metadata`, transcript "OK." | | `@google/genai` `ai.live.connect({callbacks})` | `setupComplete → modelTurn.thought → outputTranscription → generationComplete → turnComplete` | | `POST /v1alpha/auth_tokens` and `/v1beta/auth_tokens` `{uses:1, expireTime, newSessionExpireTime}` | both 200 `{name:"auth_tokens/…"}`; with `bidiGenerateContentSetup{model, generationConfig}` lock → 200 | | ephemeral token on `…v1alpha.GenerativeService.BidiGenerateContent` (query or header) | close **1008** `Method doesn't allow unregistered callers` | | ephemeral token on `…v1alpha.GenerativeService.BidiGenerateContentConstrained` | **works** with `Authorization: Token ` header and with `?access_token=`; locked token works too | Not exercised: realtime audio/video input, VAD tuning, interruptions, goAway timing (sessions lasted < 10 s), Live translate / transcribe models, half-cascade models (none remain).