Gemini transcription — gemini-3.5-transcribe (unary) and gemini-3.5-transcribe-live (Live API)
Status: DOCUMENTED · LIVE_VERIFIED 2026-09-18 (generateContent with inline WAV: speech, silence, diarization + word timestamps, SMART mode, incompatibility error) · LIVE_DISCOVERED (undocumented parts[].audioTranscription field in the generateContent response). Free tier, $0.
Sources: Audio transcription guide · Model card · discovery v1beta rev. 20260918 (GenerationConfig.audioTranscriptionConfig) · Interactions reference (transcription_config) · Pricing #gemini-3.5-transcribe.
Last verified: 2026-09-18. Twins: endpoints fragment (api_family transcription), parameters/gemini-speech.json (transcription rows), objects AudioTranscriptionConfig, GenerateContentResponse (transcription).
1. Two models, two transports
| Model | Live supportedGenerationMethods |
Transport | Max audio | Diarization | Word timestamps | Custom vocabulary | Smart mode | Price (paid) |
|---|---|---|---|---|---|---|---|---|
gemini-3.5-transcribe |
generateContent, countTokens (no batch, cache, tools, thinking) |
unary REST / Interactions | 1 h per request (30 min with diarization or timestamps) | ✔ (≤ 8 speakers, 3+ experimental) | ✔ | ✔ (≤ 1,000 terms, best ≤ 100) | ✔ | $2 /1M audio tokens (~$0.003/min) in, $12 /1M text out (~$0.002/min) → ~$0.005/min; free tier free |
gemini-3.5-transcribe-live |
bidiGenerateContent only |
WebSocket Live API (docs/gemini/live-api.md) |
10 min per session | ✘ | ✘ | ✔ | ✔ | see Live pricing |
Calling the live model with generateContent → 400 INVALID_ARGUMENT "models/gemini-3.5-transcribe-live only supports real-time bidirectional streaming via WebSocket (bidiGenerateContent). Please use the Gemini Live API…" (live). gemini-3.5-live-translate-preview belongs to the Live agent as well.
Both GA since the 2026-08 changelog entry; 85+ locales with utterance-level auto-detection and code-switching.
2. Request (generateContent)
POST /v1beta/models/gemini-3.5-transcribe:generateContent
{
"contents": [{"parts": [{"inlineData": {"mimeType": "audio/wav", "data": "<base64>"}}]}],
"generationConfig": {"audioTranscriptionConfig": {
"languageCodes": ["en-US"], "mode": "VERBATIM", "diarization": true, "wordTimestamp": true
}}
}generationConfig.audioTranscriptionConfig.* (discovery) |
Interactions generation_config.transcription_config.* |
Meaning |
|---|---|---|
languageCodes[] (also languageHints / languageAuto objects) |
language_codes[] |
BCP-47 hints; empty/omitted = auto-detect |
customVocabulary[] (alias adaptationPhrases[]) |
custom_vocabulary[] |
≤ 1,000 bias terms; incompatible with diarization/timestamps |
mode: VERBATIM (default) | SMART |
mode: "smart" or {type:"verbatim", diarization_mode:"speaker", timestamp_granularities:["word"]} |
SMART = disfluency removal, inline self-corrections, lists/dates/currency formatting; incompatible with diarization/timestamps |
diarization: true |
mode.diarization_mode: "speaker" |
speaker labels |
wordTimestamp: true |
mode.timestamp_granularities: ["word"] |
start/end offsets per word |
Accepted MIME types: audio/wav mp3 aiff aac ogg flac mpeg m4a l16 opus alaw mulaw webm. Use the Files API (fileData.fileUri) for anything longer than a few seconds. No text prompt is needed.
3. Response — observed shapes
| Input | candidates[0] |
|---|---|
| 0.5 s silence (16 kHz WAV) | {"content": {}, "finishReason": "STOP", "index": 0} — empty content, usage 13 audio tokens |
| 0.8 s TTS "OK" | {"content": {"parts": [{"audioTranscription": {"text": "Okay."}}], "role": "model"}, "finishReason": "STOP"} — 21 audio tokens |
same + diarization, wordTimestamp, languageCodes |
parts[0] = {"text": "Okay.", "audioTranscription": {"text": "Okay.", "speakerLabel": "spk:0", "words": [{"word": "Okay.", "startOffset": "0.200s", "endOffset": "0.500s"}]}} |
mode: SMART + customVocabulary + diarization |
400 INVALID_ARGUMENT "Transcription mode SMART is incompatible with diarization." |
LIVE_DISCOVERED: the audioTranscription part field (text, speakerLabel, words[] {word, startOffset, endOffset}) is not in the generate-content reference; the guide documents the Interactions shape instead (steps[].content[].annotations[] {type:"word_info", text, speaker:"spk_1", start_offset, end_offset} and interaction.output_text). Note the label formats differ (spk:0 vs spk_1). usageMetadata.promptTokensDetails[] = [{modality:"AUDIO", tokenCount}] at ≈ 25 tokens/s; no output tokens were billed for these tiny clips.
4. Live log (2026-09-18, $0)
transcribe-silence 200 · transcribe-smart-config 200 (empty) · transcribe-speech 200 "Okay." · transcribe-speech-diarization 200 · transcribe-incompatible 400 · transcribe-live-unary 400. Raw: tmp-live/gemini-media/transcribe-*.json. Examples: examples/gemini/transcription/ (transcribe.py LIVE_VERIFIED).