SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
5.2 KB

# Gemini transcription — gemini-3.5-transcribe (unary) and gemini-3.5-transcribe-live (Live API)

Status: DOCUMENTED · LIVE_VERIFIED 2026-09-18 (generateContent with inline WAV: speech, silence, diarization + word timestamps, SMART mode, incompatibility error) · LIVE_DISCOVERED (undocumented parts[].audioTranscription field in the generateContent response). Free tier, $0. Sources: Audio transcription guide · Model card · discovery v1beta rev. 20260918 (GenerationConfig.audioTranscriptionConfig) · Interactions reference (transcription_config) · Pricing #gemini-3.5-transcribe. Last verified: 2026-09-18. Twins: endpoints fragment (api_family transcription), parameters/gemini-speech.json (transcription rows), objects AudioTranscriptionConfig, GenerateContentResponse (transcription).

# 1. Two models, two transports

Model Live supportedGenerationMethods Transport Max audio Diarization Word timestamps Custom vocabulary Smart mode Price (paid)
gemini-3.5-transcribe generateContent, countTokens (no batch, cache, tools, thinking) unary REST / Interactions 1 h per request (30 min with diarization or timestamps) ✔ (≤ 8 speakers, 3+ experimental) ✔ ✔ (≤ 1,000 terms, best ≤ 100) ✔ $2 /1M audio tokens (~$0.003/min) in, $12 /1M text out (~$0.002/min) → ~$0.005/min; free tier free
gemini-3.5-transcribe-live bidiGenerateContent only WebSocket Live API (docs/gemini/live-api.md) 10 min per session ✘ ✘ ✔ ✔ see Live pricing

Calling the live model with generateContent → 400 INVALID_ARGUMENT "models/gemini-3.5-transcribe-live only supports real-time bidirectional streaming via WebSocket (bidiGenerateContent). Please use the Gemini Live API…" (live). gemini-3.5-live-translate-preview belongs to the Live agent as well.

Both GA since the 2026-08 changelog entry; 85+ locales with utterance-level auto-detection and code-switching.

# 2. Request (generateContent)

json
POST /v1beta/models/gemini-3.5-transcribe:generateContent
{
  "contents": [{"parts": [{"inlineData": {"mimeType": "audio/wav", "data": "<base64>"}}]}],
  "generationConfig": {"audioTranscriptionConfig": {
    "languageCodes": ["en-US"], "mode": "VERBATIM", "diarization": true, "wordTimestamp": true
  }}
}
generationConfig.audioTranscriptionConfig.* (discovery) Interactions generation_config.transcription_config.* Meaning
languageCodes[] (also languageHints / languageAuto objects) language_codes[] BCP-47 hints; empty/omitted = auto-detect
customVocabulary[] (alias adaptationPhrases[]) custom_vocabulary[] ≤ 1,000 bias terms; incompatible with diarization/timestamps
mode: VERBATIM (default) | SMART mode: "smart" or {type:"verbatim", diarization_mode:"speaker", timestamp_granularities:["word"]} SMART = disfluency removal, inline self-corrections, lists/dates/currency formatting; incompatible with diarization/timestamps
diarization: true mode.diarization_mode: "speaker" speaker labels
wordTimestamp: true mode.timestamp_granularities: ["word"] start/end offsets per word

Accepted MIME types: audio/wav mp3 aiff aac ogg flac mpeg m4a l16 opus alaw mulaw webm. Use the Files API (fileData.fileUri) for anything longer than a few seconds. No text prompt is needed.

# 3. Response — observed shapes

Input candidates[0]
0.5 s silence (16 kHz WAV) {"content": {}, "finishReason": "STOP", "index": 0} — empty content, usage 13 audio tokens
0.8 s TTS "OK" {"content": {"parts": [{"audioTranscription": {"text": "Okay."}}], "role": "model"}, "finishReason": "STOP"} — 21 audio tokens
same + diarization, wordTimestamp, languageCodes parts[0] = {"text": "Okay.", "audioTranscription": {"text": "Okay.", "speakerLabel": "spk:0", "words": [{"word": "Okay.", "startOffset": "0.200s", "endOffset": "0.500s"}]}}
mode: SMART + customVocabulary + diarization 400 INVALID_ARGUMENT "Transcription mode SMART is incompatible with diarization."

LIVE_DISCOVERED: the audioTranscription part field (text, speakerLabel, words[] {word, startOffset, endOffset}) is not in the generate-content reference; the guide documents the Interactions shape instead (steps[].content[].annotations[] {type:"word_info", text, speaker:"spk_1", start_offset, end_offset} and interaction.output_text). Note the label formats differ (spk:0 vs spk_1). usageMetadata.promptTokensDetails[] = [{modality:"AUDIO", tokenCount}] at ≈ 25 tokens/s; no output tokens were billed for these tiny clips.

# 4. Live log (2026-09-18, $0)

transcribe-silence 200 · transcribe-smart-config 200 (empty) · transcribe-speech 200 "Okay." · transcribe-speech-diarization 200 · transcribe-incompatible 400 · transcribe-live-unary 400. Raw: tmp-live/gemini-media/transcribe-*.json. Examples: examples/gemini/transcription/ (transcribe.py LIVE_VERIFIED).