Gemini music generation — Lyria 3 / 3.5 (single-turn) and Lyria RealTime (WebSocket)
Status: Lyria 3 / 3.5 via generateContent: DOCUMENTED · PREVIEW · ACCOUNT_RESTRICTED (429 … free_tier_requests, limit: 0, model: lyria-3-clip — no free tier). Lyria RealTime (lyria-realtime-exp, WebSocket BidiGenerateMusic): DOCUMENTED · BETA (experimental) · LIVE_VERIFIED 2026-09-18 (setup → setupComplete → 2 audio chunks → STOP).
Sources: Music generation (Lyria 3.5) · Realtime music generation · Lyria prompt guide · model cards lyria-3-clip-preview, lyria-3.5-*, lyria-realtime-exp · Pricing · python-genai LiveMusic* types + live_music.py (WebSocket URL).
Last verified: 2026-09-18. Twins: endpoints fragment (api_family music-generation), parameters/gemini-music.json, streaming-events/gemini-lyria-realtime.json, objects LiveMusicClientMessage / LiveMusicServerMessage.
1. Models
| Model id | Live supportedGenerationMethods |
Output | Duration | Price (paid) | Status |
|---|---|---|---|---|---|
lyria-3-clip-preview |
generateContent, countTokens |
MP3 (card: 48 kHz; guide: 44.1 kHz) stereo + lyrics/structure text | always 30 s | $0.04 / request | Preview (2026-03-25), "legacy" |
lyria-3-pro-preview |
same | MP3 + text | a couple of minutes | $0.08 | Preview, legacy → migrate to lyria-3.5 |
lyria-3.5 |
same | MP3 (WAV via Interactions response_format {type:"audio"}) |
minutes, steerable by prompt/timestamps | $0.08 / song | GA 2026-09-03 |
lyria-3.5-clip-preview, lyria-3.5-pro-preview |
model cards exist; not in live models.list |
DOCUMENTED only |
|||
lyria-realtime-exp |
bidiGenerateMusic |
raw 16-bit PCM 48 kHz stereo, streamed | continuous | not on pricing page | Experimental (2025-05) |
None of the Lyria models supports Batch, caching, tools or thinking. Inputs: text (+ up to 10 images for 3.5). Free tier: not available (hence our 429s).
2. Lyria 3 / 3.5 — request and response
Official samples use the Interactions API (POST /v1beta/interactions {model, input}; interaction.output_audio / output_text). The models also expose generateContent:
POST /v1beta/models/lyria-3-clip-preview:generateContent
{"contents": [{"parts": [{"text": "A short instrumental acoustic guitar piece."}]}]}Expected (documented) response: candidates[0].content.parts[] = text part(s) with the generated lyrics / JSON song structure, and an inlineData audio part (base64 MP3). Prompting: mention genre, instruments, BPM, key, mood; section tags [Verse] [Chorus] [Bridge]; timestamps [0:00 - 0:10] Intro: …; custom lyrics; prompt language = lyrics language; "Instrumental only, no vocals" for instrumental tracks. Single-turn only (no multi-turn editing), non-deterministic, SynthID audio watermark, safety filters block artist voices and copyrighted lyrics.
Live 2026-09-18: lyria-3-clip-preview and lyria-3.5 → 429 RESOURCE_EXHAUSTED (limit: 0), so the response shape above is DOCUMENTED, not observed.
3. Lyria RealTime — protocol (LIVE_VERIFIED)
Endpoint: wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateMusic with header x-goog-api-key (URL from google-genai live_music.py; v1beta worked, no query-string key needed). JSON text frames.
| # | Direction | Frame | Notes |
|---|---|---|---|
| 1 | client → server | {"setup":{"model":"models/lyria-realtime-exp"}} |
first frame only |
| 2 | server → client | {"setupComplete":{}} |
wait for it before anything else (observed) |
| 3 | client → server | {"clientContent":{"weightedPrompts":[{"text":"minimal techno","weight":1.0}]}} |
weight ≠ 0, normalized; resend any time to steer |
| 4 | client → server | {"musicGenerationConfig":{"bpm":120,"temperature":1.0, …}} |
send the whole config each time — omitted fields reset to defaults; bpm/scale need RESET_CONTEXT |
| 5 | client → server | {"playbackControl":"PLAY"} |
PLAY · PAUSE · STOP (resets context, keeps prompts) · RESET_CONTEXT |
| 6… | server → client | {"serverContent":{"audioChunks":[{"data":"<base64>","mimeType":"audio/l16;rate=48000;channels=2","sourceMetadata":{"clientContent":{…},"musicGenerationConfig":{"temperature":1,"seed":603898271,"bpm":120}}}]}} |
observed: 1 chunk per frame, 384,000 bytes = exactly 2.0 s of 48 kHz stereo PCM16; seed increments per chunk; generated faster than real time → client must buffer |
| any | server → client | {"filteredPrompt":{"text":"…","filteredReason":"…"}} |
prompt dropped by safety filters |
musicGenerationConfig fields: guidance 0–6 (4.0), bpm 60–200, density 0–1, brightness 0–1, scale (12 relative-key enums + SCALE_UNSPECIFIED), muteBass, muteDrums, onlyBassAndDrums, musicGenerationMode QUALITY (default) | DIVERSITY | VOCALIZATION, temperature 0–3 (1.1), topK 1–1000 (40), seed 0–2³¹-1. Unset bpm/density/brightness/scale → model decides from the prompts. Instrumental only; control latency ≤ 2 s.
SDK: async with client.aio.live.music.connect(model="models/lyria-realtime-exp") as s: await s.set_weighted_prompts(...); await s.set_music_generation_config(...); await s.play(); async for m in s.receive(): m.server_content.audio_chunks[0].data. Node: ai.live.music.connect({model, callbacks}).
4. Live log (2026-09-18)
| Call | Result |
|---|---|
POST …lyria-3-clip-preview:generateContent (acoustic guitar prompt) |
429 limit: 0, model: lyria-3-clip → ACCOUNT_RESTRICTED |
POST …lyria-3.5:generateContent |
429 identical |
WebSocket BidiGenerateMusic v1beta |
101; frames: setupComplete → 2 × serverContent.audioChunks (2 s each) → STOP sent; raw tmp-live/gemini-media/lyria-realtime.json |
Cost: $0 (RealTime has no published price; a few seconds were generated). Examples: examples/gemini/music/ — the RealTime Python example is LIVE_VERIFIED, the Lyria 3 ones UNVERIFIED/blocked by plan.