SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
6.2 KB

# Gemini music generation — Lyria 3 / 3.5 (single-turn) and Lyria RealTime (WebSocket)

Status: Lyria 3 / 3.5 via generateContent: DOCUMENTED · PREVIEW · ACCOUNT_RESTRICTED (429 … free_tier_requests, limit: 0, model: lyria-3-clip — no free tier). Lyria RealTime (lyria-realtime-exp, WebSocket BidiGenerateMusic): DOCUMENTED · BETA (experimental) · LIVE_VERIFIED 2026-09-18 (setup → setupComplete → 2 audio chunks → STOP). Sources: Music generation (Lyria 3.5) · Realtime music generation · Lyria prompt guide · model cards lyria-3-clip-preview, lyria-3.5-*, lyria-realtime-exp · Pricing · python-genai LiveMusic* types + live_music.py (WebSocket URL). Last verified: 2026-09-18. Twins: endpoints fragment (api_family music-generation), parameters/gemini-music.json, streaming-events/gemini-lyria-realtime.json, objects LiveMusicClientMessage / LiveMusicServerMessage.

# 1. Models

Model id Live supportedGenerationMethods Output Duration Price (paid) Status
lyria-3-clip-preview generateContent, countTokens MP3 (card: 48 kHz; guide: 44.1 kHz) stereo + lyrics/structure text always 30 s $0.04 / request Preview (2026-03-25), "legacy"
lyria-3-pro-preview same MP3 + text a couple of minutes $0.08 Preview, legacy → migrate to lyria-3.5
lyria-3.5 same MP3 (WAV via Interactions response_format {type:"audio"}) minutes, steerable by prompt/timestamps $0.08 / song GA 2026-09-03
lyria-3.5-clip-preview, lyria-3.5-pro-preview model cards exist; not in live models.list DOCUMENTED only
lyria-realtime-exp bidiGenerateMusic raw 16-bit PCM 48 kHz stereo, streamed continuous not on pricing page Experimental (2025-05)

None of the Lyria models supports Batch, caching, tools or thinking. Inputs: text (+ up to 10 images for 3.5). Free tier: not available (hence our 429s).

# 2. Lyria 3 / 3.5 — request and response

Official samples use the Interactions API (POST /v1beta/interactions {model, input}; interaction.output_audio / output_text). The models also expose generateContent:

json
POST /v1beta/models/lyria-3-clip-preview:generateContent
{"contents": [{"parts": [{"text": "A short instrumental acoustic guitar piece."}]}]}

Expected (documented) response: candidates[0].content.parts[] = text part(s) with the generated lyrics / JSON song structure, and an inlineData audio part (base64 MP3). Prompting: mention genre, instruments, BPM, key, mood; section tags [Verse] [Chorus] [Bridge]; timestamps [0:00 - 0:10] Intro: …; custom lyrics; prompt language = lyrics language; "Instrumental only, no vocals" for instrumental tracks. Single-turn only (no multi-turn editing), non-deterministic, SynthID audio watermark, safety filters block artist voices and copyrighted lyrics.

Live 2026-09-18: lyria-3-clip-preview and lyria-3.5 → 429 RESOURCE_EXHAUSTED (limit: 0), so the response shape above is DOCUMENTED, not observed.

# 3. Lyria RealTime — protocol (LIVE_VERIFIED)

Endpoint: wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateMusic with header x-goog-api-key (URL from google-genai live_music.py; v1beta worked, no query-string key needed). JSON text frames.

# Direction Frame Notes
1 client → server {"setup":{"model":"models/lyria-realtime-exp"}} first frame only
2 server → client {"setupComplete":{}} wait for it before anything else (observed)
3 client → server {"clientContent":{"weightedPrompts":[{"text":"minimal techno","weight":1.0}]}} weight ≠ 0, normalized; resend any time to steer
4 client → server {"musicGenerationConfig":{"bpm":120,"temperature":1.0, …}} send the whole config each time — omitted fields reset to defaults; bpm/scale need RESET_CONTEXT
5 client → server {"playbackControl":"PLAY"} PLAY · PAUSE · STOP (resets context, keeps prompts) · RESET_CONTEXT
6… server → client {"serverContent":{"audioChunks":[{"data":"<base64>","mimeType":"audio/l16;rate=48000;channels=2","sourceMetadata":{"clientContent":{…},"musicGenerationConfig":{"temperature":1,"seed":603898271,"bpm":120}}}]}} observed: 1 chunk per frame, 384,000 bytes = exactly 2.0 s of 48 kHz stereo PCM16; seed increments per chunk; generated faster than real time → client must buffer
any server → client {"filteredPrompt":{"text":"…","filteredReason":"…"}} prompt dropped by safety filters

musicGenerationConfig fields: guidance 0–6 (4.0), bpm 60–200, density 0–1, brightness 0–1, scale (12 relative-key enums + SCALE_UNSPECIFIED), muteBass, muteDrums, onlyBassAndDrums, musicGenerationMode QUALITY (default) | DIVERSITY | VOCALIZATION, temperature 0–3 (1.1), topK 1–1000 (40), seed 0–2³¹-1. Unset bpm/density/brightness/scale → model decides from the prompts. Instrumental only; control latency ≤ 2 s.

SDK: async with client.aio.live.music.connect(model="models/lyria-realtime-exp") as s: await s.set_weighted_prompts(...); await s.set_music_generation_config(...); await s.play(); async for m in s.receive(): m.server_content.audio_chunks[0].data. Node: ai.live.music.connect({model, callbacks}).

# 4. Live log (2026-09-18)

Call Result
POST …lyria-3-clip-preview:generateContent (acoustic guitar prompt) 429 limit: 0, model: lyria-3-clip → ACCOUNT_RESTRICTED
POST …lyria-3.5:generateContent 429 identical
WebSocket BidiGenerateMusic v1beta 101; frames: setupComplete → 2 × serverContent.audioChunks (2 s each) → STOP sent; raw tmp-live/gemini-media/lyria-realtime.json

Cost: $0 (RealTime has no published price; a few seconds were generated). Examples: examples/gemini/music/ — the RealTime Python example is LIVE_VERIFIED, the Lyria 3 ones UNVERIFIED/blocked by plan.