SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
8.6 KB

# Gemini image generation (Nano Banana) — and the end of Imagen

Status: DOCUMENTED · ACCOUNT_RESTRICTED (every image model returns 429 RESOURCE_EXHAUSTED … free_tier … limit: 0 with our key — image generation has no free tier; the request shape and error shapes are verified, a successful generation is not) · Imagen: RETIRED (404 NOT_FOUND … not supported for predict, live 2026-09-18). Sources: Image generation guide · Imagen (shut down) · GenerationConfig / ImageConfig reference · Pricing · Deprecations · model cards (gemini-3.1-flash-lite-image, gemini-3.1-flash-image, gemini-3-pro-image, gemini-2.5-flash-image) · discovery v1beta rev. 20260918 · live models.list (58 models). Last verified: 2026-09-18. Twins: generated/fragments/endpoints/gemini-media-batch-tuning.json (api_family image-generation), generated/fragments/parameters/gemini-image-generation.json, generated/fragments/objects/gemini-media-objects.json.

# 1. Two surfaces, one family

Surface Path How to ask for an image Notes
generateContent (this page) POST /v1beta/models/{model}:generateContent generationConfig.responseModalities: ["TEXT","IMAGE"] (or ["IMAGE"]), generationConfig.imageConfig {aspectRatio, imageSize} Also :streamGenerateContent, and :batchGenerateContent (50 % price, 24 h).
Interactions API POST /v1beta/interactions response_format: {type:"image", mime_type, aspect_ratio, image_size} or [{type:"text"},{type:"image"}]; previous_interaction_id for multi-turn The official guide now shows only Interactions samples; documented by the interactions agent (docs/gemini/interactions-api.md).
models.predict (Imagen) POST /v1beta/models/{model}:predict instances[{prompt}], parameters{sampleCount, aspectRatio, personGeneration, imageSize} Shut down. No Imagen id in models.list; live: 404 models/imagen-4.0-fast-generate-001 is not found for API version v1beta, or is not supported for predict.

"Nano Banana" = Gemini's native image generation. Imagen is described by Google as "legacy … shut down"; imagen-4.0-* ids were deprecated 2025-06-24 and shut down 2026-08-17 (migration target gemini-3.1-flash-image).

# 2. Models (live models.list, 2026-09-18)

Model id Marketing name Status supportedGenerationMethods In / out token limits Resolutions Reference images Search grounding Thinking
gemini-3.1-flash-lite-image Nano Banana 2 Lite stable (GA 2026-06) generateContent, countTokens, batchGenerateContent 65,536 / 65,536 (card: 4,096 out) 1K only up to 14 objects, no characters no MINIMAL / HIGH
gemini-3.1-flash-image (+ -preview, shut down 2026-06-25) Nano Banana 2 stable same 65,536 / 65,536 (card 131,072 in) 512 (0.5K), 1K, 2K, 4K 10 objects + 4 characters web and image search yes
gemini-3-pro-image (+ -preview, nano-banana-pro-preview) Nano Banana Pro stable same 131,072 / 32,768 1K, 2K, 4K 6 objects + 5 characters + 3 style web search yes (default)
gemini-2.5-flash-image Nano Banana (legacy) DEPRECATED 2025-10-02, shutdown 2026-10-02 same 32,768 / 32,768 ~1024 px, no imageSize best ≤ 3 inputs no no

Inputs: text, image (all); video and PDF only on the 3.1 Flash / Flash-Lite image models; audio never. Outputs: image + text. Every image carries a SynthID watermark (3.1 Lite also C2PA).

# 3. Request (generateContent)

json
POST /v1beta/models/gemini-3.1-flash-lite-image:generateContent
{
  "contents": [{"parts": [
    {"text": "Using the provided image of my cat, add a small knitted wizard hat."},
    {"inlineData": {"mimeType": "image/png", "data": "<base64>"}}
  ]}],
  "generationConfig": {
    "responseModalities": ["TEXT", "IMAGE"],
    "imageConfig": {"aspectRatio": "1:1", "imageSize": "1K"},
    "thinkingConfig": {"thinkingLevel": "MINIMAL"}
  }
}
Parameter Type / values Rules
generationConfig.responseModalities ["TEXT","IMAGE"] default behaviour · ["IMAGE"] images only Exact match to what comes back.
generationConfig.imageConfig.aspectRatio 1:1 1:4 4:1 1:8 8:1 2:3 3:2 3:4 4:3 4:5 5:4 9:16 16:9 21:9 1:4/4:1/1:8/8:1 only on 3.1 Flash Image. Default = input image ratio, else 1:1. Setting it on a text model → 400 INVALID_ARGUMENT "Aspect ratio is not enabled for this model" (live).
generationConfig.imageConfig.imageSize 512, 1K, 2K, 4K (uppercase K) Default 1K. Lite: 1K only. 512 only on 3.1 Flash Image.
generationConfig.thinkingConfig.thinkingLevel MINIMAL (default) / HIGH 3.1 Flash / Flash-Lite Image. Thinking cannot be disabled; interim "thought images" are not billed, thinking tokens are.
generationConfig.responseFormat.image {mimeType: IMAGE_JPEG…, delivery: INLINE|URI, aspectRatio, imageSize} Newer flat per-modality format in the reference (ResponseFormatConfig); UNVERIFIED.
contents[].parts[].inlineData / fileData images (≤ 14 references on Gemini 3), video (3.1 only) See table §2 for object / character / style budgets.
tools: [{googleSearch: {}}] grounding Not on Lite. Image Search (search_types: ["image_search"], Interactions) only on 3.1 Flash Image; display search_suggestions.

Multi-turn editing = resend the conversation (user text → model image part → new user instruction). Interactions API does the same with previous_interaction_id.

# 4. Response

candidates[0].content.parts[] mixes {text} and {inlineData:{mimeType:"image/png", data}} parts (thought images carry thought: true). usageMetadata.candidatesTokensDetails[] reports {modality:"IMAGE", tokenCount}; finishReason may be IMAGE_SAFETY. The model does not always honour a requested number of images.

Output token cost per image (from the guide's tables): 3.1 Flash Image 0.5K → 747, 1K → 1120, 2K → 1680, 4K → 2520; 3 Pro Image 1K/2K → 1120, 4K → 2000; 2.5 Flash Image → 1290.

# 5. Pricing (paid tier, USD; image output billed as tokens)

Model Input /1M Text+thinking out /1M Image out /1M tokens ≈ per image Batch
gemini-3.1-flash-lite-image $0.25 $1.50 $30 $0.0336 (1K) $0.125 / $0.75 / $15 → $0.0168
gemini-3.1-flash-image $0.50 $3 $60 $0.045 (0.5K), $0.067 (1K), $0.101 (2K), $0.151 (4K) half
gemini-3-pro-image $2 (image input 560 tok = $0.0011) $12 $120 $0.134 (1K/2K), $0.24 (4K) $1 / $6 / $0.067–0.12 per image; Flex same as batch; Priority ×1.8
gemini-2.5-flash-image $0.30 — $30 $0.039 (1290 tok) $0.15 / $0.0195

Free tier: "Not available" for all image models — which is exactly what our key hit (limit: 0 on generate_content_free_tier_requests). Search grounding: 5,000 free requests/month shared across Gemini 3.x, then $14 / 1,000.

# 6. Live verification log (2026-09-18, all free)

Call Result
gemini-3.1-flash-lite-image:generateContent "a plain white square", 1K 1:1 (×2) 429 RESOURCE_EXHAUSTED — Quota exceeded for metric: generate_content_free_tier_requests, limit: 0, model: gemini-3.1-flash-lite-image → ACCOUNT_RESTRICTED (no free tier)
same on gemini-2.5-flash-image 429 identical
imageSize: "1k" (lowercase), imageSize: "4K" on Lite 429 (quota check precedes validation — validation rules above are documented, not observed)
imageConfig.aspectRatio on gemini-3.5-flash-lite 400 INVALID_ARGUMENT "Aspect ratio is not enabled for this model"
imagen-4.0-fast-generate-001:predict, imagen-4.0-generate-001:predict, gemini-3.5-flash-lite:predict 404 NOT_FOUND "models/… is not found for API version v1beta, or is not supported for predict"

Raw bodies: tmp-live/gemini-media/image-*.json, predict-*.json. Examples in examples/gemini/image-generation/ are therefore UNVERIFIED (request shape correct, blocked by plan).

# 7. Limitations (documented)

Best languages: EN, ar-EG, de-DE, es-MX, fr-FR, hi-IN, id-ID, it-IT, ja-JP, ko-KR, pt-BR, ru-RU, ua-UA, vi-VN, zh-CN. Generate text first, then ask for the image containing it. 3.1 Flash Image search grounding does not return real people's photos. Semantic negative prompts ("an empty street") rather than "no cars".