SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
4.7 KB

# Gemini multimodal input — images, video, audio, PDF, YouTube; token math and limits

Status: DOCUMENTED + LIVE_VERIFIED for inline PNG, inline/uploaded PDF, inline text blob, per-part/global mediaResolution, YouTube URL and HTTPS image URL token counting (2026-09-18, gemini-3.5-flash-lite). Video/audio generation calls were not made (token cost) — counted only. Sources: https://ai.google.dev/gemini-api/docs/generate-content/image-understanding · …/video-understanding · …/audio · …/document-processing · …/file-input-methods · …/media-resolution · …/tokens · https://ai.google.dev/gemini-api/docs/embeddings#multimodal Machine-readable: generated/fragments/parameters/gemini-generate-content.json (parts.*), generated/fragments/compatibility/gemini-feature-model-matrix.json#media_token_rules Last verified: 2026-09-18

# 1. Input methods

Method Shape Limit Persistence
Inline parts[].inlineData {mimeType, data(base64)} 20 MB total request (image/audio/video guides); file-input-methods says 100 MB (50 MB PDF) none
Files API parts[].fileData {fileUri: File.uri, mimeType?} 2 GB/file, 20 GB/project 48 h
GCS registration files:register → fileData 2 GB/file, no cap up to 30 days
External URL fileData.fileUri: https://… (public or pre-signed) 100 MB per payload fetched per request (8 s observed for a 400 KB JPEG)
YouTube fileData.fileUri: https://www.youtube.com/watch?v=… (public; preview, free) free tier ≤ 8 h video/day; ≤ 10 videos/request (2.5+) —

Supported MIME families (discovery Blob.mimeType): images image/png, jpeg, jpg, webp, heic, heif, gif, avif; audio/*; video/*; text text/plain, html, css, javascript, x-typescript, csv, markdown, x-python, xml, rtf; application/pdf, json, x-javascript, x-typescript, x-python-code, x-ipynb+json, rtf. The declared MIME is not checked against the bytes (PNG as image/bmp → 200).

# 2. Token math

Modality Documented rule Observed (gemini-3.5-flash-lite)
Image (Gemini 3) mediaResolution: LOW 280 · MEDIUM 560 · HIGH/default 1120 · ULTRA_HIGH 2240 (per-part only) 1×1 PNG: default 1089 · LOW 256 · MEDIUM 529 · HIGH 1089 · ULTRA_HIGH 2209 (IMAGE modality); HTTPS JPEG 1080
Image (2.x rule) ≤ 384 px both sides = 258 tokens; else 768×768 tiles × 258; ≤ 3,600 images/request embedding-2 image = 258
Video (static) 263 tok/s (tokens guide) ≈ 100/s low-res, 300/s high-res at 1 fps; Gemini 3 per frame: 70 (LOW/MEDIUM) / 280 (HIGH); videoMetadata.fps scales it; agentic mode loads on demand (up to −88 %) YouTube 9hE5-98ZeCg: VIDEO 10650 + AUDIO 4777 (total 15431); 10 s clip: VIDEO 710 + AUDIO 320
Audio 32 tok/s (tokens guide) / 25 tok/s (media-resolution table, fixed across resolutions); ≤ 9.5 h per prompt; downsampled to 16 kbps mono (counted inside video above)
PDF ≤ 50 MB / 1,000 pages; pages scaled to ≤ 3072×3072, ≥ 768×768; Gemini 3: 280/560/1120 tokens per page by resolution + native text (not billed); page tokens reported under IMAGE (docs) 1-page PDF: generateContent IMAGE 520, countTokens DOCUMENT 560
Text blob counted as text text/plain inline → TEXT 17

usageMetadata.promptTokensDetails[] gives the per-modality split; agentic video splits across promptTokenCount, thoughtsTokenCount, toolUsePromptTokenCount, candidatesTokenCount.

# 3. Controls

  • generationConfig.mediaResolution (global; MEDIA_RESOLUTION_LOW|MEDIUM|HIGH; ULTRA_HIGH → 400 here) and per-part parts[].mediaResolution.level (Gemini 3; adds ULTRA_HIGH). Recommended: images HIGH, PDFs MEDIUM, video LOW (HIGH for on-screen text), audio default.
  • parts[].videoMetadata {startOffset, endOffset, fps} — clipping/sampling (static mode; also for YouTube).
  • parts[].mediaProcessing: STATIC | AGENTIC — agentic on 3.8/3.7/3.6 Flash and 3.5 Flash-Lite (not probed).
  • Timestamps in prompts: MM:SS references work for video/audio (docs).
  • Put a single image/video before the text prompt for best results (docs).

# 4. YouTube specifics

Public videos only (bogus id → 400 Request contains an invalid argument.); preview, no charge for the URL fetch but video tokens are billed normally; ≥ 2.5: up to 10 videos per request; clipping via videoMetadata. Counting with countTokens is free — use it to size requests (verified).

# 5. Bounding boxes / segmentation

Image understanding returns normalized [ymin, xmin, ymax, xmax] in 0–1000 (docs); segmentation is not supported on Gemini 3.x (use 2.5 Flash with thinking off).

Examples: examples/gemini/multimodal/ (inline 1×1 PNG + inline PDF). Tests: tests/gemini/test_count_tokens.py (image/PDF/YouTube counts).