Gemini multimodal input — images, video, audio, PDF, YouTube; token math and limits
Status: DOCUMENTED + LIVE_VERIFIED for inline PNG, inline/uploaded PDF, inline text blob, per-part/global mediaResolution, YouTube URL and HTTPS image URL token counting (2026-09-18, gemini-3.5-flash-lite). Video/audio generation calls were not made (token cost) — counted only.
Sources: https://ai.google.dev/gemini-api/docs/generate-content/image-understanding · …/video-understanding · …/audio · …/document-processing · …/file-input-methods · …/media-resolution · …/tokens · https://ai.google.dev/gemini-api/docs/embeddings#multimodal
Machine-readable: generated/fragments/parameters/gemini-generate-content.json (parts.*), generated/fragments/compatibility/gemini-feature-model-matrix.json#media_token_rules
Last verified: 2026-09-18
1. Input methods
| Method | Shape | Limit | Persistence |
|---|---|---|---|
| Inline | parts[].inlineData {mimeType, data(base64)} |
20 MB total request (image/audio/video guides); file-input-methods says 100 MB (50 MB PDF) | none |
| Files API | parts[].fileData {fileUri: File.uri, mimeType?} |
2 GB/file, 20 GB/project | 48 h |
| GCS registration | files:register → fileData |
2 GB/file, no cap | up to 30 days |
| External URL | fileData.fileUri: https://… (public or pre-signed) |
100 MB per payload | fetched per request (8 s observed for a 400 KB JPEG) |
| YouTube | fileData.fileUri: https://www.youtube.com/watch?v=… (public; preview, free) |
free tier ≤ 8 h video/day; ≤ 10 videos/request (2.5+) | — |
Supported MIME families (discovery Blob.mimeType): images image/png, jpeg, jpg, webp, heic, heif, gif, avif; audio/*; video/*; text text/plain, html, css, javascript, x-typescript, csv, markdown, x-python, xml, rtf; application/pdf, json, x-javascript, x-typescript, x-python-code, x-ipynb+json, rtf. The declared MIME is not checked against the bytes (PNG as image/bmp → 200).
2. Token math
| Modality | Documented rule | Observed (gemini-3.5-flash-lite) |
|---|---|---|
| Image (Gemini 3) | mediaResolution: LOW 280 · MEDIUM 560 · HIGH/default 1120 · ULTRA_HIGH 2240 (per-part only) |
1×1 PNG: default 1089 · LOW 256 · MEDIUM 529 · HIGH 1089 · ULTRA_HIGH 2209 (IMAGE modality); HTTPS JPEG 1080 |
| Image (2.x rule) | ≤ 384 px both sides = 258 tokens; else 768×768 tiles × 258; ≤ 3,600 images/request | embedding-2 image = 258 |
| Video (static) | 263 tok/s (tokens guide) ≈ 100/s low-res, 300/s high-res at 1 fps; Gemini 3 per frame: 70 (LOW/MEDIUM) / 280 (HIGH); videoMetadata.fps scales it; agentic mode loads on demand (up to −88 %) |
YouTube 9hE5-98ZeCg: VIDEO 10650 + AUDIO 4777 (total 15431); 10 s clip: VIDEO 710 + AUDIO 320 |
| Audio | 32 tok/s (tokens guide) / 25 tok/s (media-resolution table, fixed across resolutions); ≤ 9.5 h per prompt; downsampled to 16 kbps mono | (counted inside video above) |
≤ 50 MB / 1,000 pages; pages scaled to ≤ 3072×3072, ≥ 768×768; Gemini 3: 280/560/1120 tokens per page by resolution + native text (not billed); page tokens reported under IMAGE (docs) |
1-page PDF: generateContent IMAGE 520, countTokens DOCUMENT 560 |
|
| Text blob | counted as text | text/plain inline → TEXT 17 |
usageMetadata.promptTokensDetails[] gives the per-modality split; agentic video splits across promptTokenCount, thoughtsTokenCount, toolUsePromptTokenCount, candidatesTokenCount.
3. Controls
generationConfig.mediaResolution(global;MEDIA_RESOLUTION_LOW|MEDIUM|HIGH;ULTRA_HIGH→ 400 here) and per-partparts[].mediaResolution.level(Gemini 3; addsULTRA_HIGH). Recommended: images HIGH, PDFs MEDIUM, video LOW (HIGH for on-screen text), audio default.parts[].videoMetadata {startOffset, endOffset, fps}— clipping/sampling (static mode; also for YouTube).parts[].mediaProcessing: STATIC | AGENTIC— agentic on 3.8/3.7/3.6 Flash and 3.5 Flash-Lite (not probed).- Timestamps in prompts:
MM:SSreferences work for video/audio (docs). - Put a single image/video before the text prompt for best results (docs).
4. YouTube specifics
Public videos only (bogus id → 400 Request contains an invalid argument.); preview, no charge for the URL fetch but video tokens are billed normally; ≥ 2.5: up to 10 videos per request; clipping via videoMetadata. Counting with countTokens is free — use it to size requests (verified).
5. Bounding boxes / segmentation
Image understanding returns normalized [ymin, xmin, ymax, xmax] in 0–1000 (docs); segmentation is not supported on Gemini 3.x (use 2.5 Flash with thinking off).
Examples: examples/gemini/multimodal/ (inline 1×1 PNG + inline PDF). Tests: tests/gemini/test_count_tokens.py (image/PDF/YouTube counts).