# Gemini multimodal input — images, video, audio, PDF, YouTube; token math and limits **Status:** `DOCUMENTED` + `LIVE_VERIFIED` for inline PNG, inline/uploaded PDF, inline text blob, per-part/global `mediaResolution`, YouTube URL and HTTPS image URL token counting (2026-09-18, `gemini-3.5-flash-lite`). Video/audio generation calls were **not** made (token cost) — counted only. **Sources:** https://ai.google.dev/gemini-api/docs/generate-content/image-understanding · …/video-understanding · …/audio · …/document-processing · …/file-input-methods · …/media-resolution · …/tokens · https://ai.google.dev/gemini-api/docs/embeddings#multimodal **Machine-readable:** `generated/fragments/parameters/gemini-generate-content.json` (parts.*), `generated/fragments/compatibility/gemini-feature-model-matrix.json#media_token_rules` **Last verified:** 2026-09-18 ## 1. Input methods | Method | Shape | Limit | Persistence | |---|---|---|---| | Inline | `parts[].inlineData {mimeType, data(base64)}` | 20 MB total request (image/audio/video guides); file-input-methods says 100 MB (50 MB PDF) | none | | Files API | `parts[].fileData {fileUri: File.uri, mimeType?}` | 2 GB/file, 20 GB/project | 48 h | | GCS registration | `files:register` → `fileData` | 2 GB/file, no cap | up to 30 days | | External URL | `fileData.fileUri: https://…` (public or pre-signed) | 100 MB per payload | fetched per request (8 s observed for a 400 KB JPEG) | | YouTube | `fileData.fileUri: https://www.youtube.com/watch?v=…` (public; preview, free) | free tier ≤ 8 h video/day; ≤ 10 videos/request (2.5+) | — | Supported MIME families (discovery `Blob.mimeType`): images `image/png, jpeg, jpg, webp, heic, heif, gif, avif`; `audio/*`; `video/*`; text `text/plain, html, css, javascript, x-typescript, csv, markdown, x-python, xml, rtf`; `application/pdf, json, x-javascript, x-typescript, x-python-code, x-ipynb+json, rtf`. The declared MIME is not checked against the bytes (PNG as `image/bmp` → 200). ## 2. Token math | Modality | Documented rule | Observed (gemini-3.5-flash-lite) | |---|---|---| | Image (Gemini 3) | `mediaResolution`: LOW 280 · MEDIUM 560 · HIGH/default 1120 · ULTRA_HIGH 2240 (per-part only) | 1×1 PNG: default 1089 · LOW 256 · MEDIUM 529 · HIGH 1089 · ULTRA_HIGH 2209 (IMAGE modality); HTTPS JPEG 1080 | | Image (2.x rule) | ≤ 384 px both sides = 258 tokens; else 768×768 tiles × 258; ≤ 3,600 images/request | embedding-2 image = 258 | | Video (static) | 263 tok/s (tokens guide) ≈ 100/s low-res, 300/s high-res at 1 fps; Gemini 3 per frame: 70 (LOW/MEDIUM) / 280 (HIGH); `videoMetadata.fps` scales it; agentic mode loads on demand (up to −88 %) | YouTube 9hE5-98ZeCg: VIDEO 10650 + AUDIO 4777 (total 15431); 10 s clip: VIDEO 710 + AUDIO 320 | | Audio | 32 tok/s (tokens guide) / 25 tok/s (media-resolution table, fixed across resolutions); ≤ 9.5 h per prompt; downsampled to 16 kbps mono | (counted inside video above) | | PDF | ≤ 50 MB / 1,000 pages; pages scaled to ≤ 3072×3072, ≥ 768×768; Gemini 3: 280/560/1120 tokens per page by resolution + native text (**not billed**); page tokens reported under `IMAGE` (docs) | 1-page PDF: generateContent `IMAGE 520`, countTokens `DOCUMENT 560` | | Text blob | counted as text | `text/plain` inline → TEXT 17 | `usageMetadata.promptTokensDetails[]` gives the per-modality split; agentic video splits across `promptTokenCount`, `thoughtsTokenCount`, `toolUsePromptTokenCount`, `candidatesTokenCount`. ## 3. Controls - `generationConfig.mediaResolution` (global; `MEDIA_RESOLUTION_LOW|MEDIUM|HIGH`; `ULTRA_HIGH` → 400 here) and per-part `parts[].mediaResolution.level` (Gemini 3; adds `ULTRA_HIGH`). Recommended: images HIGH, PDFs MEDIUM, video LOW (HIGH for on-screen text), audio default. - `parts[].videoMetadata {startOffset, endOffset, fps}` — clipping/sampling (static mode; also for YouTube). - `parts[].mediaProcessing: STATIC | AGENTIC` — agentic on 3.8/3.7/3.6 Flash and 3.5 Flash-Lite (not probed). - Timestamps in prompts: `MM:SS` references work for video/audio (docs). - Put a single image/video **before** the text prompt for best results (docs). ## 4. YouTube specifics Public videos only (bogus id → `400 Request contains an invalid argument.`); preview, no charge for the URL fetch but video tokens are billed normally; ≥ 2.5: up to 10 videos per request; clipping via `videoMetadata`. Counting with `countTokens` is free — use it to size requests (verified). ## 5. Bounding boxes / segmentation Image understanding returns normalized `[ymin, xmin, ymax, xmax]` in 0–1000 (docs); segmentation is **not** supported on Gemini 3.x (use 2.5 Flash with thinking off). Examples: `examples/gemini/multimodal/` (inline 1×1 PNG + inline PDF). Tests: `tests/gemini/test_count_tokens.py` (image/PDF/YouTube counts).