Gemini token counting — models.countTokens and usageMetadata
Status: DOCUMENTED + LIVE_VERIFIED (13 countTokens probes, free; 2026-09-18).
Sources: https://ai.google.dev/api/tokens · https://ai.google.dev/gemini-api/docs/generate-content/tokens · https://ai.google.dev/gemini-api/docs/long-context · discovery CountTokensRequest/Response
Machine-readable: generated/fragments/parameters/gemini-count-tokens.json, generated/fragments/objects/gemini-core-objects.json#CountTokensResponse|UsageMetadata|ModalityTokenCount
Last verified: 2026-09-18
1. POST /v1beta/models/{model}:countTokens (free, no billing)
Two mutually exclusive bodies:
| Body | Counts | Live |
|---|---|---|
{"contents": [...]} |
the prompt turns only | "The quick brown fox jumps over the lazy dog." → 11 (flash-lite and flash) |
{"generateContentRequest": {"model": "models/<id>", "contents", "systemInstruction", "tools", "toolConfig", "cachedContent", "generationConfig"}} |
everything the model will see | + system text + 1 function declaration → 51; generationConfig.mediaResolution LOW changes the image count (1089 → 256); the wrapper's model is required in models/… form but the path model wins (mismatch accepted) |
| both | contents ignored (documented, verified) |
3 tokens = wrapper text |
| neither | 400 CountTokens requires generate_content_request or contents to be set. |
|
top-level systemInstruction |
400 Invalid JSON payload received. Unknown name "systemInstruction" — nest it in the wrapper |
SDK caveat (verified 2026-09-18): both official SDKs refuse the wrapper fields in Developer-API mode — google-genai 2.24 count_tokens(config=CountTokensConfig(system_instruction=…)) raises ValueError: system_instruction parameter is only supported in Gemini Enterprise Agent Platform mode, not in Gemini Developer API mode., and @google/genai 2.23 throws the same message. The REST endpoint accepts generateContentRequest.systemInstruction/tools fine; send that form with a raw request (see examples/gemini/count-tokens/count_tokens.py|.ts).
Response: {"totalTokens": 1092, "promptTokensDetails": [{"modality": "TEXT", "tokenCount": 3}, {"modality": "IMAGE", "tokenCount": 1089}], "cachedContentTokenCount"?, "cacheTokensDetails"?}.
Multimodal counts observed: 1×1 PNG 1089 (LOW 256, MEDIUM 529, ULTRA_HIGH 2209 per part); 1-page PDF DOCUMENT 560; YouTube video VIDEO 10650 + AUDIO 4777; 10 s clip 710 + 320; HTTPS JPEG 1080. Legacy countTextTokens / countMessageTokens → 404 / 501.
2. usageMetadata on generateContent
| Field | Meaning | Observed |
|---|---|---|
promptTokenCount |
input incl. system, tools, media, cached tokens, signatures echoed back | 5 for "Reply with OK." |
cachedContentTokenCount |
cached prefix (explicit or implicit hit) | not observed (caching restricted) |
candidatesTokenCount |
output tokens across candidates; absent when only thoughts were generated | 1–26 |
thoughtsTokenCount |
billed thinking tokens | 60 / 70 / 103 / 117 |
toolUsePromptTokenCount |
server-side tool prompts (search results, agentic video loads) | — |
totalTokenCount |
prompt + thoughts + candidates | 136 = 16 + 117 + 3 |
*TokensDetails[] |
{modality, tokenCount} for prompt / cache / candidates / toolUse |
[{TEXT 5}], [{IMAGE 1089},{TEXT 5}] |
serviceTier |
standard / flex / priority |
standard; flex when requested |
Streaming: usageMetadata is on every chunk; the last chunk is authoritative. Embeddings: usageMetadata.promptTokenCount + promptTokenDetails (embedding-2 only). Caches: usageMetadata.totalTokenCount on the CachedContent.
3. Rules of thumb (docs)
1 token ≈ 4 characters; 100 tokens ≈ 60–80 English words. Context windows: 1,048,576 input / 65,536 output for Gemini 3.x Flash/Pro (models.get → inputTokenLimit, outputTokenLimit). Long-context guidance: put the question after the data; many-shot in-context learning; explicit caching to amortize large prefixes.
Examples: examples/gemini/count-tokens/. Tests: tests/gemini/test_count_tokens.py.