SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
14 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
4.1 KB

# Gemini token counting — models.countTokens and usageMetadata

Status: DOCUMENTED + LIVE_VERIFIED (13 countTokens probes, free; 2026-09-18). Sources: https://ai.google.dev/api/tokens · https://ai.google.dev/gemini-api/docs/generate-content/tokens · https://ai.google.dev/gemini-api/docs/long-context · discovery CountTokensRequest/Response Machine-readable: generated/fragments/parameters/gemini-count-tokens.json, generated/fragments/objects/gemini-core-objects.json#CountTokensResponse|UsageMetadata|ModalityTokenCount Last verified: 2026-09-18

# 1. POST /v1beta/models/{model}:countTokens (free, no billing)

Two mutually exclusive bodies:

Body Counts Live
{"contents": [...]} the prompt turns only "The quick brown fox jumps over the lazy dog." → 11 (flash-lite and flash)
{"generateContentRequest": {"model": "models/<id>", "contents", "systemInstruction", "tools", "toolConfig", "cachedContent", "generationConfig"}} everything the model will see + system text + 1 function declaration → 51; generationConfig.mediaResolution LOW changes the image count (1089 → 256); the wrapper's model is required in models/… form but the path model wins (mismatch accepted)
both contents ignored (documented, verified) 3 tokens = wrapper text
neither 400 CountTokens requires generate_content_request or contents to be set.
top-level systemInstruction 400 Invalid JSON payload received. Unknown name "systemInstruction" — nest it in the wrapper

SDK caveat (verified 2026-09-18): both official SDKs refuse the wrapper fields in Developer-API mode — google-genai 2.24 count_tokens(config=CountTokensConfig(system_instruction=…)) raises ValueError: system_instruction parameter is only supported in Gemini Enterprise Agent Platform mode, not in Gemini Developer API mode., and @google/genai 2.23 throws the same message. The REST endpoint accepts generateContentRequest.systemInstruction/tools fine; send that form with a raw request (see examples/gemini/count-tokens/count_tokens.py|.ts).

Response: {"totalTokens": 1092, "promptTokensDetails": [{"modality": "TEXT", "tokenCount": 3}, {"modality": "IMAGE", "tokenCount": 1089}], "cachedContentTokenCount"?, "cacheTokensDetails"?}.

Multimodal counts observed: 1×1 PNG 1089 (LOW 256, MEDIUM 529, ULTRA_HIGH 2209 per part); 1-page PDF DOCUMENT 560; YouTube video VIDEO 10650 + AUDIO 4777; 10 s clip 710 + 320; HTTPS JPEG 1080. Legacy countTextTokens / countMessageTokens → 404 / 501.

# 2. usageMetadata on generateContent

Field Meaning Observed
promptTokenCount input incl. system, tools, media, cached tokens, signatures echoed back 5 for "Reply with OK."
cachedContentTokenCount cached prefix (explicit or implicit hit) not observed (caching restricted)
candidatesTokenCount output tokens across candidates; absent when only thoughts were generated 1–26
thoughtsTokenCount billed thinking tokens 60 / 70 / 103 / 117
toolUsePromptTokenCount server-side tool prompts (search results, agentic video loads) —
totalTokenCount prompt + thoughts + candidates 136 = 16 + 117 + 3
*TokensDetails[] {modality, tokenCount} for prompt / cache / candidates / toolUse [{TEXT 5}], [{IMAGE 1089},{TEXT 5}]
serviceTier standard / flex / priority standard; flex when requested

Streaming: usageMetadata is on every chunk; the last chunk is authoritative. Embeddings: usageMetadata.promptTokenCount + promptTokenDetails (embedding-2 only). Caches: usageMetadata.totalTokenCount on the CachedContent.

# 3. Rules of thumb (docs)

1 token ≈ 4 characters; 100 tokens ≈ 60–80 English words. Context windows: 1,048,576 input / 65,536 output for Gemini 3.x Flash/Pro (models.get → inputTokenLimit, outputTokenLimit). Long-context guidance: put the question after the data; many-shot in-context learning; explicit caching to amortize large prefixes.

Examples: examples/gemini/count-tokens/. Tests: tests/gemini/test_count_tokens.py.