SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
14 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
4.1 KB · 45 lines markdown
Rendered Raw Blame History
1# Gemini token counting — `models.countTokens` and `usageMetadata`23**Status:** `DOCUMENTED` + `LIVE_VERIFIED` (13 countTokens probes, free; 2026-09-18).4**Sources:** https://ai.google.dev/api/tokens · https://ai.google.dev/gemini-api/docs/generate-content/tokens · https://ai.google.dev/gemini-api/docs/long-context · discovery `CountTokensRequest/Response`5**Machine-readable:** `generated/fragments/parameters/gemini-count-tokens.json`, `generated/fragments/objects/gemini-core-objects.json#CountTokensResponse|UsageMetadata|ModalityTokenCount`6**Last verified:** 2026-09-1878## 1. `POST /v1beta/models/{model}:countTokens` (free, no billing)910Two mutually exclusive bodies:1112| Body | Counts | Live |13|---|---|---|14| `{"contents": [...]}` | the prompt turns only | "The quick brown fox jumps over the lazy dog." → 11 (flash-lite and flash) |15| `{"generateContentRequest": {"model": "models/<id>", "contents", "systemInstruction", "tools", "toolConfig", "cachedContent", "generationConfig"}}` | everything the model will see | + system text + 1 function declaration → 51; `generationConfig.mediaResolution LOW` changes the image count (1089 → 256); the wrapper's `model` is required in `models/…` form but the **path model wins** (mismatch accepted) |16| both | `contents` ignored (documented, verified) | 3 tokens = wrapper text |17| neither | `400 CountTokens requires generate_content_request or contents to be set.` | |18| top-level `systemInstruction` | `400 Invalid JSON payload received. Unknown name "systemInstruction"` — nest it in the wrapper | |1920**SDK caveat (verified 2026-09-18):** both official SDKs refuse the wrapper fields in Developer-API mode — `google-genai` 2.24 `count_tokens(config=CountTokensConfig(system_instruction=…))` raises `ValueError: system_instruction parameter is only supported in Gemini Enterprise Agent Platform mode, not in Gemini Developer API mode.`, and `@google/genai` 2.23 throws the same message. The REST endpoint accepts `generateContentRequest.systemInstruction/tools` fine; send that form with a raw request (see `examples/gemini/count-tokens/count_tokens.py|.ts`).2122Response: `{"totalTokens": 1092, "promptTokensDetails": [{"modality": "TEXT", "tokenCount": 3}, {"modality": "IMAGE", "tokenCount": 1089}], "cachedContentTokenCount"?, "cacheTokensDetails"?}`.2324Multimodal counts observed: 1×1 PNG 1089 (LOW 256, MEDIUM 529, ULTRA_HIGH 2209 per part); 1-page PDF `DOCUMENT 560`; YouTube video `VIDEO 10650 + AUDIO 4777`; 10 s clip `710 + 320`; HTTPS JPEG 1080. Legacy `countTextTokens` / `countMessageTokens` → 404 / 501.2526## 2. `usageMetadata` on generateContent2728| Field | Meaning | Observed |29|---|---|---|30| `promptTokenCount` | input incl. system, tools, media, **cached** tokens, signatures echoed back | 5 for "Reply with OK." |31| `cachedContentTokenCount` | cached prefix (explicit or implicit hit) | not observed (caching restricted) |32| `candidatesTokenCount` | output tokens across candidates; **absent** when only thoughts were generated | 1–26 |33| `thoughtsTokenCount` | billed thinking tokens | 60 / 70 / 103 / 117 |34| `toolUsePromptTokenCount` | server-side tool prompts (search results, agentic video loads) | — |35| `totalTokenCount` | prompt + thoughts + candidates | 136 = 16 + 117 + 3 |36| `*TokensDetails[]` | `{modality, tokenCount}` for prompt / cache / candidates / toolUse | `[{TEXT 5}]`, `[{IMAGE 1089},{TEXT 5}]` |37| `serviceTier` | `standard` / `flex` / `priority` | `standard`; `flex` when requested |3839Streaming: `usageMetadata` is on every chunk; the last chunk is authoritative. Embeddings: `usageMetadata.promptTokenCount` + `promptTokenDetails` (embedding-2 only). Caches: `usageMetadata.totalTokenCount` on the CachedContent.4041## 3. Rules of thumb (docs)421 token ≈ 4 characters; 100 tokens ≈ 60–80 English words. Context windows: 1,048,576 input / 65,536 output for Gemini 3.x Flash/Pro (`models.get` → `inputTokenLimit`, `outputTokenLimit`). Long-context guidance: put the question after the data; many-shot in-context learning; explicit caching to amortize large prefixes.4344Examples: `examples/gemini/count-tokens/`. Tests: `tests/gemini/test_count_tokens.py`.45