Gemini thinking — thinkingConfig, thought summaries, thought signatures
Status: DOCUMENTED + LIVE_VERIFIED (gemini-3.5-flash, gemini-3.8-flash, gemini-3.5-flash-lite; gemini-3.1-pro-preview → ACCOUNT_RESTRICTED 429 on this free-tier key). 2026-09-18.
Sources: https://ai.google.dev/gemini-api/docs/generate-content/thinking · https://ai.google.dev/gemini-api/docs/generate-content/thought-signatures · https://ai.google.dev/gemini-api/docs/generate-content/gemini-3 · https://ai.google.dev/gemini-api/docs/generate-content/whats-new-gemini-3.5 · discovery ThinkingConfig
Machine-readable: generated/fragments/parameters/gemini-generate-content.json (generationConfig.thinkingConfig.*, contents[].parts[].thought*), generated/fragments/compatibility/gemini-feature-model-matrix.json
Last verified: 2026-09-18
1. Controls
"generationConfig": {"thinkingConfig": {"thinkingLevel": "low", "includeThoughts": true}}
"generationConfig": {"thinkingConfig": {"thinkingBudget": 256, "includeThoughts": true}}| Field | Values | Rules (docs + live) |
|---|---|---|
thinkingLevel |
MINIMAL | LOW | MEDIUM | HIGH (case-insensitive; "low", "HIGH" both OK) |
Gemini 3+ only. Defaults: high (3.1 Pro, 3 Flash), medium (3.5/3.6/3.7/3.8 Flash), minimal (3.5/3.1 Flash-Lite, 3.1 Flash-Lite Image). minimal unsupported on 3.8/3.7 Flash and 3.1 Pro → 400 Thinking level MINIMAL is not supported for this model. Please retry with other thinking level. Unknown → 400 Invalid value at 'generation_config.thinking_config.thinking_level'. |
thinkingBudget |
int32 in [-1, 65535] (live: thinking_budget must be in the range [-1, 65535]) |
2.5-era control; accepted on 3.x for compatibility (not recommended, not with 3.1 Pro). -1 = dynamic, 0 = off where allowed. 2.5 Pro 128–32768 (cannot disable); 2.5 Flash 0–24576; 2.5 Flash-Lite 512–24576 (default no thinking). Live 3.x: Flash-Lite 0 → 400 Request contains an invalid argument.; Flash 0 → 200 (no thoughts); Flash-Lite -1 → dynamic (60 thought tokens for "Reply with OK."). |
includeThoughts |
bool | Returns thought summaries as parts {text, thought: true} when the model reasoned enough. Not returned on trivial prompts (verified on 3.8 Flash low and 3.5 Flash-Lite). Billing is always on full thoughts (thoughtsTokenCount), never on the summary. |
both thinkingLevel + thinkingBudget |
— | 400 You can only set only one of thinking budget and thinking level. |
| model without thinking | — | error (docs); embedding model → 404 not supported for generateContent |
2. Live numbers
| Probe | Model | Config | Result |
|---|---|---|---|
| e_flash_thoughts_budget | gemini-3.5-flash | budget 256, includeThoughts, maxOutputTokens 600, "17*23?" | parts: {text:"**My Calculation of a Multiplication Problem**…", thought:true}, {text:"391", thoughtSignature}; usage prompt 16 / candidates 3 / thoughts 117 / total 136; 16.8 s |
| e_g38_level_low | gemini-3.8-flash | level low + includeThoughts | OK. no thought part, no thoughtsTokenCount (8.1 s) |
| e_g38_level_high | gemini-3.8-flash | level high | OK., thoughtsTokenCount: 70 for a trivial prompt (3.3 s) |
| e_flash_level_minimal | gemini-3.5-flash | level minimal | 200, no thoughts (17.5 s latency) |
| e_lite_budget_minus1 | gemini-3.5-flash-lite | budget -1, maxOutputTokens 64 | only a thought part, finishReason: MAX_TOKENS, thoughts 60, no candidatesTokenCount — maxOutputTokens caps thoughts + answer |
| e_stream_thoughts | gemini-3.5-flash (SSE) | budget 128, includeThoughts | 2 chunks, thoughts 103 billed, no thought part emitted |
Pricing: output price applies to candidatesTokenCount + thoughtsTokenCount. Signatures sent back count as input tokens. Prefer lowering thinkingLevel over a small maxOutputTokens (which truncates and still bills the thinking).
3. Thought signatures (parts[].thoughtSignature)
Encrypted reasoning state; the API is stateless so you carry it. Observed: every Gemini 3.x response (3.5 Flash-Lite included, thinking minimal) puts a signature on its last part — even the plain OK answer.
Rules (docs, Gemini 3):
- Function calling: the first
functionCallpart of each step in the current turn must be echoed with its signature (parallel calls: only the first carries one; sequential steps: each). Missing →400 Function call <name> in the <n>. content block is missing a thought_signature. Enforced even atthinkingLevel: minimal. InterleavingFC1, FR1, FC2, FR2instead ofFC1, FC2, FR1, FR2→ 400. - Text/other parts: signature on the last part; returning it is recommended (quality), not validated.
- Image generation/editing: strictly validated on all parts (media domain).
- Gemini 2.5: signatures only with function declarations, on the first part, optional to return.
- Escape hatches: dummy values
"skip_thought_signature_validator"or"context_engineering_is_the_way_to_go"bypass validation for injected history (verified: accepted). - Live: echoing the received signature → 200; a random base64 →
400 Corrupted thought signature. - Since 3.5 Flash, thought context is preserved across turns when the full history with signatures is sent ("thought preservation"); this raises input tokens — clear signatures for simple follow-ups if cost matters.
- Streaming: the signature arrives in the last chunk inside an empty-text part.
- OpenAI-compat layer:
tool_calls[].extra_content.google.thought_signature(compat domain).
4. Model matrix (from docs, 2026-09)
| Model | Default | Levels | Disable? |
|---|---|---|---|
| gemini-3.8-flash / 3.7-flash | medium | low, medium, high | no (minimal → 400) |
| gemini-3.6-flash / 3.5-flash | medium | minimal, low, medium, high | via thinkingBudget: 0 (verified 3.5) |
| gemini-3.5-flash-lite / 3.1-flash-lite | minimal | minimal, low, medium, high | budget 0 → 400; use minimal |
| gemini-3.1-pro-preview | high | low, medium, high | no |
| gemini-3-flash-preview | high | minimal, low, medium, high | — |
| gemini-2.5-pro | dynamic | (budget 128–32768) | no |
| gemini-2.5-flash | dynamic | (budget 0–24576) | budget 0 |
| gemini-2.5-flash-lite | off | (budget 512–24576) | default off |
Examples: examples/gemini/thinking/. Tests: tests/gemini/test_thinking.py.