# Gemini thinking — `thinkingConfig`, thought summaries, thought signatures **Status:** `DOCUMENTED` + `LIVE_VERIFIED` (`gemini-3.5-flash`, `gemini-3.8-flash`, `gemini-3.5-flash-lite`; `gemini-3.1-pro-preview` → `ACCOUNT_RESTRICTED` 429 on this free-tier key). 2026-09-18. **Sources:** https://ai.google.dev/gemini-api/docs/generate-content/thinking · https://ai.google.dev/gemini-api/docs/generate-content/thought-signatures · https://ai.google.dev/gemini-api/docs/generate-content/gemini-3 · https://ai.google.dev/gemini-api/docs/generate-content/whats-new-gemini-3.5 · discovery `ThinkingConfig` **Machine-readable:** `generated/fragments/parameters/gemini-generate-content.json` (`generationConfig.thinkingConfig.*`, `contents[].parts[].thought*`), `generated/fragments/compatibility/gemini-feature-model-matrix.json` **Last verified:** 2026-09-18 ## 1. Controls ```json "generationConfig": {"thinkingConfig": {"thinkingLevel": "low", "includeThoughts": true}} "generationConfig": {"thinkingConfig": {"thinkingBudget": 256, "includeThoughts": true}} ``` | Field | Values | Rules (docs + live) | |---|---|---| | `thinkingLevel` | `MINIMAL` \| `LOW` \| `MEDIUM` \| `HIGH` (case-insensitive; `"low"`, `"HIGH"` both OK) | Gemini 3+ only. Defaults: `high` (3.1 Pro, 3 Flash), `medium` (3.5/3.6/3.7/3.8 Flash), `minimal` (3.5/3.1 Flash-Lite, 3.1 Flash-Lite Image). `minimal` unsupported on 3.8/3.7 Flash and 3.1 Pro → `400 Thinking level MINIMAL is not supported for this model. Please retry with other thinking level.` Unknown → `400 Invalid value at 'generation_config.thinking_config.thinking_level'`. | | `thinkingBudget` | int32 in **[-1, 65535]** (live: `thinking_budget must be in the range [-1, 65535]`) | 2.5-era control; accepted on 3.x for compatibility (not recommended, not with 3.1 Pro). `-1` = dynamic, `0` = off where allowed. 2.5 Pro 128–32768 (cannot disable); 2.5 Flash 0–24576; 2.5 Flash-Lite 512–24576 (default no thinking). **Live 3.x:** Flash-Lite `0` → `400 Request contains an invalid argument.`; Flash `0` → 200 (no thoughts); Flash-Lite `-1` → dynamic (60 thought tokens for "Reply with OK."). | | `includeThoughts` | bool | Returns thought **summaries** as parts `{text, thought: true}` when the model reasoned enough. Not returned on trivial prompts (verified on 3.8 Flash low and 3.5 Flash-Lite). Billing is always on full thoughts (`thoughtsTokenCount`), never on the summary. | | both `thinkingLevel` + `thinkingBudget` | — | `400 You can only set only one of thinking budget and thinking level.` | | model without thinking | — | error (docs); embedding model → 404 not supported for generateContent | ## 2. Live numbers | Probe | Model | Config | Result | |---|---|---|---| | e_flash_thoughts_budget | gemini-3.5-flash | budget 256, includeThoughts, maxOutputTokens 600, "17*23?" | parts: `{text:"**My Calculation of a Multiplication Problem**…", thought:true}`, `{text:"391", thoughtSignature}`; usage prompt 16 / candidates 3 / **thoughts 117** / total 136; 16.8 s | | e_g38_level_low | gemini-3.8-flash | level low + includeThoughts | `OK.` no thought part, no `thoughtsTokenCount` (8.1 s) | | e_g38_level_high | gemini-3.8-flash | level high | `OK.`, `thoughtsTokenCount: 70` for a trivial prompt (3.3 s) | | e_flash_level_minimal | gemini-3.5-flash | level minimal | 200, no thoughts (17.5 s latency) | | e_lite_budget_minus1 | gemini-3.5-flash-lite | budget -1, maxOutputTokens 64 | only a thought part, `finishReason: MAX_TOKENS`, thoughts 60, no `candidatesTokenCount` — **maxOutputTokens caps thoughts + answer** | | e_stream_thoughts | gemini-3.5-flash (SSE) | budget 128, includeThoughts | 2 chunks, thoughts 103 billed, **no thought part emitted** | Pricing: output price applies to `candidatesTokenCount + thoughtsTokenCount`. Signatures sent back count as input tokens. Prefer lowering `thinkingLevel` over a small `maxOutputTokens` (which truncates and still bills the thinking). ## 3. Thought signatures (`parts[].thoughtSignature`) Encrypted reasoning state; the API is stateless so you carry it. Observed: **every** Gemini 3.x response (3.5 Flash-Lite included, thinking minimal) puts a signature on its last part — even the plain `OK` answer. Rules (docs, Gemini 3): - Function calling: the **first `functionCall` part of each step in the current turn** must be echoed with its signature (parallel calls: only the first carries one; sequential steps: each). Missing → `400 Function call in the . content block is missing a thought_signature`. Enforced even at `thinkingLevel: minimal`. Interleaving `FC1, FR1, FC2, FR2` instead of `FC1, FC2, FR1, FR2` → 400. - Text/other parts: signature on the last part; returning it is recommended (quality), not validated. - Image generation/editing: strictly validated on all parts (media domain). - Gemini 2.5: signatures only with function declarations, on the first part, optional to return. - Escape hatches: dummy values `"skip_thought_signature_validator"` or `"context_engineering_is_the_way_to_go"` bypass validation for injected history (verified: accepted). - Live: echoing the received signature → 200; a random base64 → `400 Corrupted thought signature.` - Since 3.5 Flash, thought context is preserved across turns when the full history with signatures is sent ("thought preservation"); this raises input tokens — clear signatures for simple follow-ups if cost matters. - Streaming: the signature arrives in the last chunk inside an empty-text part. - OpenAI-compat layer: `tool_calls[].extra_content.google.thought_signature` (compat domain). ## 4. Model matrix (from docs, 2026-09) | Model | Default | Levels | Disable? | |---|---|---|---| | gemini-3.8-flash / 3.7-flash | medium | low, medium, high | no (`minimal` → 400) | | gemini-3.6-flash / 3.5-flash | medium | minimal, low, medium, high | via `thinkingBudget: 0` (verified 3.5) | | gemini-3.5-flash-lite / 3.1-flash-lite | minimal | minimal, low, medium, high | `budget 0` → 400; use `minimal` | | gemini-3.1-pro-preview | high | low, medium, high | no | | gemini-3-flash-preview | high | minimal, low, medium, high | — | | gemini-2.5-pro | dynamic | (budget 128–32768) | no | | gemini-2.5-flash | dynamic | (budget 0–24576) | `budget 0` | | gemini-2.5-flash-lite | off | (budget 512–24576) | default off | Examples: `examples/gemini/thinking/`. Tests: `tests/gemini/test_thinking.py`.