SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
6.3 KB

# Gemini thinking — thinkingConfig, thought summaries, thought signatures

Status: DOCUMENTED + LIVE_VERIFIED (gemini-3.5-flash, gemini-3.8-flash, gemini-3.5-flash-lite; gemini-3.1-pro-preview → ACCOUNT_RESTRICTED 429 on this free-tier key). 2026-09-18. Sources: https://ai.google.dev/gemini-api/docs/generate-content/thinking · https://ai.google.dev/gemini-api/docs/generate-content/thought-signatures · https://ai.google.dev/gemini-api/docs/generate-content/gemini-3 · https://ai.google.dev/gemini-api/docs/generate-content/whats-new-gemini-3.5 · discovery ThinkingConfig Machine-readable: generated/fragments/parameters/gemini-generate-content.json (generationConfig.thinkingConfig.*, contents[].parts[].thought*), generated/fragments/compatibility/gemini-feature-model-matrix.json Last verified: 2026-09-18

# 1. Controls

json
"generationConfig": {"thinkingConfig": {"thinkingLevel": "low", "includeThoughts": true}}
"generationConfig": {"thinkingConfig": {"thinkingBudget": 256, "includeThoughts": true}}
Field Values Rules (docs + live)
thinkingLevel MINIMAL | LOW | MEDIUM | HIGH (case-insensitive; "low", "HIGH" both OK) Gemini 3+ only. Defaults: high (3.1 Pro, 3 Flash), medium (3.5/3.6/3.7/3.8 Flash), minimal (3.5/3.1 Flash-Lite, 3.1 Flash-Lite Image). minimal unsupported on 3.8/3.7 Flash and 3.1 Pro → 400 Thinking level MINIMAL is not supported for this model. Please retry with other thinking level. Unknown → 400 Invalid value at 'generation_config.thinking_config.thinking_level'.
thinkingBudget int32 in [-1, 65535] (live: thinking_budget must be in the range [-1, 65535]) 2.5-era control; accepted on 3.x for compatibility (not recommended, not with 3.1 Pro). -1 = dynamic, 0 = off where allowed. 2.5 Pro 128–32768 (cannot disable); 2.5 Flash 0–24576; 2.5 Flash-Lite 512–24576 (default no thinking). Live 3.x: Flash-Lite 0 → 400 Request contains an invalid argument.; Flash 0 → 200 (no thoughts); Flash-Lite -1 → dynamic (60 thought tokens for "Reply with OK.").
includeThoughts bool Returns thought summaries as parts {text, thought: true} when the model reasoned enough. Not returned on trivial prompts (verified on 3.8 Flash low and 3.5 Flash-Lite). Billing is always on full thoughts (thoughtsTokenCount), never on the summary.
both thinkingLevel + thinkingBudget — 400 You can only set only one of thinking budget and thinking level.
model without thinking — error (docs); embedding model → 404 not supported for generateContent

# 2. Live numbers

Probe Model Config Result
e_flash_thoughts_budget gemini-3.5-flash budget 256, includeThoughts, maxOutputTokens 600, "17*23?" parts: {text:"**My Calculation of a Multiplication Problem**…", thought:true}, {text:"391", thoughtSignature}; usage prompt 16 / candidates 3 / thoughts 117 / total 136; 16.8 s
e_g38_level_low gemini-3.8-flash level low + includeThoughts OK. no thought part, no thoughtsTokenCount (8.1 s)
e_g38_level_high gemini-3.8-flash level high OK., thoughtsTokenCount: 70 for a trivial prompt (3.3 s)
e_flash_level_minimal gemini-3.5-flash level minimal 200, no thoughts (17.5 s latency)
e_lite_budget_minus1 gemini-3.5-flash-lite budget -1, maxOutputTokens 64 only a thought part, finishReason: MAX_TOKENS, thoughts 60, no candidatesTokenCount — maxOutputTokens caps thoughts + answer
e_stream_thoughts gemini-3.5-flash (SSE) budget 128, includeThoughts 2 chunks, thoughts 103 billed, no thought part emitted

Pricing: output price applies to candidatesTokenCount + thoughtsTokenCount. Signatures sent back count as input tokens. Prefer lowering thinkingLevel over a small maxOutputTokens (which truncates and still bills the thinking).

# 3. Thought signatures (parts[].thoughtSignature)

Encrypted reasoning state; the API is stateless so you carry it. Observed: every Gemini 3.x response (3.5 Flash-Lite included, thinking minimal) puts a signature on its last part — even the plain OK answer.

Rules (docs, Gemini 3):

  • Function calling: the first functionCall part of each step in the current turn must be echoed with its signature (parallel calls: only the first carries one; sequential steps: each). Missing → 400 Function call <name> in the <n>. content block is missing a thought_signature. Enforced even at thinkingLevel: minimal. Interleaving FC1, FR1, FC2, FR2 instead of FC1, FC2, FR1, FR2 → 400.
  • Text/other parts: signature on the last part; returning it is recommended (quality), not validated.
  • Image generation/editing: strictly validated on all parts (media domain).
  • Gemini 2.5: signatures only with function declarations, on the first part, optional to return.
  • Escape hatches: dummy values "skip_thought_signature_validator" or "context_engineering_is_the_way_to_go" bypass validation for injected history (verified: accepted).
  • Live: echoing the received signature → 200; a random base64 → 400 Corrupted thought signature.
  • Since 3.5 Flash, thought context is preserved across turns when the full history with signatures is sent ("thought preservation"); this raises input tokens — clear signatures for simple follow-ups if cost matters.
  • Streaming: the signature arrives in the last chunk inside an empty-text part.
  • OpenAI-compat layer: tool_calls[].extra_content.google.thought_signature (compat domain).

# 4. Model matrix (from docs, 2026-09)

Model Default Levels Disable?
gemini-3.8-flash / 3.7-flash medium low, medium, high no (minimal → 400)
gemini-3.6-flash / 3.5-flash medium minimal, low, medium, high via thinkingBudget: 0 (verified 3.5)
gemini-3.5-flash-lite / 3.1-flash-lite minimal minimal, low, medium, high budget 0 → 400; use minimal
gemini-3.1-pro-preview high low, medium, high no
gemini-3-flash-preview high minimal, low, medium, high —
gemini-2.5-pro dynamic (budget 128–32768) no
gemini-2.5-flash dynamic (budget 0–24576) budget 0
gemini-2.5-flash-lite off (budget 512–24576) default off

Examples: examples/gemini/thinking/. Tests: tests/gemini/test_thinking.py.