SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
3.9 KB

# Gemini streaming — models.streamGenerateContent

Status: DOCUMENTED + LIVE_VERIFIED (SSE and JSON-array variants, structured-output stream, thinking stream; 2026-09-18, gemini-3.5-flash-lite / gemini-3.5-flash). Sources: https://ai.google.dev/api/generate-content#method:-models.streamgeneratecontent · https://ai.google.dev/gemini-api/docs/generate-content/text-generation · https://ai.google.dev/gemini-api/docs/generate-content/thought-signatures Machine-readable: generated/fragments/streaming-events/gemini-core.json, generated/fragments/endpoints/gemini-core.json Last verified: 2026-09-18

# 1. Two framings, one chunk type

POST /v1beta/models/{model}:streamGenerateContent takes the exact generateContent body. The response framing depends on the alt query parameter:

Variant URL Content-Type Framing (observed)
SSE (recommended) …:streamGenerateContent?alt=sse text/event-stream data: <json>\r\n\r\n per chunk. No event:/id: fields, no [DONE]. 4 chunks = 8 lines.
JSON array (default) …:streamGenerateContent application/json; charset=UTF-8 A pretty-printed JSON array streamed incrementally: line 1 [{, chunks separated by a line containing only ,, closed by ] (122 lines for the same 4 chunks). Needs an incremental array parser.

Both carried byte-identical chunk objects for the same prompt. Bonus (undocumented): …:generateContent?alt=sse returns a single data: event with the full response.

# 2. Chunk = partial GenerateContentResponse

Observed sequence for "Count from 1 to 12" (26 output tokens):

# parts finishReason usageMetadata
1 {text: "1 "} — present (prompt 13, candidates 2)
2 {text: "2 3 4 5 6 7 8 9 10"} — present (running)
3 {text: " 11 12"} — present
4 {text: "", thoughtSignature: "El4K…"} STOP final (13 / 26 / 39, serviceTier: standard)

Rules:

  • usageMetadata is present in every chunk (running candidatesTokenCount); use the last one.
  • modelVersion and responseId are repeated on every chunk; responseId is constant within a stream.
  • The final chunk on Gemini 3 models carries an empty-text part with the turn's thoughtSignature — parsers must not drop empty text parts, and must keep reading until finishReason appears (docs FAQ: "the model may return the thought signature in a part with an empty text content part").
  • finishReason appears only once, on the last chunk (STOP, MAX_TOKENS observed).
  • Text deltas concatenate; with responseMimeType: application/json the deltas are partial JSON strings ({"colors": ["red", … then the rest) that concatenate into the final object.

# 3. Thinking in streams

With thinkingConfig.includeThoughts: true thought summaries stream first as parts {text, thought: true} (rolling summaries), then answer parts. Not guaranteed: on gemini-3.5-flash with thinkingBudget: 128 the stream billed thoughtsTokenCount: 103 yet emitted no thought part (2 chunks: 399 then the signature chunk). Non-streaming with budget 256 did return a thought part.

# 4. Errors

  • Pre-stream errors are plain JSON error bodies with the HTTP status (429 RESOURCE_EXHAUSTED observed).
  • Mid-stream errors are not documented for generateContent SSE (the Interactions API defines an error event; not applicable here). Treat a connection close without finishReason as an error.
  • Prompt blocked → single response with promptFeedback.blockReason and no candidates (not triggered).

# 5. Clients

  • curl: curl -N -H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" "…:streamGenerateContent?alt=sse" -d @body.json (examples/gemini/streaming/stream.sh).
  • Python: for chunk in client.models.generate_content_stream(...): chunk.text (SDK always uses alt=sse).
  • Node: for await (const chunk of await ai.models.generateContentStream({...})) chunk.text.