Gemini streaming — models.streamGenerateContent
Status: DOCUMENTED + LIVE_VERIFIED (SSE and JSON-array variants, structured-output stream, thinking stream; 2026-09-18, gemini-3.5-flash-lite / gemini-3.5-flash).
Sources: https://ai.google.dev/api/generate-content#method:-models.streamgeneratecontent · https://ai.google.dev/gemini-api/docs/generate-content/text-generation · https://ai.google.dev/gemini-api/docs/generate-content/thought-signatures
Machine-readable: generated/fragments/streaming-events/gemini-core.json, generated/fragments/endpoints/gemini-core.json
Last verified: 2026-09-18
1. Two framings, one chunk type
POST /v1beta/models/{model}:streamGenerateContent takes the exact generateContent body. The response framing depends on the alt query parameter:
| Variant | URL | Content-Type | Framing (observed) |
|---|---|---|---|
| SSE (recommended) | …:streamGenerateContent?alt=sse |
text/event-stream |
data: <json>\r\n\r\n per chunk. No event:/id: fields, no [DONE]. 4 chunks = 8 lines. |
| JSON array (default) | …:streamGenerateContent |
application/json; charset=UTF-8 |
A pretty-printed JSON array streamed incrementally: line 1 [{, chunks separated by a line containing only ,, closed by ] (122 lines for the same 4 chunks). Needs an incremental array parser. |
Both carried byte-identical chunk objects for the same prompt. Bonus (undocumented): …:generateContent?alt=sse returns a single data: event with the full response.
2. Chunk = partial GenerateContentResponse
Observed sequence for "Count from 1 to 12" (26 output tokens):
| # | parts | finishReason | usageMetadata |
|---|---|---|---|
| 1 | {text: "1 "} |
— | present (prompt 13, candidates 2) |
| 2 | {text: "2 3 4 5 6 7 8 9 10"} |
— | present (running) |
| 3 | {text: " 11 12"} |
— | present |
| 4 | {text: "", thoughtSignature: "El4K…"} |
STOP |
final (13 / 26 / 39, serviceTier: standard) |
Rules:
usageMetadatais present in every chunk (runningcandidatesTokenCount); use the last one.modelVersionandresponseIdare repeated on every chunk;responseIdis constant within a stream.- The final chunk on Gemini 3 models carries an empty-text part with the turn's
thoughtSignature— parsers must not drop empty text parts, and must keep reading untilfinishReasonappears (docs FAQ: "the model may return the thought signature in a part with an empty text content part"). finishReasonappears only once, on the last chunk (STOP,MAX_TOKENSobserved).- Text deltas concatenate; with
responseMimeType: application/jsonthe deltas are partial JSON strings ({"colors": ["red", …then the rest) that concatenate into the final object.
3. Thinking in streams
With thinkingConfig.includeThoughts: true thought summaries stream first as parts {text, thought: true} (rolling summaries), then answer parts. Not guaranteed: on gemini-3.5-flash with thinkingBudget: 128 the stream billed thoughtsTokenCount: 103 yet emitted no thought part (2 chunks: 399 then the signature chunk). Non-streaming with budget 256 did return a thought part.
4. Errors
- Pre-stream errors are plain JSON error bodies with the HTTP status (429
RESOURCE_EXHAUSTEDobserved). - Mid-stream errors are not documented for
generateContentSSE (the Interactions API defines anerrorevent; not applicable here). Treat a connection close withoutfinishReasonas an error. - Prompt blocked → single response with
promptFeedback.blockReasonand no candidates (not triggered).
5. Clients
- curl:
curl -N -H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" "…:streamGenerateContent?alt=sse" -d @body.json(examples/gemini/streaming/stream.sh). - Python:
for chunk in client.models.generate_content_stream(...): chunk.text(SDK always usesalt=sse). - Node:
for await (const chunk of await ai.models.generateContentStream({...})) chunk.text.