# Gemini streaming — `models.streamGenerateContent` **Status:** `DOCUMENTED` + `LIVE_VERIFIED` (SSE and JSON-array variants, structured-output stream, thinking stream; 2026-09-18, `gemini-3.5-flash-lite` / `gemini-3.5-flash`). **Sources:** https://ai.google.dev/api/generate-content#method:-models.streamgeneratecontent · https://ai.google.dev/gemini-api/docs/generate-content/text-generation · https://ai.google.dev/gemini-api/docs/generate-content/thought-signatures **Machine-readable:** `generated/fragments/streaming-events/gemini-core.json`, `generated/fragments/endpoints/gemini-core.json` **Last verified:** 2026-09-18 ## 1. Two framings, one chunk type `POST /v1beta/models/{model}:streamGenerateContent` takes the exact `generateContent` body. The response framing depends on the `alt` query parameter: | Variant | URL | Content-Type | Framing (observed) | |---|---|---|---| | SSE (recommended) | `…:streamGenerateContent?alt=sse` | `text/event-stream` | `data: \r\n\r\n` per chunk. No `event:`/`id:` fields, no `[DONE]`. 4 chunks = 8 lines. | | JSON array (default) | `…:streamGenerateContent` | `application/json; charset=UTF-8` | A pretty-printed JSON array streamed incrementally: line 1 `[{`, chunks separated by a line containing only `,`, closed by `]` (122 lines for the same 4 chunks). Needs an incremental array parser. | Both carried byte-identical chunk objects for the same prompt. Bonus (undocumented): `…:generateContent?alt=sse` returns a single `data:` event with the full response. ## 2. Chunk = partial `GenerateContentResponse` Observed sequence for "Count from 1 to 12" (26 output tokens): | # | parts | finishReason | usageMetadata | |---|---|---|---| | 1 | `{text: "1 "}` | — | present (prompt 13, candidates 2) | | 2 | `{text: "2 3 4 5 6 7 8 9 10"}` | — | present (running) | | 3 | `{text: " 11 12"}` | — | present | | 4 | `{text: "", thoughtSignature: "El4K…"}` | `STOP` | final (13 / 26 / 39, `serviceTier: standard`) | Rules: - `usageMetadata` is present in **every** chunk (running `candidatesTokenCount`); use the last one. - `modelVersion` and `responseId` are repeated on every chunk; `responseId` is constant within a stream. - The **final chunk on Gemini 3 models carries an empty-text part with the turn's `thoughtSignature`** — parsers must not drop empty text parts, and must keep reading until `finishReason` appears (docs FAQ: "the model may return the thought signature in a part with an empty text content part"). - `finishReason` appears only once, on the last chunk (`STOP`, `MAX_TOKENS` observed). - Text deltas concatenate; with `responseMimeType: application/json` the deltas are partial JSON strings (`{"colors": ["red", …` then the rest) that concatenate into the final object. ## 3. Thinking in streams With `thinkingConfig.includeThoughts: true` thought summaries stream first as parts `{text, thought: true}` (rolling summaries), then answer parts. Not guaranteed: on `gemini-3.5-flash` with `thinkingBudget: 128` the stream billed `thoughtsTokenCount: 103` yet emitted no thought part (2 chunks: `399` then the signature chunk). Non-streaming with budget 256 did return a thought part. ## 4. Errors - Pre-stream errors are plain JSON error bodies with the HTTP status (429 `RESOURCE_EXHAUSTED` observed). - Mid-stream errors are not documented for `generateContent` SSE (the Interactions API defines an `error` event; not applicable here). Treat a connection close without `finishReason` as an error. - Prompt blocked → single response with `promptFeedback.blockReason` and no candidates (not triggered). ## 5. Clients - curl: `curl -N -H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" "…:streamGenerateContent?alt=sse" -d @body.json` (`examples/gemini/streaming/stream.sh`). - Python: `for chunk in client.models.generate_content_stream(...): chunk.text` (SDK always uses `alt=sse`). - Node: `for await (const chunk of await ai.models.generateContentStream({...})) chunk.text`.