# Gemini Batch API (`batchGenerateContent` / `asyncBatchEmbedContent` / `batches.*`) **Status:** `DOCUMENTED` · `ACCOUNT_RESTRICTED` — every create call (valid docs body on `gemini-3.8-flash`, tiny inline batch on `gemini-3.5-flash-lite`, embedding batch, and even invalid bodies) returned `400 FAILED_PRECONDITION "Precondition check failed."`; the pricing page lists Batch as **"Free Tier: Not available"** and this key is on the free tier. `GET /v1beta/batches` → `200 {}` `LIVE_VERIFIED`. Lifecycle below is therefore documented, not observed. **Sources:** [Batch API guide](https://ai.google.dev/gemini-api/docs/batch-api) · [Batch API reference](https://ai.google.dev/api/batch-api) · [Batch Mode (older twin)](https://ai.google.dev/api/batch-mode) · [Pricing](https://ai.google.dev/gemini-api/docs/pricing) · [Rate limits #batch](https://ai.google.dev/gemini-api/docs/rate-limits) · [Webhooks](https://ai.google.dev/gemini-api/docs/webhooks) · discovery v1beta rev. 20260918 · `python-genai` `BatchJob`, `JobState`. **Last verified:** 2026-09-18. Twins: endpoints fragment (api_family `batch`), `parameters/gemini-batch.json`, objects `GenerateContentBatch`, `BatchStats`, `GenerateContentBatchOutput / InlinedResponse`, `EmbedContentBatch`, lifecycle `GenerateContentBatch / EmbedContentBatch`. ## 1. Endpoints | Method / path | SDK (`google-genai`) | Body | Returns | Live 2026-09-18 | |---|---|---|---|---| | `POST /v1beta/models/{model}:batchGenerateContent` | `client.batches.create(model, src, config)` | `{batch: {displayName, inputConfig: {requests:{requests:[{request, metadata}]}} \| {fileName}, priority?, webhookConfig?}}` | `Operation` (name `batches/{id}`, `metadata` = GenerateContentBatch with `state`) | 400 FAILED_PRECONDITION | | `POST /v1beta/models/{model}:asyncBatchEmbedContent` | `client.batches.create_embeddings` | same envelope with `EmbedContentRequest`s | `Operation` (metadata EmbedContentBatch) | 400 FAILED_PRECONDITION | | `POST /v1beta/tunedModels/{id}:batchGenerateContent` / `:asyncBatchEmbedContent` | — | same | Operation | not tested (tunedModels → 501) | | `GET /v1beta/batches/{id}` | `client.batches.get(name)` | — | `Operation` (`done`, `metadata.state`, `response`) | bogus names → 400 `Could not parse the batch name` | | `GET /v1beta/batches?filter&pageSize&pageToken&returnPartialSuccess` | `client.batches.list()` | — | `ListOperationsResponse` | **200 `{}`** | | `POST /v1beta/batches/{id}:cancel` | `client.batches.cancel(name)` | empty | `{}`; op ends with `error.code = 1` (CANCELLED) | not reachable | | `DELETE /v1beta/batches/{id}` | `client.batches.delete(name)` | — | `{}` (does **not** cancel) | not reachable | | `PATCH /v1beta/batches/{id}:updateGenerateContentBatch?updateMask=` | — | `GenerateContentBatch` (priority…) | `GenerateContentBatch` | 400 FAILED_PRECONDITION (no batch existed) | | `PATCH /v1beta/batches/{id}:updateEmbedContentBatch?updateMask=` | — | `EmbedContentBatch` | `EmbedContentBatch` | not tested | Node: `ai.batches.create / createEmbeddings / get / list / cancel / delete`. OpenAI-compatibility layer also exposes Batch (`/v1beta/openai/batches`). ## 2. Input formats **Inline** (whole request < 20 MB; output comes back inline): ```json {"batch": {"display_name": "my-batch", "input_config": {"requests": {"requests": [ {"request": {"contents": [{"parts": [{"text": "Describe photosynthesis."}]}], "generationConfig": {"temperature": 0.7}}, "metadata": {"key": "request-1"}}, {"request": {"contents": [{"parts": [{"text": "Reply with OK."}]}], "generation_config": {"responseModalities": ["TEXT","IMAGE"]}}, "metadata": {"key": "request-2"}} ]}}}} ``` **File** (JSONL uploaded with the Files API, ≤ 2 GB; output = JSONL `responsesFile`): one object per line `{"key": "request-1", "request": {}}` (the guide's shell sample also shows bare request objects per line), then `{"batch": {"display_name": "…", "input_config": {"file_name": "files/123456"}}}`. Each `request` may carry `systemInstruction`, `tools`, `generationConfig` (incl. `responseModalities` for image generation with `gemini-3-pro-image-preview`), `cachedContent` (standard caching rates apply), other modalities via `fileData`. The `model` comes from the URL. Embeddings: `{"request": {"content": {"parts": [{"text": "OK"}]}, "embedContentConfig": {"taskType", "outputDimensionality", "title"}}, "metadata": {...}}` on `gemini-embedding-2` / `-001` (live methods list `asyncBatchEmbedContent`). ## 3. States and lifecycle | REST `state` (`BatchState`) | Guide prose / SDK `JobState` | Meaning | |---|---|---| | `BATCH_STATE_PENDING` | `JOB_STATE_PENDING` (SDK also `QUEUED`) | created, waiting — initial | | `BATCH_STATE_RUNNING` | `JOB_STATE_RUNNING` (SDK also `UPDATING`, `PAUSED`) | executing; `batchStats.pendingRequestCount` decreases | | `BATCH_STATE_SUCCEEDED` | `JOB_STATE_SUCCEEDED` (SDK also `PARTIALLY_SUCCEEDED`) | terminal; `done:true`, `response.inlinedResponses[]` or `response.responsesFile` | | `BATCH_STATE_FAILED` | `JOB_STATE_FAILED` | terminal; `error` (Status) | | `BATCH_STATE_CANCELLED` | `JOB_STATE_CANCELLED` (SDK `CANCELLING` in between) | after `:cancel` | | `BATCH_STATE_EXPIRED` | `JOB_STATE_EXPIRED` | pending/running > **48 h**; no results | Poll `GET /v1beta/batches/{id}`: `done` false → keep polling; `metadata.state`; when `SUCCEEDED`, `response.inlinedResponses.inlinedResponses[] = {metadata, response | error}` (same order as input) or `response.responsesFile` → `GET https://generativelanguage.googleapis.com/download/v1beta/{responsesFile}:download?alt=media`. Results are kept **6 weeks**. Target turnaround **24 h** ("in majority of cases much quicker"). Webhooks: subscribe to `batch.succeeded` / `batch.failed` (`POST /v1/webhooks`) or set `batch.webhookConfig.uris[]` per batch. Timestamps: `createTime`, `updateTime`, `endTime`. `batchStats = {requestCount, successfulRequestCount, failedRequestCount, pendingRequestCount}` (int64 strings). ## 4. Pricing and limits * **50 %** of the interactive price for every model that lists `batchGenerateContent` (text, image, TTS models …; see each pricing section's "Batch" table — e.g. `gemini-3.1-flash-lite-image` $0.0168 per 1K image, `gemini-3.1-flash-tts-preview` $0.50 / $10). Not available on the free tier. Cached-content hits billed at standard caching rates. * Rate limits: **100 concurrent batch requests**; per-model **enqueued-token caps** (tables on the rate-limits page, separate from interactive limits). * Inline request ≤ 20 MB; input file ≤ 2 GB; creation is **not idempotent**; check `batchStats.failedRequestCount` and per-line status objects. * Supported models = those with `batchGenerateContent` / `asyncBatchEmbedContent` in `models.list` (`gemini-3.5-transcribe`, Lyria, Veo, Live models: **no**). ## 5. Live log (2026-09-18, $0) | Call | Result | |---|---| | `POST …gemini-3.5-flash-lite:batchGenerateContent` inline 2× "Reply with OK." | 400 FAILED_PRECONDITION | | `POST …gemini-3.8-flash:batchGenerateContent` (guide body, 1 request) | 400 FAILED_PRECONDITION | | body without `displayName`; body with `file_name: files/bogus` | 400 FAILED_PRECONDITION (identical → gate precedes validation) | | `POST …gemini-embedding-2:asyncBatchEmbedContent` 1 request | 400 FAILED_PRECONDITION | | `GET /v1beta/batches?pageSize=5` | 200 `{}` | | `GET /v1beta/batches/bogus-batch`, `/batches/123456789` | 400 INVALID_ARGUMENT `Could not parse the batch name` | Interpretation: account/plan restriction (Batch has no free tier), not a payload problem. Examples in `examples/gemini/batch/` are `UNVERIFIED`; tests gated by `RUN_BATCH_TESTS` will `xfail`/skip on `FAILED_PRECONDITION`.