Gemini API — errors, safety blocks and retry matrix
Status: DOCUMENTED + LIVE_VERIFIED (400 / 404 / 429 bodies captured 2026-09-19). Twin: generated/fragments/errors/gemini-errors.json.
Sources: https://ai.google.dev/gemini-api/docs/api-errors · https://ai.google.dev/gemini-api/docs/troubleshooting · https://ai.google.dev/api/generate-content#FinishReason · https://ai.google.dev/api/generate-content#BlockReason · https://ai.google.dev/gemini-api/docs/rate-limits · https://ai.google.dev/gemini-api/docs/flex-inference · https://ai.google.dev/gemini-api/docs/live-api/session
Last verified: 2026-09-18 / 2026-09-19.
1. Two error envelopes
| Surface | Shape |
|---|---|
Native REST (models.*, files, caches, batches…) |
google.rpc.Status: `{"error":{"code":,"message":"…","status":"<CANONICAL_CODE>","details":[{"@type":"type.googleapis.com/google.rpc.Help" |
Interactions API (/v1beta/interactions, /v1/interactions) |
{"error":{"code":"<snake_case>","message":"…"}}; in SSE streams an event with event_type: "error" carrying the same object; unknown codes fall back to the snake_case HTTP status |
OpenAI-compat (/v1beta/openai/*) |
google.rpc.Status, sometimes wrapped in an array: [{"error":{"code":400,"message":"Missing or invalid Authorization header.","status":"INVALID_ARGUMENT"}}] |
2. Catalogue and retry matrix
| HTTP | status (native) |
Interactions code |
Typical causes (live messages in quotes) | Retry? |
|---|---|---|---|---|
| 400 | INVALID_ARGUMENT | invalid_request, parameter_unknown | malformed body; thinking_level + thinking_budget together; invalid key "API key not valid…"; OpenAI-compat without Bearer "Missing or invalid Authorization header." |
no |
| 400 | FAILED_PRECONDITION | failed_precondition | billing required (paid-only model/feature); free tier not available in region | no (fix billing/region) |
| 401 | UNAUTHENTICATED | authentication | missing/expired OAuth or ephemeral token | no |
| 403 | PERMISSION_DENIED | permission_denied | API not enabled, key restricted to other APIs, leaked key "Your API key was reported as leaked…", foreign tuned model/corpus |
no |
| 404 | NOT_FOUND | not_found, model_not_found | "Model is not found: models/imagen-4.0-generate-001 for api version v1beta"; "…gemini-3.1-pro-preview for api version v1" (preview on v1); "This model models/gemini-2.5-flash-lite is no longer available to new users. Please update your code to use models/gemini-3.5-flash-lite … We recommend you to use the Interactions API." (GET still 200) |
no |
| 409 | ALREADY_EXISTS / ABORTED | already_exists, aborted | duplicate store/document; concurrency conflict | aborted: yes |
| 416 | OUT_OF_RANGE | out_of_range | parameter outside range | no |
| 429 | RESOURCE_EXHAUSTED | rate_limit_exceeded, quota_exceeded, too_many_requests | RPM/TPM/RPD, spend limit, Flex shed; details = Help + QuotaFailure(quotaMetric, quotaId, quotaDimensions{model,location}); "Please retry in 54.2s" in message; free tier limit: 0 on Pro |
yes, backoff + jitter, honour seconds in message; no Retry-After header |
| 499 | CANCELLED | cancelled | client closed connection | n/a |
| 500 | INTERNAL | api_error | server error / oversized input | yes (then reduce input, switch model) |
| 501 | UNIMPLEMENTED | unimplemented | feature not supported for model/version | no |
| 503 | UNAVAILABLE | service_unavailable | overloaded "The model is overloaded…", Flex capacity |
yes (backoff; consider Priority or a Flash fallback) |
| 504 | DEADLINE_EXCEEDED | deadline_exceeded | long prompt/thinking, Flex queueing | yes; raise client timeout, stream, background=true, Batch |
| 200 | — | — | soft failures: promptFeedback.blockReason, candidates[].finishReason ≠ STOP (see §3) |
depends |
| WS close | — | — | Live API: session lifetime (15 min audio / 2 min video), ~10 min connection reset, goAway |
reconnect with sessionResumption handle (2 h) |
SDK defaults (troubleshooting.md): Python retries 429/5xx/timeouts up to 4 attempts, initial ≈1 s, max 60 s; tune via HttpRetryOptions(attempts, initial_delay, max_delay, exp_base, jitter, http_status_codes).
3. Safety and generation stops (HTTP 200)
Prompt blocked → no candidates. promptFeedback.blockReason: SAFETY (see safetyRatings), OTHER (terms / unsupported), BLOCKLIST, PROHIBITED_CONTENT, IMAGE_SAFETY.
Candidate stopped. candidates[0].finishReason:
| Value | Meaning / action |
|---|---|
| STOP | natural end or stop sequence |
| MAX_TOKENS | maxOutputTokens reached — with thinking models the budget can be consumed by thoughts (live: maxOutputTokens: 8 → empty text, thoughtsTokenCount: 5); raise the limit or lower thinking_level |
| SAFETY / RECITATION / LANGUAGE / OTHER / BLOCKLIST / PROHIBITED_CONTENT / SPII | content filtered; RECITATION → make prompt unique / raise temperature; LANGUAGE → unsupported language |
| IMAGE_SAFETY / IMAGE_PROHIBITED_CONTENT / IMAGE_RECITATION / IMAGE_OTHER / NO_IMAGE | image-generation stops |
| MALFORMED_FUNCTION_CALL / MALFORMED_RESPONSE | unparsable model output — retry once |
| UNEXPECTED_TOOL_CALL / TOO_MANY_TOOL_CALLS | tool declared? loop guard |
| MISSING_THOUGHT_SIGNATURE | Gemini 3 multi-turn function calling without echoing thoughtSignature parts — fix history (stateless mode) or use Interactions stateful mode |
Interactions API mirrors these as snake_case generation codes (safety, recitation, language, prohibited_content, spii, blocklist, image_*, content_blocked, malformed_function_call, malformed_tool_call, unexpected_tool_call, no_image, too_many_tool_calls, missing_thought_signature).
4. Handler snippets
Python (google-genai 2.24)
import random, time
from google import genai
from google.genai import errors, types
client = genai.Client(http_options=types.HttpOptions(
retry_options=types.HttpRetryOptions(attempts=4, initial_delay=1.0, max_delay=60.0, exp_base=2.0, jitter=1.0)))
def generate(model: str, prompt: str) -> str:
try:
r = client.models.generate_content(model=model, contents=prompt,
config=types.GenerateContentConfig(max_output_tokens=256))
except errors.ClientError as e: # 4xx (SDK already retried 429)
if e.code == 404 and "no longer available to new users" in str(e):
return generate("gemini-3.5-flash-lite", prompt) # 2.5 family closed to new keys
if e.code == 429: # free tier limit 0 / quota: back off or upgrade tier
time.sleep(min(60, random.uniform(1, 5)))
raise
except errors.ServerError: # 5xx after retries -> fall back
return generate("gemini-3.5-flash-lite", prompt)
if r.prompt_feedback and r.prompt_feedback.block_reason:
raise RuntimeError(f"prompt blocked: {r.prompt_feedback.block_reason}")
cand = r.candidates[0]
if cand.finish_reason and cand.finish_reason.name not in ("STOP", "MAX_TOKENS"):
raise RuntimeError(f"generation stopped: {cand.finish_reason.name}")
return r.text or ""TypeScript (@google/genai 2.23)
import { GoogleGenAI, ApiError } from "@google/genai";
const ai = new GoogleGenAI({ httpOptions: { timeout: 60_000 } });
const sleep = (ms: number) => new Promise(r => setTimeout(r, ms));
export async function generate(model: string, prompt: string, attempt = 0): Promise<string> {
try {
const r = await ai.models.generateContent({ model, contents: prompt, config: { maxOutputTokens: 256 } });
if (r.promptFeedback?.blockReason) throw new Error(`prompt blocked: ${r.promptFeedback.blockReason}`);
const fr = r.candidates?.[0]?.finishReason;
if (fr && !["STOP", "MAX_TOKENS"].includes(fr)) throw new Error(`generation stopped: ${fr}`);
return r.text ?? "";
} catch (e) {
if (e instanceof ApiError) {
const retryable = [429, 500, 503, 504].includes(e.status);
if (retryable && attempt < 4) {
const m = /retry in ([\d.]+)s/.exec(e.message); // 429 message carries the delay
await sleep(m ? Number(m[1]) * 1000 : 2 ** attempt * 1000 + Math.random() * 500);
return generate(model, prompt, attempt + 1);
}
if (e.status === 404 && e.message.includes("no longer available to new users"))
return generate("gemini-3.5-flash-lite", prompt, attempt);
}
throw e;
}
}OpenAI-compat callers
Send the key as Authorization: Bearer; unwrap array-wrapped error bodies; unknown parameters are silently ignored (not 400), so validate features against openai-compatibility.