Gemini safety — safetySettings, safetyRatings, promptFeedback, finish reasons
Status: DOCUMENTED + LIVE_VERIFIED (8 harmless-prompt probes on gemini-3.5-flash-lite, 2026-09-18). Blocking behaviour itself was not exercised (no unsafe prompts sent).
Sources: https://ai.google.dev/gemini-api/docs/safety-settings · https://ai.google.dev/gemini-api/docs/safety-guidance · https://ai.google.dev/api/generate-content#safetysetting · discovery SafetySetting, SafetyRating, PromptFeedback
Machine-readable: generated/fragments/parameters/gemini-generate-content.json (safetySettings*, enableEnhancedCivicAnswers), generated/fragments/objects/gemini-core-objects.json#SafetyRating|PromptFeedback|FinishReason
Last verified: 2026-09-18
1. Request: safetySettings[]
"safetySettings": [{"category": "HARM_CATEGORY_HATE_SPEECH", "threshold": "BLOCK_LOW_AND_ABOVE"}]| Category (accepted live) | Meaning |
|---|---|
HARM_CATEGORY_HARASSMENT |
negative/harmful comments targeting identity or protected attributes |
HARM_CATEGORY_HATE_SPEECH |
rude, disrespectful, profane |
HARM_CATEGORY_SEXUALLY_EXPLICIT |
sexual acts / lewd content |
HARM_CATEGORY_DANGEROUS_CONTENT |
promotes/facilitates harmful acts |
HARM_CATEGORY_CIVIC_INTEGRITY |
deprecated (use generationConfig.enableEnhancedCivicAnswers); still accepted; response still rates only the 4 categories above |
HARM_CATEGORY_JAILBREAK |
jailbreak attempts (new in discovery 20260918); accepted live |
PaLM categories (DEROGATORY, TOXICITY, VIOLENCE, SEXUAL, MEDICAL, DANGEROUS) |
400 * GenerateContentRequest.safety_settings[0]: element predicate failed: $.category in (HATE_SPEECH, SEXUALLY_EXPLICIT, DANGEROUS_CONTENT, HARASSMENT, CIVIC_INTEGRITY, JAILBREAK) |
| unknown string | 400 Invalid value at 'safety_settings[0].category' (…HarmCategory), "HARM_CATEGORY_FOO" |
| Threshold | AI Studio label | Blocks when probability is |
|---|---|---|
OFF |
Off | never (filter disabled) |
BLOCK_NONE |
Block none | never (always show) |
BLOCK_ONLY_HIGH |
Block few | HIGH |
BLOCK_MEDIUM_AND_ABOVE |
Block some | MEDIUM, HIGH |
BLOCK_LOW_AND_ABOVE |
Block most | LOW, MEDIUM, HIGH |
HARM_BLOCK_THRESHOLD_UNSPECIFIED |
— | model default |
- Default when omitted: Off for Gemini 2.5 and 3 models. Core harms (e.g. child safety) are always blocked and not adjustable. Less restrictive settings may trigger a ToS review.
- Blocking is by probability, not severity.
- One entry per category is the rule; a duplicate category was silently accepted live.
- SDK
SafetySetting.method(SEVERITY|PROBABILITY) is Vertex-only.
2. Response
candidates[].safetyRatings[]— returned only when you sentsafetySettingswith a blocking threshold (BLOCK_LOW_AND_ABOVE → 4 ratings{category, probability: NEGLIGIBLE}); withBLOCK_NONE/OFFor no settings the field is absent. Includes"blocked": trueon the rating that blocked.probability∈NEGLIGIBLE | LOW | MEDIUM | HIGH.candidates[].finishReason == SAFETY→ content withheld; inspectsafetyRatings. Other filter-related reasons:RECITATION,LANGUAGE,BLOCKLIST,PROHIBITED_CONTENT,SPII,IMAGE_SAFETY,IMAGE_PROHIBITED_CONTENT,IMAGE_RECITATION,ESCALATION,PUP_LIMITED_DISABLED.promptFeedback.blockReason∈SAFETY | OTHER | BLOCKLIST | PROHIBITED_CONTENT | IMAGE_SAFETY(+promptFeedback.safetyRatings[]) → the prompt was blocked,candidatesabsent; rephrase.- AQA (
models/aqa:generateAnswer) returns 4 ratings onanswerand oninputFeedbackunconditionally (verified).
3. Civic integrity
generationConfig.enableEnhancedCivicAnswers: true (200 live) replaces the deprecated HARM_CATEGORY_CIVIC_INTEGRITY setting; "may not be available for all models".
Examples: examples/gemini/safety/. Tests: tests/gemini/test_generate_content.py::test_safety_ratings_present_when_settings_sent.