SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
3.9 KB

# Gemini safety — safetySettings, safetyRatings, promptFeedback, finish reasons

Status: DOCUMENTED + LIVE_VERIFIED (8 harmless-prompt probes on gemini-3.5-flash-lite, 2026-09-18). Blocking behaviour itself was not exercised (no unsafe prompts sent). Sources: https://ai.google.dev/gemini-api/docs/safety-settings · https://ai.google.dev/gemini-api/docs/safety-guidance · https://ai.google.dev/api/generate-content#safetysetting · discovery SafetySetting, SafetyRating, PromptFeedback Machine-readable: generated/fragments/parameters/gemini-generate-content.json (safetySettings*, enableEnhancedCivicAnswers), generated/fragments/objects/gemini-core-objects.json#SafetyRating|PromptFeedback|FinishReason Last verified: 2026-09-18

# 1. Request: safetySettings[]

json
"safetySettings": [{"category": "HARM_CATEGORY_HATE_SPEECH", "threshold": "BLOCK_LOW_AND_ABOVE"}]
Category (accepted live) Meaning
HARM_CATEGORY_HARASSMENT negative/harmful comments targeting identity or protected attributes
HARM_CATEGORY_HATE_SPEECH rude, disrespectful, profane
HARM_CATEGORY_SEXUALLY_EXPLICIT sexual acts / lewd content
HARM_CATEGORY_DANGEROUS_CONTENT promotes/facilitates harmful acts
HARM_CATEGORY_CIVIC_INTEGRITY deprecated (use generationConfig.enableEnhancedCivicAnswers); still accepted; response still rates only the 4 categories above
HARM_CATEGORY_JAILBREAK jailbreak attempts (new in discovery 20260918); accepted live
PaLM categories (DEROGATORY, TOXICITY, VIOLENCE, SEXUAL, MEDICAL, DANGEROUS) 400 * GenerateContentRequest.safety_settings[0]: element predicate failed: $.category in (HATE_SPEECH, SEXUALLY_EXPLICIT, DANGEROUS_CONTENT, HARASSMENT, CIVIC_INTEGRITY, JAILBREAK)
unknown string 400 Invalid value at 'safety_settings[0].category' (…HarmCategory), "HARM_CATEGORY_FOO"
Threshold AI Studio label Blocks when probability is
OFF Off never (filter disabled)
BLOCK_NONE Block none never (always show)
BLOCK_ONLY_HIGH Block few HIGH
BLOCK_MEDIUM_AND_ABOVE Block some MEDIUM, HIGH
BLOCK_LOW_AND_ABOVE Block most LOW, MEDIUM, HIGH
HARM_BLOCK_THRESHOLD_UNSPECIFIED — model default
  • Default when omitted: Off for Gemini 2.5 and 3 models. Core harms (e.g. child safety) are always blocked and not adjustable. Less restrictive settings may trigger a ToS review.
  • Blocking is by probability, not severity.
  • One entry per category is the rule; a duplicate category was silently accepted live.
  • SDK SafetySetting.method (SEVERITY|PROBABILITY) is Vertex-only.

# 2. Response

  • candidates[].safetyRatings[] — returned only when you sent safetySettings with a blocking threshold (BLOCK_LOW_AND_ABOVE → 4 ratings {category, probability: NEGLIGIBLE}); with BLOCK_NONE/OFF or no settings the field is absent. Includes "blocked": true on the rating that blocked. probability ∈ NEGLIGIBLE | LOW | MEDIUM | HIGH.
  • candidates[].finishReason == SAFETY → content withheld; inspect safetyRatings. Other filter-related reasons: RECITATION, LANGUAGE, BLOCKLIST, PROHIBITED_CONTENT, SPII, IMAGE_SAFETY, IMAGE_PROHIBITED_CONTENT, IMAGE_RECITATION, ESCALATION, PUP_LIMITED_DISABLED.
  • promptFeedback.blockReason ∈ SAFETY | OTHER | BLOCKLIST | PROHIBITED_CONTENT | IMAGE_SAFETY (+ promptFeedback.safetyRatings[]) → the prompt was blocked, candidates absent; rephrase.
  • AQA (models/aqa:generateAnswer) returns 4 ratings on answer and on inputFeedback unconditionally (verified).

# 3. Civic integrity

generationConfig.enableEnhancedCivicAnswers: true (200 live) replaces the deprecated HARM_CATEGORY_CIVIC_INTEGRITY setting; "may not be available for all models".

Examples: examples/gemini/safety/. Tests: tests/gemini/test_generate_content.py::test_safety_ratings_present_when_settings_sent.