SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
4.8 KB

# OpenAI Moderation API (POST /v1/moderations)

Status: DOCUMENTED + LIVE_VERIFIED (text and text+image with omni-moderation-latest). text-moderation-* RETIRED (removed 2025-10-27; live → 400). Sources: Moderations reference · Moderation guide · Pricing (Free) · Deprecations · Changelog (inline moderation scores) · model pages omni-moderation-latest, text-moderation-latest|stable · OpenAPI CreateModerationRequest / CreateModerationResponse. Last verified: 2026-09-18. Machine-readable: endpoints fragment, parameters/openai-moderations.json, objects/openai-media-objects.json.

# Models

Model Inputs Categories Price Status Live
omni-moderation-latest (default) → omni-moderation-2024-09-26 text, image_url (URL or data URL, ≤ 20 MB); no audio 13 (incl. illicit, illicit/violent) Free DOCUMENTED, LIVE_VERIFIED 200 on text; 200 on text + 1×1 PNG data URL
text-moderation-latest / -stable / -007 text 11 (illicit* = null) free RETIRED 2025-10-27 400 invalid_request_error param=model "Invalid value for 'model' = text-moderation-latest"

Both omni-* ids appear in live /v1/models. Moderation results may also be requested inline with a generation: top-level moderation: {"model": "omni-moderation-latest"} in POST /v1/responses or POST /v1/chat/completions returns scores for the input and the output (after the full output when streaming; may hold an error object) — see the Responses/Chat agents' docs.

# Request

json
{"model": "omni-moderation-latest",
 "input": "a string"  |  ["several", "strings"]  |
          [{"type":"text","text":"…"}, {"type":"image_url","image_url":{"url":"https://… | data:image/png;base64,…"}}]}

One multimodal array = one combined result; an array of strings = one result per string (same order).

# Response and score semantics

json
{"id":"modr-6765","model":"omni-moderation-latest","results":[{
  "flagged": false,
  "categories": {"harassment": false, "...": false},
  "category_scores": {"harassment": 5.5e-05, "violence": 0.00054, "...": 0},
  "category_applied_input_types": {"harassment": ["text"], "violence": ["text","image"], "...": ["text"]}}]}
  • flagged: any category flagged (model-chosen thresholds).
  • categories: per-category booleans.
  • category_scores: confidence 0–1; not calibrated probabilities — OpenAI upgrades the model over time and warns that custom thresholds on category_scores may need recalibration.
  • category_applied_input_types: which input modalities contributed. Live: text-only input → every category ["text"]; text + image → image-capable categories ["text","image"], others ["text"]. Image-only input yields score 0 for text-only categories.

# Categories

Category Description (docs) Inputs
harassment expresses, incites or promotes harassing language towards any target text
harassment/threatening harassment incl. violence or serious harm text
hate hate based on race, gender, ethnicity, religion, nationality, sexual orientation, disability, caste (non-protected groups → harassment) text
hate/threatening hateful content incl. violence/serious harm text
illicit advice/instructions for wrongdoing (e.g. "how to shoplift") text
illicit/violent illicit + violence or weapon procurement text
self-harm promotes, encourages or depicts self-harm (suicide, cutting, eating disorders) text, image
self-harm/intent speaker expresses intent/engagement in self-harm text, image
self-harm/instructions encourages or instructs self-harm text, image
sexual content meant to arouse; sexual services (excl. sex-ed/wellness) text, image
sexual/minors sexual content involving under-18s text
violence depicts death, violence, physical injury text, image
violence/graphic … in graphic detail text, image

# Usage guidance (guide + safety best practices)

Free to use; recommended to screen both user input and model output; treat scores as signals for your own policy (filter, route to review, act on accounts), not as an automatic block; pair with safety_identifier (see docs/openai/safety.md). Images: URL or base64 data URL, ≤ 20 MB; audio not classified.

# Examples and tests

examples/openai/moderation/ — moderate.sh|py|ts (text) and moderate_image.py (multimodal data URL), all LIVE_VERIFIED. tests/openai/test_moderation.py (cheap, always runs).