OpenAI Safety surface: identifiers, alerts/cases APIs, content provenance, platform safeguards
Status: safety_identifier / OpenAI-Safety-Identifier DOCUMENTED · GET /v1/safety/alerts/{id} DOCUMENTED (guide + spec) + LIVE_VERIFIED reachable (404 safety_alert_not_found) · GET /v1/safety/cases/{id} LIVE_DISCOVERED (spec-only) + ACCOUNT_RESTRICTED (403, missing scope api.safety.read) · POST /v1/content_provenance_checks DOCUMENTED + LIVE_VERIFIED · safety classifiers / cybersecurity / misalignment monitoring DOCUMENTED (behavioural, not testable).
Sources: Safety best practices · Safety classifiers · Misalignment monitoring · Cybersecurity checks · Under-18 guidance · Content provenance · Moderation · OpenAPI SafetyAlertResource, SafetyCaseResource, ProvenanceResource, WebhookSafety*.
Last verified: 2026-09-18.
Machine-readable: endpoints fragment, parameters/openai-safety.json, objects/openai-media-objects.json.
1. Safety identifiers
| Surface | How | Constraint |
|---|---|---|
POST /v1/responses, POST /v1/chat/completions |
body safety_identifier: "<hash>" |
string ≤ 64 chars; hash username/email; session id for anonymous previews; replaces legacy user (with prompt_cache_key for caching) |
Realtime (POST /v1/realtime/client_secrets, direct WS/WebRTC from a trusted backend) |
header OpenAI-Safety-Identifier: <hash> |
identifiers do not carry over between APIs/sessions — send on each |
| Images / Embeddings / Audio … | legacy user body field |
"help OpenAI monitor and detect abuse" |
Purpose: lets OpenAI throttle/block a single abusive end user (error identifier blocked on future GPT-5 requests; cannot currently be unblocked) instead of the whole organization, and gives you actionable abuse feedback. Recommended, not required.
2. Platform safeguards (what OpenAI applies)
| Safeguard | Scope | Developer-visible effect | Source |
|---|---|---|---|
| Safety classifiers (bio/chem risk) | GPT-5 family | requests classified into risk thresholds; repeated high-risk → error + warning e-mail → org access to GPT-5 stopped after ~7 days; possible delayed streaming while extra checks run (add a spinner); per-safety_identifier blocks |
safety-checks guide |
| Cybersecurity checks | GPT-5.3-Codex and newer (5.4, 5.5) — "High Cybersecurity Capability" | error code cyber_policy; access temporarily revoked for the identifier or the whole org (if no safety_identifier); ZDR orgs also get request-level rejections (can appear mid-stream); appeal via support within the 7-day window; Trusted Access for Cyber aliases gpt-daybreak-blue-latest → gpt-5.6-sol, gpt-daybreak-red-latest → gpt-5.6-cyber |
cybersecurity guide |
| Misalignment monitoring | Responses API (persisted reasoning / WebSockets / compaction = monitored and auto-stopped; other Responses requests = monitored, webhook only); Chat Completions not covered | HTTP 403 invalid_request_error code misalignment_policy_violation before streaming starts; no resume; earlier actions are not undone; details via safety alerts |
misalignment guide |
| Computer-use safety checks | computer_use_preview tool |
pending_safety_checks on the computer_call (malicious_instructions, irrelevant_domain, sensitive_domain) must be echoed back as acknowledged_safety_checks after human confirmation — owned by the tools agent, see docs/tools/ |
tools-computer-use guide |
| Fine-tuning safety checks | supervised fine-tuning | training data screened; see fine-tuning agent | supervised-fine-tuning guide |
| Image / video content filters | GPT image (`moderation: low | auto`), Sora (under-18 content only, no real people, no copyrighted characters/music) | image_generation_user_error; failed video jobs with error.misalignment |
3. Safety alerts and cases APIs
Webhooks → ids
| Event | Sent when | data.id → endpoint |
|---|---|---|
safety.alert.created |
approved misalignment alert for an API project | alert_<32 hex> → GET /v1/safety/alerts/{id} |
safety.org_alert.created |
same, enterprise workspace | GET /v1/safety/alerts/{id} |
safety.warning_issued |
warning for a safety identifier in your org | C-… → GET /v1/safety/cases/{id} |
safety.deactivation_issued |
deactivation of a safety identifier | GET /v1/safety/cases/{id} |
Payload carries only the id; verify the signature, ack, then fetch in background with a key of the same project.
GET /v1/safety/alerts/{id} → safety.alert
| Field | Type | Meaning |
|---|---|---|
id, object: "safety.alert", created_at |
||
request_id, response_id, model |
string | the flagged Responses request |
request_paused |
bool | block registration succeeded (does not confirm execution stopped) |
error_type |
potentially_unintended_data_transfer | potentially_unintended_data_access | potentially_unintended_destructive_activity | other (accept future values) |
|
reason |
string | null | customer-safe description; null for ZDR requests |
Live: GET /v1/safety/alerts/salert_atlasprobe → 404 {"error":{"message":"Safety alert not found.","type":"invalid_request_error","param":null,"code":"safety_alert_not_found"}} — endpoint usable with a standard project key. SDK: client.safety.alerts.retrieve(id) (Python/Node).
GET /v1/safety/cases/{id} → safety.case (spec-only)
Fields: id, object: "safety.case", created_at, entity_identifier (the safety identifier), reason (string|null), notice: {type: "warning" | "deactivation"}. Not in the downloaded guides nor in the Python/Node SDK API lists.
Live: GET /v1/safety/cases/C-atlasprobe → 403 with a non-standard body {"error": "You have insufficient permissions for this operation. Missing scopes: api.safety.read. Check that you have the correct role in your organization, and if you're using a restricted API key, that it has the necessary scopes."} (error is a plain string). Status ACCOUNT_RESTRICTED; the scope name api.safety.read is LIVE_DISCOVERED.
4. Content provenance — POST /v1/content_provenance_checks
Synchronous check of one file for OpenAI provenance signals (not a general AI detector; other vendors' content is not detected).
| Aspect | Value |
|---|---|
| Request | multipart/form-data, single part file with its media type (-F "file=@x.png;type=image/png", Opus → audio/ogg); no extra type field |
| Formats | images PNG, JPEG, WebP; audio MP3, Opus, AAC, FLAC, WAV, PCM (≤ 60 s decoded); ≤ 50 MiB |
| Signals | C2PA Content Credentials (images; metadata, lost on edit/convert) · SynthID watermark (images + audio; survives some transforms) |
| Response | {object: "content_provenance_check", created_at, results: [C2PA?, SynthID]}; inapplicable checks omitted; no top-level outcome |
| C2PA result | `outcome detected |
| SynthID result | outcome, model, generated_at (nullable) |
| Errors | 400 malformed/unsupported/blocked file · 404 org without access · 429 rate_limit_exceeded (strict limits, honor Retry-After; higher limits by application) |
| Policy | not ZDR-eligible; don't use to reverse-engineer/evade watermarks or infer creators; pair with human review |
| SDK | Python ≥ 2.52 client.content_provenance_checks.create(file=(name, fh, "image/png")); Node client.contentProvenanceChecks.create({file: toStreamingFile(...)}) |
Live (2026-09-18): fresh gpt-image-1-mini JPEG → c2pa {outcome: detected, validation_state: trusted, issuer: "OpenAI OpCo, LLC", model: "API / gpt-image", generated_at: "2026-09-19T01:43:43Z"}, synthid {outcome: not_detected}. Hand-made 1×1 PNG → c2pa not_detected / not_present, synthid not_detected. So GPT image outputs carry C2PA (even JPEG) but no SynthID watermark was detected on this model.
5. Safety best practices (checklist from the guide)
Use the free Moderation API (or inline moderation scores) · adversarial/red-team testing · human in the loop for high-stakes and code · prompt engineering to constrain topic/tone · KYC / login · constrain input length and output tokens, prefer validated inputs/outputs · let users report issues · communicate limitations · implement safety identifiers · revoke compromised keys promptly · follow CSAM guidance (NCMEC/Thorn) and the Under-18 API guidance if serving minors · report vulnerabilities via the Coordinated Vulnerability Disclosure Program.
Examples and tests
examples/openai/images/provenance_check.sh|py (LIVE_VERIFIED). Safety alert/case retrieval is exercised (free, error shapes only) by tests/openai/test_moderation.py::test_safety_alert_404_shape and ::test_safety_case_scope_shape. No example creates alerts: they are produced by OpenAI monitoring, not by developers.