# OpenAI Safety surface: identifiers, alerts/cases APIs, content provenance, platform safeguards **Status:** `safety_identifier` / `OpenAI-Safety-Identifier` DOCUMENTED · `GET /v1/safety/alerts/{id}` DOCUMENTED (guide + spec) + LIVE_VERIFIED reachable (404 `safety_alert_not_found`) · `GET /v1/safety/cases/{id}` LIVE_DISCOVERED (spec-only) + **ACCOUNT_RESTRICTED** (403, missing scope `api.safety.read`) · `POST /v1/content_provenance_checks` DOCUMENTED + LIVE_VERIFIED · safety classifiers / cybersecurity / misalignment monitoring DOCUMENTED (behavioural, not testable). **Sources:** [Safety best practices](https://developers.openai.com/api/docs/guides/safety-best-practices) · [Safety classifiers](https://developers.openai.com/api/docs/guides/safety-checks) · [Misalignment monitoring](https://developers.openai.com/api/docs/guides/safety-checks/misalignment-monitoring) · [Cybersecurity checks](https://developers.openai.com/api/docs/guides/safety-checks/cybersecurity) · [Under-18 guidance](https://developers.openai.com/api/docs/guides/safety-checks/under-18-api-guidance) · [Content provenance](https://developers.openai.com/api/docs/guides/content-provenance) · [Moderation](https://developers.openai.com/api/docs/guides/moderation) · OpenAPI `SafetyAlertResource`, `SafetyCaseResource`, `ProvenanceResource`, `WebhookSafety*`. **Last verified:** 2026-09-18. **Machine-readable:** endpoints fragment, `parameters/openai-safety.json`, `objects/openai-media-objects.json`. ## 1. Safety identifiers | Surface | How | Constraint | |---|---|---| | `POST /v1/responses`, `POST /v1/chat/completions` | body `safety_identifier: ""` | string ≤ 64 chars; hash username/email; session id for anonymous previews; replaces legacy `user` (with `prompt_cache_key` for caching) | | Realtime (`POST /v1/realtime/client_secrets`, direct WS/WebRTC from a trusted backend) | header `OpenAI-Safety-Identifier: ` | identifiers do **not** carry over between APIs/sessions — send on each | | Images / Embeddings / Audio … | legacy `user` body field | "help OpenAI monitor and detect abuse" | Purpose: lets OpenAI throttle/block a single abusive end user (error *identifier blocked* on future GPT-5 requests; cannot currently be unblocked) instead of the whole organization, and gives you actionable abuse feedback. Recommended, not required. ## 2. Platform safeguards (what OpenAI applies) | Safeguard | Scope | Developer-visible effect | Source | |---|---|---|---| | **Safety classifiers** (bio/chem risk) | GPT-5 family | requests classified into risk thresholds; repeated high-risk → error + warning e-mail → org access to GPT-5 stopped after ~7 days; possible **delayed streaming** while extra checks run (add a spinner); per-`safety_identifier` blocks | safety-checks guide | | **Cybersecurity checks** | GPT-5.3-Codex and newer (5.4, 5.5) — "High Cybersecurity Capability" | error code **`cyber_policy`**; access temporarily revoked for the identifier or the whole org (if no `safety_identifier`); ZDR orgs also get request-level rejections (can appear mid-stream); appeal via support within the 7-day window; Trusted Access for Cyber aliases `gpt-daybreak-blue-latest` → `gpt-5.6-sol`, `gpt-daybreak-red-latest` → `gpt-5.6-cyber` | cybersecurity guide | | **Misalignment monitoring** | Responses API (persisted reasoning / WebSockets / compaction = monitored **and** auto-stopped; other Responses requests = monitored, webhook only); Chat Completions **not covered** | HTTP **403** `invalid_request_error` code **`misalignment_policy_violation`** before streaming starts; no resume; earlier actions are not undone; details via safety alerts | misalignment guide | | **Computer-use safety checks** | `computer_use_preview` tool | `pending_safety_checks` on the computer_call (malicious_instructions, irrelevant_domain, sensitive_domain) must be echoed back as `acknowledged_safety_checks` after human confirmation — owned by the tools agent, see `docs/tools/` | tools-computer-use guide | | **Fine-tuning safety checks** | supervised fine-tuning | training data screened; see fine-tuning agent | supervised-fine-tuning guide | | **Image / video content filters** | GPT image (`moderation: low|auto`), Sora (under-18 content only, no real people, no copyrighted characters/music) | `image_generation_user_error`; failed video jobs with `error.misalignment` | see `images.md`, `video.md` | ## 3. Safety alerts and cases APIs ### Webhooks → ids | Event | Sent when | `data.id` → endpoint | |---|---|---| | `safety.alert.created` | approved misalignment alert for an API project | `alert_<32 hex>` → `GET /v1/safety/alerts/{id}` | | `safety.org_alert.created` | same, enterprise workspace | `GET /v1/safety/alerts/{id}` | | `safety.warning_issued` | warning for a safety identifier in your org | `C-…` → `GET /v1/safety/cases/{id}` | | `safety.deactivation_issued` | deactivation of a safety identifier | `GET /v1/safety/cases/{id}` | Payload carries only the id; verify the signature, ack, then fetch in background with a key of the same project. ### `GET /v1/safety/alerts/{id}` → `safety.alert` | Field | Type | Meaning | |---|---|---| | `id`, `object: "safety.alert"`, `created_at` | | | | `request_id`, `response_id`, `model` | string | the flagged Responses request | | `request_paused` | bool | block registration succeeded (does **not** confirm execution stopped) | | `error_type` | `potentially_unintended_data_transfer` \| `potentially_unintended_data_access` \| `potentially_unintended_destructive_activity` \| `other` (accept future values) | | | `reason` | string \| null | customer-safe description; null for ZDR requests | Live: `GET /v1/safety/alerts/salert_atlasprobe` → **404** `{"error":{"message":"Safety alert not found.","type":"invalid_request_error","param":null,"code":"safety_alert_not_found"}}` — endpoint usable with a standard project key. SDK: `client.safety.alerts.retrieve(id)` (Python/Node). ### `GET /v1/safety/cases/{id}` → `safety.case` (spec-only) Fields: `id`, `object: "safety.case"`, `created_at`, `entity_identifier` (the safety identifier), `reason` (string|null), `notice: {type: "warning" | "deactivation"}`. Not in the downloaded guides nor in the Python/Node SDK API lists. Live: `GET /v1/safety/cases/C-atlasprobe` → **403** with a **non-standard body** `{"error": "You have insufficient permissions for this operation. Missing scopes: api.safety.read. Check that you have the correct role in your organization, and if you're using a restricted API key, that it has the necessary scopes."}` (error is a plain string). Status ACCOUNT_RESTRICTED; the scope name `api.safety.read` is LIVE_DISCOVERED. ## 4. Content provenance — `POST /v1/content_provenance_checks` Synchronous check of one file for **OpenAI** provenance signals (not a general AI detector; other vendors' content is not detected). | Aspect | Value | |---|---| | Request | `multipart/form-data`, single part `file` with its media type (`-F "file=@x.png;type=image/png"`, Opus → `audio/ogg`); no extra `type` field | | Formats | images PNG, JPEG, WebP; audio MP3, Opus, AAC, FLAC, WAV, PCM (≤ 60 s decoded); ≤ 50 MiB | | Signals | **C2PA Content Credentials** (images; metadata, lost on edit/convert) · **SynthID** watermark (images + audio; survives some transforms) | | Response | `{object: "content_provenance_check", created_at, results: [C2PA?, SynthID]}`; inapplicable checks omitted; no top-level outcome | | C2PA result | `outcome detected|not_detected`, `validation_state trusted|valid|invalid|not_present`, `issuer`, `model`, `generated_at` — `detected` only for a trusted/valid **OpenAI-issued** manifest with an AI-generation action | | SynthID result | `outcome`, `model`, `generated_at` (nullable) | | Errors | 400 malformed/unsupported/blocked file · 404 org without access · 429 `rate_limit_exceeded` (strict limits, honor `Retry-After`; higher limits by application) | | Policy | not ZDR-eligible; don't use to reverse-engineer/evade watermarks or infer creators; pair with human review | | SDK | Python ≥ 2.52 `client.content_provenance_checks.create(file=(name, fh, "image/png"))`; Node `client.contentProvenanceChecks.create({file: toStreamingFile(...)})` | Live (2026-09-18): fresh `gpt-image-1-mini` JPEG → `c2pa {outcome: detected, validation_state: trusted, issuer: "OpenAI OpCo, LLC", model: "API / gpt-image", generated_at: "2026-09-19T01:43:43Z"}`, `synthid {outcome: not_detected}`. Hand-made 1×1 PNG → `c2pa not_detected / not_present`, `synthid not_detected`. So GPT image outputs carry C2PA (even JPEG) but no SynthID watermark was detected on this model. ## 5. Safety best practices (checklist from the guide) Use the free Moderation API (or inline `moderation` scores) · adversarial/red-team testing · human in the loop for high-stakes and code · prompt engineering to constrain topic/tone · KYC / login · constrain input length and output tokens, prefer validated inputs/outputs · let users report issues · communicate limitations · **implement safety identifiers** · revoke compromised keys promptly · follow CSAM guidance (NCMEC/Thorn) and the Under-18 API guidance if serving minors · report vulnerabilities via the Coordinated Vulnerability Disclosure Program. ## Examples and tests `examples/openai/images/provenance_check.sh|py` (LIVE_VERIFIED). Safety alert/case retrieval is exercised (free, error shapes only) by `tests/openai/test_moderation.py::test_safety_alert_404_shape` and `::test_safety_case_scope_shape`. No example creates alerts: they are produced by OpenAI monitoring, not by developers.