SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
9.4 KB

# OpenAI Safety surface: identifiers, alerts/cases APIs, content provenance, platform safeguards

Status: safety_identifier / OpenAI-Safety-Identifier DOCUMENTED · GET /v1/safety/alerts/{id} DOCUMENTED (guide + spec) + LIVE_VERIFIED reachable (404 safety_alert_not_found) · GET /v1/safety/cases/{id} LIVE_DISCOVERED (spec-only) + ACCOUNT_RESTRICTED (403, missing scope api.safety.read) · POST /v1/content_provenance_checks DOCUMENTED + LIVE_VERIFIED · safety classifiers / cybersecurity / misalignment monitoring DOCUMENTED (behavioural, not testable). Sources: Safety best practices · Safety classifiers · Misalignment monitoring · Cybersecurity checks · Under-18 guidance · Content provenance · Moderation · OpenAPI SafetyAlertResource, SafetyCaseResource, ProvenanceResource, WebhookSafety*. Last verified: 2026-09-18. Machine-readable: endpoints fragment, parameters/openai-safety.json, objects/openai-media-objects.json.

# 1. Safety identifiers

Surface How Constraint
POST /v1/responses, POST /v1/chat/completions body safety_identifier: "<hash>" string ≤ 64 chars; hash username/email; session id for anonymous previews; replaces legacy user (with prompt_cache_key for caching)
Realtime (POST /v1/realtime/client_secrets, direct WS/WebRTC from a trusted backend) header OpenAI-Safety-Identifier: <hash> identifiers do not carry over between APIs/sessions — send on each
Images / Embeddings / Audio … legacy user body field "help OpenAI monitor and detect abuse"

Purpose: lets OpenAI throttle/block a single abusive end user (error identifier blocked on future GPT-5 requests; cannot currently be unblocked) instead of the whole organization, and gives you actionable abuse feedback. Recommended, not required.

# 2. Platform safeguards (what OpenAI applies)

Safeguard Scope Developer-visible effect Source
Safety classifiers (bio/chem risk) GPT-5 family requests classified into risk thresholds; repeated high-risk → error + warning e-mail → org access to GPT-5 stopped after ~7 days; possible delayed streaming while extra checks run (add a spinner); per-safety_identifier blocks safety-checks guide
Cybersecurity checks GPT-5.3-Codex and newer (5.4, 5.5) — "High Cybersecurity Capability" error code cyber_policy; access temporarily revoked for the identifier or the whole org (if no safety_identifier); ZDR orgs also get request-level rejections (can appear mid-stream); appeal via support within the 7-day window; Trusted Access for Cyber aliases gpt-daybreak-blue-latest → gpt-5.6-sol, gpt-daybreak-red-latest → gpt-5.6-cyber cybersecurity guide
Misalignment monitoring Responses API (persisted reasoning / WebSockets / compaction = monitored and auto-stopped; other Responses requests = monitored, webhook only); Chat Completions not covered HTTP 403 invalid_request_error code misalignment_policy_violation before streaming starts; no resume; earlier actions are not undone; details via safety alerts misalignment guide
Computer-use safety checks computer_use_preview tool pending_safety_checks on the computer_call (malicious_instructions, irrelevant_domain, sensitive_domain) must be echoed back as acknowledged_safety_checks after human confirmation — owned by the tools agent, see docs/tools/ tools-computer-use guide
Fine-tuning safety checks supervised fine-tuning training data screened; see fine-tuning agent supervised-fine-tuning guide
Image / video content filters GPT image (`moderation: low auto`), Sora (under-18 content only, no real people, no copyrighted characters/music) image_generation_user_error; failed video jobs with error.misalignment

# 3. Safety alerts and cases APIs

# Webhooks → ids

Event Sent when data.id → endpoint
safety.alert.created approved misalignment alert for an API project alert_<32 hex> → GET /v1/safety/alerts/{id}
safety.org_alert.created same, enterprise workspace GET /v1/safety/alerts/{id}
safety.warning_issued warning for a safety identifier in your org C-… → GET /v1/safety/cases/{id}
safety.deactivation_issued deactivation of a safety identifier GET /v1/safety/cases/{id}

Payload carries only the id; verify the signature, ack, then fetch in background with a key of the same project.

# GET /v1/safety/alerts/{id} → safety.alert

Field Type Meaning
id, object: "safety.alert", created_at
request_id, response_id, model string the flagged Responses request
request_paused bool block registration succeeded (does not confirm execution stopped)
error_type potentially_unintended_data_transfer | potentially_unintended_data_access | potentially_unintended_destructive_activity | other (accept future values)
reason string | null customer-safe description; null for ZDR requests

Live: GET /v1/safety/alerts/salert_atlasprobe → 404 {"error":{"message":"Safety alert not found.","type":"invalid_request_error","param":null,"code":"safety_alert_not_found"}} — endpoint usable with a standard project key. SDK: client.safety.alerts.retrieve(id) (Python/Node).

# GET /v1/safety/cases/{id} → safety.case (spec-only)

Fields: id, object: "safety.case", created_at, entity_identifier (the safety identifier), reason (string|null), notice: {type: "warning" | "deactivation"}. Not in the downloaded guides nor in the Python/Node SDK API lists. Live: GET /v1/safety/cases/C-atlasprobe → 403 with a non-standard body {"error": "You have insufficient permissions for this operation. Missing scopes: api.safety.read. Check that you have the correct role in your organization, and if you're using a restricted API key, that it has the necessary scopes."} (error is a plain string). Status ACCOUNT_RESTRICTED; the scope name api.safety.read is LIVE_DISCOVERED.

# 4. Content provenance — POST /v1/content_provenance_checks

Synchronous check of one file for OpenAI provenance signals (not a general AI detector; other vendors' content is not detected).

Aspect Value
Request multipart/form-data, single part file with its media type (-F "file=@x.png;type=image/png", Opus → audio/ogg); no extra type field
Formats images PNG, JPEG, WebP; audio MP3, Opus, AAC, FLAC, WAV, PCM (≤ 60 s decoded); ≤ 50 MiB
Signals C2PA Content Credentials (images; metadata, lost on edit/convert) · SynthID watermark (images + audio; survives some transforms)
Response {object: "content_provenance_check", created_at, results: [C2PA?, SynthID]}; inapplicable checks omitted; no top-level outcome
C2PA result `outcome detected
SynthID result outcome, model, generated_at (nullable)
Errors 400 malformed/unsupported/blocked file · 404 org without access · 429 rate_limit_exceeded (strict limits, honor Retry-After; higher limits by application)
Policy not ZDR-eligible; don't use to reverse-engineer/evade watermarks or infer creators; pair with human review
SDK Python ≥ 2.52 client.content_provenance_checks.create(file=(name, fh, "image/png")); Node client.contentProvenanceChecks.create({file: toStreamingFile(...)})

Live (2026-09-18): fresh gpt-image-1-mini JPEG → c2pa {outcome: detected, validation_state: trusted, issuer: "OpenAI OpCo, LLC", model: "API / gpt-image", generated_at: "2026-09-19T01:43:43Z"}, synthid {outcome: not_detected}. Hand-made 1×1 PNG → c2pa not_detected / not_present, synthid not_detected. So GPT image outputs carry C2PA (even JPEG) but no SynthID watermark was detected on this model.

# 5. Safety best practices (checklist from the guide)

Use the free Moderation API (or inline moderation scores) · adversarial/red-team testing · human in the loop for high-stakes and code · prompt engineering to constrain topic/tone · KYC / login · constrain input length and output tokens, prefer validated inputs/outputs · let users report issues · communicate limitations · implement safety identifiers · revoke compromised keys promptly · follow CSAM guidance (NCMEC/Thorn) and the Under-18 API guidance if serving minors · report vulnerabilities via the Coordinated Vulnerability Disclosure Program.

# Examples and tests

examples/openai/images/provenance_check.sh|py (LIVE_VERIFIED). Safety alert/case retrieval is exercised (free, error shapes only) by tests/openai/test_moderation.py::test_safety_alert_404_shape and ::test_safety_case_scope_shape. No example creates alerts: they are produced by OpenAI monitoring, not by developers.