SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
24.9 KB

# OpenAI Responses API — complete reference

Status: DOCUMENTED + LIVE_VERIFIED (all HTTP endpoints exercised 2026-09-18 with gpt-5.4-nano / gpt-4.1-nano; WebSocket mode DOCUMENTED/UNVERIFIED). Machine-readable twins: generated/fragments/endpoints/openai-core.json, generated/fragments/parameters/openai-responses.json, generated/fragments/objects/openai-responses-objects.json, generated/fragments/streaming-events/openai-responses*.json.

Sources

  • https://developers.openai.com/api/reference/resources/responses/methods/create (+ retrieve, delete, cancel, compact, input_items/list, input_tokens/count, streaming-events, websocket-events)
  • https://developers.openai.com/api/docs/guides/text · conversation-state · background · compaction · token-counting · websocket-mode · streaming-responses · migrate-to-responses
  • OpenAPI spec sources/openai/openapi/openapi-master.yaml (schemas CreateResponse, Response, ResponseStreamEvent, ResponsesClientEvent, ResponsesServerEvent, CompactResponseMethodPublicBody, TokenCountsBody)
  • Live probes: tmp-live/openai-core/*.json (sanitized), log reports/live-requests.jsonl

Last verified: 2026-09-18


# 1. Endpoints

Method Path Purpose Status Live note (2026-09-18)
POST /v1/responses Create a response (sync, stream=true SSE, or background=true) LIVE_VERIFIED gpt-5.4-nano, 10 in / 6 out tokens, status=completed
POST /v1/responses?beta=true Same operation on the beta OpenAPI surface (BetaCreateResponse; adds multi_agent, agent_message/multi_agent_call items, optional openai-beta header) BETA, LIVE_VERIFIED identical key set to the GA response
GET /v1/responses/{id} Retrieve a stored response; ?include[]=…, ?stream=true&starting_after=N re-attaches to a background stream LIVE_VERIFIED 200; store=false id → 404; deleted id → 404
DELETE /v1/responses/{id} Delete a stored response → {id, object:"response.deleted", deleted:true} LIVE_VERIFIED 200
POST /v1/responses/{id}/cancel Cancel a background response; idempotent while cancelled/in-flight LIVE_VERIFIED immediate cancel → status=cancelled; 2nd cancel → same object; cancel after completion → 400 "Cannot cancel a completed response."
GET /v1/responses/{id}/input_items List the rendered input context (cursor pagination limit 1–100/20, order asc|desc (default desc), after, include[]) LIVE_VERIFIED chain child lists 3 items
POST /v1/responses/input_tokens Count input tokens of a Responses payload (free) → {object:"response.input_tokens", input_tokens} LIVE_VERIFIED 10 tokens = real usage.input_tokens
POST /v1/responses/compact Stateless compaction → object:"response.compaction" LIVE_VERIFIED [message, message, compaction]
WS wss://api.openai.com/v1/responses WebSocket mode (response.create / response.steer client events) DOCUMENTED, UNVERIFIED not run (needs WS client dep)

Auth: Authorization: Bearer $OPENAI_API_KEY (+ optional OpenAI-Organization, OpenAI-Project). Observed response headers: x-request-id, openai-processing-ms, openai-version: 2020-10-01, x-ratelimit-{limit,remaining,reset}-{requests,tokens} (account-specific values, not documented limits).

SDK surface (verified installed): Python openai==3.16.2 → client.responses.{create,retrieve,delete,cancel,compact,stream,parse}, client.responses.input_items.list, client.responses.input_tokens.count, client.responses.connect() (WS). Node openai@7.18.0 → client.responses.{create,retrieve,delete,cancel,compact,stream,parse}, client.responses.inputItems.list, client.responses.inputTokens.count, ResponsesWS.


# 2. Request body — POST /v1/responses

Notation: a.b nested object, x[] array element, x[](t) union member discriminated by type (or role). Full list (511 rows incl. every input item type) in generated/fragments/parameters/openai-responses.json.

# 2.1 Top-level parameters

Parameter Type Default Constraints / enum Status Notes
model string — any Responses model id (ModelIdsResponses) LIVE used gpt-5.4-nano (snapshot echoed as gpt-5.4-nano-2026-03-17)
input string | Item[] — see §2.2 LIVE string ≡ one user message
instructions string | null null — LIVE system/developer message; not inherited through previous_response_id
previous_response_id string | null null mutually exclusive with conversation LIVE 400 mutually_exclusive_parameters if both; 400 previous_response_not_found if unknown
conversation string | {id} | null null LIVE items of the response are appended to the conversation
store boolean | null true LIVE 30-day retention; false → not retrievable (404) and reasoning items carry encrypted_content
background boolean | null false not over WebSocket LIVE returns status=queued immediately
stream boolean | null false LIVE SSE; see docs/openai/streaming-events.md
stream_options.include_obfuscation boolean true DOC deltas carry an obfuscation pad unless false
max_output_tokens integer | null null minimum 16 LIVE 8 → 400 integer_below_min_value; budget includes reasoning + formatting tokens
max_tool_calls integer | null null DOC cap on built-in tool calls per response
parallel_tool_calls boolean | null true LIVE
tools[] Tool[] — 16 tool types (function, file_search, web_search(+preview), computer(+preview), mcp, code_interpreter, image_generation, local_shell, shell, apply_patch, custom, namespace, tool_search, programmatic_tool_calling) DOC shapes owned by the tools fragments
tool_choice "none"|"auto"|"required" | object auto objects: allowed_tools{mode, tools[]}, hosted {type}, function{name}, mcp{server_label,name}, custom{name}, programmatic_tool_calling, apply_patch, shell DOC
text.format {type:text} | {type:json_schema,name,schema,strict,description} | {type:json_object} text LIVE see docs/openai/structured-outputs.md
text.verbosity low|medium|high | null medium GPT-5+ only LIVE gpt-4.1-nano + low → 400 unsupported_value ("Supported values are: 'medium'")
reasoning.effort none|minimal|low|medium|high|xhigh|max | null medium (model-dependent) reasoning models only LIVE gpt-5.4-nano echoes effort:"none" when omitted; gpt-4.1-nano → 400 unsupported_parameter
reasoning.summary auto|concise|detailed | null null LIVE auto echoed as detailed; summary items only when reasoning actually happened
reasoning.context auto|current_turn|all_turns | null auto all_turns = GPT-5.6 family default LIVE (echo) response echoes effective mode (current_turn on gpt-5.4-nano)
reasoning.mode standard|pro standard GPT-5.6 LIVE (echo)
reasoning.generate_summary same enum — deprecated → summary LEGACY
include[] string[] null file_search_call.results, web_search_call.results, web_search_call.action.sources, message.input_image.image_url, computer_call_output.output.image_url, code_interpreter_call.outputs, reasoning.encrypted_content, message.output_text.logprobs LIVE reasoning.encrypted_content is now populated by default when store=false
top_logprobs integer — 0–20 LIVE with include=["message.output_text.logprobs"] → output_text.logprobs[]
temperature / top_p number | null 1 / 1 0–2 / 0–1 LIVE live default top_p echoed = 0.98 on gpt-5.4-nano
truncation auto|disabled | null disabled deprecated in spec, still accepted LEGACY, LIVE auto drops oldest items instead of 400 on overflow; prefer context_management
context_management[] [{type:"compaction", compact_threshold≥1000}] null minItems 1 DOC server-side compaction (§6)
service_tier auto|default|flex|scale|priority|fast|ultrafast | null auto LIVE response echoes effective tier (default)
metadata map<string,string> (≤16 keys, key ≤64, value ≤512) {} LIVE
safety_identifier string | null (≤64) null LIVE replaces user for abuse detection
prompt_cache_key string | null null LIVE routing hint (< GPT-5.6) / accounting (GPT-5.6+)
prompt_cache_retention in_memory|24h | null 24h (non-ZDR) deprecated → prompt_cache_options.ttl LEGACY, LIVE echoed "24h" by default, "in_memory" when requested
prompt_cache_options {ttl:"30m", mode:implicit|explicit, prewarm, comparison_response_id} null GPT-5.6+ only DOC gpt-5.4-nano → 400 invalid_parameter "prompt_cache_options is not supported on this model"
moderation {model, policy:{input:{mode:score|block}, output:{mode}}} | null null DOC results in response.moderation
prompt {id, version, variables} | null null DOC reusable dashboard prompt
user string — deprecated LEGACY
personality friendly|pragmatic | string — only in TokenCountsBody schema LIVE_DISCOVERED accepted by POST /v1/responses on gpt-5.4-nano (200) and echoed as personality: null in the Response object — undocumented on create
multi_agent object — ?beta=true only BETA multi-agent orchestration (beta agents surface)

# 2.2 input items (union Item)

type Direction Key fields
message (type optional) in role: user|assistant|system|developer, content: string | Part[], phase: commentary|final_answer (assistant only), status
function_call / function_call_output in/out call_id, name, arguments (JSON string) / output: string | Part[]; caller, namespace, async (programmatic calling)
reasoning in/out id, summary[]{type:summary_text,text}, content[], encrypted_content, status — replay verbatim
compaction in/out encrypted_content (opaque); from /responses/compact or server-side compaction
compaction_trigger in must be the final input item; forces compaction
configuration_update in reasoning.effort change mid-conversation (gpt-6-astra only) without breaking the cached prefix
item_reference in {type:"item_reference", id} — reference a stored item
additional_tools in developer-role item adding tools mid-thread
tool items in/out file_search_call, web_search_call, computer_call(_output), code_interpreter_call, image_generation_call, local_shell_call(_output), shell_call(_output), apply_patch_call(_output), mcp_list_tools, mcp_approval_request, mcp_approval_response, mcp_call, custom_tool_call(_output), tool_search_call/tool_search_output, program/program_output — see tools docs

Message content parts: input_text{text}, input_image{image_url \| file_id, detail: low\|high\|auto\|original}, input_file{file_id \| file_url \| file_data+filename, detail: auto\|low\|high}, input_audio{input_audio:{data,format:mp3\|wav}}; assistant messages use output_text{text, annotations[], logprobs[]} / refusal{refusal}. Every part accepts prompt_cache_breakpoint: {mode:"explicit"} (GPT-5.6+). Details: docs/openai/multimodal-input.md.

Live check (p8): mixed forms in one request — {type:message, role:developer, content:"…"}, {type:message, role:user, content:[input_text]}, {type:message, role:assistant, phase:"commentary", content:"…"} and bare {role:user, content:"…"} → 200.


# 3. The Response object

Top-level fields (spec Response + live). * = returned live but absent from the OpenAPI Response schema (LIVE_DISCOVERED).

Field Type Notes
id, object:"response", created_at, completed_at completed_at null until terminal
status completed|failed|in_progress|cancelled|queued|incomplete
error {code, message} | null
incomplete_details {reason: max_output_tokens|max_messages|content_filter|steered} | null live: max_output_tokens when reasoning ate the budget
model string snapshot id
output[] OutputItem[] messages, reasoning, tool calls, compaction…
output_text SDK-only helper not in the HTTP body
usage {input_tokens, input_tokens_details:{cached_tokens, cache_write_tokens}, output_tokens, output_tokens_details:{reasoning_tokens}, total_tokens}
instructions, previous_response_id, conversation{id}, store, background, max_output_tokens, max_tool_calls, parallel_tool_calls, tools, tool_choice, text{format,verbosity}, reasoning{effort,summary,context,mode}, truncation, service_tier, metadata, safety_identifier, user, prompt_cache_key, prompt_cache_retention, prompt_cache_options, prompt_cache_diagnostics, moderation, temperature, top_p, top_logprobs echo of effective request settings
billing* {payer:"developer"} undocumented
tool_usage* {image_gen:{input_tokens, input_tokens_details:{image_tokens,text_tokens}, output_tokens, output_tokens_details, total_tokens}, web_search:{num_requests}} undocumented tool accounting
frequency_penalty, presence_penalty number (0.0) echoed although not Responses parameters
personality* null echoed (see §2.1)
output[].phase (message) final_answer / commentary in spec for input/output messages; live always final_answer

Live example (a, trimmed, ids shortened):

json
{"id":"resp_0d1a…","object":"response","created_at":1789782269,"status":"completed","background":false,
 "billing":{"payer":"developer"},"completed_at":1789782270,"error":null,"incomplete_details":null,
 "instructions":null,"max_output_tokens":32,"model":"gpt-5.4-nano-2026-03-17","moderation":null,
 "output":[{"id":"msg_0d1a…","type":"message","status":"completed","role":"assistant","phase":"final_answer",
           "content":[{"type":"output_text","annotations":[],"logprobs":[],"text":"OK."}]}],
 "parallel_tool_calls":true,"previous_response_id":null,"prompt_cache_key":null,"prompt_cache_retention":"24h",
 "reasoning":{"context":"current_turn","effort":"none","mode":"standard","summary":null},"safety_identifier":null,
 "service_tier":"default","store":true,"temperature":1.0,"text":{"format":{"type":"text"},"verbosity":"medium"},
 "tool_choice":"auto","tool_usage":{"image_gen":{"input_tokens":0,"output_tokens":0,"total_tokens":0,"…":"…"},"web_search":{"num_requests":0}},
 "tools":[],"top_logprobs":0,"top_p":0.98,"truncation":"disabled",
 "usage":{"input_tokens":10,"input_tokens_details":{"cache_write_tokens":0,"cached_tokens":0},"output_tokens":6,
          "output_tokens_details":{"reasoning_tokens":0},"total_tokens":16},"user":null,"metadata":{}}

# 4. Conversation-state strategies

Strategy How Billing Retention / ZDR Live
Manual replay (stateless) append all output items (incl. reasoning with encrypted_content, assistant phase) + new user message to input; store:false full context re-billed each turn (cache helps) ZDR-safe; encrypted reasoning decrypted in memory only p8/example previous_response_id.py
previous_response_id pass the last response id; send only new items; instructions are not carried over prior turns still billed as input tokens (27 vs 10 on turn 2) needs store:true (default); chain broken if the parent is deleted/expired (30 days) c, g2
Conversations API conversation: "conv_…"; server prepends conversation items and appends the response's items same no 30-day TTL for conversation items k1–k11
WebSocket continuation response.create with previous_response_id on the same socket; connection-local cache same works with store:false/ZDR while the parent is cached; previous_response_not_found otherwise docs only
Compaction context_management (server-side) or /responses/compact (standalone) fewer input tokens after compaction; may lower cache hits ZDR-friendly with store:false i, i2

Rules that bite:

  • previous_response_id xor conversation (400 mutually_exclusive_parameters).
  • Reasoning models: keep every item between the last user message and your function_call_output untouched; reasoning.context: all_turns (GPT-5.6) re-renders earlier reasoning only when the request has access to it.
  • instructions are per-request; to change the system prompt mid-chain just send a new instructions.

# 5. Background mode

  • background:true → HTTP 200 immediately with status:"queued", output:[], usage:null. Poll GET /v1/responses/{id} while queued|in_progress.
  • Cancel: POST …/cancel → Response with status:"cancelled" (usage zeros); repeated cancel on a cancelled response returns the same object (idempotent, documented + observed). Cancelling an already completed response fails: HTTP 400 invalid_request_error "Cannot cancel a completed response." (observed in examples/openai/responses/background.ts, first run) — so "idempotent" only holds for in-flight/cancelled states. Cancelling a foreground response = drop the connection.
  • Streaming a background response: create with background:true, stream:true; keep the last sequence_number; resume with GET /v1/responses/{id}?stream=true&starting_after=N (live: after seq 1 the server replayed response.in_progress(2)…response.completed(9)). Only possible if created with stream:true.
  • Background streams emit an extra response.queued event (seq 1) after response.created.
  • Retention: ZDR projects run background with store=false, data kept ~10 min for polling; Modified-Abuse-Monitoring projects keep background responses only with explicit store=true.
  • Documented latency caveat: time-to-first-token is higher than synchronous.

# 6. Compaction

Server-side: context_management: [{"type":"compaction","compact_threshold": N}] (N ≥ 1000 tokens). When rendered context crosses N the server compacts mid-response, emits a compaction output item (and response.compaction.compacting SSE event) and continues. Continue by appending outputs (stateless) or via previous_response_id (never prune manually in that mode). Latency tip: in stateless mode you may drop items before the latest compaction item.

Standalone: POST /v1/responses/compact {model, input | previous_response_id, instructions?, tools?, prompt_cache_*?, service_tier?} →

json
{"id":"resp_0ed5…","object":"response.compaction","created_at":1789782277,
 "output":[{"id":"msg_…","type":"message","status":"completed","role":"user","content":[{"type":"input_text","text":"Reply with OK."}]},
           {"id":"msg_…","type":"message","role":"user","content":[{"type":"input_text","text":"Reply with OK again."}]},
           {"id":"cmp_…","type":"compaction","encrypted_content":"gAAAAABq…"}],
 "usage":{"input_tokens":119,"input_tokens_details":{"cache_write_tokens":0,"cached_tokens":0},"output_tokens":…,"output_tokens_details":{"reasoning_tokens":…},"total_tokens":…}}

Observations: the assistant message was folded into the compaction item while user messages were retained; previous_response_id alone (no input) also works (i2). Pass the returned output as-is as the next input (do not prune). Compaction can lower prompt-cache reuse (prompt_cache_diagnostics.reason = context_compacted on GPT-5.6+).


# 7. Token counting — POST /v1/responses/input_tokens

Body = the create body minus generation settings (model, input, instructions, tools, tool_choice, text, reasoning, previous_response_id, conversation, truncation, parallel_tool_calls, personality). Response {object:"response.input_tokens", input_tokens}. Live: "Reply with OK." → 10 (equals the real call's usage.input_tokens); with instructions + 1×1 image + one function tool → 49. Free; ?beta=true variant works. Counts include role/boundary formatting tokens; output-side max_output_tokens also covers non-visible formatting tokens.


# 8. WebSocket mode (wss://api.openai.com/v1/responses) — DOCUMENTED

  • Client events: response.create (same body as HTTP create + stream_id; stream implicit, background unsupported, generate:false = warm-up), response.steer (previous_response_id, input of messages/function outputs — queues mid-turn steering), beta response.inject (response_id, input).
  • Server events: the 59 SSE events + stream_id echo, plus response.steer.accepted|pending|failed, beta response.inject.created|failed, and error (previous_response_not_found, invalid_stream_id, websocket_stream_limit_reached, websocket_connection_limit_reached).
  • stream_id = ordered lane (FIFO within a lane, concurrent across lanes); previous_response_id = lineage → fork a lane from another lane's response. Limits: 16 in-flight responses, 32 named lanes, 60-minute connections. Reconnect = new cache; continue with previous_response_id only if store=true.
  • SDKs: Python client.responses.connect() (pip install "openai[realtime]"), Node ResponsesWS.

# 9. Retrieve / list / delete semantics

  • GET /v1/responses/{id} returns the same object as create (status may have progressed). include[] works on retrieve (live: message.output_text.logprobs).
  • GET …/input_items on a previous_response_id child returned (desc): new user message, prior assistant message (phase:"final_answer"), prior user message — i.e. the rendered context, not only the literal input.
  • DELETE is allowed on any stored response including chain parents (child kept its copy of the context). After delete → 404 Response with id '…' not found.
  • Responses attached to a conversation persist through the conversation (no 30-day TTL).

# 10. Errors observed (all type: invalid_request_error)

HTTP code param Trigger
400 integer_below_min_value max_output_tokens value 8 (< 16)
400 mutually_exclusive_parameters null conversation + previous_response_id
400 previous_response_not_found previous_response_id unknown id
400 unsupported_parameter reasoning.effort reasoning on gpt-4.1-nano
400 unsupported_value text.verbosity low on gpt-4.1-nano (only medium)
400 invalid_parameter prompt_cache_options on gpt-5.4-nano (GPT-5.6+ only)
400 null input text.format.type=json_object without the word "json" in input
400 null null POST …/cancel on a response whose status is already completed ("Cannot cancel a completed response.")
404 null null GET/DELETE unknown, unstored or deleted response

# 11. Live verification log (2026-09-18, total ≈ $0.0015)

# Call Result
a POST minimal (gpt-5.4-nano) 200 completed, "OK.", 16 tokens; extra fields billing, tool_usage, phase
a2 POST ?beta=true 200, same keys
b POST stream:true 10 events: created, in_progress, output_item.added, content_part.added, output_text.delta ×2, output_text.done, content_part.done, output_item.done, completed
c previous_response_id 200, input_tokens 27 (prior turn re-billed)
d/d2 store:false → GET 200 / 404
e text.format json_schema strict {"answer":"OK."}
f/f2/f3 reasoning effort:low, summary:auto (32 tok) → 0 reasoning tokens; effort:medium 32 tok → incomplete (32 reasoning tokens); 256 tok → completed, 118 reasoning tokens, 1 summary_text, encrypted_content present
g1–g6 GET, input_items, include[], DELETE ×2, GET after delete 200/200/200/200/200/404
h/h2/h3 input_tokens (plain / rich / beta) 10 / 49 / 10
i/i2 compact (input / previous_response_id) response.compaction, output [message, message, compaction]
j1–j6 background create → cancel → poll → cancel again; background+stream → resume starting_after=1 queued → cancelled → cancelled → cancelled(idempotent); stream has response.queued; resume replays seq 2–9
k1–k11 conversations create/items/list/retrieve item/response with conversation/list/update/retrieve/delete item/delete/retrieve deleted all 200, final GET 404
n3/n4 prompt caching (1555-token instructions, prompt_cache_key, 1.5 s apart) cached_tokens: 0 on both calls (no hit observed on Responses)
o1 1×1 PNG input_image on gpt-4.1-nano 200, 15 input tokens
p7 many params (verbosity, top_logprobs+logprobs include, truncation:auto, service_tier:auto, safety_identifier, metadata, prompt_cache_retention:in_memory, parallel_tool_calls:false) 200, all echoed
p10 personality:"friendly" 200, echoed personality: null