OpenAI Responses API — complete reference
Status: DOCUMENTED + LIVE_VERIFIED (all HTTP endpoints exercised 2026-09-18 with gpt-5.4-nano / gpt-4.1-nano; WebSocket mode DOCUMENTED/UNVERIFIED). Machine-readable twins: generated/fragments/endpoints/openai-core.json, generated/fragments/parameters/openai-responses.json, generated/fragments/objects/openai-responses-objects.json, generated/fragments/streaming-events/openai-responses*.json.
Sources
- https://developers.openai.com/api/reference/resources/responses/methods/create (+ retrieve, delete, cancel, compact, input_items/list, input_tokens/count, streaming-events, websocket-events)
- https://developers.openai.com/api/docs/guides/text · conversation-state · background · compaction · token-counting · websocket-mode · streaming-responses · migrate-to-responses
- OpenAPI spec
sources/openai/openapi/openapi-master.yaml(schemasCreateResponse,Response,ResponseStreamEvent,ResponsesClientEvent,ResponsesServerEvent,CompactResponseMethodPublicBody,TokenCountsBody) - Live probes:
tmp-live/openai-core/*.json(sanitized), logreports/live-requests.jsonl
Last verified: 2026-09-18
1. Endpoints
| Method | Path | Purpose | Status | Live note (2026-09-18) |
|---|---|---|---|---|
| POST | /v1/responses |
Create a response (sync, stream=true SSE, or background=true) |
LIVE_VERIFIED | gpt-5.4-nano, 10 in / 6 out tokens, status=completed |
| POST | /v1/responses?beta=true |
Same operation on the beta OpenAPI surface (BetaCreateResponse; adds multi_agent, agent_message/multi_agent_call items, optional openai-beta header) |
BETA, LIVE_VERIFIED | identical key set to the GA response |
| GET | /v1/responses/{id} |
Retrieve a stored response; ?include[]=…, ?stream=true&starting_after=N re-attaches to a background stream |
LIVE_VERIFIED | 200; store=false id → 404; deleted id → 404 |
| DELETE | /v1/responses/{id} |
Delete a stored response → {id, object:"response.deleted", deleted:true} |
LIVE_VERIFIED | 200 |
| POST | /v1/responses/{id}/cancel |
Cancel a background response; idempotent while cancelled/in-flight | LIVE_VERIFIED | immediate cancel → status=cancelled; 2nd cancel → same object; cancel after completion → 400 "Cannot cancel a completed response." |
| GET | /v1/responses/{id}/input_items |
List the rendered input context (cursor pagination limit 1–100/20, order asc|desc (default desc), after, include[]) |
LIVE_VERIFIED | chain child lists 3 items |
| POST | /v1/responses/input_tokens |
Count input tokens of a Responses payload (free) → {object:"response.input_tokens", input_tokens} |
LIVE_VERIFIED | 10 tokens = real usage.input_tokens |
| POST | /v1/responses/compact |
Stateless compaction → object:"response.compaction" |
LIVE_VERIFIED | [message, message, compaction] |
| WS | wss://api.openai.com/v1/responses |
WebSocket mode (response.create / response.steer client events) |
DOCUMENTED, UNVERIFIED | not run (needs WS client dep) |
Auth: Authorization: Bearer $OPENAI_API_KEY (+ optional OpenAI-Organization, OpenAI-Project). Observed response headers: x-request-id, openai-processing-ms, openai-version: 2020-10-01, x-ratelimit-{limit,remaining,reset}-{requests,tokens} (account-specific values, not documented limits).
SDK surface (verified installed): Python openai==3.16.2 → client.responses.{create,retrieve,delete,cancel,compact,stream,parse}, client.responses.input_items.list, client.responses.input_tokens.count, client.responses.connect() (WS). Node openai@7.18.0 → client.responses.{create,retrieve,delete,cancel,compact,stream,parse}, client.responses.inputItems.list, client.responses.inputTokens.count, ResponsesWS.
2. Request body — POST /v1/responses
Notation: a.b nested object, x[] array element, x[](t) union member discriminated by type (or role). Full list (511 rows incl. every input item type) in generated/fragments/parameters/openai-responses.json.
2.1 Top-level parameters
| Parameter | Type | Default | Constraints / enum | Status | Notes |
|---|---|---|---|---|---|
model |
string | — | any Responses model id (ModelIdsResponses) |
LIVE | used gpt-5.4-nano (snapshot echoed as gpt-5.4-nano-2026-03-17) |
input |
string | Item[] | — | see §2.2 | LIVE | string ≡ one user message |
instructions |
string | null | null | — | LIVE | system/developer message; not inherited through previous_response_id |
previous_response_id |
string | null | null | mutually exclusive with conversation |
LIVE | 400 mutually_exclusive_parameters if both; 400 previous_response_not_found if unknown |
conversation |
string | {id} | null |
null | LIVE | items of the response are appended to the conversation | |
store |
boolean | null | true | LIVE | 30-day retention; false → not retrievable (404) and reasoning items carry encrypted_content |
|
background |
boolean | null | false | not over WebSocket | LIVE | returns status=queued immediately |
stream |
boolean | null | false | LIVE | SSE; see docs/openai/streaming-events.md |
|
stream_options.include_obfuscation |
boolean | true | DOC | deltas carry an obfuscation pad unless false |
|
max_output_tokens |
integer | null | null | minimum 16 | LIVE | 8 → 400 integer_below_min_value; budget includes reasoning + formatting tokens |
max_tool_calls |
integer | null | null | DOC | cap on built-in tool calls per response | |
parallel_tool_calls |
boolean | null | true | LIVE | ||
tools[] |
Tool[] | — | 16 tool types (function, file_search, web_search(+preview), computer(+preview), mcp, code_interpreter, image_generation, local_shell, shell, apply_patch, custom, namespace, tool_search, programmatic_tool_calling) | DOC | shapes owned by the tools fragments |
tool_choice |
"none"|"auto"|"required" | object |
auto | objects: allowed_tools{mode, tools[]}, hosted {type}, function{name}, mcp{server_label,name}, custom{name}, programmatic_tool_calling, apply_patch, shell |
DOC | |
text.format |
{type:text} | {type:json_schema,name,schema,strict,description} | {type:json_object} |
text | LIVE | see docs/openai/structured-outputs.md |
|
text.verbosity |
low|medium|high | null |
medium | GPT-5+ only | LIVE | gpt-4.1-nano + low → 400 unsupported_value ("Supported values are: 'medium'") |
reasoning.effort |
none|minimal|low|medium|high|xhigh|max | null |
medium (model-dependent) | reasoning models only | LIVE | gpt-5.4-nano echoes effort:"none" when omitted; gpt-4.1-nano → 400 unsupported_parameter |
reasoning.summary |
auto|concise|detailed | null |
null | LIVE | auto echoed as detailed; summary items only when reasoning actually happened |
|
reasoning.context |
auto|current_turn|all_turns | null |
auto | all_turns = GPT-5.6 family default |
LIVE (echo) | response echoes effective mode (current_turn on gpt-5.4-nano) |
reasoning.mode |
standard|pro |
standard | GPT-5.6 | LIVE (echo) | |
reasoning.generate_summary |
same enum | — | deprecated → summary |
LEGACY | |
include[] |
string[] | null | file_search_call.results, web_search_call.results, web_search_call.action.sources, message.input_image.image_url, computer_call_output.output.image_url, code_interpreter_call.outputs, reasoning.encrypted_content, message.output_text.logprobs |
LIVE | reasoning.encrypted_content is now populated by default when store=false |
top_logprobs |
integer | — | 0–20 | LIVE | with include=["message.output_text.logprobs"] → output_text.logprobs[] |
temperature / top_p |
number | null | 1 / 1 | 0–2 / 0–1 | LIVE | live default top_p echoed = 0.98 on gpt-5.4-nano |
truncation |
auto|disabled | null |
disabled | deprecated in spec, still accepted | LEGACY, LIVE | auto drops oldest items instead of 400 on overflow; prefer context_management |
context_management[] |
[{type:"compaction", compact_threshold≥1000}] |
null | minItems 1 | DOC | server-side compaction (§6) |
service_tier |
auto|default|flex|scale|priority|fast|ultrafast | null |
auto | LIVE | response echoes effective tier (default) |
|
metadata |
map<string,string> (≤16 keys, key ≤64, value ≤512) | {} | LIVE | ||
safety_identifier |
string | null (≤64) | null | LIVE | replaces user for abuse detection |
|
prompt_cache_key |
string | null | null | LIVE | routing hint (< GPT-5.6) / accounting (GPT-5.6+) | |
prompt_cache_retention |
in_memory|24h | null |
24h (non-ZDR) | deprecated → prompt_cache_options.ttl |
LEGACY, LIVE | echoed "24h" by default, "in_memory" when requested |
prompt_cache_options |
{ttl:"30m", mode:implicit|explicit, prewarm, comparison_response_id} |
null | GPT-5.6+ only | DOC | gpt-5.4-nano → 400 invalid_parameter "prompt_cache_options is not supported on this model" |
moderation |
{model, policy:{input:{mode:score|block}, output:{mode}}} | null |
null | DOC | results in response.moderation |
|
prompt |
{id, version, variables} | null |
null | DOC | reusable dashboard prompt | |
user |
string | — | deprecated | LEGACY | |
personality |
friendly|pragmatic | string |
— | only in TokenCountsBody schema |
LIVE_DISCOVERED | accepted by POST /v1/responses on gpt-5.4-nano (200) and echoed as personality: null in the Response object — undocumented on create |
multi_agent |
object | — | ?beta=true only |
BETA | multi-agent orchestration (beta agents surface) |
2.2 input items (union Item)
type |
Direction | Key fields |
|---|---|---|
message (type optional) |
in | role: user|assistant|system|developer, content: string | Part[], phase: commentary|final_answer (assistant only), status |
function_call / function_call_output |
in/out | call_id, name, arguments (JSON string) / output: string | Part[]; caller, namespace, async (programmatic calling) |
reasoning |
in/out | id, summary[]{type:summary_text,text}, content[], encrypted_content, status — replay verbatim |
compaction |
in/out | encrypted_content (opaque); from /responses/compact or server-side compaction |
compaction_trigger |
in | must be the final input item; forces compaction |
configuration_update |
in | reasoning.effort change mid-conversation (gpt-6-astra only) without breaking the cached prefix |
item_reference |
in | {type:"item_reference", id} — reference a stored item |
additional_tools |
in | developer-role item adding tools mid-thread |
| tool items | in/out | file_search_call, web_search_call, computer_call(_output), code_interpreter_call, image_generation_call, local_shell_call(_output), shell_call(_output), apply_patch_call(_output), mcp_list_tools, mcp_approval_request, mcp_approval_response, mcp_call, custom_tool_call(_output), tool_search_call/tool_search_output, program/program_output — see tools docs |
Message content parts: input_text{text}, input_image{image_url \| file_id, detail: low\|high\|auto\|original}, input_file{file_id \| file_url \| file_data+filename, detail: auto\|low\|high}, input_audio{input_audio:{data,format:mp3\|wav}}; assistant messages use output_text{text, annotations[], logprobs[]} / refusal{refusal}. Every part accepts prompt_cache_breakpoint: {mode:"explicit"} (GPT-5.6+). Details: docs/openai/multimodal-input.md.
Live check (p8): mixed forms in one request — {type:message, role:developer, content:"…"}, {type:message, role:user, content:[input_text]}, {type:message, role:assistant, phase:"commentary", content:"…"} and bare {role:user, content:"…"} → 200.
3. The Response object
Top-level fields (spec Response + live). * = returned live but absent from the OpenAPI Response schema (LIVE_DISCOVERED).
| Field | Type | Notes |
|---|---|---|
id, object:"response", created_at, completed_at |
completed_at null until terminal |
|
status |
completed|failed|in_progress|cancelled|queued|incomplete |
|
error |
{code, message} | null |
|
incomplete_details |
{reason: max_output_tokens|max_messages|content_filter|steered} | null |
live: max_output_tokens when reasoning ate the budget |
model |
string | snapshot id |
output[] |
OutputItem[] | messages, reasoning, tool calls, compaction… |
output_text |
SDK-only helper | not in the HTTP body |
usage |
{input_tokens, input_tokens_details:{cached_tokens, cache_write_tokens}, output_tokens, output_tokens_details:{reasoning_tokens}, total_tokens} |
|
instructions, previous_response_id, conversation{id}, store, background, max_output_tokens, max_tool_calls, parallel_tool_calls, tools, tool_choice, text{format,verbosity}, reasoning{effort,summary,context,mode}, truncation, service_tier, metadata, safety_identifier, user, prompt_cache_key, prompt_cache_retention, prompt_cache_options, prompt_cache_diagnostics, moderation, temperature, top_p, top_logprobs |
echo of effective request settings | |
billing* |
{payer:"developer"} |
undocumented |
tool_usage* |
{image_gen:{input_tokens, input_tokens_details:{image_tokens,text_tokens}, output_tokens, output_tokens_details, total_tokens}, web_search:{num_requests}} |
undocumented tool accounting |
frequency_penalty, presence_penalty |
number (0.0) | echoed although not Responses parameters |
personality* |
null | echoed (see §2.1) |
output[].phase (message) |
final_answer / commentary |
in spec for input/output messages; live always final_answer |
Live example (a, trimmed, ids shortened):
{"id":"resp_0d1a…","object":"response","created_at":1789782269,"status":"completed","background":false,
"billing":{"payer":"developer"},"completed_at":1789782270,"error":null,"incomplete_details":null,
"instructions":null,"max_output_tokens":32,"model":"gpt-5.4-nano-2026-03-17","moderation":null,
"output":[{"id":"msg_0d1a…","type":"message","status":"completed","role":"assistant","phase":"final_answer",
"content":[{"type":"output_text","annotations":[],"logprobs":[],"text":"OK."}]}],
"parallel_tool_calls":true,"previous_response_id":null,"prompt_cache_key":null,"prompt_cache_retention":"24h",
"reasoning":{"context":"current_turn","effort":"none","mode":"standard","summary":null},"safety_identifier":null,
"service_tier":"default","store":true,"temperature":1.0,"text":{"format":{"type":"text"},"verbosity":"medium"},
"tool_choice":"auto","tool_usage":{"image_gen":{"input_tokens":0,"output_tokens":0,"total_tokens":0,"…":"…"},"web_search":{"num_requests":0}},
"tools":[],"top_logprobs":0,"top_p":0.98,"truncation":"disabled",
"usage":{"input_tokens":10,"input_tokens_details":{"cache_write_tokens":0,"cached_tokens":0},"output_tokens":6,
"output_tokens_details":{"reasoning_tokens":0},"total_tokens":16},"user":null,"metadata":{}}4. Conversation-state strategies
| Strategy | How | Billing | Retention / ZDR | Live |
|---|---|---|---|---|
| Manual replay (stateless) | append all output items (incl. reasoning with encrypted_content, assistant phase) + new user message to input; store:false |
full context re-billed each turn (cache helps) | ZDR-safe; encrypted reasoning decrypted in memory only | p8/example previous_response_id.py |
previous_response_id |
pass the last response id; send only new items; instructions are not carried over |
prior turns still billed as input tokens (27 vs 10 on turn 2) | needs store:true (default); chain broken if the parent is deleted/expired (30 days) |
c, g2 |
| Conversations API | conversation: "conv_…"; server prepends conversation items and appends the response's items |
same | no 30-day TTL for conversation items | k1–k11 |
| WebSocket continuation | response.create with previous_response_id on the same socket; connection-local cache |
same | works with store:false/ZDR while the parent is cached; previous_response_not_found otherwise |
docs only |
| Compaction | context_management (server-side) or /responses/compact (standalone) |
fewer input tokens after compaction; may lower cache hits | ZDR-friendly with store:false |
i, i2 |
Rules that bite:
previous_response_idxorconversation(400mutually_exclusive_parameters).- Reasoning models: keep every item between the last user message and your
function_call_outputuntouched;reasoning.context: all_turns(GPT-5.6) re-renders earlier reasoning only when the request has access to it. instructionsare per-request; to change the system prompt mid-chain just send a newinstructions.
5. Background mode
background:true→ HTTP 200 immediately withstatus:"queued",output:[],usage:null. PollGET /v1/responses/{id}whilequeued|in_progress.- Cancel:
POST …/cancel→ Response withstatus:"cancelled"(usage zeros); repeated cancel on a cancelled response returns the same object (idempotent, documented + observed). Cancelling an alreadycompletedresponse fails: HTTP 400invalid_request_error"Cannot cancel a completed response." (observed inexamples/openai/responses/background.ts, first run) — so "idempotent" only holds for in-flight/cancelled states. Cancelling a foreground response = drop the connection. - Streaming a background response: create with
background:true, stream:true; keep the lastsequence_number; resume withGET /v1/responses/{id}?stream=true&starting_after=N(live: after seq 1 the server replayedresponse.in_progress(2)…response.completed(9)). Only possible if created withstream:true. - Background streams emit an extra
response.queuedevent (seq 1) afterresponse.created. - Retention: ZDR projects run background with
store=false, data kept ~10 min for polling; Modified-Abuse-Monitoring projects keep background responses only with explicitstore=true. - Documented latency caveat: time-to-first-token is higher than synchronous.
6. Compaction
Server-side: context_management: [{"type":"compaction","compact_threshold": N}] (N ≥ 1000 tokens). When rendered context crosses N the server compacts mid-response, emits a compaction output item (and response.compaction.compacting SSE event) and continues. Continue by appending outputs (stateless) or via previous_response_id (never prune manually in that mode). Latency tip: in stateless mode you may drop items before the latest compaction item.
Standalone: POST /v1/responses/compact {model, input | previous_response_id, instructions?, tools?, prompt_cache_*?, service_tier?} →
{"id":"resp_0ed5…","object":"response.compaction","created_at":1789782277,
"output":[{"id":"msg_…","type":"message","status":"completed","role":"user","content":[{"type":"input_text","text":"Reply with OK."}]},
{"id":"msg_…","type":"message","role":"user","content":[{"type":"input_text","text":"Reply with OK again."}]},
{"id":"cmp_…","type":"compaction","encrypted_content":"gAAAAABq…"}],
"usage":{"input_tokens":119,"input_tokens_details":{"cache_write_tokens":0,"cached_tokens":0},"output_tokens":…,"output_tokens_details":{"reasoning_tokens":…},"total_tokens":…}}Observations: the assistant message was folded into the compaction item while user messages were retained; previous_response_id alone (no input) also works (i2). Pass the returned output as-is as the next input (do not prune). Compaction can lower prompt-cache reuse (prompt_cache_diagnostics.reason = context_compacted on GPT-5.6+).
7. Token counting — POST /v1/responses/input_tokens
Body = the create body minus generation settings (model, input, instructions, tools, tool_choice, text, reasoning, previous_response_id, conversation, truncation, parallel_tool_calls, personality). Response {object:"response.input_tokens", input_tokens}. Live: "Reply with OK." → 10 (equals the real call's usage.input_tokens); with instructions + 1×1 image + one function tool → 49. Free; ?beta=true variant works. Counts include role/boundary formatting tokens; output-side max_output_tokens also covers non-visible formatting tokens.
8. WebSocket mode (wss://api.openai.com/v1/responses) — DOCUMENTED
- Client events:
response.create(same body as HTTP create +stream_id;streamimplicit,backgroundunsupported,generate:false= warm-up),response.steer(previous_response_id,inputof messages/function outputs — queues mid-turn steering), betaresponse.inject(response_id,input). - Server events: the 59 SSE events +
stream_idecho, plusresponse.steer.accepted|pending|failed, betaresponse.inject.created|failed, anderror(previous_response_not_found,invalid_stream_id,websocket_stream_limit_reached,websocket_connection_limit_reached). stream_id= ordered lane (FIFO within a lane, concurrent across lanes);previous_response_id= lineage → fork a lane from another lane's response. Limits: 16 in-flight responses, 32 named lanes, 60-minute connections. Reconnect = new cache; continue withprevious_response_idonly ifstore=true.- SDKs: Python
client.responses.connect()(pip install "openai[realtime]"), NodeResponsesWS.
9. Retrieve / list / delete semantics
GET /v1/responses/{id}returns the same object as create (status may have progressed).include[]works on retrieve (live:message.output_text.logprobs).GET …/input_itemson aprevious_response_idchild returned (desc): new user message, prior assistant message (phase:"final_answer"), prior user message — i.e. the rendered context, not only the literalinput.DELETEis allowed on any stored response including chain parents (child kept its copy of the context). After delete → 404Response with id '…' not found.- Responses attached to a conversation persist through the conversation (no 30-day TTL).
10. Errors observed (all type: invalid_request_error)
| HTTP | code |
param |
Trigger |
|---|---|---|---|
| 400 | integer_below_min_value |
max_output_tokens |
value 8 (< 16) |
| 400 | mutually_exclusive_parameters |
null | conversation + previous_response_id |
| 400 | previous_response_not_found |
previous_response_id |
unknown id |
| 400 | unsupported_parameter |
reasoning.effort |
reasoning on gpt-4.1-nano |
| 400 | unsupported_value |
text.verbosity |
low on gpt-4.1-nano (only medium) |
| 400 | invalid_parameter |
prompt_cache_options |
on gpt-5.4-nano (GPT-5.6+ only) |
| 400 | null | input |
text.format.type=json_object without the word "json" in input |
| 400 | null | null | POST …/cancel on a response whose status is already completed ("Cannot cancel a completed response.") |
| 404 | null | null | GET/DELETE unknown, unstored or deleted response |
11. Live verification log (2026-09-18, total ≈ $0.0015)
| # | Call | Result |
|---|---|---|
| a | POST minimal (gpt-5.4-nano) |
200 completed, "OK.", 16 tokens; extra fields billing, tool_usage, phase |
| a2 | POST ?beta=true |
200, same keys |
| b | POST stream:true |
10 events: created, in_progress, output_item.added, content_part.added, output_text.delta ×2, output_text.done, content_part.done, output_item.done, completed |
| c | previous_response_id |
200, input_tokens 27 (prior turn re-billed) |
| d/d2 | store:false → GET |
200 / 404 |
| e | text.format json_schema strict |
{"answer":"OK."} |
| f/f2/f3 | reasoning effort:low, summary:auto (32 tok) → 0 reasoning tokens; effort:medium 32 tok → incomplete (32 reasoning tokens); 256 tok → completed, 118 reasoning tokens, 1 summary_text, encrypted_content present |
|
| g1–g6 | GET, input_items, include[], DELETE ×2, GET after delete |
200/200/200/200/200/404 |
| h/h2/h3 | input_tokens (plain / rich / beta) | 10 / 49 / 10 |
| i/i2 | compact (input / previous_response_id) | response.compaction, output [message, message, compaction] |
| j1–j6 | background create → cancel → poll → cancel again; background+stream → resume starting_after=1 |
queued → cancelled → cancelled → cancelled(idempotent); stream has response.queued; resume replays seq 2–9 |
| k1–k11 | conversations create/items/list/retrieve item/response with conversation/list/update/retrieve/delete item/delete/retrieve deleted | all 200, final GET 404 |
| n3/n4 | prompt caching (1555-token instructions, prompt_cache_key, 1.5 s apart) |
cached_tokens: 0 on both calls (no hit observed on Responses) |
| o1 | 1×1 PNG input_image on gpt-4.1-nano |
200, 15 input tokens |
| p7 | many params (verbosity, top_logprobs+logprobs include, truncation:auto, service_tier:auto, safety_identifier, metadata, prompt_cache_retention:in_memory, parallel_tool_calls:false) |
200, all echoed |
| p10 | personality:"friendly" |
200, echoed personality: null |