OpenAI reasoning parameters (Responses reasoning.*, Chat reasoning_effort)
Status: DOCUMENTED + LIVE_VERIFIED (effort/summary/context echo, reasoning tokens, encrypted reasoning items, incomplete-on-budget behaviour — 2026-09-18, gpt-5.4-nano). Machine-readable: reasoning.* rows in generated/fragments/parameters/openai-responses.json, reasoning_effort in openai-chat-completions.json, ReasoningItem in generated/fragments/objects/openai-responses-objects.json.
Sources
- https://developers.openai.com/api/docs/guides/reasoning (effort, mode, context, summaries,
phase,configuration_update) · https://developers.openai.com/api/docs/guides/reasoning-best-practices - https://developers.openai.com/api/docs/guides/conversation-state · https://developers.openai.com/api/docs/guides/token-counting#understand-output-token-counts
- OpenAPI
Reasoning,ReasoningEffort,ReasoningModeEnum,ReasoningItem,SummaryTextContent
Last verified: 2026-09-18
1. Parameters
| Responses | Chat Completions | Values | Default | Notes |
|---|---|---|---|---|
reasoning.effort |
reasoning_effort |
none, minimal, low, medium, high, xhigh, max |
model-dependent (medium for gpt-5.5; none echoed for gpt-5.4-nano when omitted) |
not every model supports every value; gpt-6-astra rejects none (400); non-reasoning models reject the parameter (live: gpt-4.1-nano → 400 unsupported_parameter, param reasoning.effort) |
reasoning.mode |
— | standard, pro |
standard | GPT-5.6 only; pro = more model work billed at standard rates; independent of effort |
reasoning.summary |
— | auto, concise, detailed |
null (no summary) | auto = most detailed available; live echo summary:"detailed" for auto; concise on computer-use models and GPT-5+ |
reasoning.context |
— | auto, current_turn, all_turns |
auto (→ all_turns on GPT-5.6, current_turn before) |
response echoes the effective mode; all_turns needs access to earlier items (previous_response_id, conversation or full replay); reasoning reusable only within a model family |
reasoning.generate_summary |
— | same as summary | — | deprecated alias |
max_output_tokens |
max_completion_tokens |
≥16 | — | caps visible + reasoning + formatting tokens |
input item configuration_update |
— | {type:"configuration_update", reasoning:{effort}} |
— | gpt-6-astra only; changes effort for subsequent responses without changing request-level reasoning.effort (keeps cache prefix) |
assistant message phase |
— | commentary, final_answer |
— | GPT-5.4/5.5: label intermediate vs final assistant messages; round-trip when replaying history |
include: ["reasoning.encrypted_content"] |
— | legacy; encrypted content is now returned by default when store:false/ZDR |
Chat Completions limitation: from GPT-5.4, tool calling on Chat only works with reasoning_effort:"none"; Responses is the recommended surface for reasoning + tools.
2. Reasoning items and usage
output[]contains{type:"reasoning", id:"rs_…", summary:[{type:"summary_text", text}], content:[], encrypted_content, status}before the assistant message.content[](raw reasoning text) is not exposed for GPT models.usage.output_tokens_details.reasoning_tokenscounts hidden reasoning; billed as output tokens; included inoutput_tokensand in the context window.- Budget exhaustion →
status:"incomplete",incomplete_details.reason:"max_output_tokens", and possibly no visible output at all while input + reasoning tokens are billed. Reserve ≥25k tokens when experimenting (guide). - Replay: pass reasoning items back verbatim (
previous_response_iddoes it for you). Withstore:false,encrypted_contentis decrypted in memory and discarded (ZDR-compatible). - Streaming:
response.reasoning_summary_part.added/done,response.reasoning_summary_text.delta/done,response.reasoning_text.delta/done; use theencrypted_contentfromresponse.output_item.done(theaddedcopy may be partial).
3. Live observations (gpt-5.4-nano, 2026-09-18)
| Request | Result |
|---|---|
no reasoning param, "Reply with OK." |
response echoes reasoning:{context:"current_turn", effort:"none", mode:"standard", summary:null}, reasoning_tokens:0, no reasoning item |
effort:low, summary:auto, "What is 2+2?", 32 tokens |
completed, summary:"detailed" echoed, 0 reasoning tokens, no reasoning item (model skipped reasoning) |
effort:medium, summary:auto, strawberry-r question, 32 tokens |
status:"incomplete", reason max_output_tokens, output_tokens:32 = reasoning_tokens:32, output = one reasoning item (encrypted_content present, empty summary) — no answer |
same with 256 tokens, store:false |
completed: reasoning_tokens:118, output_tokens:125; reasoning item with encrypted_content and one summary_text ("Counting letters in 'strawberry' … total for 'r' is indeed 3"); message phase:"final_answer", text "3" |
store:false, effort:low "Reply with OK." |
completed, no reasoning item (nothing to encrypt) |
Chat gpt-5.4-nano + reasoning_effort:"low" |
200, completion_tokens_details.reasoning_tokens:0, system_fingerprint:null |
gpt-4.1-nano + reasoning.effort |
400 unsupported_parameter |
Take-aways: summary is only materialised when the model actually reasons; small nano models frequently answer trivial prompts with zero reasoning tokens even at low; the reasoning object is always echoed (useful to learn the model's defaults); a 32-token cap is unsafe for any prompt that triggers reasoning.
4. Prompting reasoning models (guide digest)
Keep prompts simple and direct; avoid "think step by step" (built in); use delimiters; provide only relevant context; for latency-sensitive flows ask for a short preamble; start at medium, compare high/xhigh only with evals; prefer Responses for tool loops so reasoning items are preserved between calls; keep every item between the last user message and your function outputs intact.