SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
6.0 KB

# OpenAI reasoning parameters (Responses reasoning.*, Chat reasoning_effort)

Status: DOCUMENTED + LIVE_VERIFIED (effort/summary/context echo, reasoning tokens, encrypted reasoning items, incomplete-on-budget behaviour — 2026-09-18, gpt-5.4-nano). Machine-readable: reasoning.* rows in generated/fragments/parameters/openai-responses.json, reasoning_effort in openai-chat-completions.json, ReasoningItem in generated/fragments/objects/openai-responses-objects.json.

Sources

Last verified: 2026-09-18

# 1. Parameters

Responses Chat Completions Values Default Notes
reasoning.effort reasoning_effort none, minimal, low, medium, high, xhigh, max model-dependent (medium for gpt-5.5; none echoed for gpt-5.4-nano when omitted) not every model supports every value; gpt-6-astra rejects none (400); non-reasoning models reject the parameter (live: gpt-4.1-nano → 400 unsupported_parameter, param reasoning.effort)
reasoning.mode — standard, pro standard GPT-5.6 only; pro = more model work billed at standard rates; independent of effort
reasoning.summary — auto, concise, detailed null (no summary) auto = most detailed available; live echo summary:"detailed" for auto; concise on computer-use models and GPT-5+
reasoning.context — auto, current_turn, all_turns auto (→ all_turns on GPT-5.6, current_turn before) response echoes the effective mode; all_turns needs access to earlier items (previous_response_id, conversation or full replay); reasoning reusable only within a model family
reasoning.generate_summary — same as summary — deprecated alias
max_output_tokens max_completion_tokens ≥16 — caps visible + reasoning + formatting tokens
input item configuration_update — {type:"configuration_update", reasoning:{effort}} — gpt-6-astra only; changes effort for subsequent responses without changing request-level reasoning.effort (keeps cache prefix)
assistant message phase — commentary, final_answer — GPT-5.4/5.5: label intermediate vs final assistant messages; round-trip when replaying history
include: ["reasoning.encrypted_content"] — legacy; encrypted content is now returned by default when store:false/ZDR

Chat Completions limitation: from GPT-5.4, tool calling on Chat only works with reasoning_effort:"none"; Responses is the recommended surface for reasoning + tools.

# 2. Reasoning items and usage

  • output[] contains {type:"reasoning", id:"rs_…", summary:[{type:"summary_text", text}], content:[], encrypted_content, status} before the assistant message. content[] (raw reasoning text) is not exposed for GPT models.
  • usage.output_tokens_details.reasoning_tokens counts hidden reasoning; billed as output tokens; included in output_tokens and in the context window.
  • Budget exhaustion → status:"incomplete", incomplete_details.reason:"max_output_tokens", and possibly no visible output at all while input + reasoning tokens are billed. Reserve ≥25k tokens when experimenting (guide).
  • Replay: pass reasoning items back verbatim (previous_response_id does it for you). With store:false, encrypted_content is decrypted in memory and discarded (ZDR-compatible).
  • Streaming: response.reasoning_summary_part.added/done, response.reasoning_summary_text.delta/done, response.reasoning_text.delta/done; use the encrypted_content from response.output_item.done (the added copy may be partial).

# 3. Live observations (gpt-5.4-nano, 2026-09-18)

Request Result
no reasoning param, "Reply with OK." response echoes reasoning:{context:"current_turn", effort:"none", mode:"standard", summary:null}, reasoning_tokens:0, no reasoning item
effort:low, summary:auto, "What is 2+2?", 32 tokens completed, summary:"detailed" echoed, 0 reasoning tokens, no reasoning item (model skipped reasoning)
effort:medium, summary:auto, strawberry-r question, 32 tokens status:"incomplete", reason max_output_tokens, output_tokens:32 = reasoning_tokens:32, output = one reasoning item (encrypted_content present, empty summary) — no answer
same with 256 tokens, store:false completed: reasoning_tokens:118, output_tokens:125; reasoning item with encrypted_content and one summary_text ("Counting letters in 'strawberry' … total for 'r' is indeed 3"); message phase:"final_answer", text "3"
store:false, effort:low "Reply with OK." completed, no reasoning item (nothing to encrypt)
Chat gpt-5.4-nano + reasoning_effort:"low" 200, completion_tokens_details.reasoning_tokens:0, system_fingerprint:null
gpt-4.1-nano + reasoning.effort 400 unsupported_parameter

Take-aways: summary is only materialised when the model actually reasons; small nano models frequently answer trivial prompts with zero reasoning tokens even at low; the reasoning object is always echoed (useful to learn the model's defaults); a 32-token cap is unsafe for any prompt that triggers reasoning.

# 4. Prompting reasoning models (guide digest)

Keep prompts simple and direct; avoid "think step by step" (built in); use delimiters; provide only relevant context; for latency-sensitive flows ask for a short preamble; start at medium, compare high/xhigh only with evals; prefer Responses for tool loops so reasoning items are preserved between calls; keep every item between the last user message and your function outputs intact.