SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
9.5 KB

# xAI Responses API (/v1/responses) — stateful text + agentic tools

Status: DOCUMENTED + LIVE_VERIFIED (2026-09-19, grok-4.3: create, stream, previous_response_id, store=false + encrypted reasoning replay, GET/DELETE/input_items, compact, json_schema, function calls, input_image, input_file, server tools web_search/x_search/code_interpreter/mcp/file_search). Machine-readable: generated/fragments/parameters/xai-responses.json, endpoints/xai-inference.json, objects/xai-inference-objects.json, streaming-events/xai-inference.json, tools/xai-tools.json.

Sources

Last verified: 2026-09-19

# 1. Endpoints

Method Path Purpose Live
POST /v1/responses Create (sync / stream:true) 200
GET /v1/responses/{id} Retrieve stored response (30 days) 200 — also 200 for a store:false response (LIVE_DISCOVERED)
DELETE /v1/responses/{id} {id, object:"response", deleted:true}; later GET → 404 200
GET /v1/responses/{id}/input_items {object:"list", data:[{id:"item_0", type:"message", role, content}], first_id, last_id, has_more}; query limit (1–100, 20), order (asc), after 200
POST /v1/responses/compact Compact a conversation into one compaction item 200

# 2. Request body (ModelRequest)

Parameter Type / values Default Live Notes
model string required ✅
input string | item[] required ✅ see §3
instructions string ✅ 400 "instructions and previous_response_id together"
previous_response_id string ✅ server rehydrates full agentic state (reasoning, tool calls)
store bool true ✅ echoed; response still retrievable when false
include[] reasoning.encrypted_content, web_search_call.action.sources, code_interpreter_call.outputs, file_search_call.results, no_inline_citations, message.output_text.logprobs (ignored) ✅
max_output_tokens int 128 000 ✅ docs: includes reasoning; live 32 → status completed, 138 reasoning tokens, text "OK" (not enforced on reasoning)
max_turns int server cap ✅ agentic turns per request; resets after each client-side tool call
reasoning {effort: low|medium|high|xhigh, summary: auto|concise|detailed, generate_summary} echo {effort:"low", summary:"detailed"} on grok-4.3 ✅ see reasoning
reasoning_effort string ✅ non-standard alias, used only if reasoning unset
text.format {type:text} | {type:json_object} | {type:json_schema, name, schema, strict, description} text ✅
tools[] see tools index ✅ accepted type strings live: function, web_search, x_search, image_generation, collections_search, file_search, code_execution, code_interpreter, mcp, shell, tool_search (403 alpha)
tool_choice auto | none | required | {type:"function", name} auto ✅
parallel_tool_calls bool true doc
temperature / top_p 0–2 / ≤1 spec 1 / 1, echo 0.7 / 0.95 doc
top_k (≥1), min_p (0–1) int / number off ✅ xAI-specific
stream bool false ✅
service_tier default | priority default ✅
prompt_cache_key string ✅ echoed routing key (= x-grok-conv-id)
safety_identifier, user string ✅
logprobs, top_logprobs bool / 0–8 accepted, ignored
background bool ❌ 400 "Argument not supported: background"
metadata object ❌ 400 "Argument not supported: metadata" (chat completions ignores it instead)
truncation string disabled echoed doc compat only
context_management[] array doc "parsed but not yet executed"
search_parameters object RETIRED Live Search → tools

# 3. Input items

Item Shape Live
message {role: user|assistant|system|developer, content: string | part[], name?, type?:"message"} ✅ (developer ok)
input_text part {type:"input_text", text} ✅
input_image part {type:"input_image", image_url: "https://…" | "data:image/png;base64,…", detail?: low|high|auto, file_id?} ✅ data URL (32×32 → image_tokens 3; ≥512 px required); file_id + empty image_url → 400 "image_url must either be a base64-encoded image or a URL"
input_file part {type:"input_file", file_id | file_url | file_data, filename?, mime_type?} ✅ file_id → agentic attachment search answered from the file ($10/1k calls) — see files
output_text part (assistant history) {type:"output_text", text} doc
replayed output items any output[] item from a previous response (reasoning incl. encrypted_content, message, function_call, web_search_call, …) ✅ reasoning + message replay
function_call_output {type, call_id, output: string | part[]} ✅
shell_call_output {type, call_id, output:[{stdout, stderr, outcome:{type:exit, exit_code}|{type:timeout}}], max_output_length} doc
compaction {type:"compaction", id:"cmp_…", encrypted_content} from /v1/responses/compact ✅

# 4. Response object (Response, object:"response")

Live minimal body (abridged):

json
{"id":"b778ce56-…","object":"response","created_at":1789789464,"completed_at":1789789464,"model":"grok-4.3","status":"completed","store":true,
 "output":[{"type":"reasoning","id":"rs_b778ce56-…","summary":[{"type":"summary_text","text":"The user requested a reply of \"OK.\""}],"status":"completed"},
           {"type":"message","id":"msg_b778ce56-…","role":"assistant","status":"completed","content":[{"type":"output_text","text":"OK.","logprobs":[],"annotations":[]}]}],
 "usage":{"input_tokens":196,"input_tokens_details":{"cached_tokens":192},"output_tokens":98,"output_tokens_details":{"reasoning_tokens":96},"total_tokens":294,
          "num_sources_used":0,"num_server_side_tools_used":0,"cost_in_usd_ticks":2884000,"context_details":{"input_tokens":196,"output_tokens":106}},
 "reasoning":{"effort":"low","summary":"detailed"},"temperature":0.7,"top_p":0.95,"text":{"format":{"type":"text"}},"tool_choice":"auto","tools":[],"parallel_tool_calls":true,
 "previous_response_id":null,"metadata":{"system_fingerprint":"fp_eb3c003fc66c14ed"},"background":false,"service_tier":"default","truncation":"disabled","top_logprobs":0,
 "presence_penalty":0.0,"frequency_penalty":0.0,"prompt_cache_key":null,"max_tool_calls":null,"safety_identifier":null,"error":null,"instructions":null,"incomplete_details":null,"max_output_tokens":2000,"user":null}

Quirks: metadata holds {system_fingerprint} (not user metadata); the reasoning output item is sometimes omitted (3 of 8 plain calls had only message); output_tokens = visible + reasoning (98 = 2 + 96); usage.server_side_tool_usage_details appears only when a server tool ran; status values completed | in_progress | incomplete (incomplete_details.reason ∈ max_output_tokens | max_prompt_tokens | max_time_limit).

Output item types observed: reasoning, message, function_call, web_search_call, custom_tool_call (x_search sub-tools — docs say x_search_call), code_interpreter_call, file_search_call, mcp_call. Documented only: image_generation_call, tool_search_call, tool_search_output, shell_call.

# 5. State management

  1. Server-side (default) — store:true + previous_response_id. Follow-ups may change tools/model; the whole agentic trajectory is rehydrated. Live: second turn saw input_tokens 312, previous_response_id echoed.
  2. Client-side / ZDR — store:false + include:["reasoning.encrypted_content"], then send back ...response.output verbatim before the new user turn. Live: replay → 200 (input_tokens 380, cached_tokens 192).
  3. Compaction — POST /v1/responses/compact {model, input} → {object:"response.compaction", id:"cmp_…", output:[{type:"compaction", id, encrypted_content}], usage:{input_tokens 588, output_tokens 138, reasoning_tokens 711, total_tokens 1437, dropped_message_count 3}}. Put output first in the next input; do not edit or merge blobs; re-compaction allowed; the pre-compaction conversation must still fit in context. Live follow-up answered "A and B" from the compacted history.

Retention: 30 days, then deleted. DELETE for early removal. Images: docs advise store:false when sending images ("the request may fail").

# 6. Tools & the agentic loop

See tool-loop and the per-tool pages under docs/tools/xai/. Summary: server-side tools (web_search, x_search, code_interpreter, file_search, mcp, image_generation) execute inside one request (bounded by max_turns); client-side function/shell calls pause the request; you answer with function_call_output/shell_call_output + previous_response_id (or replayed items).

# 7. Batch

/v1/responses bodies are accepted in the Batch API (batch_request.responses), but live the result came back as chat_get_completion — see batches.