xAI Responses API (/v1/responses) — stateful text + agentic tools
Status: DOCUMENTED + LIVE_VERIFIED (2026-09-19, grok-4.3: create, stream, previous_response_id, store=false + encrypted reasoning replay, GET/DELETE/input_items, compact, json_schema, function calls, input_image, input_file, server tools web_search/x_search/code_interpreter/mcp/file_search). Machine-readable: generated/fragments/parameters/xai-responses.json, endpoints/xai-inference.json, objects/xai-inference-objects.json, streaming-events/xai-inference.json, tools/xai-tools.json.
Sources
- https://docs.x.ai/developers/rest-api-reference/inference/responses · https://docs.x.ai/developers/model-capabilities/text/generate-text · https://docs.x.ai/developers/tools/overview · https://docs.x.ai/developers/tools/advanced-usage · https://docs.x.ai/developers/advanced-api-usage/context-compaction
- OpenAPI
ModelRequest,ModelResponse,ModelInputPart,ModelOutput,ModelTool,ModelUsage,CompactRequest/Response,ListInputItemsResponse
Last verified: 2026-09-19
1. Endpoints
| Method | Path | Purpose | Live |
|---|---|---|---|
| POST | /v1/responses |
Create (sync / stream:true) |
200 |
| GET | /v1/responses/{id} |
Retrieve stored response (30 days) | 200 — also 200 for a store:false response (LIVE_DISCOVERED) |
| DELETE | /v1/responses/{id} |
{id, object:"response", deleted:true}; later GET → 404 |
200 |
| GET | /v1/responses/{id}/input_items |
{object:"list", data:[{id:"item_0", type:"message", role, content}], first_id, last_id, has_more}; query limit (1–100, 20), order (asc), after |
200 |
| POST | /v1/responses/compact |
Compact a conversation into one compaction item |
200 |
2. Request body (ModelRequest)
| Parameter | Type / values | Default | Live | Notes |
|---|---|---|---|---|
model |
string | required | ✅ | |
input |
string | item[] | required | ✅ | see §3 |
instructions |
string | ✅ | 400 "instructions and previous_response_id together" | |
previous_response_id |
string | ✅ | server rehydrates full agentic state (reasoning, tool calls) | |
store |
bool | true | ✅ | echoed; response still retrievable when false |
include[] |
reasoning.encrypted_content, web_search_call.action.sources, code_interpreter_call.outputs, file_search_call.results, no_inline_citations, message.output_text.logprobs (ignored) |
✅ | ||
max_output_tokens |
int | 128 000 | ✅ | docs: includes reasoning; live 32 → status completed, 138 reasoning tokens, text "OK" (not enforced on reasoning) |
max_turns |
int | server cap | ✅ | agentic turns per request; resets after each client-side tool call |
reasoning |
{effort: low|medium|high|xhigh, summary: auto|concise|detailed, generate_summary} |
echo {effort:"low", summary:"detailed"} on grok-4.3 |
✅ | see reasoning |
reasoning_effort |
string | ✅ | non-standard alias, used only if reasoning unset |
|
text.format |
{type:text} | {type:json_object} | {type:json_schema, name, schema, strict, description} |
text | ✅ | |
tools[] |
see tools index | ✅ | accepted type strings live: function, web_search, x_search, image_generation, collections_search, file_search, code_execution, code_interpreter, mcp, shell, tool_search (403 alpha) |
|
tool_choice |
auto | none | required | {type:"function", name} |
auto | ✅ | |
parallel_tool_calls |
bool | true | doc | |
temperature / top_p |
0–2 / ≤1 | spec 1 / 1, echo 0.7 / 0.95 | doc | |
top_k (≥1), min_p (0–1) |
int / number | off | ✅ | xAI-specific |
stream |
bool | false | ✅ | |
service_tier |
default | priority |
default | ✅ | |
prompt_cache_key |
string | ✅ echoed | routing key (= x-grok-conv-id) |
|
safety_identifier, user |
string | ✅ | ||
logprobs, top_logprobs |
bool / 0–8 | accepted, ignored | ||
background |
bool | ❌ 400 "Argument not supported: background" | ||
metadata |
object | ❌ 400 "Argument not supported: metadata" | (chat completions ignores it instead) | |
truncation |
string | disabled echoed |
doc | compat only |
context_management[] |
array | doc | "parsed but not yet executed" | |
search_parameters |
object | RETIRED | Live Search → tools |
3. Input items
| Item | Shape | Live |
|---|---|---|
| message | {role: user|assistant|system|developer, content: string | part[], name?, type?:"message"} |
✅ (developer ok) |
input_text part |
{type:"input_text", text} |
✅ |
input_image part |
{type:"input_image", image_url: "https://…" | "data:image/png;base64,…", detail?: low|high|auto, file_id?} |
✅ data URL (32×32 → image_tokens 3; ≥512 px required); file_id + empty image_url → 400 "image_url must either be a base64-encoded image or a URL" |
input_file part |
{type:"input_file", file_id | file_url | file_data, filename?, mime_type?} |
✅ file_id → agentic attachment search answered from the file ($10/1k calls) — see files |
output_text part (assistant history) |
{type:"output_text", text} |
doc |
| replayed output items | any output[] item from a previous response (reasoning incl. encrypted_content, message, function_call, web_search_call, …) |
✅ reasoning + message replay |
function_call_output |
{type, call_id, output: string | part[]} |
✅ |
shell_call_output |
{type, call_id, output:[{stdout, stderr, outcome:{type:exit, exit_code}|{type:timeout}}], max_output_length} |
doc |
compaction |
{type:"compaction", id:"cmp_…", encrypted_content} from /v1/responses/compact |
✅ |
4. Response object (Response, object:"response")
Live minimal body (abridged):
{"id":"b778ce56-…","object":"response","created_at":1789789464,"completed_at":1789789464,"model":"grok-4.3","status":"completed","store":true,
"output":[{"type":"reasoning","id":"rs_b778ce56-…","summary":[{"type":"summary_text","text":"The user requested a reply of \"OK.\""}],"status":"completed"},
{"type":"message","id":"msg_b778ce56-…","role":"assistant","status":"completed","content":[{"type":"output_text","text":"OK.","logprobs":[],"annotations":[]}]}],
"usage":{"input_tokens":196,"input_tokens_details":{"cached_tokens":192},"output_tokens":98,"output_tokens_details":{"reasoning_tokens":96},"total_tokens":294,
"num_sources_used":0,"num_server_side_tools_used":0,"cost_in_usd_ticks":2884000,"context_details":{"input_tokens":196,"output_tokens":106}},
"reasoning":{"effort":"low","summary":"detailed"},"temperature":0.7,"top_p":0.95,"text":{"format":{"type":"text"}},"tool_choice":"auto","tools":[],"parallel_tool_calls":true,
"previous_response_id":null,"metadata":{"system_fingerprint":"fp_eb3c003fc66c14ed"},"background":false,"service_tier":"default","truncation":"disabled","top_logprobs":0,
"presence_penalty":0.0,"frequency_penalty":0.0,"prompt_cache_key":null,"max_tool_calls":null,"safety_identifier":null,"error":null,"instructions":null,"incomplete_details":null,"max_output_tokens":2000,"user":null}Quirks: metadata holds {system_fingerprint} (not user metadata); the reasoning output item is sometimes omitted (3 of 8 plain calls had only message); output_tokens = visible + reasoning (98 = 2 + 96); usage.server_side_tool_usage_details appears only when a server tool ran; status values completed | in_progress | incomplete (incomplete_details.reason ∈ max_output_tokens | max_prompt_tokens | max_time_limit).
Output item types observed: reasoning, message, function_call, web_search_call, custom_tool_call (x_search sub-tools — docs say x_search_call), code_interpreter_call, file_search_call, mcp_call. Documented only: image_generation_call, tool_search_call, tool_search_output, shell_call.
5. State management
- Server-side (default) —
store:true+previous_response_id. Follow-ups may change tools/model; the whole agentic trajectory is rehydrated. Live: second turn sawinput_tokens 312,previous_response_idechoed. - Client-side / ZDR —
store:false+include:["reasoning.encrypted_content"], then send back...response.outputverbatim before the new user turn. Live: replay → 200 (input_tokens 380,cached_tokens 192). - Compaction —
POST /v1/responses/compact {model, input}→{object:"response.compaction", id:"cmp_…", output:[{type:"compaction", id, encrypted_content}], usage:{input_tokens 588, output_tokens 138, reasoning_tokens 711, total_tokens 1437, dropped_message_count 3}}. Putoutputfirst in the nextinput; do not edit or merge blobs; re-compaction allowed; the pre-compaction conversation must still fit in context. Live follow-up answered "A and B" from the compacted history.
Retention: 30 days, then deleted. DELETE for early removal. Images: docs advise store:false when sending images ("the request may fail").
6. Tools & the agentic loop
See tool-loop and the per-tool pages under docs/tools/xai/. Summary: server-side tools (web_search, x_search, code_interpreter, file_search, mcp, image_generation) execute inside one request (bounded by max_turns); client-side function/shell calls pause the request; you answer with function_call_output/shell_call_output + previous_response_id (or replayed items).
7. Batch
/v1/responses bodies are accepted in the Batch API (batch_request.responses), but live the result came back as chat_get_completion — see batches.