# xAI Responses API (`/v1/responses`) — stateful text + agentic tools **Status:** `DOCUMENTED` + `LIVE_VERIFIED` (2026-09-19, `grok-4.3`: create, stream, previous_response_id, store=false + encrypted reasoning replay, GET/DELETE/input_items, compact, json_schema, function calls, input_image, input_file, server tools web_search/x_search/code_interpreter/mcp/file_search). Machine-readable: `generated/fragments/parameters/xai-responses.json`, `endpoints/xai-inference.json`, `objects/xai-inference-objects.json`, `streaming-events/xai-inference.json`, `tools/xai-tools.json`. **Sources** - https://docs.x.ai/developers/rest-api-reference/inference/responses · https://docs.x.ai/developers/model-capabilities/text/generate-text · https://docs.x.ai/developers/tools/overview · https://docs.x.ai/developers/tools/advanced-usage · https://docs.x.ai/developers/advanced-api-usage/context-compaction - OpenAPI `ModelRequest`, `ModelResponse`, `ModelInputPart`, `ModelOutput`, `ModelTool`, `ModelUsage`, `CompactRequest/Response`, `ListInputItemsResponse` **Last verified:** 2026-09-19 ## 1. Endpoints | Method | Path | Purpose | Live | |---|---|---|---| | POST | `/v1/responses` | Create (sync / `stream:true`) | 200 | | GET | `/v1/responses/{id}` | Retrieve stored response (30 days) | 200 — **also 200 for a `store:false` response (LIVE_DISCOVERED)** | | DELETE | `/v1/responses/{id}` | `{id, object:"response", deleted:true}`; later GET → 404 | 200 | | GET | `/v1/responses/{id}/input_items` | `{object:"list", data:[{id:"item_0", type:"message", role, content}], first_id, last_id, has_more}`; query `limit` (1–100, 20), `order` (asc), `after` | 200 | | POST | `/v1/responses/compact` | Compact a conversation into one `compaction` item | 200 | ## 2. Request body (`ModelRequest`) | Parameter | Type / values | Default | Live | Notes | |---|---|---|---|---| | `model` | string | required | ✅ | | | `input` | string \| item[] | required | ✅ | see §3 | | `instructions` | string | | ✅ | **400 "instructions and previous_response_id together"** | | `previous_response_id` | string | | ✅ | server rehydrates full agentic state (reasoning, tool calls) | | `store` | bool | true | ✅ | echoed; response still retrievable when false | | `include[]` | `reasoning.encrypted_content`, `web_search_call.action.sources`, `code_interpreter_call.outputs`, `file_search_call.results`, `no_inline_citations`, `message.output_text.logprobs` (ignored) | | ✅ | | | `max_output_tokens` | int | 128 000 | ✅ | docs: includes reasoning; live `32` → status `completed`, 138 reasoning tokens, text "OK" (not enforced on reasoning) | | `max_turns` | int | server cap | ✅ | agentic turns per request; resets after each client-side tool call | | `reasoning` | `{effort: low\|medium\|high\|xhigh, summary: auto\|concise\|detailed, generate_summary}` | echo `{effort:"low", summary:"detailed"}` on grok-4.3 | ✅ | see [reasoning](reasoning.md) | | `reasoning_effort` | string | | ✅ | non-standard alias, used only if `reasoning` unset | | `text.format` | `{type:text}` \| `{type:json_object}` \| `{type:json_schema, name, schema, strict, description}` | text | ✅ | | | `tools[]` | see [tools index](../tools/xai/index.md) | | ✅ | accepted type strings live: `function`, `web_search`, `x_search`, `image_generation`, `collections_search`, `file_search`, `code_execution`, `code_interpreter`, `mcp`, `shell`, `tool_search` (403 alpha) | | `tool_choice` | `auto` \| `none` \| `required` \| `{type:"function", name}` | auto | ✅ | | | `parallel_tool_calls` | bool | true | doc | | | `temperature` / `top_p` | 0–2 / ≤1 | spec 1 / 1, **echo 0.7 / 0.95** | doc | | | `top_k` (≥1), `min_p` (0–1) | int / number | off | ✅ | xAI-specific | | `stream` | bool | false | ✅ | | | `service_tier` | `default` \| `priority` | default | ✅ | | | `prompt_cache_key` | string | | ✅ echoed | routing key (= `x-grok-conv-id`) | | `safety_identifier`, `user` | string | | ✅ | | | `logprobs`, `top_logprobs` | bool / 0–8 | | accepted, ignored | | | `background` | bool | | ❌ 400 "Argument not supported: background" | | | `metadata` | object | | ❌ 400 "Argument not supported: metadata" | (chat completions ignores it instead) | | `truncation` | string | `disabled` echoed | doc | compat only | | `context_management[]` | array | | doc | "parsed but not yet executed" | | `search_parameters` | object | | RETIRED | Live Search → tools | ## 3. Input items | Item | Shape | Live | |---|---|---| | message | `{role: user\|assistant\|system\|developer, content: string \| part[], name?, type?:"message"}` | ✅ (developer ok) | | `input_text` part | `{type:"input_text", text}` | ✅ | | `input_image` part | `{type:"input_image", image_url: "https://…" \| "data:image/png;base64,…", detail?: low\|high\|auto, file_id?}` | ✅ data URL (32×32 → image_tokens 3; ≥512 px required); `file_id` + empty `image_url` → 400 "image_url must either be a base64-encoded image or a URL" | | `input_file` part | `{type:"input_file", file_id \| file_url \| file_data, filename?, mime_type?}` | ✅ `file_id` → agentic attachment search answered from the file ($10/1k calls) — see [files](files.md) | | `output_text` part (assistant history) | `{type:"output_text", text}` | doc | | replayed output items | any `output[]` item from a previous response (`reasoning` incl. `encrypted_content`, `message`, `function_call`, `web_search_call`, …) | ✅ reasoning + message replay | | `function_call_output` | `{type, call_id, output: string \| part[]}` | ✅ | | `shell_call_output` | `{type, call_id, output:[{stdout, stderr, outcome:{type:exit, exit_code}\|{type:timeout}}], max_output_length}` | doc | | `compaction` | `{type:"compaction", id:"cmp_…", encrypted_content}` from `/v1/responses/compact` | ✅ | ## 4. Response object (`Response`, `object:"response"`) Live minimal body (abridged): ```json {"id":"b778ce56-…","object":"response","created_at":1789789464,"completed_at":1789789464,"model":"grok-4.3","status":"completed","store":true, "output":[{"type":"reasoning","id":"rs_b778ce56-…","summary":[{"type":"summary_text","text":"The user requested a reply of \"OK.\""}],"status":"completed"}, {"type":"message","id":"msg_b778ce56-…","role":"assistant","status":"completed","content":[{"type":"output_text","text":"OK.","logprobs":[],"annotations":[]}]}], "usage":{"input_tokens":196,"input_tokens_details":{"cached_tokens":192},"output_tokens":98,"output_tokens_details":{"reasoning_tokens":96},"total_tokens":294, "num_sources_used":0,"num_server_side_tools_used":0,"cost_in_usd_ticks":2884000,"context_details":{"input_tokens":196,"output_tokens":106}}, "reasoning":{"effort":"low","summary":"detailed"},"temperature":0.7,"top_p":0.95,"text":{"format":{"type":"text"}},"tool_choice":"auto","tools":[],"parallel_tool_calls":true, "previous_response_id":null,"metadata":{"system_fingerprint":"fp_eb3c003fc66c14ed"},"background":false,"service_tier":"default","truncation":"disabled","top_logprobs":0, "presence_penalty":0.0,"frequency_penalty":0.0,"prompt_cache_key":null,"max_tool_calls":null,"safety_identifier":null,"error":null,"instructions":null,"incomplete_details":null,"max_output_tokens":2000,"user":null} ``` Quirks: `metadata` holds `{system_fingerprint}` (not user metadata); the `reasoning` output item is **sometimes omitted** (3 of 8 plain calls had only `message`); `output_tokens` = visible + reasoning (98 = 2 + 96); `usage.server_side_tool_usage_details` appears only when a server tool ran; `status` values `completed | in_progress | incomplete` (`incomplete_details.reason ∈ max_output_tokens | max_prompt_tokens | max_time_limit`). Output item types observed: `reasoning`, `message`, `function_call`, `web_search_call`, `custom_tool_call` (x_search sub-tools — docs say `x_search_call`), `code_interpreter_call`, `file_search_call`, `mcp_call`. Documented only: `image_generation_call`, `tool_search_call`, `tool_search_output`, `shell_call`. ## 5. State management 1. **Server-side (default)** — `store:true` + `previous_response_id`. Follow-ups may change tools/model; the whole agentic trajectory is rehydrated. Live: second turn saw `input_tokens 312`, `previous_response_id` echoed. 2. **Client-side / ZDR** — `store:false` + `include:["reasoning.encrypted_content"]`, then send back `...response.output` verbatim before the new user turn. Live: replay → 200 (input_tokens 380, `cached_tokens 192`). 3. **Compaction** — `POST /v1/responses/compact {model, input}` → `{object:"response.compaction", id:"cmp_…", output:[{type:"compaction", id, encrypted_content}], usage:{input_tokens 588, output_tokens 138, reasoning_tokens 711, total_tokens 1437, dropped_message_count 3}}`. Put `output` first in the next `input`; do not edit or merge blobs; re-compaction allowed; the pre-compaction conversation must still fit in context. Live follow-up answered "A and B" from the compacted history. Retention: 30 days, then deleted. `DELETE` for early removal. Images: docs advise `store:false` when sending images ("the request may fail"). ## 6. Tools & the agentic loop See [tool-loop](tool-loop.md) and the per-tool pages under `docs/tools/xai/`. Summary: server-side tools (`web_search`, `x_search`, `code_interpreter`, `file_search`, `mcp`, `image_generation`) execute inside one request (bounded by `max_turns`); client-side `function`/`shell` calls pause the request; you answer with `function_call_output`/`shell_call_output` + `previous_response_id` (or replayed items). ## 7. Batch `/v1/responses` bodies are accepted in the Batch API (`batch_request.responses`), but live the result came back as `chat_get_completion` — see [batches](batches.md).