xAI Batch API — asynchronous discounted processing (/v1/batches)
Status: DOCUMENTED + LIVE_VERIFIED (2026-09-19: create → add 2 requests → poll (11 s) → requests → results → cancel; DELETE → 405). Not in the OpenAPI spec (REST reference + guide only). Machine-readable: parameters/xai-files-collections-batches.json, endpoints/xai-inference.json, objects/xai-inference-objects.json (Batch, BatchRequestMetadata, BatchResult).
Sources: https://docs.x.ai/developers/rest-api-reference/inference/batches · https://docs.x.ai/developers/advanced-api-usage/batch-api · https://docs.x.ai/developers/pricing#batch-api-pricing Last verified: 2026-09-19
Lifecycle
POST /v1/batches {"name"}→Batch(batch_idbatch_<uuid>,create_time/expire_timeasYYYY-MM-DD, live expiry +30 days,statecounters all 0).POST /v1/batches/{id}/requests {"batch_requests":[{"batch_request_id":"r1","batch_request":{"chat_get_completion":{…ChatRequest}}}, …]}→ 200 with empty body. Other request kinds:responses(ModelRequest — accepted, result returned aschat_get_completionlive),image_generation,image_edit,video_generation,video_extension. Uniquebatch_request_idrecommended (UUID generated otherwise); ≤25 MB per request; ≤1 000 add calls / 30 s / team.- Poll
GET /v1/batches/{id}untilstate.num_pending == 0(live: 2 → 0 in 11 s; docs: typically ≤24 h, best effort). Individual states viaGET /v1/batches/{id}/requests(pending | succeeded | failed | cancelled,endpointxai_api.Chat/GetCompletion,model, times). GET /v1/batches/{id}/results?limit=…&pagination_token=…→{results:[{batch_request_id, batch_result:{response:{chat_get_completion:ChatCompletion}} | {error}}], pagination_token}; available as requests finish; cost per result inusage.cost_in_usd_ticks.POST /v1/batches/{id}:cancel→ Batch withcancel_time(pending requests dropped, completed results kept; works after completion too).GET /v1/batches?limit&pagination_tokenlists. No delete (405).
Alternative: JSONL file — lines {"custom_id","method":"POST","url":"/v1/chat/completions"|"/v1/responses"|"/v1/images/generations"|"/v1/images/edits"|"/v1/videos/generations"|"/v1/videos/edits"|"/v1/videos/extensions","body":{…}}, upload via Files API, then POST /v1/batches {name, input_file_id}. ≤200 MB, ≤50 000 lines, sealed after creation; an invalid line cancels the batch with cancel_by_xai_message.
Pricing / limits
20 % off all token types (input, cached, output, reasoning) for grok-4.3, grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning, grok-4.20-multi-agent-0309; other models accept batch at standard price or reject ("not supported for batch processing"). Batch requests do not count towards per-minute rate limits. Not combinable with priority processing. Image/video results are signed URLs valid 1 h. Both server-side tools and client-side function tools work inside batch requests (multi-turn tool calling = new batch request).
Live result sample
{"batch_request_id":"r1","batch_result":{"response":{"chat_get_completion":{"id":"445efed6-…","object":"chat.completion","model":"grok-4.3",
"choices":[{"index":0,"message":{"role":"assistant","content":"OK.","reasoning_content":"The user requested a reply of \"OK.\"","refusal":null},"finish_reason":"stop"}],
"usage":{"prompt_tokens":196,"completion_tokens":2,"total_tokens":373,"prompt_tokens_details":{…"cached_tokens":192},"completion_tokens_details":{"reasoning_tokens":175,…},"num_sources_used":0,"cost_in_usd_ticks":3887200},"system_fingerprint":"fp_eb3c003fc66c14ed","service_tier":"default"}}}}SDK: xai-sdk client.batch.create(batch_name=…, input_file_id=…), .add(batch_id, batch_requests=[chat, image_req, …]), .get, .list, .list_batch_requests, .list_batch_results (.succeeded / .failed), .cancel; batch.cost_breakdown.total_cost_usd_ticks (SDK/gRPC only — not in the REST body live). Examples: examples/xai/batches/. Test: tests/xai/test_files.py::test_batch_lifecycle (gated by RUN_BATCH_TESTS).