SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
4.1 KB

# xAI Batch API — asynchronous discounted processing (/v1/batches)

Status: DOCUMENTED + LIVE_VERIFIED (2026-09-19: create → add 2 requests → poll (11 s) → requests → results → cancel; DELETE → 405). Not in the OpenAPI spec (REST reference + guide only). Machine-readable: parameters/xai-files-collections-batches.json, endpoints/xai-inference.json, objects/xai-inference-objects.json (Batch, BatchRequestMetadata, BatchResult).

Sources: https://docs.x.ai/developers/rest-api-reference/inference/batches · https://docs.x.ai/developers/advanced-api-usage/batch-api · https://docs.x.ai/developers/pricing#batch-api-pricing Last verified: 2026-09-19

# Lifecycle

  1. POST /v1/batches {"name"} → Batch (batch_id batch_<uuid>, create_time/expire_time as YYYY-MM-DD, live expiry +30 days, state counters all 0).
  2. POST /v1/batches/{id}/requests {"batch_requests":[{"batch_request_id":"r1","batch_request":{"chat_get_completion":{…ChatRequest}}}, …]} → 200 with empty body. Other request kinds: responses (ModelRequest — accepted, result returned as chat_get_completion live), image_generation, image_edit, video_generation, video_extension. Unique batch_request_id recommended (UUID generated otherwise); ≤25 MB per request; ≤1 000 add calls / 30 s / team.
  3. Poll GET /v1/batches/{id} until state.num_pending == 0 (live: 2 → 0 in 11 s; docs: typically ≤24 h, best effort). Individual states via GET /v1/batches/{id}/requests (pending | succeeded | failed | cancelled, endpoint xai_api.Chat/GetCompletion, model, times).
  4. GET /v1/batches/{id}/results?limit=…&pagination_token=… → {results:[{batch_request_id, batch_result:{response:{chat_get_completion:ChatCompletion}} | {error}}], pagination_token}; available as requests finish; cost per result in usage.cost_in_usd_ticks.
  5. POST /v1/batches/{id}:cancel → Batch with cancel_time (pending requests dropped, completed results kept; works after completion too). GET /v1/batches?limit&pagination_token lists. No delete (405).

Alternative: JSONL file — lines {"custom_id","method":"POST","url":"/v1/chat/completions"|"/v1/responses"|"/v1/images/generations"|"/v1/images/edits"|"/v1/videos/generations"|"/v1/videos/edits"|"/v1/videos/extensions","body":{…}}, upload via Files API, then POST /v1/batches {name, input_file_id}. ≤200 MB, ≤50 000 lines, sealed after creation; an invalid line cancels the batch with cancel_by_xai_message.

# Pricing / limits

20 % off all token types (input, cached, output, reasoning) for grok-4.3, grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning, grok-4.20-multi-agent-0309; other models accept batch at standard price or reject ("not supported for batch processing"). Batch requests do not count towards per-minute rate limits. Not combinable with priority processing. Image/video results are signed URLs valid 1 h. Both server-side tools and client-side function tools work inside batch requests (multi-turn tool calling = new batch request).

# Live result sample

json
{"batch_request_id":"r1","batch_result":{"response":{"chat_get_completion":{"id":"445efed6-…","object":"chat.completion","model":"grok-4.3",
 "choices":[{"index":0,"message":{"role":"assistant","content":"OK.","reasoning_content":"The user requested a reply of \"OK.\"","refusal":null},"finish_reason":"stop"}],
 "usage":{"prompt_tokens":196,"completion_tokens":2,"total_tokens":373,"prompt_tokens_details":{…"cached_tokens":192},"completion_tokens_details":{"reasoning_tokens":175,…},"num_sources_used":0,"cost_in_usd_ticks":3887200},"system_fingerprint":"fp_eb3c003fc66c14ed","service_tier":"default"}}}}

SDK: xai-sdk client.batch.create(batch_name=…, input_file_id=…), .add(batch_id, batch_requests=[chat, image_req, …]), .get, .list, .list_batch_requests, .list_batch_results (.succeeded / .failed), .cancel; batch.cost_breakdown.total_cost_usd_ticks (SDK/gRPC only — not in the REST body live). Examples: examples/xai/batches/. Test: tests/xai/test_files.py::test_batch_lifecycle (gated by RUN_BATCH_TESTS).