OpenAI Batch API
Status: DOCUMENTED · LIVE_VERIFIED (create / list / retrieve, full lifecycle to completed, output download) · LIVE_DISCOVERED (cancel on a terminal batch → 409; bogus input_file_id accepted then failed).
Sources: Batch guide · Batches reference · Pricing · OpenAPI openapi-master.yaml.
Last verified: 2026-09-18. Twins: generated/fragments/endpoints/openai-files-vectorstores-batch-finetuning-evals.json, generated/fragments/parameters/openai-batch.json, generated/fragments/status-lifecycles/openai-lifecycles.json (object batch).
Asynchronous execution of a JSONL file of requests with a 50 % price discount, a separate rate-limit pool, and a 24 h completion window.
1. Endpoints
| Method / path | SDK | Body / query | 2026-09-18 |
|---|---|---|---|
POST /v1/batches |
client.batches.create |
input_file_id (req), endpoint (req), completion_window: "24h" (req, only value), metadata?, output_expires_after? {anchor:"created_at", seconds 3600–2592000} |
200 → status: validating |
GET /v1/batches |
batches.list |
after, limit (1–100, default 20) |
200 |
GET /v1/batches/{id} |
batches.retrieve |
200; 404 No batch found with id '…' |
|
POST /v1/batches/{id}/cancel |
batches.cancel |
409 when the batch was already terminal (failed); documented flow cancelling → cancelled (≤ 10 min) not exercised |
Batches cannot be deleted; they stay in the list (two test batches remain in this project's history: one completed, one failed).
2. Supported endpoints (endpoint enum, one per batch, one model per file)
/v1/responses, /v1/chat/completions, /v1/embeddings, /v1/completions, /v1/moderations, /v1/images/generations, /v1/images/edits, /v1/videos (POST only, JSON not multipart, assets via file_id / image_url; batch-generated videos downloadable 24 h).
3. Input JSONL format
Each line: {"custom_id": "<unique>", "method": "POST", "url": "<same as endpoint>", "body": {…endpoint body…}}. stream is rejected. Order of output lines is not guaranteed — join on custom_id.
{"custom_id":"r-1","method":"POST","url":"/v1/responses","body":{"model":"gpt-5.4-nano","input":"Reply with OK.","max_output_tokens":16}}
{"custom_id":"c-1","method":"POST","url":"/v1/chat/completions","body":{"model":"gpt-4.1-nano","messages":[{"role":"user","content":"Reply with OK."}],"max_tokens":8}}
{"custom_id":"e-1","method":"POST","url":"/v1/embeddings","body":{"model":"text-embedding-3-small","input":"hello world"}}
{"custom_id":"m-1","method":"POST","url":"/v1/moderations","body":{"model":"omni-moderation-latest","input":"This is a harmless test sentence."}}
{"custom_id":"i-1","method":"POST","url":"/v1/images/generations","body":{"model":"gpt-image-1-mini","prompt":"a red square","size":"1024x1024"}}
{"custom_id":"v-1","method":"POST","url":"/v1/videos","body":{"model":"sora-2","prompt":"a paper plane","seconds":"4"}}(Each example above must live in its own file — a file targets a single endpoint and a single model. Only the /v1/responses line was executed live; the others are DOCUMENTED shapes taken from the guide/spec.)
4. Output / error JSONL format (verified)
{"id":"batch_req_6aade973…","custom_id":"atlas-platform-agent-req-1","response":{"status_code":200,"request_id":"cbbd7532-…","body":{"id":"resp_…","object":"response","status":"completed","model":"gpt-5.4-nano-2026-03-17","output":[{"type":"message","content":[{"type":"output_text","text":"OK"}]}],"usage":{…}}},"error":null}Error-file lines have "response": null and "error": {"code": "…", "message": "…"}; expired requests use code: "batch_expired". response.body is the normal endpoint object. Output file purpose is batch_output; its expires_at honoured output_expires_after (observed created_at + 3600). Default deletion: 30 days after completion.
5. Lifecycle (observed timings)
validating → in_progress → finalizing → completed | expired; validating → failed; any active state → cancelling → cancelled.
| t (s) | status | request_counts |
|---|---|---|
| 0 | validating | {total 0} |
2 (in_progress_at) |
in_progress | {total 1, completed 0} |
29 (finalizing_at) |
finalizing | |
30 (completed_at) |
completed | {total 1, completed 1, failed 0}; usage {input_tokens 10, output_tokens 5, total 15, cached 0, reasoning 0} |
A batch created with a non-existent input_file_id was accepted (200, validating) and moved to failed within a minute with errors.data[0] = {code:"invalid_request", message:"Cannot find file …, or organization … does not have access to it.", param:"file_id", line:null}. Cancelling it afterwards returned 409.
6. Batch object
id, object:"batch", endpoint, model (resolved snapshot, e.g. gpt-5.4-nano-2026-03-17), errors {object:"list", data[] {code,message,param,line}}, input_file_id, completion_window, status, output_file_id, error_file_id, timestamps created_at, in_progress_at, expires_at (= created + 24 h), finalizing_at, completed_at, failed_at, expired_at, cancelling_at, cancelled_at, request_counts {total, completed, failed}, usage (batches created after 2025-09-07), metadata.
7. Limits (documented — separate from synchronous limits)
| Limit | Value |
|---|---|
| Requests per batch | 50 000 (embeddings: also ≤ 50 000 inputs across the batch) |
| Input file size | 200 MB |
| Batch creations | 2 000 per hour |
| Queued prompt tokens per model | org-specific (Platform → Settings → Limits); not exposed by the API |
| Output tokens | no limit |
8. Pricing (cite: pricing page, 2026-09-18)
Batch = 50 % of the standard per-token price; the pricing page publishes a dedicated "Batch" table (e.g. gpt-5.4-nano $0.10 input / $0.01 cached / $0.625 output per 1M; gpt-5.4 $1.25 / $0.13 / $7.50; gpt-5.5 $2.50 / $0.25 / $15.00 for < 272K context). Fine-tuned models also have Batch rates (e.g. gpt-4.1-nano-2025-04-14 $0.10 / $0.025 / $0.40). Requests that complete in an expired batch are billed. Test cost 2026-09-18: 15 tokens on gpt-5.4-nano ≈ $0.000004.
9. Errors and edge cases
| Situation | HTTP / state |
|---|---|
Unknown input_file_id |
200 then status: failed, errors.data[].code = invalid_request |
| Cancel on terminal batch | 409 (undocumented) |
| Unknown batch id | 404 invalid_request_error |
| Line-level failure | appears in error_file_id, batch still completed |
| Window exceeded | expired; unfinished lines in error file with batch_expired |
10. Best practices
- One model per file; pre-validate JSON lines and unique
custom_ids locally (validation failures fail the whole batch). - Keep bodies small (reference images by URL/file_id; no base64 blobs).
- Poll with back-off (a 1-line batch finished in 30 s; large ones can take hours); or use webhooks
batch.completed | failed | expired | cancelled(see the webhooks agent's docs). - Set
output_expires_afterif you do not want outputs retained 30 days; delete input files after completion (purpose=batchfiles auto-expire in 30 d). - Prefer Batch for evals/classification/embedding backfills; use Flex/Priority service tiers for latency-tolerant synchronous traffic instead.