OpenAI Evals API
Status: DOCUMENTED · DEPRECATED (Evals platform: read-only 2026-10-31, API shutdown 2026-11-30; migration path: Promptfoo cookbook / Datasets) · LIVE_VERIFIED for all 12 endpoints on 2026-09-18 (evals CRUD, runs create/list/retrieve/cancel/delete, output items list/retrieve) with custom and logs configs and string_check/text_similarity/python/label_model criteria; no model sampling was run (cost).
Sources: Evals guide · Evals reference · Graders guide · Evaluation best practices · Deprecations.
Last verified: 2026-09-18. Twins: generated/fragments/parameters/openai-evals.json (124 rows), objects (Eval, EvalRun, EvalRunOutputItem, data-source configs/sources), lifecycles (eval.run).
1. Architecture
Eval (data_source_config: schema of `item` [+ `sample`]; testing_criteria: graders)
└── Run (data_source: where items come from + optional model sampling; name, metadata)
└── Output item (per row: datasource_item, sample, results[] per grader, status pass|fail)- item namespace = one dataset row; sample namespace = model output (
model,choices,output_text,output_json,output_tools,output_reasoning_summary,output_audio,input_tools). - The Eval fixes the schema and the graders; each Run supplies data and optionally a model + prompt template to generate samples. Runs of the same eval are comparable (
report_urldashboard).
2. Endpoints
| Method / path | SDK (python) | Notes | 2026-09-18 |
|---|---|---|---|
POST /v1/evals |
evals.create |
name, metadata, data_source_config (req), testing_criteria[] (req) |
201 |
GET /v1/evals |
.list |
after, limit (20), order asc/desc, `order_by created_at |
updated_at` |
GET /v1/evals/{id} |
.retrieve |
200; 404 after delete | |
POST /v1/evals/{id} |
.update |
name, metadata |
200 |
DELETE /v1/evals/{id} |
.delete |
{object:"eval.deleted", deleted:true, eval_id} |
200 |
POST /v1/evals/{id}/runs |
.runs.create |
name, metadata, data_source (req) |
201 queued |
GET /v1/evals/{id}/runs |
.runs.list |
after, limit, order, `status ∈ queued |
in_progress |
GET /v1/evals/{id}/runs/{run} |
.runs.retrieve |
200 | |
POST /v1/evals/{id}/runs/{run} |
.runs.cancel |
200 on a completed run (status unchanged) | 200 |
DELETE /v1/evals/{id}/runs/{run} |
.runs.delete |
{object:"eval.run.deleted", deleted:true, run_id} |
200 |
GET /v1/evals/{id}/runs/{run}/output_items |
.runs.output_items.list |
after, limit, order, `status ∈ pass |
fail` |
GET /v1/evals/{id}/runs/{run}/output_items/{item} |
.runs.output_items.retrieve |
200 |
3. data_source_config (on the Eval)
| type | Fields | Resulting schema |
|---|---|---|
custom |
item_schema (JSON Schema of a row, req), include_sample_schema (bool) |
{properties:{item:<yours>, sample:<platform sample schema>}, required:[item, sample]} — sample.required = ["model","choices"] (verified) |
logs |
metadata filters on stored Responses/Completions logs |
platform-defined item (the stored request/response) + sample |
stored_completions |
metadata |
deprecated alias of logs |
Response adds max_items (null) alongside type and schema.
4. testing_criteria[] = graders (see graders.md)
string_check, text_similarity (+pass_threshold), python (+pass_threshold, image_tag), score_model (+pass_threshold, range, sampling_params), label_model (labels, passing_labels). Each criterion receives an id = <name>-<uuid> used in per_testing_criteria_results[].testing_criteria; the response also carries grdr_id, inactive_at.
5. data_source (on the Run)
| type | source |
Sampling | Use |
|---|---|---|---|
jsonl |
{type:"file_content", content:[{item, sample?}]} or {type:"file_id", id} (file purpose=evals) |
none — sample must be supplied and must satisfy the sample schema (model + choices) |
grade pre-computed outputs (free, verified) |
completions |
file_content | file_id | {type:"stored_completions", metadata, model, created_after, created_before, limit} |
model + input_messages {type:"template", template:[{role, content}]} or {type:"item_reference", item_reference:"item.input_trajectory"}; sampling_params {temperature 1, top_p 1, seed 42, max_completion_tokens, reasoning_effort, response_format, tools} |
Chat Completions sampling |
responses |
file_content | file_id | {type:"responses", metadata, model, instructions_search, created_after, created_before, reasoning_effort, temperature, top_p, users, tools} |
model + input_messages (template / item_reference); sampling_params {…, text.format, tools} |
Responses API sampling (verified 400 Missing required parameter: 'data_source.model' or 'data_source.input_messages' when omitted) |
6. Objects (verified shapes)
- Eval
object:"eval":id(eval_…),name,data_source_config,testing_criteria[],metadata,created_at. - EvalRun
object:"eval.run":id(evalrun_…),eval_id,status,model(null for jsonl),name,report_url,result_counts {total, errored, failed, passed},per_testing_criteria_results[] {testing_criteria, testing_criteria_id, passed, failed},per_model_usage[] {model_name, invocation_count, prompt_tokens, completion_tokens, total_tokens, cached_tokens}(empty without sampling),data_source,metadata,error,created_at. - EvalRunOutputItem
object:"eval.run.output_item":id(outputitem_…),run_id,eval_id,status pass|fail,datasource_item_id(0-based row index),datasource_item(item+sampleas submitted),results[] {name (= criterion id), type (null live), score, passed, sample?},sample {input[], output[], finish_reason, model, usage, error, temperature, max_completion_tokens, top_p, seed},created_at; live extras_datasource_item_content_hash,available_includes→LIVE_DISCOVERED.
Run lifecycle (observed, 3 items, no sampling): queued → in_progress (≈0.3 s) → completed (≈3.6 s). Webhooks: eval.run.succeeded | failed | canceled.
7. Three eval designs
A. Exact-match classification on a custom dataset (verified, $0 without sampling)
POST /v1/evals
{"name":"ticket-categorization",
"data_source_config":{"type":"custom","include_sample_schema":true,
"item_schema":{"type":"object","properties":{"ticket_text":{"type":"string"},"correct_label":{"type":"string"}},"required":["ticket_text","correct_label"]}},
"testing_criteria":[{"type":"string_check","name":"label_match","input":"{{sample.output_text}}","operation":"eq","reference":"{{item.correct_label}}"}]}Run with sampling (costs tokens): {"data_source":{"type":"responses","model":"gpt-5.4-nano","input_messages":{"type":"template","template":[{"role":"developer","content":"Classify as Hardware, Software or Other. One word."},{"role":"user","content":"{{item.ticket_text}}"}]},"source":{"type":"file_id","id":"file-…"}}}.
Run without sampling (free): {"data_source":{"type":"jsonl","source":{"type":"file_content","content":[{"item":{"ticket_text":"Monitor is dead","correct_label":"Hardware"},"sample":{"model":"my-app-v1","output_text":"Hardware","choices":[{"index":0,"message":{"role":"assistant","content":"Hardware"},"finish_reason":"stop"}]}}]}}}.
B. Model-graded quality on production logs (logs config, DOCUMENTED)
{"name":"support-answer-quality","data_source_config":{"type":"logs","metadata":{"usecase":"support-bot"}},
"testing_criteria":[{"type":"label_model","name":"helpful","model":"gpt-4.1-mini","labels":["helpful","unhelpful"],"passing_labels":["helpful"],
"input":[{"role":"developer","content":"Label the assistant answer as helpful or unhelpful."},{"role":"user","content":"Question: {{item.input}}\nAnswer: {{sample.output_text}}"}]}]}Run: {"data_source":{"type":"responses","source":{"type":"responses","metadata":{"usecase":"support-bot"},"created_after":1789000000,"limit":200}}} (grades stored responses; cost = grader model tokens). Creation of this eval was verified (201); the run was not executed.
C. Python grader with partial credit (verified in a 3-criteria eval)
{"type":"python","name":"py","pass_threshold":0.5,
"source":"from rapidfuzz import fuzz\ndef grade(sample, item):\n return fuzz.WRatio(sample['output_text'], item['expected']) / 100.0"}Observed with 3 rows (OK/ok/Nope vs expected OK): result_counts {total 3, passed 1, failed 2}; per-criterion pass counts identical for string_check (case-sensitive), text_similarity fuzzy_match ≥ 0.8 and the python grader.
8. Cost control
- Grade pre-computed samples with
jsonlruns: no model calls, $0 (verified). Onlyscore_model/label_modelcriteria andcompletions/responsessampling cost tokens;text_similarity cosinecallstext-embedding-3-large. - Use
stored_completions/responsessources withlimitand time windows; start with ≤ 50 items. - Pick the cheapest sampled model (
gpt-5.4-nano) and a cheap grader (gpt-4.1-nano) while iterating on prompts;per_model_usagereports exact tokens. - Delete runs/evals when done (they are free to keep but the platform shuts down 2026-11-30).
9. Errors observed
| Call | HTTP | Message |
|---|---|---|
run create, sample lacking model |
400 | Error validating file against schema: 'model' is a required property. Please check the datasource and try again. |
run create, responses source without model |
400 | Missing required parameter: 'data_source.model' or 'data_source.input_messages'. |
| GET deleted eval | 404 | Eval eval_… cannot be found. Please confirm the eval_id or permissions to view it. |
| DELETE already-deleted run | 404 | Run evalrun_… cannot be found… |