SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
10.2 KB

# OpenAI Evals API

Status: DOCUMENTED · DEPRECATED (Evals platform: read-only 2026-10-31, API shutdown 2026-11-30; migration path: Promptfoo cookbook / Datasets) · LIVE_VERIFIED for all 12 endpoints on 2026-09-18 (evals CRUD, runs create/list/retrieve/cancel/delete, output items list/retrieve) with custom and logs configs and string_check/text_similarity/python/label_model criteria; no model sampling was run (cost). Sources: Evals guide · Evals reference · Graders guide · Evaluation best practices · Deprecations. Last verified: 2026-09-18. Twins: generated/fragments/parameters/openai-evals.json (124 rows), objects (Eval, EvalRun, EvalRunOutputItem, data-source configs/sources), lifecycles (eval.run).

# 1. Architecture

text
Eval  (data_source_config: schema of `item` [+ `sample`];  testing_criteria: graders)
 └── Run (data_source: where items come from + optional model sampling; name, metadata)
      └── Output item (per row: datasource_item, sample, results[] per grader, status pass|fail)
  • item namespace = one dataset row; sample namespace = model output (model, choices, output_text, output_json, output_tools, output_reasoning_summary, output_audio, input_tools).
  • The Eval fixes the schema and the graders; each Run supplies data and optionally a model + prompt template to generate samples. Runs of the same eval are comparable (report_url dashboard).

# 2. Endpoints

Method / path SDK (python) Notes 2026-09-18
POST /v1/evals evals.create name, metadata, data_source_config (req), testing_criteria[] (req) 201
GET /v1/evals .list after, limit (20), order asc/desc, `order_by created_at updated_at`
GET /v1/evals/{id} .retrieve 200; 404 after delete
POST /v1/evals/{id} .update name, metadata 200
DELETE /v1/evals/{id} .delete {object:"eval.deleted", deleted:true, eval_id} 200
POST /v1/evals/{id}/runs .runs.create name, metadata, data_source (req) 201 queued
GET /v1/evals/{id}/runs .runs.list after, limit, order, `status ∈ queued in_progress
GET /v1/evals/{id}/runs/{run} .runs.retrieve 200
POST /v1/evals/{id}/runs/{run} .runs.cancel 200 on a completed run (status unchanged) 200
DELETE /v1/evals/{id}/runs/{run} .runs.delete {object:"eval.run.deleted", deleted:true, run_id} 200
GET /v1/evals/{id}/runs/{run}/output_items .runs.output_items.list after, limit, order, `status ∈ pass fail`
GET /v1/evals/{id}/runs/{run}/output_items/{item} .runs.output_items.retrieve 200

# 3. data_source_config (on the Eval)

type Fields Resulting schema
custom item_schema (JSON Schema of a row, req), include_sample_schema (bool) {properties:{item:<yours>, sample:<platform sample schema>}, required:[item, sample]} — sample.required = ["model","choices"] (verified)
logs metadata filters on stored Responses/Completions logs platform-defined item (the stored request/response) + sample
stored_completions metadata deprecated alias of logs

Response adds max_items (null) alongside type and schema.

# 4. testing_criteria[] = graders (see graders.md)

string_check, text_similarity (+pass_threshold), python (+pass_threshold, image_tag), score_model (+pass_threshold, range, sampling_params), label_model (labels, passing_labels). Each criterion receives an id = <name>-<uuid> used in per_testing_criteria_results[].testing_criteria; the response also carries grdr_id, inactive_at.

# 5. data_source (on the Run)

type source Sampling Use
jsonl {type:"file_content", content:[{item, sample?}]} or {type:"file_id", id} (file purpose=evals) none — sample must be supplied and must satisfy the sample schema (model + choices) grade pre-computed outputs (free, verified)
completions file_content | file_id | {type:"stored_completions", metadata, model, created_after, created_before, limit} model + input_messages {type:"template", template:[{role, content}]} or {type:"item_reference", item_reference:"item.input_trajectory"}; sampling_params {temperature 1, top_p 1, seed 42, max_completion_tokens, reasoning_effort, response_format, tools} Chat Completions sampling
responses file_content | file_id | {type:"responses", metadata, model, instructions_search, created_after, created_before, reasoning_effort, temperature, top_p, users, tools} model + input_messages (template / item_reference); sampling_params {…, text.format, tools} Responses API sampling (verified 400 Missing required parameter: 'data_source.model' or 'data_source.input_messages' when omitted)

# 6. Objects (verified shapes)

  • Eval object:"eval": id (eval_…), name, data_source_config, testing_criteria[], metadata, created_at.
  • EvalRun object:"eval.run": id (evalrun_…), eval_id, status, model (null for jsonl), name, report_url, result_counts {total, errored, failed, passed}, per_testing_criteria_results[] {testing_criteria, testing_criteria_id, passed, failed}, per_model_usage[] {model_name, invocation_count, prompt_tokens, completion_tokens, total_tokens, cached_tokens} (empty without sampling), data_source, metadata, error, created_at.
  • EvalRunOutputItem object:"eval.run.output_item": id (outputitem_…), run_id, eval_id, status pass|fail, datasource_item_id (0-based row index), datasource_item (item + sample as submitted), results[] {name (= criterion id), type (null live), score, passed, sample?}, sample {input[], output[], finish_reason, model, usage, error, temperature, max_completion_tokens, top_p, seed}, created_at; live extras _datasource_item_content_hash, available_includes → LIVE_DISCOVERED.

Run lifecycle (observed, 3 items, no sampling): queued → in_progress (≈0.3 s) → completed (≈3.6 s). Webhooks: eval.run.succeeded | failed | canceled.

# 7. Three eval designs

# A. Exact-match classification on a custom dataset (verified, $0 without sampling)

json
POST /v1/evals
{"name":"ticket-categorization",
 "data_source_config":{"type":"custom","include_sample_schema":true,
   "item_schema":{"type":"object","properties":{"ticket_text":{"type":"string"},"correct_label":{"type":"string"}},"required":["ticket_text","correct_label"]}},
 "testing_criteria":[{"type":"string_check","name":"label_match","input":"{{sample.output_text}}","operation":"eq","reference":"{{item.correct_label}}"}]}

Run with sampling (costs tokens): {"data_source":{"type":"responses","model":"gpt-5.4-nano","input_messages":{"type":"template","template":[{"role":"developer","content":"Classify as Hardware, Software or Other. One word."},{"role":"user","content":"{{item.ticket_text}}"}]},"source":{"type":"file_id","id":"file-…"}}}. Run without sampling (free): {"data_source":{"type":"jsonl","source":{"type":"file_content","content":[{"item":{"ticket_text":"Monitor is dead","correct_label":"Hardware"},"sample":{"model":"my-app-v1","output_text":"Hardware","choices":[{"index":0,"message":{"role":"assistant","content":"Hardware"},"finish_reason":"stop"}]}}]}}}.

# B. Model-graded quality on production logs (logs config, DOCUMENTED)

json
{"name":"support-answer-quality","data_source_config":{"type":"logs","metadata":{"usecase":"support-bot"}},
 "testing_criteria":[{"type":"label_model","name":"helpful","model":"gpt-4.1-mini","labels":["helpful","unhelpful"],"passing_labels":["helpful"],
   "input":[{"role":"developer","content":"Label the assistant answer as helpful or unhelpful."},{"role":"user","content":"Question: {{item.input}}\nAnswer: {{sample.output_text}}"}]}]}

Run: {"data_source":{"type":"responses","source":{"type":"responses","metadata":{"usecase":"support-bot"},"created_after":1789000000,"limit":200}}} (grades stored responses; cost = grader model tokens). Creation of this eval was verified (201); the run was not executed.

# C. Python grader with partial credit (verified in a 3-criteria eval)

json
{"type":"python","name":"py","pass_threshold":0.5,
 "source":"from rapidfuzz import fuzz\ndef grade(sample, item):\n    return fuzz.WRatio(sample['output_text'], item['expected']) / 100.0"}

Observed with 3 rows (OK/ok/Nope vs expected OK): result_counts {total 3, passed 1, failed 2}; per-criterion pass counts identical for string_check (case-sensitive), text_similarity fuzzy_match ≥ 0.8 and the python grader.

# 8. Cost control

  • Grade pre-computed samples with jsonl runs: no model calls, $0 (verified). Only score_model/label_model criteria and completions/responses sampling cost tokens; text_similarity cosine calls text-embedding-3-large.
  • Use stored_completions/responses sources with limit and time windows; start with ≤ 50 items.
  • Pick the cheapest sampled model (gpt-5.4-nano) and a cheap grader (gpt-4.1-nano) while iterating on prompts; per_model_usage reports exact tokens.
  • Delete runs/evals when done (they are free to keep but the platform shuts down 2026-11-30).

# 9. Errors observed

Call HTTP Message
run create, sample lacking model 400 Error validating file against schema: 'model' is a required property. Please check the datasource and try again.
run create, responses source without model 400 Missing required parameter: 'data_source.model' or 'data_source.input_messages'.
GET deleted eval 404 Eval eval_… cannot be found. Please confirm the eval_id or permissions to view it.
DELETE already-deleted run 404 Run evalrun_… cannot be found…