OpenAI Graders (/v1/fine_tuning/alpha/graders)
Status: DOCUMENTED · BETA (alpha path) · DEPRECATED (tied to the Evals shutdown 2026-11-30 and fine-tuning wind-down 2027-01-06) · LIVE_VERIFIED for validate and run with string_check, text_similarity, python, multi graders (2026-09-18). Model graders (score_model, label_model) not run (cost) → DOCUMENTED.
Sources: Graders guide · Graders reference · RFT guide · Evals guide · Deprecations.
Last verified: 2026-09-18. Twins: generated/fragments/parameters/openai-graders.json (102 rows), objects fragment (Grader*, RunGraderResponse, ValidateGraderResponse).
Graders score a model sample against a dataset item and return a number in [0, 1] (or a custom range). The same JSON objects are used as testing_criteria in Evals and as method.reinforcement.grader in RFT; the two alpha endpoints let you test them standalone.
1. Endpoints
| Method / path | Body | Response | 2026-09-18 |
|---|---|---|---|
POST /v1/fine_tuning/alpha/graders/validate |
{grader} |
{grader} (normalised echo) |
200; invalid operation → 400 {"type":"invalid_value","param":"operation","message":"Invalid value: 'contains'. Supported values are: 'eq', 'ne', 'like', and 'ilike'."} |
POST /v1/fine_tuning/alpha/graders/run |
{grader, model_sample (req, string), item?} |
{reward, metadata{…}, sub_rewards{}, model_grader_token_usage_per_model{}} |
200 |
SDK: client.fine_tuning.alpha.graders.validate/run · client.fineTuning.alpha.graders.validate/run. Free for non-model graders; model graders bill the grader model's tokens.
2. Templating
{{ item.<path> }} = dataset row (eval item / RFT training line, JSON-path style nesting), {{ sample.<var> }} = model output: output_text, output_json (only with response_format/structured output), output_tools (chat-style tool_calls), choices, output_audio {data, transcript}. In /run, model_sample fills sample.output_text (and output_json when it parses as JSON).
3. Grader types
type |
Required fields | Score | Notes |
|---|---|---|---|
string_check |
name, input, reference, `operation ∈ eq |
ne | like |
text_similarity |
name, input, reference, `evaluation_metric ∈ fuzzy_match |
bleu | gleu |
score_model |
name, model, input[] (messages with `role ∈ user |
assistant | system |
label_model |
name, model (structured-outputs capable), input[], labels[], passing_labels[] ⊆ labels |
pass/fail | evals testing_criteria and multi sub-graders; verified accepted at eval creation |
python |
name, source (must define grade(sample: dict, item: dict) -> float), image_tag? (e.g. 2025-05-08), evals add pass_threshold |
float | sandbox: no network, 2 min, 2 GB RAM, 1 GB disk, 2 CPUs, source < 256 kB; packages numpy, scipy, sympy, pandas, rapidfuzz, scikit-learn, rouge-score, deepdiff, jsonschema, pydantic, pyyaml, nltk (+corpora), sqlparse, rdkit, scikit-bio, ast-grep-py. Verified: reward 1.0, execution_time 2.3 s |
multi |
name, graders {key: grader} (no nested multi), calculate_output formula over keys |
formula result | operators + - * / ^, functions min max abs floor ceil exp sqrt log. Verified: 0.5*exact + 0.5*fuzzy → 1.0 with sub_rewards {exact:{reward, metadata}, fuzzy:{…}}. Guide: RFT only, but accepted by /run |
Model grader constraints (guide): model ∈ gpt-4o-2024-08-06, gpt-4o-mini-2024-07-18, gpt-4.1-2025-04-14, gpt-4.1-mini-2025-04-14, gpt-4.1-nano-2025-04-14, o1-2024-12-17, o3-mini-2025-01-31, o3-2025-04-16, o4-mini-2025-04-16; temperature unsupported on reasoning models; reasoning_effort unsupported on non-reasoning models.
4. Run response (verified shape)
{"reward": 1.0,
"metadata": {"name": "atlas_exact", "type": "string_check",
"errors": {"formula_parse_error": false, "sample_parse_error": false, "sample_parse_error_details": null,
"truncated_observation_error": false, "unresponsive_reward_error": false, "invalid_variable_error": false,
"invalid_variable_error_details": null, "other_error": false,
"python_grader_server_error": false, "python_grader_server_error_type": null,
"python_grader_runtime_error": false, "python_grader_runtime_error_details": null,
"model_grader_server_error": false, "model_grader_refusal_error": false, "model_grader_refusal_error_details": null,
"model_grader_parse_error": false, "model_grader_parse_error_details": null,
"model_grader_exceeded_max_tokens_error": false, "model_grader_server_error_details": null,
"endpoint_grader_internal_error": false, "endpoint_grader_internal_error_details": null,
"endpoint_grader_server_error": false, "endpoint_grader_server_error_details": null,
"endpoint_grader_safety_check_error": false},
"execution_time": 0.00045, "metadata": {}, "scores": {}, "token_usage": null, "sampled_model_name": null},
"sub_rewards": {}, "model_grader_token_usage_per_model": {}}The live errors object has more keys than the OpenAPI spec (*_details, model_grader_exceeded_max_tokens_error, endpoint_grader_*) → LIVE_DISCOVERED fields.
5. Examples
{"type":"string_check","name":"exact","input":"{{sample.output_text}}","reference":"{{item.expected}}","operation":"eq"}
{"type":"text_similarity","name":"fuzzy","input":"{{sample.output_text}}","reference":"{{item.expected}}","evaluation_metric":"fuzzy_match","pass_threshold":0.8}
{"type":"python","name":"py","source":"def grade(sample, item):\n return 1.0 if sample['output_text'].strip() == item['expected'] else 0.0"}
{"type":"multi","name":"combo","graders":{"name":{"type":"text_similarity","name":"n","input":"{{sample.output_json.name}}","reference":"{{item.name}}","evaluation_metric":"fuzzy_match"},"email":{"type":"string_check","name":"e","input":"{{sample.output_json.email}}","reference":"{{item.email}}","operation":"eq"}},"calculate_output":"(name + email) / 2"}
{"type":"score_model","name":"judge","model":"gpt-4.1-mini-2025-04-14","range":[0,1],"pass_threshold":0.5,"input":[{"role":"system","content":"Score 1 if the answer matches the reference, 0.5 if similar, else 0."},{"role":"user","content":"Reference: {{item.reference_answer}} Answer: {{sample.output_text}}"}],"sampling_params":{"temperature":0,"max_completions_tokens":256}}6. Design tips (guide)
Prefer smooth scores over pass/fail for RFT; combine exact-match on critical fields with similarity on free text via multi; guard against reward hacking by cross-checking model graders with human grades; balance label distributions; for tool-calling grade sample.output_tools[0].function.name and .arguments separately.