# OpenAI Graders (`/v1/fine_tuning/alpha/graders`) **Status:** `DOCUMENTED` · `BETA` (alpha path) · `DEPRECATED` (tied to the Evals shutdown 2026-11-30 and fine-tuning wind-down 2027-01-06) · `LIVE_VERIFIED` for `validate` and `run` with `string_check`, `text_similarity`, `python`, `multi` graders (2026-09-18). Model graders (`score_model`, `label_model`) not run (cost) → `DOCUMENTED`. **Sources:** [Graders guide](https://developers.openai.com/api/docs/guides/graders) · [Graders reference](https://developers.openai.com/api/reference/resources/graders) · [RFT guide](https://developers.openai.com/api/docs/guides/reinforcement-fine-tuning) · [Evals guide](https://developers.openai.com/api/docs/guides/evals) · [Deprecations](https://developers.openai.com/api/docs/deprecations#2026-06-03-evals-platform). **Last verified:** 2026-09-18. Twins: `generated/fragments/parameters/openai-graders.json` (102 rows), objects fragment (`Grader*`, `RunGraderResponse`, `ValidateGraderResponse`). Graders score a model sample against a dataset item and return a number in [0, 1] (or a custom `range`). The same JSON objects are used as **`testing_criteria`** in Evals and as **`method.reinforcement.grader`** in RFT; the two alpha endpoints let you test them standalone. ## 1. Endpoints | Method / path | Body | Response | 2026-09-18 | |---|---|---|---| | `POST /v1/fine_tuning/alpha/graders/validate` | `{grader}` | `{grader}` (normalised echo) | 200; invalid `operation` → 400 `{"type":"invalid_value","param":"operation","message":"Invalid value: 'contains'. Supported values are: 'eq', 'ne', 'like', and 'ilike'."}` | | `POST /v1/fine_tuning/alpha/graders/run` | `{grader, model_sample (req, string), item?}` | `{reward, metadata{…}, sub_rewards{}, model_grader_token_usage_per_model{}}` | 200 | SDK: `client.fine_tuning.alpha.graders.validate/run` · `client.fineTuning.alpha.graders.validate/run`. Free for non-model graders; model graders bill the grader model's tokens. ## 2. Templating `{{ item. }}` = dataset row (eval item / RFT training line, JSON-path style nesting), `{{ sample. }}` = model output: `output_text`, `output_json` (only with `response_format`/structured output), `output_tools` (chat-style `tool_calls`), `choices`, `output_audio {data, transcript}`. In `/run`, `model_sample` fills `sample.output_text` (and `output_json` when it parses as JSON). ## 3. Grader types | `type` | Required fields | Score | Notes | |---|---|---|---| | `string_check` | `name`, `input`, `reference`, `operation ∈ eq | ne | like | ilike` | 0/1 | `like` = contains (case-sensitive), `ilike` = contains (case-insensitive). Verified: `eq` → `reward 1.0`, exec 0.4 ms | | `text_similarity` | `name`, `input`, `reference`, `evaluation_metric ∈ fuzzy_match | bleu | gleu | meteor | cosine | rouge_1…rouge_5 | rouge_l`; evals add `pass_threshold` | 0–1 | `cosine` uses `text-embedding-3-large`, **evals only**. Verified: fuzzy_match("fuzzy wuzzy was a bear" vs "…had no hair") → 0.7556 | | `score_model` | `name`, `model`, `input[]` (messages with `role ∈ user|assistant|system|developer`, `content` text/image/audio parts, `type:"message"`), optional `range` (default [0,1]), `sampling_params {seed, top_p, temperature, max_completions_tokens, reasoning_effort}`, evals add `pass_threshold` | number in range | model must return `{result: float, steps:[{description, conclusion}]}`; non-numeric → 0 | | `label_model` | `name`, `model` (structured-outputs capable), `input[]`, `labels[]`, `passing_labels[] ⊆ labels` | pass/fail | evals `testing_criteria` and `multi` sub-graders; verified accepted at eval creation | | `python` | `name`, `source` (must define `grade(sample: dict, item: dict) -> float`), `image_tag?` (e.g. `2025-05-08`), evals add `pass_threshold` | float | sandbox: no network, 2 min, 2 GB RAM, 1 GB disk, 2 CPUs, source < 256 kB; packages numpy, scipy, sympy, pandas, rapidfuzz, scikit-learn, rouge-score, deepdiff, jsonschema, pydantic, pyyaml, nltk (+corpora), sqlparse, rdkit, scikit-bio, ast-grep-py. Verified: reward 1.0, `execution_time` 2.3 s | | `multi` | `name`, `graders {key: grader}` (no nested multi), `calculate_output` formula over keys | formula result | operators `+ - * / ^`, functions `min max abs floor ceil exp sqrt log`. Verified: `0.5*exact + 0.5*fuzzy` → 1.0 with `sub_rewards {exact:{reward, metadata}, fuzzy:{…}}`. Guide: RFT only, but accepted by `/run` | Model grader constraints (guide): `model` ∈ gpt-4o-2024-08-06, gpt-4o-mini-2024-07-18, gpt-4.1-2025-04-14, gpt-4.1-mini-2025-04-14, gpt-4.1-nano-2025-04-14, o1-2024-12-17, o3-mini-2025-01-31, o3-2025-04-16, o4-mini-2025-04-16; `temperature` unsupported on reasoning models; `reasoning_effort` unsupported on non-reasoning models. ## 4. Run response (verified shape) ```json {"reward": 1.0, "metadata": {"name": "atlas_exact", "type": "string_check", "errors": {"formula_parse_error": false, "sample_parse_error": false, "sample_parse_error_details": null, "truncated_observation_error": false, "unresponsive_reward_error": false, "invalid_variable_error": false, "invalid_variable_error_details": null, "other_error": false, "python_grader_server_error": false, "python_grader_server_error_type": null, "python_grader_runtime_error": false, "python_grader_runtime_error_details": null, "model_grader_server_error": false, "model_grader_refusal_error": false, "model_grader_refusal_error_details": null, "model_grader_parse_error": false, "model_grader_parse_error_details": null, "model_grader_exceeded_max_tokens_error": false, "model_grader_server_error_details": null, "endpoint_grader_internal_error": false, "endpoint_grader_internal_error_details": null, "endpoint_grader_server_error": false, "endpoint_grader_server_error_details": null, "endpoint_grader_safety_check_error": false}, "execution_time": 0.00045, "metadata": {}, "scores": {}, "token_usage": null, "sampled_model_name": null}, "sub_rewards": {}, "model_grader_token_usage_per_model": {}} ``` The live `errors` object has more keys than the OpenAPI spec (`*_details`, `model_grader_exceeded_max_tokens_error`, `endpoint_grader_*`) → `LIVE_DISCOVERED` fields. ## 5. Examples ```json {"type":"string_check","name":"exact","input":"{{sample.output_text}}","reference":"{{item.expected}}","operation":"eq"} {"type":"text_similarity","name":"fuzzy","input":"{{sample.output_text}}","reference":"{{item.expected}}","evaluation_metric":"fuzzy_match","pass_threshold":0.8} {"type":"python","name":"py","source":"def grade(sample, item):\n return 1.0 if sample['output_text'].strip() == item['expected'] else 0.0"} {"type":"multi","name":"combo","graders":{"name":{"type":"text_similarity","name":"n","input":"{{sample.output_json.name}}","reference":"{{item.name}}","evaluation_metric":"fuzzy_match"},"email":{"type":"string_check","name":"e","input":"{{sample.output_json.email}}","reference":"{{item.email}}","operation":"eq"}},"calculate_output":"(name + email) / 2"} {"type":"score_model","name":"judge","model":"gpt-4.1-mini-2025-04-14","range":[0,1],"pass_threshold":0.5,"input":[{"role":"system","content":"Score 1 if the answer matches the reference, 0.5 if similar, else 0."},{"role":"user","content":"Reference: {{item.reference_answer}} Answer: {{sample.output_text}}"}],"sampling_params":{"temperature":0,"max_completions_tokens":256}} ``` ## 6. Design tips (guide) Prefer smooth scores over pass/fail for RFT; combine exact-match on critical fields with similarity on free text via `multi`; guard against reward hacking by cross-checking model graders with human grades; balance label distributions; for tool-calling grade `sample.output_tools[0].function.name` and `.arguments` separately.