OpenAI Fine-tuning API (SFT · DPO · RFT · vision · distillation)
Status: DOCUMENTED · DEPRECATED (self-serve fine-tuning is winding down) · ACCOUNT_RESTRICTED for job creation with our key (403 training_not_available) · LIVE_VERIFIED for list/retrieve/events/checkpoints/pause reachability (404 shapes on bogus ids) · checkpoint permissions ACCOUNT_RESTRICTED (401 missing scope).
Sources: Fine-tuning reference · Model optimization · Supervised fine-tuning · Vision fine-tuning · DPO · Reinforcement fine-tuning · RFT use cases · Fine-tuning best practices · Deprecations · Pricing.
Last verified: 2026-09-18. Twins: generated/fragments/parameters/openai-fine-tuning.json (95 rows incl. training-line pseudo-parameters), objects + lifecycles fragments.
0. Deprecation timeline (read first)
| Date | Effect |
|---|---|
| 2026-05-07 | Orgs that never fine-tuned can no longer create jobs |
| 2026-07-02 | Orgs without fine-tuned-model inference in the last 60 days can no longer create jobs |
| 2027-01-06 | No new fine-tuning jobs for anyone. Inference on existing ft: models continues until the base model is deprecated |
Live confirmation (2026-09-18): POST /v1/fine_tuning/jobs with a placeholder file → 403 {"code":"training_not_available","message":"OpenAI is winding down the fine-tuning platform and your organization is no longer able to create new fine-tuning training jobs…"}. Nothing below about job creation could be exercised; it is DOCUMENTED/UNVERIFIED.
1. Endpoints
| Method / path | SDK (python) | Notes | 2026-09-18 |
|---|---|---|---|
POST /v1/fine_tuning/jobs |
fine_tuning.jobs.create |
see §3 | 403 training_not_available |
GET /v1/fine_tuning/jobs |
.list |
after, limit (20), metadata[k]=v filter (or metadata=null) |
200 {object:"list", data:[], has_more} |
GET /v1/fine_tuning/jobs/{id} |
.retrieve |
404 fine_tune_not_found (bogus id) |
|
POST /v1/fine_tuning/jobs/{id}/cancel |
.cancel |
any non-terminal state | not called |
POST /v1/fine_tuning/jobs/{id}/pause |
.pause |
RFT: snapshot + stop | 404 (bogus id) |
POST /v1/fine_tuning/jobs/{id}/resume |
.resume |
RFT: continue from last checkpoint | not called |
GET /v1/fine_tuning/jobs/{id}/events |
.list_events |
after, limit (20); events `type: message |
metrics(+moderation_checks` per guides) |
GET /v1/fine_tuning/jobs/{id}/checkpoints |
.checkpoints.list |
after, limit (10) |
404 (bogus id) |
GET /v1/fine_tuning/checkpoints/{ckpt}/permissions |
.checkpoints.permissions.retrieve/list |
project_id, after, limit (10), order ascending/descending |
401 Missing scopes: api.fine_tuning.checkpoints.read (Owner role / admin key) |
POST /v1/fine_tuning/checkpoints/{ckpt}/permissions |
.permissions.create |
{project_ids: [...]} — share a checkpoint with other projects of the org |
not called (admin write) |
DELETE /v1/fine_tuning/checkpoints/{ckpt}/permissions/{permission_id} |
.permissions.delete |
not called | |
POST /v1/fine_tuning/alpha/graders/{run,validate} |
see graders.md |
200 |
2. Methods and base-model matrix
Method (method.type) |
What you supply | Base models (guide, pinned snapshots) | Notes |
|---|---|---|---|
supervised (default) |
prompt + ideal completion | gpt-4.1-2025-04-14, gpt-4.1-mini-2025-04-14, gpt-4.1-nano-2025-04-14 (also gpt-4o-2024-08-06, gpt-4o-mini-2024-07-18, legacy gpt-4, gpt-3.5-turbo) |
≥ 10 examples; 50–100 recommended |
| vision SFT | SFT lines with image_url parts |
gpt-4o-2024-08-06 |
≤ 50 000 image examples, ≤ 10 images/example, ≤ 10 MB/image, JPEG/PNG/WEBP RGB(A); no images in assistant turns; detail: low = 85 tokens |
dpo |
prompt + preferred + non-preferred | gpt-4.1-* snapshots (guide example also shows gpt-4o-2024-08-06) |
text only, one-turn |
reinforcement |
prompt + reference fields + grader | o4-mini-2025-04-16 only (reasoning models) |
expert graders required; response_format optional |
| distillation | = SFT on captured outputs of a stronger model | any SFT base | Responses API stores responses 30 d (store: true) → export → SFT (see §4.4) |
Model pages marking Fine-tuning | v1/fine-tuning | Supported (2026-09-18): gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, gpt-4o, gpt-4o-mini, o4-mini, gpt-4, gpt-3.5-turbo. No gpt-5.x / gpt-6 model supports fine-tuning. A fine-tuned model can itself be the model of a new job (continued fine-tuning).
3. POST /v1/fine_tuning/jobs — parameters
| Parameter | Type | Notes |
|---|---|---|
model (req) |
string | base or ft: model |
training_file (req) |
file id (purpose=fine-tune, .jsonl) |
|
validation_file |
file id | periodic validation metrics; also used for full_valid_* checkpoint metrics |
method.type |
`supervised | dpo |
method.supervised.hyperparameters |
n_epochs, batch_size, learning_rate_multiplier (each "auto" | number) |
top-level hyperparameters is deprecated |
method.dpo.hyperparameters |
above + beta ("auto" | 0–2) |
higher β = stay closer to the reference policy |
method.reinforcement.grader (req for RFT) |
grader object (string_check, text_similarity, python, score_model, label_model, multi) | see graders.md |
method.reinforcement.hyperparameters |
n_epochs, batch_size, learning_rate_multiplier, reasoning_effort (`default |
low |
method.reinforcement.response_format |
{type:"json_schema", json_schema:{name, strict, schema}} |
in guide, not in spec → UNVERIFIED; populates sample.output_json |
suffix |
≤ 64 chars | → ft:gpt-4.1-nano-2025-04-14:<org>:<suffix>:<id> |
seed |
int ≤ 2 147 483 647 | reproducibility (best effort) |
integrations[] |
{type:"wandb", wandb:{project (req), name, entity, tags[]}} |
metrics streamed to Weights & Biases; default tags openai/finetune, openai/{base}, openai/{ftjob} |
metadata |
≤ 16 pairs | filterable in list |
4. Training data formats (JSONL, one example per line, ≥ 10 lines)
4.1 Supervised (chat format)
{"messages":[{"role":"system","content":"You are a terse assistant."},{"role":"user","content":"Capital of France?"},{"role":"assistant","content":"Paris"}]}
{"messages":[{"role":"user","content":"Weather in Boston?"},{"role":"assistant","tool_calls":[{"id":"call_1","type":"function","function":{"name":"get_weather","arguments":"{\"city\":\"Boston\"}"}}]}],"tools":[{"type":"function","function":{"name":"get_weather","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}}],"parallel_tool_calls":false}
{"messages":[{"role":"user","content":"Hi"},{"role":"assistant","content":"Hello!","weight":0},{"role":"user","content":"Bye"},{"role":"assistant","content":"Goodbye!","weight":1}]}weight (0|1) on assistant messages excludes a turn from the loss. Legacy functions and completions (prompt/completion) formats exist for legacy models.
4.2 Vision SFT — a user content array with {"type":"image_url","image_url":{"url":"https://… or data:image/png;base64,…","detail":"low"}} parts inside the chat format.
4.3 DPO (preference)
{"input":{"messages":[{"role":"user","content":"How cold is SF today?"}],"tools":[],"parallel_tool_calls":true},"preferred_output":[{"role":"assistant","content":"Mild: high 68°F, low 57°F."}],"non_preferred_output":[{"role":"assistant","content":"Not very cold."}]}4.4 RFT
{"messages":[{"role":"user","content":"Do you have a dedicated security team?"}],"compliant":"yes","explanation":"A dedicated security team follows strict protocols."}Every non-messages key is available to the grader as {{item.<key>}}; the model output as {{sample.output_text}} / sample.output_json / sample.output_tools.
4.5 Distillation recipe
- Prompt a strong model (
gpt-4.1, or any Responses-capable model) withstore: true+metadatatags; 2. pull good responses (Evalslogsdata source or the Responses list/GET /v1/responses/{id}); 3. write them as §4.1 lines; 4. SFT a small base (gpt-4.1-mini/-nano); 5. compare base vs. fine-tuned with an eval on held-out data.
5. Job object and lifecycle
fine_tuning.job: id (ftjob-…), model, fine_tuned_model (null until success), status (validating_files → queued → running → succeeded | failed | cancelled; RFT adds paused), error {code,message,param}, hyperparameters (SFT echo), method, training_file, validation_file, result_files[] (CSV metrics, purpose fine-tune-results), trained_tokens, seed, estimated_finish, integrations[], metadata, organization_id, created_at, finished_at; guide examples also show user_provided_suffix, usage_metrics, shared_with_openai.
Events (fine_tuning.job.event): level info|warn|error, message, type message|metrics, data (per-step metrics: train_loss, train_mean_token_accuracy, valid_*, full_valid_*; RFT adds train_reward_mean, valid_reward_mean, reasoning/grader token counts, per-grader scores).
Checkpoints (fine_tuning.job.checkpoint): fine_tuned_model_checkpoint = ft:<base>:<org>:<suffix>:<id>:ckpt-step-<n> (usable as a model id), step_number, metrics {step, train_loss, train_mean_token_accuracy, valid_loss, valid_mean_token_accuracy, full_valid_loss, full_valid_mean_token_accuracy}. SFT keeps the last 3 epochs' checkpoints; RFT keeps the 3 best by valid_reward_mean. Checkpoint permissions (checkpoint.permission {id, project_id, created_at}) let the org owner share a checkpoint with other projects — admin scope api.fine_tuning.checkpoints.read/write.
6. Safety checks
After training, 13 moderation categories (advice, harassment/threatening, hate, hate/threatening, highly-sensitive, illicit, propaganda, self-harm/instructions, self-harm/intent, sensitive, sexual/minors, sexual, violence) are evaluated; failing thresholds block deployment; details via events of type moderation_checks.
7. Pricing (pricing page, 2026-09-18 — cite only)
| Model | Training | Input | Cached input | Output |
|---|---|---|---|---|
o4-mini-2025-04-16 (RFT) |
$100 / hour of core training time (+ grader model tokens at that model's rate) | $4.00 | $1.00 | $16.00 (halved with data sharing) |
gpt-4.1-2025-04-14 |
$25 / 1M training tokens | $3.00 | $0.75 | $12.00 |
gpt-4.1-mini-2025-04-14 |
$5.00 | $0.80 | $0.20 | $3.20 |
gpt-4.1-nano-2025-04-14 |
$1.50 | $0.20 | $0.05 | $0.80 |
gpt-4o-2024-08-06 |
$25.00 | $3.75 | $1.875 | $15.00 |
gpt-4o-mini-2024-07-18 |
$3.00 | $0.30 | $0.15 | $1.20 |
Batch inference on fine-tuned models is 50 % of the above. Training tokens ≈ tokens in the file × epochs.
8. Evaluation loop (recommended)
Build an eval first (evals.md), hold out 10–20 % of the data, train, then run the same eval against the base model, the final ft: model and 1–2 checkpoints; watch full_valid_loss for overfitting; for RFT compare valid_reward_mean with a human-graded sample to detect grader hacking.
9. Observed error shapes
| Call | HTTP | Body |
|---|---|---|
| create (any) | 403 | {"error":{"message":"OpenAI is winding down the fine-tuning platform and your organization is no longer able to create new fine-tuning training jobs. Learn more …","type":"invalid_request_error","param":null,"code":"training_not_available"}} |
| retrieve/events/checkpoints/pause on bogus id | 404 | {"error":{"message":"Could not find fine tune: ftjob-…","type":"invalid_request_error","param":"fine_tune_id","code":"fine_tune_not_found"}} |
| checkpoint permissions list | 401 | You have insufficient permissions for this operation. Missing scopes: api.fine_tuning.checkpoints.read… |