SPB Git forge

spb/doc-api

Public
2commits 1branches 0releases
15.7 MBsize
maindefault branch
13 days agolast push
Python 88.3% TypeScript 7.6% Shell 4.1%
12.5 KB

# OpenAI Fine-tuning API (SFT · DPO · RFT · vision · distillation)

Status: DOCUMENTED · DEPRECATED (self-serve fine-tuning is winding down) · ACCOUNT_RESTRICTED for job creation with our key (403 training_not_available) · LIVE_VERIFIED for list/retrieve/events/checkpoints/pause reachability (404 shapes on bogus ids) · checkpoint permissions ACCOUNT_RESTRICTED (401 missing scope). Sources: Fine-tuning reference · Model optimization · Supervised fine-tuning · Vision fine-tuning · DPO · Reinforcement fine-tuning · RFT use cases · Fine-tuning best practices · Deprecations · Pricing. Last verified: 2026-09-18. Twins: generated/fragments/parameters/openai-fine-tuning.json (95 rows incl. training-line pseudo-parameters), objects + lifecycles fragments.

# 0. Deprecation timeline (read first)

Date Effect
2026-05-07 Orgs that never fine-tuned can no longer create jobs
2026-07-02 Orgs without fine-tuned-model inference in the last 60 days can no longer create jobs
2027-01-06 No new fine-tuning jobs for anyone. Inference on existing ft: models continues until the base model is deprecated

Live confirmation (2026-09-18): POST /v1/fine_tuning/jobs with a placeholder file → 403 {"code":"training_not_available","message":"OpenAI is winding down the fine-tuning platform and your organization is no longer able to create new fine-tuning training jobs…"}. Nothing below about job creation could be exercised; it is DOCUMENTED/UNVERIFIED.

# 1. Endpoints

Method / path SDK (python) Notes 2026-09-18
POST /v1/fine_tuning/jobs fine_tuning.jobs.create see §3 403 training_not_available
GET /v1/fine_tuning/jobs .list after, limit (20), metadata[k]=v filter (or metadata=null) 200 {object:"list", data:[], has_more}
GET /v1/fine_tuning/jobs/{id} .retrieve 404 fine_tune_not_found (bogus id)
POST /v1/fine_tuning/jobs/{id}/cancel .cancel any non-terminal state not called
POST /v1/fine_tuning/jobs/{id}/pause .pause RFT: snapshot + stop 404 (bogus id)
POST /v1/fine_tuning/jobs/{id}/resume .resume RFT: continue from last checkpoint not called
GET /v1/fine_tuning/jobs/{id}/events .list_events after, limit (20); events `type: message metrics(+moderation_checks` per guides)
GET /v1/fine_tuning/jobs/{id}/checkpoints .checkpoints.list after, limit (10) 404 (bogus id)
GET /v1/fine_tuning/checkpoints/{ckpt}/permissions .checkpoints.permissions.retrieve/list project_id, after, limit (10), order ascending/descending 401 Missing scopes: api.fine_tuning.checkpoints.read (Owner role / admin key)
POST /v1/fine_tuning/checkpoints/{ckpt}/permissions .permissions.create {project_ids: [...]} — share a checkpoint with other projects of the org not called (admin write)
DELETE /v1/fine_tuning/checkpoints/{ckpt}/permissions/{permission_id} .permissions.delete not called
POST /v1/fine_tuning/alpha/graders/{run,validate} see graders.md 200

# 2. Methods and base-model matrix

Method (method.type) What you supply Base models (guide, pinned snapshots) Notes
supervised (default) prompt + ideal completion gpt-4.1-2025-04-14, gpt-4.1-mini-2025-04-14, gpt-4.1-nano-2025-04-14 (also gpt-4o-2024-08-06, gpt-4o-mini-2024-07-18, legacy gpt-4, gpt-3.5-turbo) ≥ 10 examples; 50–100 recommended
vision SFT SFT lines with image_url parts gpt-4o-2024-08-06 ≤ 50 000 image examples, ≤ 10 images/example, ≤ 10 MB/image, JPEG/PNG/WEBP RGB(A); no images in assistant turns; detail: low = 85 tokens
dpo prompt + preferred + non-preferred gpt-4.1-* snapshots (guide example also shows gpt-4o-2024-08-06) text only, one-turn
reinforcement prompt + reference fields + grader o4-mini-2025-04-16 only (reasoning models) expert graders required; response_format optional
distillation = SFT on captured outputs of a stronger model any SFT base Responses API stores responses 30 d (store: true) → export → SFT (see §4.4)

Model pages marking Fine-tuning | v1/fine-tuning | Supported (2026-09-18): gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, gpt-4o, gpt-4o-mini, o4-mini, gpt-4, gpt-3.5-turbo. No gpt-5.x / gpt-6 model supports fine-tuning. A fine-tuned model can itself be the model of a new job (continued fine-tuning).

# 3. POST /v1/fine_tuning/jobs — parameters

Parameter Type Notes
model (req) string base or ft: model
training_file (req) file id (purpose=fine-tune, .jsonl)
validation_file file id periodic validation metrics; also used for full_valid_* checkpoint metrics
method.type `supervised dpo
method.supervised.hyperparameters n_epochs, batch_size, learning_rate_multiplier (each "auto" | number) top-level hyperparameters is deprecated
method.dpo.hyperparameters above + beta ("auto" | 0–2) higher β = stay closer to the reference policy
method.reinforcement.grader (req for RFT) grader object (string_check, text_similarity, python, score_model, label_model, multi) see graders.md
method.reinforcement.hyperparameters n_epochs, batch_size, learning_rate_multiplier, reasoning_effort (`default low
method.reinforcement.response_format {type:"json_schema", json_schema:{name, strict, schema}} in guide, not in spec → UNVERIFIED; populates sample.output_json
suffix ≤ 64 chars → ft:gpt-4.1-nano-2025-04-14:<org>:<suffix>:<id>
seed int ≤ 2 147 483 647 reproducibility (best effort)
integrations[] {type:"wandb", wandb:{project (req), name, entity, tags[]}} metrics streamed to Weights & Biases; default tags openai/finetune, openai/{base}, openai/{ftjob}
metadata ≤ 16 pairs filterable in list

# 4. Training data formats (JSONL, one example per line, ≥ 10 lines)

# 4.1 Supervised (chat format)

jsonl
{"messages":[{"role":"system","content":"You are a terse assistant."},{"role":"user","content":"Capital of France?"},{"role":"assistant","content":"Paris"}]}
{"messages":[{"role":"user","content":"Weather in Boston?"},{"role":"assistant","tool_calls":[{"id":"call_1","type":"function","function":{"name":"get_weather","arguments":"{\"city\":\"Boston\"}"}}]}],"tools":[{"type":"function","function":{"name":"get_weather","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}}],"parallel_tool_calls":false}
{"messages":[{"role":"user","content":"Hi"},{"role":"assistant","content":"Hello!","weight":0},{"role":"user","content":"Bye"},{"role":"assistant","content":"Goodbye!","weight":1}]}

weight (0|1) on assistant messages excludes a turn from the loss. Legacy functions and completions (prompt/completion) formats exist for legacy models.

# 4.2 Vision SFT — a user content array with {"type":"image_url","image_url":{"url":"https://… or data:image/png;base64,…","detail":"low"}} parts inside the chat format.

# 4.3 DPO (preference)

jsonl
{"input":{"messages":[{"role":"user","content":"How cold is SF today?"}],"tools":[],"parallel_tool_calls":true},"preferred_output":[{"role":"assistant","content":"Mild: high 68°F, low 57°F."}],"non_preferred_output":[{"role":"assistant","content":"Not very cold."}]}

# 4.4 RFT

jsonl
{"messages":[{"role":"user","content":"Do you have a dedicated security team?"}],"compliant":"yes","explanation":"A dedicated security team follows strict protocols."}

Every non-messages key is available to the grader as {{item.<key>}}; the model output as {{sample.output_text}} / sample.output_json / sample.output_tools.

# 4.5 Distillation recipe

  1. Prompt a strong model (gpt-4.1, or any Responses-capable model) with store: true + metadata tags; 2. pull good responses (Evals logs data source or the Responses list/GET /v1/responses/{id}); 3. write them as §4.1 lines; 4. SFT a small base (gpt-4.1-mini/-nano); 5. compare base vs. fine-tuned with an eval on held-out data.

# 5. Job object and lifecycle

fine_tuning.job: id (ftjob-…), model, fine_tuned_model (null until success), status (validating_files → queued → running → succeeded | failed | cancelled; RFT adds paused), error {code,message,param}, hyperparameters (SFT echo), method, training_file, validation_file, result_files[] (CSV metrics, purpose fine-tune-results), trained_tokens, seed, estimated_finish, integrations[], metadata, organization_id, created_at, finished_at; guide examples also show user_provided_suffix, usage_metrics, shared_with_openai.

Events (fine_tuning.job.event): level info|warn|error, message, type message|metrics, data (per-step metrics: train_loss, train_mean_token_accuracy, valid_*, full_valid_*; RFT adds train_reward_mean, valid_reward_mean, reasoning/grader token counts, per-grader scores).

Checkpoints (fine_tuning.job.checkpoint): fine_tuned_model_checkpoint = ft:<base>:<org>:<suffix>:<id>:ckpt-step-<n> (usable as a model id), step_number, metrics {step, train_loss, train_mean_token_accuracy, valid_loss, valid_mean_token_accuracy, full_valid_loss, full_valid_mean_token_accuracy}. SFT keeps the last 3 epochs' checkpoints; RFT keeps the 3 best by valid_reward_mean. Checkpoint permissions (checkpoint.permission {id, project_id, created_at}) let the org owner share a checkpoint with other projects — admin scope api.fine_tuning.checkpoints.read/write.

# 6. Safety checks

After training, 13 moderation categories (advice, harassment/threatening, hate, hate/threatening, highly-sensitive, illicit, propaganda, self-harm/instructions, self-harm/intent, sensitive, sexual/minors, sexual, violence) are evaluated; failing thresholds block deployment; details via events of type moderation_checks.

# 7. Pricing (pricing page, 2026-09-18 — cite only)

Model Training Input Cached input Output
o4-mini-2025-04-16 (RFT) $100 / hour of core training time (+ grader model tokens at that model's rate) $4.00 $1.00 $16.00 (halved with data sharing)
gpt-4.1-2025-04-14 $25 / 1M training tokens $3.00 $0.75 $12.00
gpt-4.1-mini-2025-04-14 $5.00 $0.80 $0.20 $3.20
gpt-4.1-nano-2025-04-14 $1.50 $0.20 $0.05 $0.80
gpt-4o-2024-08-06 $25.00 $3.75 $1.875 $15.00
gpt-4o-mini-2024-07-18 $3.00 $0.30 $0.15 $1.20

Batch inference on fine-tuned models is 50 % of the above. Training tokens ≈ tokens in the file × epochs.

Build an eval first (evals.md), hold out 10–20 % of the data, train, then run the same eval against the base model, the final ft: model and 1–2 checkpoints; watch full_valid_loss for overfitting; for RFT compare valid_reward_mean with a human-graded sample to detect grader hacking.

# 9. Observed error shapes

Call HTTP Body
create (any) 403 {"error":{"message":"OpenAI is winding down the fine-tuning platform and your organization is no longer able to create new fine-tuning training jobs. Learn more …","type":"invalid_request_error","param":null,"code":"training_not_available"}}
retrieve/events/checkpoints/pause on bogus id 404 {"error":{"message":"Could not find fine tune: ftjob-…","type":"invalid_request_error","param":"fine_tune_id","code":"fine_tune_not_found"}}
checkpoint permissions list 401 You have insufficient permissions for this operation. Missing scopes: api.fine_tuning.checkpoints.read…