OpenAI's fine-tuning API takes training data as JSON Lines: a text file where every line is one complete JSON object describing one training example. The shape of that object depends on the training method. Supervised fine-tuning uses a list of chat messages, the same structure you send to the chat API; preference tuning uses a prompt with a preferred and a non-preferred answer; reinforcement fine-tuning uses a prompt plus whatever reference fields your grader needs.

First, the status, because it changes what you should do with this page. OpenAI is winding down its self-serve fine-tuning platform. According to its deprecations page, organisations that had never run fine-tuning lost the ability to create jobs on 7 May 2026; from 2 July 2026 job creation also stopped for organisations without inference on a fine-tuned model in the previous 60 days; and on 6 January 2027 the remaining active customers lose the ability to create new jobs. Inference on existing fine-tuned models continues until their base models are deprecated.

Why learn the format, then? Because messages JSONL is the common currency for chat fine-tuning: open-weight trainers accept it almost unchanged, and if you still have access, a correct file is the difference between a useful model and a wasted job. This article covers the format field by field, a validator, a worked example and the move to open models.

Advertisement

The pipeline and where the format sits

From JSONL lines to a fine-tuned model: one data contract, three training methodstrain.jsonlone example per linevalid.jsonlheld-out examplesLocal validatorschema, roles, tokensFiles APIpurpose=fine-tuneFine-tuning jobmethod + base modelsupervisedmessagesdpopref. pairsreinforcementgraderft: model idevents, metrics, checkpointsSame messages JSONLalso feeds open-weight trainers (TRL, etc.)portableSelf-serve job creation is being wound down; the data format remains the common currency.
Training and validation files are checked locally, uploaded through the Files API with purpose fine-tune, and referenced by a job that names the base model and the method. The same messages files can be used to train an open-weight model.

Everything the model learns comes from the examples; the method decides only how they become a loss. The data file is the real specification of your fine-tune, so treat it like code: versioned, reviewed and tested.

Supervised fine-tuning: the messages format

Each line holds an object with a messages array. Each message has a role and content. Here are two examples for a ticket router whose target output is a small JSON object:

{"messages": [{"role": "system", "content": "Classify the support ticket. Reply with JSON."}, {"role": "user", "content": "My invoice for March was charged twice."}, {"role": "assistant", "content": "{\"queue\": \"billing\", \"priority\": \"high\"}"}]}
{"messages": [{"role": "system", "content": "Classify the support ticket. Reply with JSON."}, {"role": "user", "content": "How do I export my data to CSV?"}, {"role": "assistant", "content": "{\"queue\": \"how_to\", \"priority\": \"low\"}"}]}

Points that matter in practice:

  • One object per line. No pretty-printing across lines in the real file, no trailing commas, UTF-8. The examples on this page are wrapped for reading only.
  • The assistant message is the target. The model is trained to produce the assistant turns given everything before them. System and user turns are context, not targets.
  • Train with the production system prompt. Otherwise you serve something you did not train. If you shorten the prompt, train with the short one.
  • JSON output is still a string, serialised inside content with quotes escaped.
  • Minimum and useful size. The guide sets a hard minimum of 10 examples and reports clear improvements from 50 to 100 well-crafted ones. Quality and coverage of edge cases beat raw volume.
Advertisement

Multi-turn examples and the weight key

A multi-turn conversation contains several assistant messages. By default they are all training targets. The optional weight key on an assistant message, set to 0 or 1, controls that: weight 0 keeps the message as context but excludes it from training. This is how you include a realistic bad turn, followed by a correction, without teaching the model to produce the bad turn.

{"messages": [
  {"role": "system", "content": "You are a terse SQL tutor."},
  {"role": "user", "content": "What does COALESCE do?"},
  {"role": "assistant", "content": "COALESCE returns the first non-null argument, I think maybe.", "weight": 0},
  {"role": "user", "content": "Be precise."},
  {"role": "assistant", "content": "COALESCE(a, b, ...) returns the first argument that is not NULL, or NULL if all are.", "weight": 1}
]}

Use weights sparingly: zeroing every turn but the last throws away most of the signal you paid for. Zero only turns you do not want imitated.

Tool calling examples

To teach a model when and how to call your functions, the example includes a tools list in the same JSON-schema shape as the chat API, assistant messages with tool_calls, tool messages carrying the results, and a final assistant answer. parallel_tool_calls can be set per example.

{"messages": [
  {"role": "user", "content": "What's the order status for A-1042?"},
  {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function",
     "function": {"name": "get_order", "arguments": "{\"order_id\": \"A-1042\"}"}}]},
  {"role": "tool", "tool_call_id": "call_1", "content": "{\"status\": \"shipped\", \"eta\": \"2026-10-03\"}"},
  {"role": "assistant", "content": "Order A-1042 has shipped and should arrive on 3 October."}
 ],
 "tools": [{"type": "function", "function": {"name": "get_order",
   "description": "Look up an order by id",
   "parameters": {"type": "object", "properties": {"order_id": {"type": "string"}},
                  "required": ["order_id"]}}}],
 "parallel_tool_calls": false}

Three things go wrong often. The arguments field is a JSON-encoded string, not an object. Every tool_call_id in a tool message must match an id in a preceding call. And the tool definitions in training should match those you will send in production; a model trained with one parameter name and served with another will fail quietly. More on designing tool schemas in function calling with small models.

Preference tuning: the DPO format

Direct preference optimisation teaches the model to prefer one answer over another for the same prompt. Each line has an input (messages, and optionally tools and parallel_tool_calls), a preferred_output and a non_preferred_output, each a list containing the final assistant message.

{"input": {"messages": [{"role": "user", "content": "Summarise this incident in one sentence: ..."}],
           "tools": [], "parallel_tool_calls": true},
 "preferred_output": [{"role": "assistant", "content": "A bad config push at 14:02 dropped 3% of API requests for 11 minutes."}],
 "non_preferred_output": [{"role": "assistant", "content": "There was an incident that affected some users for a while."}]}

The guide states that training currently uses one-turn conversations: the preferred and non-preferred messages must be the last assistant message. The job uses method={"type": "dpo", ...} with a beta hyperparameter, either a number up to 2 or "auto"; a higher beta keeps the model closer to its previous behaviour. Good pairs differ in the property you care about and as little as possible otherwise. If the preferred answers are also systematically longer, the model learns length. A common sequence is supervised fine-tuning first, then DPO on top. The mechanics of the loss are covered in DPO alignment.

Reinforcement fine-tuning and vision examples

Reinforcement fine-tuning (RFT) is supported only on o-series reasoning models; the guide lists o4-mini-2025-04-16. Each line has a messages array holding the prompt, plus any extra fields your grader needs. The grader, defined in the job's method, scores each sampled answer, and templates reference the training item as {{item.field}} and the model's answer as {{sample.output_text}} or {{sample.output_json.field}}.

{"messages": [{"role": "user", "content": "Is this clause compliant with policy 4.2? <clause text>"}],
 "compliant": "no", "explanation": "Retention exceeds 90 days."}

grader (inside method.reinforcement):
{"type": "string_check", "name": "verdict", "operation": "eq",
 "input": "{{sample.output_json.compliant}}", "reference": "{{item.compliant}}"}

Grader types include string_check, text_similarity, score_model, python and multi. RFT is worth considering only when answers can be checked automatically and reliably; a grader that can be gamed will be gamed.

Vision fine-tuning puts images in user messages as content parts of type image_url, with either an HTTPS URL or a base64 data URL. The guide limits each example to 10 images of at most 10 MB, in JPEG, PNG or WEBP, and skips images containing people, faces, children or CAPTCHAs. Assistant messages cannot contain images. The guide lists gpt-4o-2024-08-06 for vision.

Validate locally before you upload

The API rejects malformed files, but only after upload, and some problems, such as a truncated example, are not errors at all. A local validator catches both. The token count below is approximate because it ignores the few framing tokens added per message, which is fine for spotting outliers.

import json, sys
import tiktoken   # o200k_base is the encoding used by the gpt-4o and gpt-4.1 families

enc = tiktoken.get_encoding("o200k_base")
ROLES = {"system", "user", "assistant", "tool"}

def check(path, max_tokens=65_536):
    errors, sizes = [], []
    for n, line in enumerate(open(path, encoding="utf-8"), 1):
        try:
            ex = json.loads(line)
        except json.JSONDecodeError as e:
            errors.append(f"{n}: not JSON ({e})"); continue
        msgs = ex.get("messages")
        if not isinstance(msgs, list) or not msgs:
            errors.append(f"{n}: missing messages"); continue
        open_calls = set()
        for m in msgs:
            if m.get("role") not in ROLES:
                errors.append(f"{n}: bad role {m.get('role')!r}")
            if "weight" in m and (m["role"] != "assistant" or m["weight"] not in (0, 1)):
                errors.append(f"{n}: weight must be 0/1 on assistant messages")
            for call in m.get("tool_calls", []):
                open_calls.add(call["id"])
                json.loads(call["function"]["arguments"])   # arguments are a JSON *string*
            if m.get("role") == "tool":
                open_calls.discard(m.get("tool_call_id"))
        if open_calls:
            errors.append(f"{n}: tool calls without results {open_calls}")
        if msgs[-1].get("role") != "assistant":
            errors.append(f"{n}: last message should be the assistant target")
        toks = sum(len(enc.encode(m["content"])) for m in msgs
                   if isinstance(m.get("content"), str))   # image parts not counted
        sizes.append(toks)            # approximate: ignores per-message framing tokens
        if toks > max_tokens:
            errors.append(f"{n}: ~{toks} tokens, will be truncated")
    print(f"{len(sizes)} examples, {sum(sizes)} content tokens, max {max(sizes, default=0)}")
    return errors

if __name__ == "__main__":
    problems = check(sys.argv[1])
    print("\n".join(problems[:50]) or "ok")
    sys.exit(1 if problems else 0)

The limit reflects the best-practices guide's 65,536 tokens per example for gpt-4o-mini and gpt-4.1-mini; check your model's limit. Over-long examples are truncated, and if that removes the target, the example teaches nothing. Also check for duplicates, train/validation leakage and personal data.

Creating and monitoring a job

Upload both files with purpose fine-tune, then create a job naming the base model, the method and optional hyperparameters. The supervised and DPO guides list the dated gpt-4.1, gpt-4.1-mini and gpt-4.1-nano snapshots.

from openai import OpenAI
client = OpenAI()

train = client.files.create(file=open("train.jsonl", "rb"), purpose="fine-tune")
valid = client.files.create(file=open("valid.jsonl", "rb"), purpose="fine-tune")

job = client.fine_tuning.jobs.create(
    model="gpt-4.1-mini-2025-04-14",
    training_file=train.id,
    validation_file=valid.id,
    suffix="ticket-router",
    method={"type": "supervised",
            "supervised": {"hyperparameters": {"n_epochs": 3}}},
)

for ev in client.fine_tuning.jobs.list_events(job.id, limit=20):
    print(ev.created_at, ev.message)

# when job.status == "succeeded", job.fine_tuned_model holds an "ft:..." model id

Hyperparameters such as n_epochs, batch_size and learning_rate_multiplier can be left automatic; move epochs by one or two if the model under-fits (ignores your format) or over-fits (copies training answers). The job reports training and validation loss, and checkpoints let you pick an earlier model. Loss is a proxy: run your own task evaluation, as in evaluating small models, before switching traffic.

Worked example: a ticket router

A support team routes tickets into 12 queues with two priorities. A prompted base model reaches 81% queue accuracy on a 500-ticket labelled set, mostly failing on rare queues and returning prose instead of JSON about 2% of the time.

  1. Data. Sample 1,200 historical tickets stratified by queue so rare queues have at least 40 examples each. Two agents relabel disagreements. Hold out 200 as validation and keep the original 500 as the test set, de-duplicated against training.
  2. Format. Build SFT lines with the production system prompt and the JSON target, as in the first example. Run the validator: three lines failed because labels contained raw newlines; one ticket of 90,000 tokens (a pasted log) was trimmed.
  3. Train. Three epochs on a mini model. Validation loss flattened after the second epoch.
  4. Evaluate. On the test set, queue accuracy rose to 93% and invalid JSON fell to zero. Rare-queue accuracy was the weakest bucket, so 150 more rare-queue examples went into the next version.
  5. Plan for the wind-down. Because job creation ends on 6 January 2027, the team trained the same files on a small open model, following fine-tuning small language models and LoRA, and kept whichever scored better on the same test set.

Porting the same data to open-weight models

The messages format maps directly onto open-weight tooling. Hugging Face TRL's SFTTrainer accepts a conversational dataset with a messages column and applies the target model's chat template, and its DPO trainer expects prompt, chosen and rejected columns, a mechanical conversion from the OpenAI DPO lines.

# The same files feed an open-weight trainer. TRL's SFTTrainer accepts a
# conversational dataset with a "messages" column and applies the model's chat template.
from datasets import load_dataset
from trl import SFTConfig, SFTTrainer

ds = load_dataset("json", data_files={"train": "train.jsonl", "test": "valid.jsonl"})
trainer = SFTTrainer(model="Qwen/Qwen2.5-1.5B-Instruct",
                     train_dataset=ds["train"], eval_dataset=ds["test"],
                     args=SFTConfig(output_dir="ticket-router", num_train_epochs=3))
trainer.train()

# DPO pairs convert to TRL's prompt / chosen / rejected columns:
def to_trl(ex):
    return {"prompt": ex["input"]["messages"],
            "chosen": ex["preferred_output"],
            "rejected": ex["non_preferred_output"]}

Two adjustments are usually needed. Chat templates differ between model families, so let the tokenizer apply the template rather than hand-writing special tokens; chat templates explains why. And per-message weights have no universal equivalent: many trainers compute loss on all assistant tokens, so drop weight-0 turns or mask them yourself.

Failure modes

  • Train-serve prompt mismatch. Different system prompt or tool schemas in production. Generate training lines from the same code path that builds production requests.
  • Format learned, facts not. Fine-tuning teaches style and narrow decisions, not changing knowledge; use retrieval.
  • Silent truncation. Over-long examples lose their targets. Validate token counts.
  • Label noise. A few percent of contradictory labels caps achievable accuracy. Relabel disagreements before adding volume.
  • Leakage. Test examples in training inflate scores. De-duplicate across splits.
  • Platform dependency. A model you cannot retrain is a liability once job creation ends. Keep data and evaluation independent of any one provider.

What to do next

  1. Check whether your organisation can still create fine-tuning jobs, and note 6 January 2027 if it can.
  2. Export your training data to the messages JSONL format and put it under version control with a data card.
  3. Run the validator, and fix every role, tool-call and token-limit problem it reports.
  4. Build a fixed test set and an evaluation script before training anything.
  5. Train once on the hosted API if available, and once on a small open model with the same files; compare on the same test set.
  6. Convert any preference data to both the DPO format and TRL's prompt, chosen, rejected columns.
  7. Inventory existing fine-tuned models and their base models' deprecation dates, and plan replacements.
Key takeaway: OpenAI fine-tuning data is JSONL with one example per line. Supervised examples are chat messages whose assistant turns are the targets, optionally weighted 0 or 1, with tool calls and tool results in chat-API shape. DPO lines pair a one-turn input with preferred and non-preferred outputs, and RFT lines carry reference fields for a grader. Self-serve job creation ends for remaining customers on 6 January 2027, but the format remains portable: validate it locally, evaluate on a fixed test set, and keep the same files ready for open-weight trainers.