Almost every fine-tuning tool you will meet, hosted or open source, wants its training data as JSONL: one JSON object per line, one training example per object. The format looks too simple to deserve an article. In practice it is where a surprising share of failed fine-tunes go wrong, and they rarely fail loudly. A byte order mark makes the first example unreadable, a stray pretty-printed object shifts every later line, a conversation whose content is a list in one row and a string in the next breaks schema inference, and a template that truncates from the right silently removes the very assistant turn you wanted to teach.
This article treats JSONL as a data-engineering artifact rather than an upload form. It covers the format rules, the record shapes that open-source trainers expect, what the loader does to turn a line into tokens and a loss mask, how to write and validate files as streams, how to deduplicate and split without leaking, and the failure modes worth testing for. The hosted-API schemas are covered field by field in the OpenAI fine-tuning format guide; everything here applies whichever trainer reads the file.
What JSONL actually is
JSON Lines, also called newline-delimited JSON, has exactly three rules. The file is UTF-8 and must not start with a byte order mark. Each line is one valid JSON value; objects are usual, but a blank line is not a valid value. The line terminator is \n, and \r\n is tolerated only because whitespace around a JSON value is ignored. A newline after the last value is strongly recommended but not required, and the conventional extensions are .jsonl and, compressed, .jsonl.gz.
Two consequences follow. First, a record may never contain a raw newline. JSON already forbids unescaped control characters inside strings, so a correct serializer writes a newline inside a message as the two characters backslash and n, and the record stays on one physical line. Second, the file is splittable and appendable: you can count, sample, shard, concatenate and stream it with line-oriented tools without parsing it, and a reader can resume at any line boundary. Those two properties are why trainers prefer it to a single JSON array, which must be parsed whole before the first example is available.
The record shapes trainers expect
JSONL only fixes the framing. The keys inside each object are a contract with the trainer, and the most widely copied contract in open source is the one used by Hugging Face TRL. It separates the format of a record, standard (plain strings) or conversational (lists of role and content messages), from its type, which depends on the training method.
| Type | Keys | Used by |
|---|---|---|
| Language modeling | text or messages | SFT on whole sequences, continued pretraining |
| Prompt-completion | prompt, completion | SFT with loss on the completion only |
| Prompt-only | prompt | Online methods such as GRPO that generate their own completions |
| Preference | prompt, chosen, rejected | DPO, ORPO, reward models |
| Unpaired preference | prompt, completion, label | KTO |
A conversational SFT line and a conversational preference line look like this. Note that the preference record keeps the prompt separate from both answers; TRL calls this an explicit prompt and recommends it, because the trainer then knows exactly where the shared prefix ends.
{"messages": [{"role": "system", "content": "You route support tickets."}, {"role": "user", "content": "Card charged twice for order 1182"}, {"role": "assistant", "content": "{\"queue\": \"billing\", \"priority\": \"high\"}"}]}
{"prompt": [{"role": "user", "content": "Summarise: the build failed on step 3"}], "chosen": [{"role": "assistant", "content": "Step 3 failed."}], "rejected": [{"role": "assistant", "content": "Everything passed."}]}Tool-calling data adds a tools key holding JSON Schema function definitions, and assistant messages carry tool_calls instead of text. Vision data replaces a string content with a list of typed parts. Each extension is another way for two rows in the same file to disagree in shape, which matters in the loader section below.
From a line to tokens and a loss mask
The trainer never sees your JSON. For conversational records it renders the message list through the model's chat template, which inserts the model-specific role markers such as <|im_start|>assistant, then tokenizes the resulting string. That is why the same JSONL file can train a Llama, a Qwen and a Phi model: the roles are portable, the markers are not. The Llama chat template article walks through one template in detail.
The second thing the loader builds is a loss mask: which token positions contribute to the loss. If every token counts, the model is also trained to produce the user's words and the system prompt, which wastes capacity and can teach it to echo instructions. In TRL, a prompt-completion record computes loss on the completion by default (completion_only_loss), and a conversational record can be restricted to assistant turns with assistant_only_loss=True, which requires a chat template that marks assistant spans with generation tags. The shape of your record therefore decides the mask, which is a strong reason to choose prompt-completion when each example has exactly one answer to learn.
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct")
record = {"messages": [
{"role": "user", "content": "Card charged twice for order 1182"},
{"role": "assistant", "content": "billing / high"},
]}
text = tok.apply_chat_template(record["messages"], tokenize=False)
ids = tok(text, add_special_tokens=False)["input_ids"]
print(len(ids), repr(text[-60:])) # inspect the rendered tail and the token countRender a handful of records like this before every run. It is the cheapest way to catch a missing end-of-turn token, a duplicated system prompt or a template that does not match the base model.
Writing JSONL that cannot be malformed
Never build JSONL with string formatting. Serialize each record with a real JSON encoder, refuse non-finite floats, and write to a temporary file that you rename only when the whole shard is complete, so a crashed job never leaves a truncated last line that a later job appends to.
import json, os, hashlib
def write_jsonl(path, records):
tmp = path + ".tmp"
n = 0
with open(tmp, "w", encoding="utf-8", newline="\n") as f:
for rec in records:
line = json.dumps(rec, ensure_ascii=False, allow_nan=False, separators=(",", ":"))
f.write(line + "\n")
n += 1
f.flush()
os.fsync(f.fileno())
os.replace(tmp, path) # atomic on the same filesystem
return nThree details carry weight. allow_nan=False matters because Python's default emits the bare tokens NaN and Infinity, which are not JSON and which stricter parsers reject. newline="\n" stops Windows from writing carriage returns. ensure_ascii=False keeps text readable and the file smaller; the escaped form is equally valid if you prefer pure ASCII files. The compact separators save a few percent on large datasets and change nothing semantically.
Reading and validating as a stream
A validator should read one line at a time, report the line number of every problem, and never stop at the first error, because a person fixing data wants the full list. Check framing, then shape, then content, then length.
import json, sys
ROLES = {"system", "user", "assistant", "tool"}
def validate(path, tok, max_tokens=4096):
errors, lengths = [], []
with open(path, "rb") as f:
for no, raw in enumerate(f, start=1):
if no == 1 and raw.startswith(b"\xef\xbb\xbf"):
errors.append((no, "byte order mark")); raw = raw[3:]
try:
rec = json.loads(raw.decode("utf-8"))
except (UnicodeDecodeError, json.JSONDecodeError) as e:
errors.append((no, f"not JSON: {e}")); continue
msgs = rec.get("messages")
if not isinstance(msgs, list) or not msgs:
errors.append((no, "messages missing or empty")); continue
if any(m.get("role") not in ROLES or not isinstance(m.get("content"), str) for m in msgs):
errors.append((no, "bad role or non-string content")); continue
if msgs[-1]["role"] != "assistant":
errors.append((no, "last turn is not the assistant"))
text = tok.apply_chat_template(msgs, tokenize=False)
n = len(tok(text, add_special_tokens=False)["input_ids"])
lengths.append(n)
if n > max_tokens:
errors.append((no, f"{n} tokens exceeds {max_tokens}"))
return errors, lengthsOpening the file in binary mode and splitting only on the newline byte is deliberate. Python's str.splitlines() also splits on U+2028, U+2029 and U+0085, which JSON allows unescaped inside strings, so a reader built on it can cut a valid record in half. Keep the token-length histogram the validator returns; it drives the packing and truncation decisions below.
Deduplication, splits and sharding
Duplicates in SFT data are not harmless repetition: a template answer that appears four hundred times will be memorised and over-produced. Hash a normalised form of each record, for example lower-cased, whitespace-collapsed message text joined with role names, and keep the first occurrence. Near-duplicates need shingling or MinHash, but exact hashing after normalisation usually removes the bulk.
Split by group, not by line. If one customer conversation was cut into five training records, all five must land in the same split, otherwise validation loss measures memorisation of neighbouring turns. Assign the split from a hash of the group key so the assignment is stable when you add data later.
def split_of(group_key, val_pct=5):
h = int(hashlib.sha256(group_key.encode()).hexdigest(), 16)
return "val" if h % 100 < val_pct else "train"For anything past a few hundred megabytes, write shards of roughly equal size, named train-00000-of-00016.jsonl.gz and so on, so data loaders can read them in parallel and resume at shard boundaries. Gzip is the conventional compressor and most loaders read it transparently, but a gzip stream cannot be split mid-file, so it is the shard count, not the compressor, that gives you parallelism.
Worked example: forty thousand support conversations
A team wants a 1.5-billion-parameter model to route support tickets. The export holds 41,800 conversations from a ticketing system. Normalisation maps agents to the assistant role, drops internal notes and produces one conversational record per ticket. The validator flags 312 lines: 190 with empty assistant replies, 87 whose last turn is the customer, 31 with HTML fragments that turned content into a list during export and 4 that are not valid UTF-8. Exact dedup on normalised text removes 6,140 records, almost all of them the same three canned replies. Group-aware splitting by customer id puts 1,770 records in validation.
The token histogram shows a median of 410 tokens and a 99th percentile of 3,900. Setting the maximum length to 2,048 would truncate about four percent of records, and because truncation cuts from the end, those records would lose some or all of the assistant answer, which is the only part that carries loss. The team drops records over 2,048 tokens instead of truncating them, then confirms by rendering twenty random records that every one still ends with the routing decision. After the length filter removes 1,414 records, the final training set is 32,164 lines in eight gzipped shards, and the fine-tune itself is the routine part described in the SLM fine-tuning guide.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| First example rejected or key named with a stray prefix | Byte order mark from an editor or spreadsheet export | Strip it in the writer and reject it in the validator |
| Every later line misparsed | One pretty-printed object spanning several lines | Always serialize with a JSON encoder, never by hand |
| Loader raises a cast or schema error | Arrow infers one type per column and content is a string in some rows and a list in others | Normalise to one shape; newer datasets releases (4.7 and later, per the TRL docs) add a Json feature type for arbitrary tool arguments |
| Model replies with escaped JSON inside JSON | Assistant content was json.dumps-ed twice | Decode once in the validator and compare to the expected structure |
| Model never stops generating | Template rendered without the end-of-turn token for that model | Render samples with the target tokenizer before training |
| Great training loss, poor eval | Duplicates and turn-level leakage across splits | Hash dedup and group-key splits |
| Silent loss of answers | Right truncation of long records | Filter by token length or truncate the prompt side only |
Trade-offs: JSONL, Parquet or Arrow
JSONL wins on transparency: you can read it, grep it, diff it and fix it with a text editor, and every trainer accepts it. It loses on size and speed, because every key is repeated on every line and every load reparses text. Parquet stores columns with compression and a schema, which makes it far smaller and faster to scan, but nested message lists are awkward to inspect and the schema must be uniform. A sensible rule is to author, review and version datasets as JSONL, and let the training framework convert to Arrow or Parquet in its cache. Keep the JSONL as the source of truth, because it is the form a human can audit when a model misbehaves months later.
For adapter-based training, the format choices are identical; only the training step changes, as described in the LoRA article. For preference data, the pair construction matters more than the framing, and the DPO alignment guide covers it.
What to do next
- Pick one record type per file and write it down as a schema next to the data.
- Replace any hand-built serialization with a writer that uses
allow_nan=False, a fixed newline and an atomic rename. - Run a line-numbered validator on every file in CI and fail the build on any error.
- Render twenty random records through the target chat template and read them.
- Compute the token-length histogram and decide between filtering and prompt-side truncation.
- Deduplicate on normalised text and split by a stable hash of the group key.
- Shard large datasets, keep the JSONL under version control and record its hash with every model you train.