Alignment data is the small, expensive slice of a training corpus that teaches a pretrained model how to behave: follow instructions, refuse what it should refuse, prefer a careful answer to a careless one. By token count it is tiny next to pretraining text, often well under one percent, yet it decides much of what users experience. Because it is small, every record carries weight, and every mistake in how a record is formatted, masked or truncated is repeated across thousands of gradient steps.
This article is about the engineering of that data once it exists: the three record shapes, how a chat template turns a record into tokens, which tokens the loss is allowed to see, how to pack variable-length examples so a GPU is not multiplying padding, and why preference training costs roughly twice the forward passes of supervised fine-tuning unless you cache the reference model. How humans should judge and collect preferences is covered in preference data collection; here we start where that page ends, with a file of records and a GPU budget.
Three record shapes
Almost every alignment dataset reduces to one of three shapes, and each one feeds a different objective.
| Shape | Fields | Objective | What the loss sees |
|---|---|---|---|
| Demonstration (SFT) | messages: [{role, content}, ...] | Next-token cross-entropy | Assistant turns only, usually |
| Preference pair | prompt, chosen, rejected | DPO, IPO, reward-model training | Log-prob gap between the two responses |
| Unpaired label | prompt, completion, label: bool | KTO-style objectives | One completion, pushed up or down |
The shapes are not interchangeable. A preference pair only carries signal through the difference between chosen and rejected, so if both responses share a long identical prefix, that prefix contributes nothing to the gradient and only costs compute. A demonstration carries signal on every assistant token, so a single sloppy demonstration teaches sloppiness directly. Unpaired labels are the cheapest to collect from production thumbs-up and thumbs-down signals, but they are noisier because the model never sees what the better answer would have looked like.
Store all three in one normalised schema with an explicit source, license, created_at and split_hint field. Those metadata fields look like bureaucracy until the day you need to remove one vendor's data, rerun an ablation without synthetic records, or prove that an evaluation prompt never entered training.
Chat templates and token offsets
A model does not see roles; it sees tokens. The chat template is the function that serialises [{role, content}] into a single string with special tokens marking where each turn begins and ends. The single most common alignment-data bug is training with one template and serving with another: a missing end-of-turn token at training time produces a model that never stops generating, and an extra system preamble at serving time produces a model that behaves subtly differently from the one you evaluated.
The rule is to render with the tokenizer's own template, the same object the inference server will use, and to build labels from token offsets rather than by re-tokenising fragments. Tokenisers merge across boundaries, so tokenising the prompt and response separately and concatenating the ids does not always equal tokenising the whole string. The robust approach is to render the conversation incrementally and record where each assistant span starts and ends in token space.
IGNORE = -100 # PyTorch cross_entropy skips this label
def build_sft_example(tok, messages, max_len):
ids, labels = [], []
for i in range(len(messages)):
# Render the conversation up to and including turn i with the real template.
prefix = tok.apply_chat_template(messages[: i + 1], tokenize=True,
add_generation_prompt=False)
new = prefix[len(ids):] # tokens contributed by turn i
if prefix[: len(ids)] != ids:
raise ValueError("template is not prefix-stable; mask by offsets instead")
ids.extend(new)
train_on = messages[i]["role"] == "assistant"
labels.extend(new if train_on else [IGNORE] * len(new))
if len(ids) > max_len:
return None # drop, never silently cut the answer
return {"input_ids": ids, "labels": labels}The prefix-stability check matters: some templates rewrite earlier turns when a later one is appended (for example, moving a system prompt or inserting a date), and if that happens your offsets are wrong. Failing loudly is better than training on a shifted mask. Note too that the assistant span should include the end-of-turn token; that token is how the model learns to stop.
Loss masking
Loss masking decides what behaviour the model is rewarded for producing. Training on user turns teaches the model to predict user text, which wastes capacity and can make it imitate the user's style or continue writing the next user message. Training on the system prompt teaches it to regurgitate the system prompt. The standard choice is to compute loss only on assistant tokens, including tool calls the assistant emits, and to mask tool results, because those are produced by the environment, not by the model.
Masking also changes how you should average the loss. If you average per token across a batch, long responses dominate the gradient; if you average per example, short responses get the same weight as long ones. Neither is wrong, but switching between them changes effective learning rate by a factor that depends on your length distribution. Pick one, write it down, and keep it constant across ablations. With gradient accumulation, normalise by the total count of unmasked tokens over all micro-batches, not per micro-batch, or accumulation steps with short examples are silently up-weighted.
Packing without cross-contamination
Alignment examples vary from a dozen tokens to many thousands. Naive batching pads every example to the longest one in the batch, and the GPU spends its tensor-core time multiplying zeros. On a typical SFT mix, padding can easily exceed half the tokens processed. There are two remedies.
Length bucketing sorts examples into buckets of similar length and draws each batch from one bucket. It is simple and keeps examples independent, but it correlates batch composition with length, which can interact badly with learning-rate schedules if long examples cluster in time. Shuffle the order of buckets.
Packing concatenates several examples into one fixed-length row. Done naively, tokens in example two can attend to example one, which leaks unrelated context and teaches the model that conversations bleed into each other. Done properly, the attention kernel receives the boundary offsets and treats each segment as its own sequence. Variable-length attention kernels such as FlashAttention's varlen path take cumulative sequence lengths for exactly this purpose, and position ids must restart at zero for each segment.
def pack(examples, row_len):
rows, cur, cu = [], {"input_ids": [], "labels": [], "position_ids": []}, [0]
for ex in sorted(examples, key=lambda e: -len(e["input_ids"])): # first-fit decreasing
n = len(ex["input_ids"])
if len(cur["input_ids"]) + n > row_len:
rows.append((cur, cu)); cur, cu = {"input_ids": [], "labels": [], "position_ids": []}, [0]
cur["input_ids"] += ex["input_ids"]
cur["labels"] += [IGNORE] + ex["labels"][1:] # never predict across a boundary
cur["position_ids"] += list(range(n)) # restart positions per segment
cu.append(cu[-1] + n) # cu_seqlens boundary for the kernel
if cur["input_ids"]:
rows.append((cur, cu))
return rowsTwo details are easy to miss. First, the label of the last token of one segment must not be asked to predict the first token of the next; because labels are shifted by one inside the loss, mask the first label of every segment or shift per segment. Second, packing changes how many examples a step contains, so log examples per step and unmasked tokens per step separately; a jump in one without the other is a data-pipeline bug, not a modelling result.
Preference data and the reference model
Direct preference optimisation needs, for every pair, the log-probability of chosen and rejected under the policy being trained and under a frozen reference model, usually the SFT checkpoint. Done naively that is four forward passes per pair and two copies of the model in GPU memory. The loss for one pair is:
# logp_* are summed log-probs over response tokens only (prompt masked)
margin = beta * ((logp_pol_chosen - logp_ref_chosen) - (logp_pol_rejected - logp_ref_rejected))
loss = -torch.nn.functional.logsigmoid(margin).mean()The reference terms never change during training, so compute them once in an offline pass, store two floats per pair next to the record, and drop the reference model from the training job entirely. That halves forward compute and frees the memory a second model copy would occupy, which is often the difference between fitting a model on one node and needing two. The cache is only valid for one exact tokenisation and one reference checkpoint, so key it by a hash of both; a changed template with a stale cache produces a margin computed on different tokens, which trains confidently toward nonsense.
Concatenate chosen and rejected into one batch so that a single forward pass computes both, which keeps the kernels busy with a larger batch. Sum log-probs over response tokens only; including the shared prompt adds identical terms to both sides that cancel mathematically but still cost compute and add numerical noise. The GPU mechanics of the trainer itself are covered in DPO training on GPUs.
Quality controls that change results
Deduplication. Exact-hash dedup removes copies; near-duplicate dedup with MinHash over word shingles removes templated variants that differ by a name or number. Without it, a generator that produced the same answer skeleton two thousand times will dominate SFT and the model will adopt that skeleton everywhere.
Decontamination. Before every training run, compare prompts and responses against every evaluation set you will report on, using long n-gram overlap (13-grams are a common choice) and embedding similarity for paraphrases. Remove matches from training, not from evaluation. Contamination does not just inflate a benchmark; it destroys your ability to tell whether a data change helped.
Length bias. In many preference datasets the chosen response is longer than the rejected one more often than not, because raters and judge models both favour thoroughness. A DPO model learns that length is the cheapest way to raise the margin. Measure the fraction of pairs where chosen is longer; if it is far from half, rebalance, add pairs where the shorter answer wins, or use a length-normalised objective.
Degenerate pairs. Drop pairs where chosen and rejected are identical after normalisation, where either side was truncated, or where both sides are refusals. Each of these yields a gradient that is zero, random or actively misleading.
Synthetic share. Synthetic demonstrations from a stronger model are cheap and useful, but they transfer that model's habits and errors. Tag them, cap their share per task, and ablate with and without; see synthetic data generation for generation-side controls.
Worked example: auditing a 100,000-record mix
Suppose you have 60,000 SFT conversations and 40,000 preference pairs for a 7B-parameter model, a context length of 4,096 and eight GPUs. A first audit run of the pipeline above reports the following, and each line maps to an action.
- Mean SFT length 610 tokens, p99 3,900. Padded batching at 4,096 would process about 6.7 times more tokens than it trains on, so pack to 4,096 with first-fit decreasing; the audit shows rows about 95 percent full.
- 1.8 percent of SFT records exceed 4,096 tokens. Drop them rather than truncating, because truncation would cut the answer and its end-of-turn token.
- Near-duplicate clusters: one synthetic task contributes 9,000 records that collapse to 700 clusters. Keep at most a few per cluster.
- Decontamination finds 212 prompts that overlap a reported benchmark. Remove them and record the removal list.
- Chosen is longer than rejected in 71 percent of pairs. Add 3,000 pairs where a concise answer is preferred and track response length on a held-out prompt set during training.
- Reference log-probs: one offline pass over 40,000 pairs, stored as two float32 values per pair keyed by template and checkpoint hash. The DPO job now loads only the policy.
None of these steps changed a hyperparameter, yet together they decide whether the run trains on the data you think it trains on.
Failure modes
| Symptom | Likely data cause | Check |
|---|---|---|
| Model never stops generating | End-of-turn token masked or missing from labels | Decode one example's unmasked labels |
| Model repeats the user | Loss computed on user turns | Count unmasked tokens by role |
| Answers keep growing after DPO | Length bias in pairs | Fraction chosen-longer; length on held-out prompts |
| DPO loss falls instantly to near zero | Stale reference cache or chosen leaked into prompt | Recompute margins on 100 pairs |
| Benchmark jumps, live quality does not | Evaluation contamination | n-gram overlap report |
| Throughput far below peak | Padding or unpacked short examples | Unmasked tokens / processed tokens |
Make every check in this table a script that runs in the data pipeline and writes a report next to the dataset version. Reading the decoded text of a few random examples, with masked tokens rendered differently, catches more bugs per minute than any metric.
Trade-offs
More data is not automatically better. A smaller, deduplicated, carefully masked set routinely beats a larger noisy one, and it trains faster. Packing gains throughput but makes per-example debugging harder, so keep an unpacked debug path. Caching reference log-probs saves compute but couples the dataset to one checkpoint and template. Synthetic data scales cheaply but narrows style. Strict decontamination removes some genuinely useful prompts. Choose consciously, and record the choice in the dataset card so the next run can repeat or reverse it. For how this data feeds the full pipeline, see the RLHF pipeline and SFT fine-tuning on GPUs.
What to do next
- Normalise every source into one schema with source, license, created_at and split_hint fields.
- Render with the serving tokenizer's chat template and assert prefix stability on a sample.
- Mask everything except assistant tokens, including the end-of-turn token in the trained span.
- Drop over-length examples instead of truncating; log how many you dropped.
- Pack with per-segment position ids and cu_seqlens; mask the first label of each segment.
- Run MinHash dedup and n-gram decontamination against every reported evaluation set.
- Measure chosen-longer rate and degenerate pairs before any preference run.
- Precompute reference log-probs keyed by template and checkpoint hash.