Direct Preference Optimization (Rafailov and colleagues, 2023) is the most common way to align a small language model after supervised fine-tuning. It takes pairs of answers to the same prompt, one preferred and one not, and trains the model to make the preferred answer more likely relative to the other, without training a separate reward model and without sampling during training. That simplicity is why it fits small-model teams: one trainer, one dataset, one GPU for a model of a billion or so parameters.
Simple to run is not the same as simple to get right. Most failed DPO runs fail because of the data or because nobody read the metrics, not because of the loss. This page walks the workflow a practitioner actually follows: what the loss computes, how to build pairs, how to configure TRL, what each logged number means and how to evaluate the result. The derivation is in DPO math and the GPU memory analysis in DPO training on GPU. TRL parameter names were checked against its documentation in October 2026; they have changed across releases, so check yours.
What DPO optimises
Classic RLHF trains a reward model on preference pairs, then uses reinforcement learning to maximise that reward while a KL penalty keeps the policy close to a reference model. The DPO paper showed that the optimal policy for that objective has a closed form, and that rearranging it expresses the reward in terms of the policy itself: the reward of an answer is beta times the log of how much more likely the policy makes it than the reference does. Substituting that into the Bradley-Terry model of preferences gives a loss that needs only log-probabilities from two models.
In words: for each pair, compute how much the trained policy has raised the chosen answer's log-probability above the reference, and the same for the rejected answer. The difference, scaled by beta, is a margin. The loss is the negative log-sigmoid of that margin, so it is large when the policy does not prefer the chosen answer and approaches zero as it comes to prefer it strongly. The reference model, usually a frozen copy of the SFT model, is what stops the policy from drifting into degenerate text that games the margin.
The loss in code
The whole algorithm fits in a few lines. Each answer's log-probability is the sum of per-token log-probabilities over the completion only; prompt tokens are masked out, because both answers share the prompt and it carries no preference signal.
import torch
import torch.nn.functional as F
def sequence_logp(model, input_ids, attention_mask, completion_mask):
"""Sum of log-probabilities of the completion tokens only (prompt tokens masked out)."""
logits = model(input_ids=input_ids, attention_mask=attention_mask).logits[:, :-1]
targets = input_ids[:, 1:]
logp = torch.gather(logits.log_softmax(-1), 2, targets.unsqueeze(-1)).squeeze(-1)
return (logp * completion_mask[:, 1:]).sum(-1)
def dpo_loss(policy, reference, chosen, rejected, beta=0.1):
pi_c = sequence_logp(policy, **chosen)
pi_r = sequence_logp(policy, **rejected)
with torch.no_grad(): # the reference is frozen
ref_c = sequence_logp(reference, **chosen)
ref_r = sequence_logp(reference, **rejected)
reward_c = beta * (pi_c - ref_c) # implicit reward of the chosen answer
reward_r = beta * (pi_r - ref_r)
loss = -F.logsigmoid(reward_c - reward_r).mean()
return loss, reward_c.detach(), reward_r.detach()Each step therefore needs four forward passes: policy and reference, on chosen and rejected. Only the two policy passes need gradients. Libraries batch chosen and rejected together and may precompute the reference log-probabilities once for the whole dataset, which TRL offers as precompute_ref_log_probs, trading a preprocessing pass for not holding the reference in memory.
A worked example of the numbers
Take one pair. The reference assigns the chosen answer a log-probability of -42.0 and the rejected answer -49.0. At the first step the policy is identical to the reference, so both log-ratios are zero, the margin is zero and the loss is -log(0.5), about 0.693. Every DPO run starts at that loss, which makes it a useful sanity check: if your first logged loss is not near 0.693, something is wrong with masking or the reference.
After some training the policy gives the chosen answer -40.0 and the rejected -50.0. The chosen log-ratio is +2.0 and the rejected is -1.0. With beta 0.1 the implicit rewards are +0.2 and -0.1, so the margin is 0.3. The loss is -log(sigmoid(0.3)) = log(1 + e^-0.3), about 0.554. In TRL's logs these appear as rewards/chosen 0.2, rewards/rejected -0.1 and rewards/margins 0.3, and this pair counts towards rewards/accuracies because the chosen reward is higher.
Notice what beta does. With beta 0.5 the same log-probabilities give a margin of 1.5 and a loss near 0.20, so the gradient fades sooner and the policy stays closer to the reference. A small beta needs much larger log-ratio changes before the loss is satisfied, which lets the policy move further.
Where DPO sits in the pipeline
DPO is a second stage. It assumes a model that already follows the format and the task, which is what supervised fine-tuning provides. Run DPO on a base model and the pairs mostly teach format rather than preference. The SFT checkpoint becomes both the starting policy and the frozen reference.
How DPO compares with reward-model-plus-RL pipelines, and when the extra machinery is worth it, is covered in the RLHF training pipeline. The component view for small models, including where gates sit, is in DPO alignment architecture for SLMs.
Building the pairs
Data quality decides the outcome. Three properties matter most.
First, pairs should be on-policy where possible: answers sampled from your own SFT model, not from a different, stronger model. DPO raises and lowers the likelihood of the answers it is shown; if those answers are text your model would never produce, the gradient is spent moving probability mass in regions it never samples from, and much of the improvement does not show up in generation.
Second, the preference must be clear and about the right thing. Score answers with a rubric, a verifier for tasks with checkable outputs such as code that must pass tests or JSON that must validate, a strong model as judge, or humans. Drop pairs whose scores are close; near-ties are mostly label noise.
Third, control shortcuts. If the chosen answer is usually longer, the model learns that length is preferred and verbosity grows. Check the length ratio across the dataset and balance it, or use a length-normalised loss.
def build_pairs(prompts, sft_generate, score, k=4, min_gap=1.0):
"""On-policy pairs: sample from the SFT model, score, keep clear best-vs-worst pairs."""
pairs = []
for prompt in prompts:
answers = [sft_generate(prompt, temperature=0.8) for _ in range(k)]
scored = sorted(((score(prompt, a), a) for a in answers), key=lambda t: t[0])
(lo, worst), (hi, best) = scored[0], scored[-1]
if hi - lo < min_gap: # ties and near-ties teach noise
continue
if len(best) > 1.5 * len(worst): # watch for a length shortcut
flag_for_review(prompt, best, worst)
pairs.append({
"prompt": [{"role": "user", "content": prompt}],
"chosen": [{"role": "assistant", "content": best}],
"rejected": [{"role": "assistant", "content": worst}],
})
return pairsA few thousand clean pairs from your own domain often do more than a large generic preference dataset. Hold out 5 to 10 percent of pairs for evaluation and deduplicate prompts across the split.
Configuring TRL
TRL's DPOTrainer accepts preference data in standard or conversational form with prompt, chosen and rejected columns, and applies the chat template to conversational data. An explicit prompt column is recommended. The tokenizer must have a padding token and pad on the left. When no reference model is passed, the trainer uses the initial state of the policy; with a PEFT adapter, the frozen base weights serve that role without a second copy in memory.
from datasets import load_dataset
from peft import LoraConfig
from trl import DPOConfig, DPOTrainer
train = load_dataset("json", data_files="pairs_train.jsonl", split="train")
held_out = load_dataset("json", data_files="pairs_eval.jsonl", split="train")
args = DPOConfig(
output_dir="slm-dpo",
beta=0.1, # strength of the pull toward the reference
loss_type=["sigmoid"], # the original DPO loss; a list, so losses can be combined
learning_rate=1e-5, # LoRA; full fine-tuning defaults to 1e-6
num_train_epochs=1,
per_device_train_batch_size=4,
gradient_accumulation_steps=8, # effective batch of 32 pairs
max_length=1024, # prompt + completion; check how many pairs get truncated
eval_strategy="steps",
eval_steps=50,
logging_steps=10,
)
trainer = DPOTrainer(
model="my-org/slm-1b-sft", # the SFT checkpoint; with PEFT the frozen base serves as reference
args=args,
train_dataset=train,
eval_dataset=held_out,
peft_config=LoraConfig(r=16, lora_alpha=32, target_modules="all-linear"),
)
trainer.train()Defaults worth knowing in current TRL: beta is 0.1, learning_rate is 1e-6 (the docs suggest around 1e-5 for adapters), max_length is 1024 tokens with truncation keeping the start, bf16 and gradient checkpointing are on, and dropout is disabled in both models. Measure how many pairs exceed max_length: a truncated chosen answer loses exactly the ending that made it better.
Reading the metrics
TRL logs the implicit rewards and the raw log-probabilities. Read them together; each pattern below points at a specific cause.
| Pattern | What it means | What to do |
|---|---|---|
| loss starts far from 0.693 | Masking, template or reference bug | Stop and check that prompt tokens are masked and the reference matches the start policy |
| rewards/accuracies near 1.0 within a few dozen steps | Pairs are trivially separable, often by length or format | Inspect pairs, balance length, add harder pairs |
| margins rise but logps/chosen falls steadily | The model is pushing both answers down, the rejected faster | Common and not fatal alone; if generations degrade, raise beta, lower the learning rate or add an SFT term |
| margins grow without bound, eval accuracy flat | Overfitting the training pairs | Fewer epochs, more data, higher beta |
| grad_norm spikes, loss jumps | Learning rate too high for full fine-tuning | Lower the learning rate, add warmup |
| mean output length rises across checkpoints | Length shortcut in the data | Length-balance pairs or try sigmoid_norm |
The falling chosen log-probability deserves emphasis because it surprises people. The DPO loss depends only on the margin, so the cheapest way to satisfy it is often to lower both answers' likelihood, the rejected one more. TRL's documentation notes that in practice DPO tends to work by suppressing dispreferred completions rather than raising preferred ones. A moderate drop is normal; a large one means probability mass is going to text in neither answer, and that shows up as odd generations.
Choosing a loss type
TRL implements many published variants through loss_type, which takes a list so losses can be combined with loss_weights. Start with sigmoid, the original. Move only when the metrics show a specific problem: ipo for overfitting on near-deterministic preferences, since it regresses the margin towards a target instead of pushing it ever higher; robust with label_smoothing set to your estimated label-flip rate (the documented range is 0 to below 0.5) when labels are noisy; sigmoid_norm, the SimPO-style length normalisation, when length bias appears; and adding sft with a weight to keep the chosen answers' likelihood up. Change one thing at a time and compare on the same held-out set.
Evaluating the result
Training metrics show the model separates the pairs it saw. They do not show it got better. Evaluate three ways. Measure accuracy on held-out pairs, the same implicit-reward comparison on data the model never trained on. Run a head-to-head win rate against the SFT model on fresh prompts, judged by the same rubric or judge that scored the pairs, swapping answer order to cancel position bias and reporting length alongside the win rate. And run your existing task regression suite, since preference tuning can quietly cost accuracy on capabilities the pairs did not cover. Ship only if the win rate improves without regressions, then consider another round: sample from the new model, score, pair and train again.
Failure modes
- Skipping SFT. DPO on a base model learns format, not preference.
- Off-policy pairs. Answers from a much stronger model give a nice margin and little change in what your model generates.
- Length hacking. Verbosity rises because chosen answers were longer.
- Reference mismatch. A reference that is not the starting checkpoint makes every reward meaningless from step one.
- Truncation. Long chosen answers cut at max_length teach the wrong lesson.
- Too many epochs. One to three passes are typical; more mostly memorises pairs.
- Judge leakage. Evaluating with the exact judge used to label pairs rewards the judge's quirks; spot-check with humans.
What to do next
- Start from a solid SFT checkpoint and record its scores on your task suite as the baseline.
- Write a scoring rubric or verifier for your task before generating any pairs.
- Sample four or more answers per prompt from the SFT model, score them, and keep clear best-versus-worst pairs.
- Check length ratios and truncation, then hold out 5 to 10 percent of pairs.
- Train one epoch with sigmoid loss and beta 0.1 using LoRA, and confirm the first loss is near 0.693.
- Watch margins, accuracies and both log-probabilities; act on the patterns in the table.
- Evaluate held-out pair accuracy, order-swapped win rate against SFT and your regression suite before shipping.