Direct Preference Optimization turns preference pairs into a classification loss, and most teams can get a DPO run to train in an afternoon. What takes longer is getting one that is better: a checkpoint that wins against the starting model on fresh prompts without losing anything else. Most of the gap comes down to one hyperparameter people set by habit, beta, and to choosing a checkpoint by training metrics rather than by a measurement that can tell two models apart.
This article treats a DPO fine-tune as a controlled experiment on a fixed GPU budget. It derives what beta controls from the objective DPO solves, shows the gradient weight that explains why beta and learning rate must be tuned together, lays out a sweep that reuses one reference computation across runs, shows how to measure drift from the reference on sampled outputs, and works through how many evaluation prompts you need before a win rate means anything. Per-step memory and compute are covered in DPO training on GPU and the end-to-end workflow and metric reading in the DPO practitioner guide; this piece is about deciding.
What beta controls
DPO starts from the objective that RLHF optimises: maximise expected reward while paying beta times the KL divergence from a reference policy. That objective has a closed-form optimum: the reference distribution reweighted by exp(reward / beta) and renormalised. DPO's insight is to invert that relation, so the policy's own log-ratio against the reference acts as an implicit reward, and fit it to preference pairs with a logistic loss. The beta in the DPO loss is that same KL price.
A four-answer toy makes the effect concrete. The reference puts probabilities 0.40, 0.30, 0.20 and 0.10 on answers whose true rewards are 0, 1, 0.5 and -1. Computing the optimum for several betas gives:
| beta | Optimal policy | KL from reference | Expected reward |
|---|---|---|---|
| reference | 0.40, 0.30, 0.20, 0.10 | 0 | 0.30 |
| 3.0 | 0.355, 0.372, 0.210, 0.064 | 0.018 | 0.41 |
| 1.0 | 0.253, 0.515, 0.208, 0.023 | 0.138 | 0.60 |
| 0.3 | 0.041, 0.852, 0.107, 0.000 | 0.727 | 0.91 |
| 0.1 | 0.000, 0.995, 0.004, 0.000 | 1.176 | 1.00 |
Low beta is greedy: in the toy it collapses onto one answer. In a language model the rewards are not known; they are inferred from a finite set of pairs, so a low beta lets the policy move far on the strength of noisy, narrow evidence. High beta keeps the model close to the SFT checkpoint and learns less. TRL's default of 0.1 is a reasonable starting point, not a recommendation for your data.
The gradient, and why beta and learning rate move together
Write the log-ratio margin for a pair as delta = [log pi(y+) - log ref(y+)] - [log pi(y-) - log ref(y-)], summed over the completion tokens. The loss is -log sigmoid(beta * delta) and its gradient is
grad L = -beta * sigmoid(-beta * delta) * (grad log pi(y+) - grad log pi(y-))Beta appears twice. It scales the gradient directly, and it sets where the weight sigmoid(-beta * delta) dies off. The table shows the weight for two betas.
| log-ratio margin delta | weight at beta 0.1 | weight at beta 0.5 |
|---|---|---|
| -10 | 0.73 | 0.99 |
| 0 | 0.50 | 0.50 |
| 5 | 0.38 | 0.08 |
| 10 | 0.27 | 0.007 |
| 20 | 0.12 | 0.000 |
| 40 | 0.018 | 0.000 |
At beta 0.5 a pair stops contributing once its margin passes about 10 nats; at beta 0.1 it keeps pulling until roughly 40. A low beta therefore keeps pushing log-ratios much further: the margin at which pairs go quiet scales roughly with 1 / beta. The beta prefactor on the gradient matters less than it looks, because Adam normalises gradient scale away. The learning rate sets how fast the log-ratios travel, beta sets where they stop, and the drift you get after a fixed number of steps depends on both, so a beta change without a learning-rate check is two confounded experiments. Tune them as a grid. Note also that the margin is a sum over tokens, so long completions can reach large margins cheaply, one reason length effects interact with beta.
A sweep on a fixed GPU budget
On a fixed budget the most useful structure is a small grid run with LoRA adapters on a frozen SFT checkpoint, which keeps each run cheap and lets several share a GPU node. A practical first grid is three betas times three learning rates. For full fine-tuning TRL's default learning rate is 1e-6; its documentation suggests around 1e-5 for adapters. Bracket those rather than trusting them.
The reference model is the same SFT checkpoint for every run, so its log-probabilities on the training pairs are the same too. TRL can compute them up front with precompute_ref_log_probs=True, which frees the memory of a second model during training; within a sweep, each run still pays that pass unless you cache it yourself and feed it to a custom loop. Either way it is one forward pass over the pairs per run, small next to training. The configuration below uses only fields documented for TRL v1.14; check your installed version before copying it.
from itertools import product
import torch
from peft import LoraConfig
from trl import DPOConfig, DPOTrainer
GRID = list(product([0.05, 0.1, 0.3], [2e-6, 5e-6, 1e-5])) # (beta, lr)
for beta, lr in GRID:
args = DPOConfig(
output_dir=f"runs/dpo_b{beta}_lr{lr}",
model_init_kwargs={"dtype": torch.bfloat16}, # string models load in fp32 otherwise
beta=beta,
learning_rate=lr,
loss_type=["sigmoid"],
num_train_epochs=1,
per_device_train_batch_size=4,
gradient_accumulation_steps=8,
max_length=1024,
precompute_ref_log_probs=True,
eval_strategy="steps", eval_steps=100,
save_strategy="steps", save_steps=100,
seed=0,
)
trainer = DPOTrainer(
model="checkpoints/sft", # also the implicit reference
args=args,
train_dataset=train_pairs,
eval_dataset=val_pairs,
peft_config=LoraConfig(r=16, lora_alpha=32, target_modules="all-linear"),
)
trainer.train()Keep everything else fixed: data order seed, pair set, max length and evaluation prompts. Change one axis at a time after the first grid, for example adding an SFT term with loss_type=["sigmoid", "sft"] and loss_weights to the best cell, rather than widening the grid in every direction.
Measuring drift on generated outputs
The logged reward margins are measured on the training or validation pairs, which are texts the model did not write. They cannot tell you how far the model's own generations moved. For that, sample completions from the policy on held-out prompts and score the same tokens under both models; the mean of the summed log-ratio is a Monte Carlo estimate of the sequence-level KL from policy to reference.
import torch
@torch.no_grad()
def seq_logprob(model, input_ids, attn, prompt_len):
logits = model(input_ids=input_ids, attention_mask=attn).logits[:, :-1]
lp = torch.log_softmax(logits.float(), -1)
tok = lp.gather(-1, input_ids[:, 1:, None]).squeeze(-1)
mask = attn[:, 1:].clone()
mask[:, : prompt_len - 1] = 0 # score completion tokens only
return (tok * mask).sum(-1)
@torch.no_grad()
def kl_estimate(policy, ref, tok, prompts, max_new=256):
vals = []
for p in prompts:
enc = tok(p, return_tensors="pt").to(policy.device)
out = policy.generate(**enc, do_sample=True, temperature=1.0, top_p=1.0,
max_new_tokens=max_new) # also disable top-k; see below
attn = torch.ones_like(out)
n = enc.input_ids.shape[1]
vals.append((seq_logprob(policy, out, attn, n) - seq_logprob(ref, out, attn, n)).item())
return sum(vals) / len(vals)Sample at temperature 1 with no truncation so the estimate matches the distribution the KL is defined on. Many models ship a generation_config that sets top-k or top-p; override both (top-p of 1.0, and top-k disabled in the way your transformers version documents), or the estimate is biased. With LoRA the reference is the same weights with the adapter disabled, so PEFT's disable_adapter() context manager gives you the reference without a second copy. Track this number per checkpoint alongside mean output length. A checkpoint whose KL is several times its neighbours' for the same win rate has moved somewhere the pairs did not ask it to go.
How many prompts a win rate needs
Final selection uses a head-to-head win rate against the SFT model on prompts neither model trained on, judged by humans or by a judge model with both answer orders tried. The question is how many prompts you need. A win rate is a proportion, and its 95 percent interval half-width near 50 percent is about 0.98 / sqrt(n). Wilson intervals for an observed 55 percent win rate:
| Prompts | Wilson 95% interval at 55% observed | Half-width near 50% |
|---|---|---|
| 100 | 45.2% to 64.4% | 9.8 points |
| 300 | 49.3% to 60.5% | 5.7 points |
| 500 | 50.6% to 59.3% | 4.4 points |
| 1,000 | 51.9% to 58.1% | 3.1 points |
| 2,000 | 52.8% to 57.2% | 2.2 points |
import math
def wilson(wins, n, z=1.96):
p = wins / n
d = 1 + z * z / n
centre = (p + z * z / (2 * n)) / d
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
return centre - half, centre + half
# ties: count as half a win, or drop them and report the tie rate separatelyWith 100 prompts a 55 percent win is indistinguishable from a coin flip; you need roughly 500 before it clears 50 percent. When comparing two DPO checkpoints with each other rather than with the baseline, the differences are smaller and the required count larger. Use the same prompts for every candidate, so comparisons are paired, and budget generation and judging time up front; it can easily cost more GPU hours than the LoRA training runs themselves.
Worked example: choosing among nine runs
A team fine-tunes an 8B assistant on 20,000 pairs with a three-by-three grid, LoRA rank 16, one epoch, checkpoints every 100 steps. Validation pair accuracy alone would pick beta 0.05 at the highest learning rate, which separates the pairs best. That cell also shows the largest sampled KL and the longest outputs, so they carry the top three cells by a combined filter (validation accuracy above the median, KL below the grid's 75th percentile) into the expensive stage.
On 600 shared prompts with order-swapped judging, the three survivors score 57, 55 and 53 percent against SFT. The 57 percent run (beta 0.05, middle learning rate) has an interval of about 53 to 61 percent; the 55 percent run (beta 0.1) about 51 to 59; the 53 percent run's interval includes 50 and it is dropped. Both remaining runs pass the regression suite and their intervals overlap, so the team ships the beta 0.1 run for its lower KL and shorter outputs, and records why. These figures are illustrative of the decision process, not results from a specific model; the intervals follow from the formula above.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Picking by train reward accuracy | Best training metrics, worst generations | Select on held-out win rate plus regressions |
| Beta changed alone | Contradictory sweep results | Grid over beta and learning rate together |
| Too few eval prompts | Rankings flip between evaluations | Size n from the interval width you need |
| Judge position bias | First answer wins too often | Judge both orders; count disagreements as ties |
| Length drift | Win rate rises with output length | Track length; try sigmoid_norm or length-balanced pairs |
| Reference mismatch | Loss far from 0.693 at step 0 | Reference must be the exact starting checkpoint |
| Stale cached reference | Silent wrong rewards after changing tokenizer or template | Invalidate the cache on any tokenisation change |
| Unseeded comparisons | Differences within run-to-run noise | Repeat the best cell with a second seed |
Trade-offs
LoRA sweeps are cheap but may not transfer exactly to full fine-tuning; if you plan to ship a full fine-tune, confirm the chosen beta with one full run. Larger grids find better cells but spend evaluation budget, which is usually the scarcer resource; successive halving, covered in hyperparameter search, lets you start many cells and stop most early on the cheap filter. Lower beta buys larger wins on the target behaviour at the price of drift on everything else, so the regression suite, not the win rate, is often what caps it. When offline pairs stop producing gains, the next step is fresher pairs sampled from the current model, as in online DPO, rather than more tuning.
What to do next
- Freeze the SFT checkpoint, split pairs into train, validation and test, and fix seeds.
- Run a three-by-three grid over beta and learning rate with LoRA and periodic checkpoints.
- Log validation pair accuracy, sampled KL and mean output length per checkpoint.
- Prune to the top few cells on that cheap filter, never on training metrics alone.
- Size the win-rate set from the interval you need, at least several hundred shared prompts, judged in both orders.
- Run the regression suite on every finalist and drop any that regresses.
- Repeat the winner with a second seed, record the decision, then consider fresher pairs.