The usual recipe for aligning a language model to preferences has two stages. Supervised fine-tuning (SFT) teaches the model to produce responses in the right format and style. A preference stage, such as DPO, then teaches it to prefer better responses over worse ones, and DPO needs a frozen copy of the SFT model as a reference. That means two training runs, and during the second, two models in memory or a precomputation pass.
ORPO, Odds Ratio Preference Optimization (Hong, Lee and Thorne, 2024), merges the two stages into one. It trains on preference pairs from the base model, adding a small odds-ratio penalty to the ordinary SFT loss. There is no reference model at all. This article derives the objective, works through the numbers, implements it, and explains what that buys you on a GPU and where it goes wrong.
Why a single stage can work
The observation behind ORPO is about what SFT does to rejected responses. When you fine-tune on good responses, the probability of similar but bad responses also rises, because they share most of their tokens and style with the good ones. SFT teaches the domain and the format, but nothing in its loss says this response is worse than that one. Preference methods add that contrast afterwards.
ORPO adds the contrast during SFT. Every training example is a triple: a prompt, a chosen response and a rejected response. The model learns the chosen response with the usual negative log-likelihood, and at the same time a second term pushes the chosen response's odds above the rejected one's. The authors argue that a mild penalty on the disfavoured style is enough, and that the odds ratio provides one that is mild enough to coexist with SFT.
The objective from first principles
Start with the probability of a whole response. A product of token probabilities shrinks with length, so ORPO uses the length-normalised version: the exponential of the mean per-token log-probability over the response tokens. Call it P(y|x). It is always between 0 and 1.
The odds of an event with probability P are P / (1 - P). Odds run from 0 to infinity and grow sharply as P approaches 1. The odds ratio between the chosen response yw and the rejected response yl is odds(yw|x) / odds(yl|x). Taking its log and passing it through a log-sigmoid gives a loss that is small when chosen is much more likely than rejected:
P(y|x) = exp( (1/|y|) * sum_t log p(y_t | x, y_<t) )
odds(y|x) = P(y|x) / (1 - P(y|x))
log_OR = log odds(y_w|x) - log odds(y_l|x)
= (logP_w - logP_l) - (log(1 - P_w) - log(1 - P_l))
L_OR = -log sigmoid(log_OR)
L_ORPO = L_SFT(y_w) + lambda * L_ORLSFT is the usual cross-entropy on the chosen response's tokens, with prompt tokens masked out. λ sets how hard the preference term pushes. In TRL it is called beta and defaults to 0.1. The second line of log_OR is how implementations compute it: from mean log-probabilities, using a numerically stable log(1 - exp(x)).
What the gradient does
Differentiating LOR gives a product of two factors. The first is σ(−log OR), a weight that is close to 1 when the model still prefers the rejected response and fades towards 0 as the preference is learned. Pairs the model already gets right stop contributing, as in DPO.
The second factor is the direction: the gradient of log Pw scaled by 1 / (1 − Pw), minus the gradient of log Pl scaled by 1 / (1 − Pl). The 1 / (1 − P) scaling is the odds at work. When the model assigns the rejected response a high probability, its gradient is amplified and the push against it is strong. When the rejected response is already unlikely, the push is weak. Meanwhile the SFT term keeps pulling the chosen response up, which prevents the degenerate solution of lowering both.
Worked example: the numbers for one pair
Suppose the model's mean token log-probability is −0.9 on the chosen response and −1.1 on the rejected one. This snippet computes the terms exactly as TRL does, with λ = 0.1:
import math
def log1mexp(x):
# log(1 - exp(x)) for x < 0, stable
return math.log(-math.expm1(x)) if x > -0.693 else math.log1p(-math.exp(x))
def orpo_terms(chosen_mean_logp, rejected_mean_logp, beta=0.1):
log_odds = (chosen_mean_logp - rejected_mean_logp) - (
log1mexp(chosen_mean_logp) - log1mexp(rejected_mean_logp))
log_sig = -math.log1p(math.exp(-log_odds))
nll = -chosen_mean_logp
return log_odds, -log_sig, nll + beta * (-log_sig)
for c, r in [(-0.9, -1.1), (-0.9, -0.9), (-0.9, -1.6), (-0.3, -0.35), (-0.05, -0.06)]:
lo, lor, total = orpo_terms(c, r)
print(f"chosen {c:6.2f} rejected {r:6.2f} | log_odds {lo:7.4f} | L_OR {lor:.4f} | total {total:.4f}")chosen -0.90 rejected -1.10 | log_odds 0.3171 | L_OR 0.5471 | total 0.9547
chosen -0.90 rejected -0.90 | log_odds 0.0000 | L_OR 0.6931 | total 0.9693
chosen -0.90 rejected -1.60 | log_odds 0.9963 | L_OR 0.3143 | total 0.9314
chosen -0.30 rejected -0.35 | log_odds 0.1805 | L_OR 0.6070 | total 0.3607
chosen -0.05 rejected -0.06 | log_odds 0.1874 | L_OR 0.6038 | total 0.1104Three things stand out. When the two responses are equally likely, LOR is log 2, about 0.693. Widening the gap from 0.2 to 0.7 nats more than doubles the log odds. And in the last two rows, a gap of only 0.01 near certainty produces more log odds than a gap of 0.05 further away. Odds magnify differences between responses the model finds likely, which is where style preferences live. Notice also that the SFT term dominates the total: with λ = 0.1, the preference term is a correction, not the main objective.
The loss in PyTorch
In practice chosen and rejected sequences are padded to one length and concatenated into a single batch, so one forward pass serves both. Labels mask prompt and padding positions with −100.
import torch
import torch.nn.functional as F
def mean_logps(logits, labels):
# logits [B, T, V]; labels [B, T] with -100 on prompt and padding
logits, labels = logits[:, :-1], labels[:, 1:]
mask = labels != -100
safe = labels.masked_fill(~mask, 0)
tok = torch.gather(logits.log_softmax(-1), 2, safe.unsqueeze(-1)).squeeze(-1)
return (tok * mask).sum(-1) / mask.sum(-1).clamp(min=1)
def log1mexp(x):
# log(1 - exp(x)) for x < 0
return torch.where(x > -0.693, torch.log(-torch.expm1(x)), torch.log1p(-torch.exp(x)))
def orpo_step(model, chosen_ids, chosen_labels, rejected_ids, rejected_labels, beta=0.1):
ids = torch.cat([chosen_ids, rejected_ids]) # same padded length
labels = torch.cat([chosen_labels, rejected_labels])
logits = model(input_ids=ids).logits.float()
n = chosen_ids.shape[0]
lp = mean_logps(logits, labels)
lp_w, lp_l = lp[:n], lp[n:]
nll = F.cross_entropy(logits[:n, :-1].reshape(-1, logits.shape[-1]),
chosen_labels[:, 1:].reshape(-1), ignore_index=-100)
log_odds = (lp_w - lp_l) - (log1mexp(lp_w) - log1mexp(lp_l))
loss = nll - beta * F.logsigmoid(log_odds).mean()
return loss, {"nll": nll.item(), "log_odds": log_odds.mean().item(),
"acc": (lp_w > lp_l).float().mean().item()}Two details matter. The NLL is the token-averaged cross-entropy over the chosen batch, matching TRL, which computes it on the chosen half of the concatenated logits. Casting logits to float32 before the log-softmax avoids precision loss in bf16, at the cost of a large temporary tensor; at long context and large vocabulary that tensor is often the peak of activation memory.
What a step costs on the GPU
The win is in what is not there. DPO runs four forward passes per pair: policy on chosen and rejected, with gradients, and reference on both, without. The reference model's weights also sit in memory, unless you precompute its log-probabilities in a separate pass. The DPO on GPU article breaks that bill down. ORPO runs two forward passes per pair and holds one model.
| Per step, 8B model, bf16 weights | SFT | SFT then DPO (DPO stage) | ORPO |
|---|---|---|---|
| Model copies in memory | 1 | 2 (policy + frozen reference) | 1 |
| Extra weight memory | none | about 16 GB for the reference | none |
| Forward passes per example | 1 | 4 per pair (2 without grad) | 2 per pair |
| Backward passes | 1 | 2 per pair | 2 per pair |
| Training runs end to end | 1 | 2 | 1 |
Two caveats keep this honest. First, ORPO is not cheaper than SFT: a pair costs roughly twice an SFT example because both sequences are processed with gradients. Second, the optimizer dominates memory for full fine-tuning. Mixed-precision AdamW keeps around 16 bytes per parameter for weights, gradients, master weights and two moments, about 128 GB for 8B parameters, so sharding with FSDP or using LoRA is still required. Activation memory doubles relative to SFT at the same number of examples, so halve the per-device pairs or turn on gradient checkpointing; TRL enables checkpointing by default for ORPO.
Training with TRL
TRL ships an ORPOTrainer. On current main it lives under trl.experimental.orpo; older releases exported it from the top-level trl package, so check the version you have installed. The dataset needs prompt, chosen and rejected columns, in plain or conversational format.
from datasets import load_dataset
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl.experimental.orpo import ORPOConfig, ORPOTrainer
name = "Qwen/Qwen2-0.5B-Instruct"
model = AutoModelForCausalLM.from_pretrained(name)
tok = AutoTokenizer.from_pretrained(name)
data = load_dataset("trl-lib/ultrafeedback_binarized", split="train")
args = ORPOConfig(
output_dir="qwen-orpo",
beta=0.1, # lambda in the paper
max_length=1024, # prompt + completion
learning_rate=1e-6, # TRL's ORPO default
per_device_train_batch_size=4, # pairs, so 8 sequences per forward
gradient_accumulation_steps=8,
num_train_epochs=1,
logging_steps=10,
)
ORPOTrainer(model=model, args=args, processing_class=tok, train_dataset=data).train()Launch it with accelerate launch train_orpo.py for multiple GPUs. The batch-size comment matters for capacity planning: four pairs per device is eight sequences in each forward pass.
Reading the metrics
TRL logs nll_loss, log_odds_chosen (the mean log odds ratio), log_odds_ratio (the mean log-sigmoid of it), and rewards/chosen, rewards/rejected, rewards/accuracies and rewards/margins, where rewards are the mean log-probabilities scaled by beta. A healthy run shows nll_loss falling like an SFT run, log_odds_chosen rising from about zero, and accuracy climbing above one half. If rewards/chosen falls steadily alongside rewards/rejected, the preference term is winning by lowering everything; reduce beta or the learning rate. If accuracy stays near one half while NLL falls, the run is behaving like plain SFT; check that chosen and rejected actually differ after truncation.
Failure modes
- Truncation erases the signal. Chosen and rejected often share a long prefix and differ near the end. If
max_lengthcuts both before they diverge, the pair contributes zero preference signal. Measure how many pairs survive truncation with distinct completions. - Chosen responses that are not SFT quality. ORPO trains on every chosen response with full NLL weight. A chosen response that is merely less bad than the rejected one becomes a target to imitate. Filter pairs by absolute quality, not only by preference margin.
- Chat template mismatch. Training on one template and serving with another breaks the format the NLL term learned. Apply the tokenizer's chat template in both places and make sure the end-of-turn token is in the labels.
- Numerical edge cases. As a mean log-probability approaches 0, log(1 − P) goes to minus infinity. Use the stable log1mexp form and compute in float32.
- Learning rate set for SFT. Rates that are comfortable for SFT can make the preference term unstable; TRL's ORPO default is 1e-6. Sweep a small range on a held-out preference set.
- Evaluation by training loss only. Loss can fall while behaviour does not improve. Keep a held-out preference set and a small generation benchmark, and read samples.
ORPO, DPO and SFT-then-DPO compared
| Question | SFT then DPO | ORPO |
|---|---|---|
| Stages | Two runs, each tunable | One run |
| Reference model | Required, or precomputed log-probabilities | None |
| Data | SFT data plus preference pairs | Preference pairs whose chosen side is SFT quality |
| Control | Separate knobs per stage; can reuse an existing SFT model | One beta trades imitation against contrast |
| Starting point | Works well from an instruct model | Designed to start from a base model |
| Compute | SFT cost plus about 2x per pair for DPO | About 2x SFT per example |
Choose ORPO when you have one preference dataset with good chosen responses, are starting from a base model, and want one run with one model in memory. Choose SFT then DPO when you already have a strong SFT model, when your SFT data and preference data come from different sources, or when you want to tune the two behaviours independently. For the wider pipeline, see the RLHF pipeline on GPUs.
What to do next
- Reproduce the worked example table, then change the mean log-probabilities to build intuition for how odds magnify gaps near certainty.
- Audit your preference data: count pairs whose completions are identical after truncation at your
max_length, and drop or re-truncate them. - Score chosen responses for absolute quality and remove those you would not accept as SFT targets.
- Run a short ORPO job from a small base model with TRL, logging
nll_loss,log_odds_chosenandrewards/accuracies. - Sweep beta over a few values and learning rate around 1e-6, selecting by held-out preference accuracy and a generation benchmark.
- Compare against an SFT-then-DPO baseline at equal GPU hours before committing.
- Size memory with mixed-precision accounting, remembering each pair is two sequences.