The usual recipe for aligning a language model to preferences has two stages. Supervised fine-tuning (SFT) teaches the model to produce responses in the right format and style. A preference stage, such as DPO, then teaches it to prefer better responses over worse ones, and DPO needs a frozen copy of the SFT model as a reference. That means two training runs, and during the second, two models in memory or a precomputation pass.

ORPO, Odds Ratio Preference Optimization (Hong, Lee and Thorne, 2024), merges the two stages into one. It trains on preference pairs from the base model, adding a small odds-ratio penalty to the ordinary SFT loss. There is no reference model at all. This article derives the objective, works through the numbers, implements it, and explains what that buys you on a GPU and where it goes wrong.

Advertisement

Why a single stage can work

The observation behind ORPO is about what SFT does to rejected responses. When you fine-tune on good responses, the probability of similar but bad responses also rises, because they share most of their tokens and style with the good ones. SFT teaches the domain and the format, but nothing in its loss says this response is worse than that one. Preference methods add that contrast afterwards.

ORPO adds the contrast during SFT. Every training example is a triple: a prompt, a chosen response and a rejected response. The model learns the chosen response with the usual negative log-likelihood, and at the same time a second term pushes the chosen response's odds above the rejected one's. The authors argue that a mild penalty on the disfavoured style is enough, and that the odds ratio provides one that is mild enough to coexist with SFT.

Two-stage SFT + DPO versus single-stage ORPO: models resident and forward passes per preference pairSFT then DPOBase modelpretrainedSFT runNLL on chosenDPO runpolicy + frozen refAligned modelcopy = refper pair: 2 policy passes with grad + 2 reference passesORPOBase modelpretrainedOne ORPO runNLL(chosen) + λ · odds-ratio term, no referenceAligned modelper pair: 2 policy passes with grad (chosen and rejected, usually one concatenated batch)Policy computes, per response y:P(y|x) = exp(mean token log-prob)odds(y|x) = P / (1 - P)log OR = log odds(chosen) - log odds(rejected)Loss per pair:L = NLL(chosen) - λ · log σ(log OR)first term: learn the chosen stylesecond term: push rejected below chosen
ORPO removes the separate SFT stage and the frozen reference model; each step runs the policy on the chosen and rejected responses only.

The objective from first principles

Start with the probability of a whole response. A product of token probabilities shrinks with length, so ORPO uses the length-normalised version: the exponential of the mean per-token log-probability over the response tokens. Call it P(y|x). It is always between 0 and 1.

The odds of an event with probability P are P / (1 - P). Odds run from 0 to infinity and grow sharply as P approaches 1. The odds ratio between the chosen response yw and the rejected response yl is odds(yw|x) / odds(yl|x). Taking its log and passing it through a log-sigmoid gives a loss that is small when chosen is much more likely than rejected:

P(y|x)   = exp( (1/|y|) * sum_t log p(y_t | x, y_<t) )
odds(y|x) = P(y|x) / (1 - P(y|x))
log_OR   = log odds(y_w|x) - log odds(y_l|x)
         = (logP_w - logP_l) - (log(1 - P_w) - log(1 - P_l))
L_OR     = -log sigmoid(log_OR)
L_ORPO   = L_SFT(y_w) + lambda * L_OR

LSFT is the usual cross-entropy on the chosen response's tokens, with prompt tokens masked out. λ sets how hard the preference term pushes. In TRL it is called beta and defaults to 0.1. The second line of log_OR is how implementations compute it: from mean log-probabilities, using a numerically stable log(1 - exp(x)).

Advertisement

What the gradient does

Differentiating LOR gives a product of two factors. The first is σ(−log OR), a weight that is close to 1 when the model still prefers the rejected response and fades towards 0 as the preference is learned. Pairs the model already gets right stop contributing, as in DPO.

The second factor is the direction: the gradient of log Pw scaled by 1 / (1 − Pw), minus the gradient of log Pl scaled by 1 / (1 − Pl). The 1 / (1 − P) scaling is the odds at work. When the model assigns the rejected response a high probability, its gradient is amplified and the push against it is strong. When the rejected response is already unlikely, the push is weak. Meanwhile the SFT term keeps pulling the chosen response up, which prevents the degenerate solution of lowering both.

Worked example: the numbers for one pair

Suppose the model's mean token log-probability is −0.9 on the chosen response and −1.1 on the rejected one. This snippet computes the terms exactly as TRL does, with λ = 0.1:

import math

def log1mexp(x):
    # log(1 - exp(x)) for x < 0, stable
    return math.log(-math.expm1(x)) if x > -0.693 else math.log1p(-math.exp(x))

def orpo_terms(chosen_mean_logp, rejected_mean_logp, beta=0.1):
    log_odds = (chosen_mean_logp - rejected_mean_logp) - (
        log1mexp(chosen_mean_logp) - log1mexp(rejected_mean_logp))
    log_sig = -math.log1p(math.exp(-log_odds))
    nll = -chosen_mean_logp
    return log_odds, -log_sig, nll + beta * (-log_sig)

for c, r in [(-0.9, -1.1), (-0.9, -0.9), (-0.9, -1.6), (-0.3, -0.35), (-0.05, -0.06)]:
    lo, lor, total = orpo_terms(c, r)
    print(f"chosen {c:6.2f} rejected {r:6.2f} | log_odds {lo:7.4f} | L_OR {lor:.4f} | total {total:.4f}")
chosen  -0.90 rejected  -1.10 | log_odds  0.3171 | L_OR 0.5471 | total 0.9547
chosen  -0.90 rejected  -0.90 | log_odds  0.0000 | L_OR 0.6931 | total 0.9693
chosen  -0.90 rejected  -1.60 | log_odds  0.9963 | L_OR 0.3143 | total 0.9314
chosen  -0.30 rejected  -0.35 | log_odds  0.1805 | L_OR 0.6070 | total 0.3607
chosen  -0.05 rejected  -0.06 | log_odds  0.1874 | L_OR 0.6038 | total 0.1104

Three things stand out. When the two responses are equally likely, LOR is log 2, about 0.693. Widening the gap from 0.2 to 0.7 nats more than doubles the log odds. And in the last two rows, a gap of only 0.01 near certainty produces more log odds than a gap of 0.05 further away. Odds magnify differences between responses the model finds likely, which is where style preferences live. Notice also that the SFT term dominates the total: with λ = 0.1, the preference term is a correction, not the main objective.

The loss in PyTorch

In practice chosen and rejected sequences are padded to one length and concatenated into a single batch, so one forward pass serves both. Labels mask prompt and padding positions with −100.

import torch
import torch.nn.functional as F

def mean_logps(logits, labels):
    # logits [B, T, V]; labels [B, T] with -100 on prompt and padding
    logits, labels = logits[:, :-1], labels[:, 1:]
    mask = labels != -100
    safe = labels.masked_fill(~mask, 0)
    tok = torch.gather(logits.log_softmax(-1), 2, safe.unsqueeze(-1)).squeeze(-1)
    return (tok * mask).sum(-1) / mask.sum(-1).clamp(min=1)

def log1mexp(x):
    # log(1 - exp(x)) for x < 0
    return torch.where(x > -0.693, torch.log(-torch.expm1(x)), torch.log1p(-torch.exp(x)))

def orpo_step(model, chosen_ids, chosen_labels, rejected_ids, rejected_labels, beta=0.1):
    ids = torch.cat([chosen_ids, rejected_ids])        # same padded length
    labels = torch.cat([chosen_labels, rejected_labels])
    logits = model(input_ids=ids).logits.float()
    n = chosen_ids.shape[0]
    lp = mean_logps(logits, labels)
    lp_w, lp_l = lp[:n], lp[n:]
    nll = F.cross_entropy(logits[:n, :-1].reshape(-1, logits.shape[-1]),
                          chosen_labels[:, 1:].reshape(-1), ignore_index=-100)
    log_odds = (lp_w - lp_l) - (log1mexp(lp_w) - log1mexp(lp_l))
    loss = nll - beta * F.logsigmoid(log_odds).mean()
    return loss, {"nll": nll.item(), "log_odds": log_odds.mean().item(),
                  "acc": (lp_w > lp_l).float().mean().item()}

Two details matter. The NLL is the token-averaged cross-entropy over the chosen batch, matching TRL, which computes it on the chosen half of the concatenated logits. Casting logits to float32 before the log-softmax avoids precision loss in bf16, at the cost of a large temporary tensor; at long context and large vocabulary that tensor is often the peak of activation memory.

What a step costs on the GPU

The win is in what is not there. DPO runs four forward passes per pair: policy on chosen and rejected, with gradients, and reference on both, without. The reference model's weights also sit in memory, unless you precompute its log-probabilities in a separate pass. The DPO on GPU article breaks that bill down. ORPO runs two forward passes per pair and holds one model.

Per step, 8B model, bf16 weightsSFTSFT then DPO (DPO stage)ORPO
Model copies in memory12 (policy + frozen reference)1
Extra weight memorynoneabout 16 GB for the referencenone
Forward passes per example14 per pair (2 without grad)2 per pair
Backward passes12 per pair2 per pair
Training runs end to end121

Two caveats keep this honest. First, ORPO is not cheaper than SFT: a pair costs roughly twice an SFT example because both sequences are processed with gradients. Second, the optimizer dominates memory for full fine-tuning. Mixed-precision AdamW keeps around 16 bytes per parameter for weights, gradients, master weights and two moments, about 128 GB for 8B parameters, so sharding with FSDP or using LoRA is still required. Activation memory doubles relative to SFT at the same number of examples, so halve the per-device pairs or turn on gradient checkpointing; TRL enables checkpointing by default for ORPO.

Training with TRL

TRL ships an ORPOTrainer. On current main it lives under trl.experimental.orpo; older releases exported it from the top-level trl package, so check the version you have installed. The dataset needs prompt, chosen and rejected columns, in plain or conversational format.

from datasets import load_dataset
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl.experimental.orpo import ORPOConfig, ORPOTrainer

name = "Qwen/Qwen2-0.5B-Instruct"
model = AutoModelForCausalLM.from_pretrained(name)
tok = AutoTokenizer.from_pretrained(name)
data = load_dataset("trl-lib/ultrafeedback_binarized", split="train")

args = ORPOConfig(
    output_dir="qwen-orpo",
    beta=0.1,                        # lambda in the paper
    max_length=1024,                 # prompt + completion
    learning_rate=1e-6,              # TRL's ORPO default
    per_device_train_batch_size=4,   # pairs, so 8 sequences per forward
    gradient_accumulation_steps=8,
    num_train_epochs=1,
    logging_steps=10,
)
ORPOTrainer(model=model, args=args, processing_class=tok, train_dataset=data).train()

Launch it with accelerate launch train_orpo.py for multiple GPUs. The batch-size comment matters for capacity planning: four pairs per device is eight sequences in each forward pass.

Reading the metrics

TRL logs nll_loss, log_odds_chosen (the mean log odds ratio), log_odds_ratio (the mean log-sigmoid of it), and rewards/chosen, rewards/rejected, rewards/accuracies and rewards/margins, where rewards are the mean log-probabilities scaled by beta. A healthy run shows nll_loss falling like an SFT run, log_odds_chosen rising from about zero, and accuracy climbing above one half. If rewards/chosen falls steadily alongside rewards/rejected, the preference term is winning by lowering everything; reduce beta or the learning rate. If accuracy stays near one half while NLL falls, the run is behaving like plain SFT; check that chosen and rejected actually differ after truncation.

Failure modes

  • Truncation erases the signal. Chosen and rejected often share a long prefix and differ near the end. If max_length cuts both before they diverge, the pair contributes zero preference signal. Measure how many pairs survive truncation with distinct completions.
  • Chosen responses that are not SFT quality. ORPO trains on every chosen response with full NLL weight. A chosen response that is merely less bad than the rejected one becomes a target to imitate. Filter pairs by absolute quality, not only by preference margin.
  • Chat template mismatch. Training on one template and serving with another breaks the format the NLL term learned. Apply the tokenizer's chat template in both places and make sure the end-of-turn token is in the labels.
  • Numerical edge cases. As a mean log-probability approaches 0, log(1 − P) goes to minus infinity. Use the stable log1mexp form and compute in float32.
  • Learning rate set for SFT. Rates that are comfortable for SFT can make the preference term unstable; TRL's ORPO default is 1e-6. Sweep a small range on a held-out preference set.
  • Evaluation by training loss only. Loss can fall while behaviour does not improve. Keep a held-out preference set and a small generation benchmark, and read samples.

ORPO, DPO and SFT-then-DPO compared

QuestionSFT then DPOORPO
StagesTwo runs, each tunableOne run
Reference modelRequired, or precomputed log-probabilitiesNone
DataSFT data plus preference pairsPreference pairs whose chosen side is SFT quality
ControlSeparate knobs per stage; can reuse an existing SFT modelOne beta trades imitation against contrast
Starting pointWorks well from an instruct modelDesigned to start from a base model
ComputeSFT cost plus about 2x per pair for DPOAbout 2x SFT per example

Choose ORPO when you have one preference dataset with good chosen responses, are starting from a base model, and want one run with one model in memory. Choose SFT then DPO when you already have a strong SFT model, when your SFT data and preference data come from different sources, or when you want to tune the two behaviours independently. For the wider pipeline, see the RLHF pipeline on GPUs.

What to do next

  1. Reproduce the worked example table, then change the mean log-probabilities to build intuition for how odds magnify gaps near certainty.
  2. Audit your preference data: count pairs whose completions are identical after truncation at your max_length, and drop or re-truncate them.
  3. Score chosen responses for absolute quality and remove those you would not accept as SFT targets.
  4. Run a short ORPO job from a small base model with TRL, logging nll_loss, log_odds_chosen and rewards/accuracies.
  5. Sweep beta over a few values and learning rate around 1e-6, selecting by held-out preference accuracy and a generation benchmark.
  6. Compare against an SFT-then-DPO baseline at equal GPU hours before committing.
  7. Size memory with mixed-precision accounting, remembering each pair is two sequences.
Key takeaway: ORPO folds preference alignment into supervised fine-tuning. The loss is the usual NLL on the chosen response plus λ times −log σ of the log odds ratio between chosen and rejected, where each response's probability is the exponential of its mean token log-probability. There is no reference model, so each pair costs two policy passes and memory holds one model; that is cheaper than DPO but about twice SFT per example. The odds term pushes hardest against rejected responses the model finds likely, while the NLL term keeps the chosen response rising. Success depends on chosen responses being SFT quality, on truncation preserving the difference between pairs, and on a small beta and learning rate.