ORPO (odds ratio preference optimization) folds supervised fine-tuning and preference alignment into a single training stage with no reference model. That makes it look like the cheapest way to get a model that both follows your format and prefers your better answers. In practice, most ORPO projects succeed or fail on things the loss does not see: how the preference pairs were built, what the chat template and truncation did to them, how much memory two sequences per example really costs, and whether anyone checked the result against the model it is replacing.
This page is the practitioner's recipe: decide whether ORPO fits, build pairs, size the job, run it with Hugging Face TRL, sweep the one hyperparameter that matters most, and promote the adapter only through explicit gates. The derivation, gradient analysis and per-step GPU cost are covered in ORPO, in depth: the odds-ratio objective; here the math gets one paragraph.
The objective in one paragraph
For a prompt x, ORPO scores a response y by its length-normalised likelihood: the exponential of the mean per-token log-probability. The odds of y are that probability divided by one minus it. The loss is the ordinary negative log-likelihood on the chosen response plus lambda times minus the log-sigmoid of the log odds ratio between chosen and rejected. The first term teaches the domain and the format; the second pushes the rejected style down relative to the chosen one. Because both terms use only the policy's own probabilities, there is no frozen reference model to load or to precompute, which is the practical difference from DPO. In TRL the weight lambda is called beta and defaults to 0.1.
When ORPO is the right tool
ORPO is a good fit when you have a base or lightly tuned model, a few thousand to a few hundred thousand prompts with a better and a worse answer each, and you would otherwise run SFT followed by DPO. One stage means one set of hyperparameters, one checkpoint lineage and roughly half the pipeline to maintain.
It is a weaker fit in three cases. If your model is already a strong instruction-tuned release, the NLL term keeps pulling it toward the chosen texts, which may be worse written than what it already produces; DPO or GRPO on top of the existing model usually disturbs less. If your signal is a verifiable reward (unit tests pass, the answer matches) rather than a pairwise judgment, an online RL method uses it more directly. And if you have only demonstrations, with no rejected answers, plain SFT is the honest choice: inventing rejected answers by corrupting good ones teaches the model to avoid corruption, not to avoid your real failure modes.
Building preference pairs
The rejected answer should be a mistake your model actually makes. The most reliable way to get such answers is to sample several completions per prompt from the model you are about to train, score them against a written rubric, and pair the best with a clearly worse one. Pairs taken from a public dataset generated by other models teach the difference between those models' styles, which may have little to do with your traffic.
Four filters matter more than volume. Keep a pair only if the judge's score gap is large, because near-ties are mostly judge noise. Deduplicate prompts by normalised text so a few popular questions do not dominate. Watch the length ratio: if chosen answers are systematically longer, the model learns length rather than quality. And hold out a prompt-disjoint evaluation split before any training run sees the data.
import hashlib, json, random, re
def norm(prompt: str) -> str:
return re.sub(r"\s+", " ", prompt.strip().lower())
def build_pairs(samples, min_gap=2.0, max_len_ratio=1.8, seed=0):
"""samples: {"prompt", "completions": [{"text", "score"}]} with rubric scores 1-10."""
rng, seen, pairs = random.Random(seed), set(), []
for s in samples:
key = hashlib.sha1(norm(s["prompt"]).encode()).hexdigest()
if key in seen:
continue # one pair per distinct prompt
ranked = sorted(s["completions"], key=lambda c: c["score"], reverse=True)
best, worst = ranked[0], ranked[-1]
if best["score"] - worst["score"] < min_gap:
continue # near-tie: judge noise, not signal
lb, lw = len(best["text"]), len(worst["text"])
if max(lb, lw) / max(1, min(lb, lw)) > max_len_ratio:
continue # length would explain the preference
seen.add(key)
pairs.append({"prompt": [{"role": "user", "content": s["prompt"]}],
"chosen": [{"role": "assistant", "content": best["text"]}],
"rejected": [{"role": "assistant", "content": worst["text"]}]})
rng.shuffle(pairs)
cut = max(1, len(pairs) // 20) # 5% prompt-disjoint eval split
return pairs[cut:], pairs[:cut]
train, held_out = build_pairs(json.load(open("scored_samples.json")))
The pipeline
Templates, masking and truncation
The pairs above use the conversational format with an explicit prompt, which TRL's trainer accepts and renders through the tokenizer's chat template. Three details decide whether the tokens the loss sees are the tokens you meant.
First, the template. The chat template stored with the tokenizer is the one the model will be served with; if your serving stack uses a different template, the model is trained on one format and queried in another. Pin the tokenizer revision and render one example by hand before training. Second, masking: the NLL term should cover only the chosen completion, not the prompt, and the odds-ratio term compares completion likelihoods. Confirm on a decoded batch that prompt positions carry the ignore label. Third, truncation: max_length (default 1024) caps prompt plus completion. When a long pair is cut, the end of the answer, often including the end-of-turn token, disappears, and the model is trained on answers that never stop. Measure the token length distribution first and either raise the cap or drop the longest few per cent of pairs.
Memory and batch planning
ORPO saves the reference model, but every example still has two sequences: TRL concatenates chosen and rejected into one forward pass. Activation memory therefore behaves like an SFT batch of twice the size at the pair's longest length. Plan with this rule of thumb, and then measure one step, because kernels, checkpointing and sequence padding move the numbers.
| Component | Full fine-tune (bf16, AdamW) | LoRA on bf16 base |
|---|---|---|
| Weights | 2 bytes x params | 2 bytes x params (frozen) |
| Gradients | 2 bytes x params | adapter params only |
| Optimizer state | about 12 bytes x params (fp32 master + two moments) | adapter params only |
| Activations | grows with 2 x batch x length | same, 2 x batch x length |
| Reference model | none (DPO would add one) | none |
For a 7B model, full fine-tuning needs on the order of 16 bytes per parameter before activations, about 112 GB, so it is a multi-GPU job with sharded optimizer state. LoRA keeps the frozen base at about 14 GB and makes a single 80 GB card workable for moderate lengths, with gradient checkpointing on (TRL's ORPO default). If memory is short, reduce the per-device batch and raise gradient accumulation rather than shortening max_length below the length your answers need. LoRA, in depth covers rank and target choices.
The training script
The script below is the whole training stage. Note what is absent: no reference model, no precomputed log-probabilities, no separate SFT checkpoint. Launch it with accelerate launch train_orpo.py on one or more GPUs. The trainer and its config are under trl.experimental, which means their arguments can change between TRL releases; pin the TRL version in the run manifest.
# pip install trl peft -- ORPO lives under trl.experimental in TRL 1.x
import torch
from datasets import Dataset
from peft import LoraConfig
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl.experimental.orpo import ORPOConfig, ORPOTrainer
BASE = "your-org/base-7b" # pin a revision in real runs
tok = AutoTokenizer.from_pretrained(BASE)
model = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16)
args = ORPOConfig(
output_dir="runs/orpo-l0.1-lr5e-6",
beta=0.1, # lambda in the paper
learning_rate=5e-6, # TRL default is 1e-6; sweep it
lr_scheduler_type="cosine",
warmup_steps=50,
num_train_epochs=1,
per_device_train_batch_size=2,
gradient_accumulation_steps=16,
max_length=2048, # prompt + completion; measure first
gradient_checkpointing=True,
bf16=True,
logging_steps=10,
eval_strategy="steps",
eval_steps=200,
save_steps=200,
report_to="none",
)
trainer = ORPOTrainer(
model=model,
args=args,
processing_class=tok,
train_dataset=Dataset.from_list(train),
eval_dataset=Dataset.from_list(held_out),
peft_config=LoraConfig(r=16, lora_alpha=32, lora_dropout=0.0,
target_modules="all-linear", task_type="CAUSAL_LM"),
)
trainer.train()
trainer.save_model()
Sweeping lambda and learning rate
Lambda sets how hard the model is pushed away from rejected answers relative to how hard it imitates chosen ones. Too small and the run is SFT on the chosen half. Too large and the model learns to make both responses unlikely, which shows up as the NLL term rising and outputs getting short, hedged or repetitive. The learning rate interacts with it, so sweep the two together on a small grid, one epoch each, and compare on held-out generations, not on training loss.
| Signal (TRL metric) | Healthy run | Lambda or LR too high |
|---|---|---|
nll_loss | falls, then flattens | falls, then climbs back up |
rewards/accuracies | rises well above 0.5 on eval | near 1.0 on train, flat on eval |
log_odds_chosen | grows steadily | grows fast while generations degrade |
rewards/margins | widens | widens because chosen and rejected both fall |
A practical grid is lambda in {0.05, 0.1, 0.25, 0.5} by learning rate in two or three values around the TRL default scaled for LoRA (adapters usually tolerate larger rates than full fine-tuning). Eight to twelve one-epoch runs on a fixed seed are cheap next to a bad release.
Promotion gates
Promotion needs two kinds of evidence. A pairwise comparison against the incumbent on held-out prompts, with the judge run in both orders to cancel position bias, shows whether the new adapter is preferred. A fixed regression suite shows it did not break something the pairs never covered: refusals that must stay, output formats downstream parsers rely on, stop behaviour, and a few general-capability probes. A model can win the pairwise vote and still fail a regression; both gates must pass.
def promote(candidate, incumbent, judge, prompts, regressions, min_win=0.55):
"""Return (ok, report). candidate/incumbent: callables prompt -> text."""
wins = ties = 0
for pr in prompts: # held-out, prompt-disjoint
a, b = candidate(pr), incumbent(pr)
verdict = judge(pr, a, b, swap_order=True) # judge both orders, average
wins += verdict == "A"
ties += verdict == "tie"
win_rate = (wins + 0.5 * ties) / len(prompts)
failed = [t["name"] for t in regressions if not t["check"](candidate(t["prompt"]))]
report = {"win_rate": round(win_rate, 3), "regressions": failed}
return win_rate >= min_win and not failed, report
Worked example: a support assistant
A support team has a 7B model that answers product questions. Complaints cluster on two patterns: inventing configuration flags and burying the answer under boilerplate. They take 9,000 recent prompts, sample four answers each from the current model, and have an LLM judge score each against a rubric that names both patterns, with humans spot-checking 200 pairs. After the gap, dedup and length filters, 5,800 pairs remain; 290 are held out.
The token histogram shows 4 per cent of pairs above 2,048 tokens; they are dropped rather than truncated. A LoRA run on one 80 GB GPU with batch 2 and accumulation 16 fits with checkpointing on. The sweep's lambda 0.5 run has the best training margins but its NLL climbs after 300 steps and its answers start to omit steps, so it is rejected. Lambda 0.1 at the middle learning rate wins 61 per cent of held-out comparisons and passes every regression except one: it stopped emitting the JSON summary block a dashboard parses. Twenty pairs whose chosen answers include that block are added, the run is repeated, and the second candidate passes both gates. These figures are illustrative; your own will differ, but the shape of the process will not.
Failure modes
- Chosen answers worse than the model's own. The NLL term imitates whatever is labelled chosen. If pairs come from a weaker model, quality drops even as margins grow.
- Length learning. Systematically longer chosen answers produce a verbose model. Filter on length ratio and report mean output length at every gate.
- Truncated completions. Pairs cut at
max_lengthlose their end-of-turn token, and the model learns not to stop. Drop over-long pairs. - Template mismatch. Training and serving templates differ, and quality looks fine in the notebook and poor in production.
- Over-pushing. High lambda or learning rate makes both responses unlikely; NLL rises and generations collapse into short or repetitive text.
- Judge leakage. Evaluating with the same judge and rubric that labelled the pairs rewards agreement with the judge. Use a second judge or human review for promotion.
- Evaluation contamination. Prompts near-duplicated between train and eval inflate win rates; split by normalised prompt, not by row.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| ORPO vs SFT then DPO | one stage, no reference model, simpler lineage | cannot tune imitation and preference separately; NLL can pull a strong model toward weaker texts |
| LoRA vs full fine-tune | single-GPU runs, cheap sweeps, swappable adapters | lower ceiling on large behavioural shifts |
| LLM judge vs human labels | scale and speed | judge bias becomes model bias; needs spot checks |
| Own-model samples vs public pairs | targets your real failure modes | costs generation and judging compute |
| Larger lambda | stronger preference signal | risk of likelihood collapse |
For the operational side of running these jobs repeatedly (queues, checkpoints, cost tracking), see fine-tuning operations.
What to do next
- Write the rubric that defines a better answer for your product, naming the failure patterns.
- Sample several answers per prompt from the model you will train and score them.
- Filter pairs on score gap, prompt duplicates and length ratio; hold out a prompt-disjoint split.
- Render one example through the chat template and check masking and the end-of-turn token.
- Measure token lengths; set
max_lengthto cover your answers and drop outliers. - Run a LoRA job with TRL's experimental ORPO trainer and pin the TRL version.
- Sweep lambda and learning rate together; reject runs where NLL climbs back.
- Promote only when the pairwise win rate and the regression suite both pass.