Plain DPO trains once on a fixed set of preference pairs, usually collected from some other model. Iterative DPO turns that into a loop. Sample fresh completions from the current policy, score them, turn the scores into pairs, run a DPO pass, check the result and repeat. The Llama 3 paper describes six rounds of reward modelling, rejection sampling, SFT and DPO. Self-Rewarding Language Models ran three iterations on Llama 2 70B with the model judging its own outputs. Iterative Reasoning Preference Optimization used correct and incorrect chains of thought as pairs and reported GSM8K rising from 55.6 to 81.6 percent on Llama-2-70B-Chat.

This article is about the loop, not the single step. What one DPO step computes, where its memory goes and how to host the reference model are covered in DPO training on GPU, in depth, and the loss itself is derived in DPO math. Here the questions are what a round costs, how to build pairs, which reference to use, how to share GPUs and what goes wrong across many rounds.

Advertisement

Why a single offline round goes stale

DPO's loss only sees the pairs you give it. If those pairs were written by humans or by a different model, they come from a different distribution to what your policy actually produces. As training moves the policy away from that data, it makes mistakes no pair covers and sees more pairs whose rejected answers it would never write. Those easy pairs get a gradient weight, sigmoid of minus the margin, near zero, so the run gets less informative the longer it goes.

Iterative DPO fixes the distribution problem by sampling the training data from the policy itself, once per round. Its pairs describe the errors the current model actually makes. The cost is an inference workload back inside the training pipeline, which DPO was meant to remove from RLHF.

Iterative DPO sits between offline DPO (one dataset, one pass) and fully online DPO, where every batch is freshly sampled. Data is on-policy at the start of each round and slightly stale by its end, in exchange for large batch jobs that are easy to schedule, evaluate and roll back.

The round loop as a data flow

Each round has six phases. The prompt pool mixes new prompts with a replayed slice of earlier ones. The inference engine samples K completions per prompt from the current policy. A scorer (a reward model, an LLM judge or a verifier such as unit tests) rates every completion, and pair construction turns the K scores into zero or one pair per prompt. DPO training produces a candidate policy, and a gate decides whether that candidate becomes the next round's starting point.

Prompt poolfresh + replayed promptsGenerate K samplesinference engine, policy tScorereward model / judge / verifierBuild pairsbest vs worst, margin filterDPO trainingpolicy t -> policy t+1Reference choiceSFT or policy tRound gateheld-out evals, drift checksCheckpoint hand-offweights to inference enginepromptsK completionsscorespairsref log-probscandidatepassnext roundOne round: sample on-policy, score, pair, train, gate, redeploy.Generation and scoring are inference; only the DPO box needs gradients.A failed gate discards the candidate and the next round restarts from policy t.
The iterative DPO round. Only the training box needs gradients; everything else is batch inference, and the round gate is what stops a bad round from compounding.

The phases are strictly sequential within a round, because generation needs the latest weights and training needs all pairs.

Advertisement

A worked budget for one round

Take an 8B model, 20,000 prompts per round, K = 8 samples per prompt and an average completion of 400 tokens. Generation produces 20,000 x 8 x 400 = 64 million tokens. After scoring, suppose 30 percent of prompts produce only ties (every sample equally good or bad) and are dropped, leaving 14,000 pairs. With 200-token prompts, training processes 14,000 x 2 x 600 = 16.8 million tokens.

Now compare the FLOPs. Training one token costs roughly 6P for forward and backward, plus about 2P for the reference forward, so 8 x 8.03e9 x 16.8e6 comes to about 1.08e18 FLOPs. Decoding costs about 2P per generated token, so 2 x 8.03e9 x 64e6 is about 1.03e18 FLOPs. Prompt prefill adds little, because an engine sampling eight completions per prompt processes the shared prompt once. Generation costs about as many FLOPs as training.

Wall-clock time is worse, because decoding is limited by memory bandwidth, not arithmetic, and runs far below peak unless batches are very large. If an 8-GPU node sustains an aggregate 20,000 tokens per second on this model, which is an assumption you must measure yourself, then 64 million tokens take 3,200 seconds, just over 53 minutes, before any scoring. Measure all three phases before choosing K or the prompt count; in most setups the loop is not training-bound.

PhaseWork per round (this example)Bound byMain lever
Generate64M tokens decodedHBM bandwidth, KV cache sizeContinuous batching, larger batches, smaller K
Score160,000 completionsScorer size and prompt lengthSmaller reward model, verifier where possible
Train16.8M tokens, 1.08e18 FLOPsCompute, logits memoryPrecomputed reference, fused loss
GateHeld-out evalsEval suite sizeFixed, cached eval set

Building pairs from K scored samples

Pair construction decides what the model learns, and it is where most iterative runs go wrong. The common rule is best versus worst. Among K samples, pair the highest-scored as chosen with the lowest-scored as rejected. That maximises the score gap, but the worst sample is often degenerate, such as a truncated or empty answer or one in the wrong language. The model then learns the trivial lesson that it should not do that, and learns little else. Two refinements help: pair the best with a random lower-half sample, and drop pairs whose score gap is below a threshold, since a noisy scorer turns small gaps into coin flips.

  • Verifiable tasks (maths answers, code with tests) give clean binary labels. Pair a correct sample with an incorrect one, and skip prompts where all K samples agree. Iterative RPO added a negative log-likelihood term on the chosen answer, and TRL's multi-loss option (loss_type=["sigmoid", "sft"]) gives a similar effect.
  • Reward models produce continuous scores with a scale that drifts as the policy moves. Normalize within a prompt, for example with a gap in standard deviations across the K samples, rather than using a global threshold.
  • LLM judges are flexible but biased towards length, position and their own style; see LLM-as-judge calibration and bias. Randomize answer order and record the judge's version in the round's metadata.
  • Length control. Log the chosen-minus-rejected length difference each round. If it keeps growing, the scorer is rewarding verbosity and every round compounds it.

K trades generation cost against pair quality. K = 2 gives many ties, while very large K mostly finds more extreme outliers. For comparison, the Llama 3 paper reports sampling 10 to 30 outputs per prompt for rejection sampling.

Which reference each round uses

In round one the policy and the reference are both the SFT checkpoint, and the first logged loss is ln 2 = 0.693. From round two onward you must choose.

  1. Move the reference. Round t trains from policy t with policy t as reference. Each round is a fresh, small KL-constrained step, the first loss is again 0.693, and the ln 2 sanity check still works. The catch is that nothing anchors the model to SFT, so small per-round drifts accumulate over many rounds.
  2. Anchor to SFT. Train from policy t but keep SFT as reference. The implicit reward now measures total drift from SFT, which limits long-run wandering. The first loss of round t is no longer 0.693: pairs the previous rounds already learned start with positive margins, and a first loss of, say, 0.55 is expected, not a bug.
  3. Restart from SFT on accumulated data. Each round trains from SFT on the union of all rounds' pairs. It is the most stable option and the most expensive.

Whichever you choose, the reference is frozen within a round. That makes precomputing reference log-probabilities a clear win: one inference pass per round, then the trainer never loads the reference at all.

Round driver and trainer

The driver below runs one round per invocation. In production, run each phase in its own process so that the inference engine and the trainer never hold GPU memory at the same time. It uses vLLM's offline LLM.generate with SamplingParams(n=K) and writes pairs in TRL's conversational preference format:

# iterate.py <round>
import json, random, subprocess, sys
from pathlib import Path

ROUND = int(sys.argv[1])
PREV = Path(f"ckpt/round{ROUND - 1}")        # round 0 is the SFT checkpoint
OUT = Path(f"ckpt/round{ROUND}")
K, TEMP, MIN_GAP = 8, 0.8, 0.5               # samples per prompt, temperature, score gap

def phase_generate(prompts):
    from vllm import LLM, SamplingParams
    from transformers import AutoTokenizer
    tok = AutoTokenizer.from_pretrained(PREV)
    llm = LLM(model=str(PREV), dtype="bfloat16")
    texts = [tok.apply_chat_template([{"role": "user", "content": q}],
                                     tokenize=False, add_generation_prompt=True)
             for q in prompts]
    params = SamplingParams(n=K, temperature=TEMP, top_p=0.95, max_tokens=1024, seed=ROUND)
    outs = llm.generate(texts, params)
    return [[c.text for c in o.outputs] for o in outs]

def phase_score(prompts, samples):
    from scorer import score_batch            # your code; returns one float per completion
    return [score_batch(q, cands) for q, cands in zip(prompts, samples)]

def build_pairs(prompts, samples, scores):
    rng = random.Random(ROUND)
    pairs = []
    for q, cands, s in zip(prompts, samples, scores):
        order = sorted(range(len(cands)), key=lambda i: s[i])
        hi = order[-1]
        lo = rng.choice(order[:len(order) // 2])  # lower half, not the degenerate worst
        if s[hi] - s[lo] < MIN_GAP:           # a tie teaches nothing but costs a full step
            continue
        pairs.append({"prompt":   [{"role": "user", "content": q}],
                      "chosen":   [{"role": "assistant", "content": cands[hi]}],
                      "rejected": [{"role": "assistant", "content": cands[lo]}]})
    return pairs

if __name__ == "__main__":
    prompts = json.load(open(f"prompts/round{ROUND}.json"))
    samples = phase_generate(prompts)          # run as its own process in production
    scores = phase_score(prompts, samples)
    pairs = build_pairs(prompts, samples, scores)
    Path(f"pairs/round{ROUND}.jsonl").write_text("\n".join(json.dumps(x) for x in pairs))
    print(f"round {ROUND}: {len(pairs)} pairs from {len(prompts)} prompts")
    subprocess.run([sys.executable, "train_dpo.py", str(PREV), f"pairs/round{ROUND}.jsonl",
                    str(OUT)], check=True)
    subprocess.run([sys.executable, "gate.py", str(PREV), str(OUT)], check=True)

train_dpo.py is an ordinary TRL DPOTrainer run, with ckpt/round0 (SFT) as ref_model and three iterative-specific choices: one epoch per round, since the next round brings fresh data; a learning rate at or below round one's; and precompute_ref_log_probs=True, because the reference is frozen within a round. Everything else follows the single-round guidance.

Scheduling GPUs between generation and training

There are two basic layouts. In a time-shared layout, one pool of GPUs runs the inference engine, shuts it down, runs the trainer, shuts it down and repeats. It is simple, but each switch reloads weights and rebuilds the engine. In a split-pool layout, dedicated inference GPUs generate round t+1's samples while training GPUs are still finishing round t. That overlap only works if you accept one round of staleness, sampling from policy t-1 while t trains. That moves you back towards off-policy data, which is the thing iterative DPO exists to avoid.

Time-sharing is the right default for rounds measured in hours. Split pools make sense when generation dominates and rounds are short, and in that case the architecture starts to look like the RLHF pipelines on GPU with a weight-sync channel between trainer and engine. The engine's own throughput is set by its batching and KV-cache management, covered in vLLM continuous batching and KV cache.

Checkpoint hand-off is a common source of silent bugs. With LoRA, merge the adapter into the base weights before loading into the engine, or confirm that your engine version serves the adapter natively. Carry the tokenizer and chat template with every checkpoint. An engine that applies a different template from the trainer produces completions the trainer then scores under different token boundaries.

Failure modes that appear only across rounds

  • Reward hacking compounds. A scorer weakness exploited slightly in round one is exploited heavily by round four, because every round samples from a model that already leans towards the exploit. Watch for rising length, formatting tics and flattery, and keep a scorer-independent eval in the gate.
  • Diversity collapse. The K samples become near-duplicates, more prompts yield ties, the pair count falls and the rounds get cheaper and less useful. Track the tie fraction and a diversity metric per round.
  • Self-judging echo. When the policy is also the judge, as in self-rewarding setups, both drift together. Keep an external judge or human spot checks on a fixed sample.

Operational guidance and trade-offs

Treat each round as a release. Version each round's prompts, samples, pairs, config and checkpoint together. The gate should compare candidate t+1 with policy t on a fixed held-out preference set, a capability suite that the scorer does not see, and length and refusal statistics. A failing candidate is discarded and the round is re-run from policy t with a new seed or pair rule.

The trade-offs are straightforward to state. Iterative DPO keeps DPO's simple, stable loss and adds on-policy data, at the cost of an inference phase that often dominates wall-clock time. It is cheaper and more predictable than PPO-style RLHF, but less on-policy, and it depends heavily on the scorer's quality. For small models and domain tasks the pair and gate design matters more than the loop; DPO alignment for small models covers that side. Most gains usually arrive in the first two or three rounds. Stop when gated improvements shrink to noise rather than running a fixed number of rounds.

What to do next

  1. Get one offline DPO round working first, with the first loss at about 0.693 and the chosen log-prob stable.
  2. Measure generation throughput for your model and engine in tokens per second, then choose prompts per round and K so that one round fits your wall-clock budget.
  3. Write the scorer and pair builder, with a minimum score gap, tie skipping and chosen-versus-rejected length logging.
  4. Decide the reference policy (move it, anchor to SFT or restart) and record each round's expected starting loss.
  5. Build the round gate: a held-out preference set, a scorer-independent capability suite and length and diversity metrics. Discard failing candidates.
  6. Run three rounds in a time-shared layout, and consider split pools only if generation clearly dominates.
Key takeaway: Iterative DPO turns a one-shot preference pass into a loop. Each round samples K completions from the current policy, scores them, builds pairs from the spread and runs a short DPO pass, so the training data keeps tracking the model's own mistakes. The price is an inference phase that costs about as many FLOPs as training and usually more wall-clock time. Budget rounds by measured generation throughput, filter ties and degenerate pairs, choose deliberately between a moving and an anchored reference, and put a round gate with scorer-independent evals between rounds so one bad round can't compound into the next.