Every preference-tuned model, whether trained with a reward model and PPO or directly with DPO, learns from records that say: for this prompt, response A is better than response B. The training algorithm gets most of the attention, but the records decide what the model learns. If the chosen responses are systematically longer, the model learns to be long. If raters disagree on half the items, the model learns noise. If the candidates came from a model you no longer train, the pairs teach distinctions your policy no longer makes.
This page is about the collection operation itself: where prompts come from, how candidate responses are produced, how judgements are captured, how you know the judgements are any good, when an AI judge is acceptable, and how to release the result as a dataset you can version and audit. It assumes you know where preference data sits in the training loop; that loop is covered end to end in the RLHF pipeline, and the DPO side in DPO alignment for small models.
What a preference record is
At minimum a record is a triple: a prompt, a chosen response and a rejected response. That is what DPO and pairwise reward-model losses consume. Everything else in the record exists so you can debug, filter and reweight later, and in practice that metadata is what separates a dataset you can trust from one you cannot. Store which model and sampling settings produced each response, how long each one is, who or what judged it under which guideline version, how long they took, which order the responses were shown in, and how strong the preference was.
{
"pair_id": "r3-000412-b",
"prompt_id": "p-88213",
"prompt": "Summarise this refund policy for a customer in two sentences: ...",
"chosen": "...",
"rejected": "...",
"chosen_meta": {"model": "sft-v7", "temperature": 0.9, "tokens": 61},
"rejected_meta": {"model": "sft-v7", "temperature": 0.9, "tokens": 148},
"strength": "better", // much_better | better | slightly_better
"axis": "helpfulness",
"judge": {"type": "human", "rater_id": "r-0193", "guideline": "v4.2", "seconds": 74},
"shown_order": ["rejected", "chosen"],
"round": 3,
"source": "support_tickets_2026q3"
}The shown_order field looks fussy until you find a rater who picks whichever response is on the left. Response lengths let you measure length bias before it is baked into a model. The guideline version lets you drop or relabel everything judged under a rule you later changed. None of these can be reconstructed after the fact, so capture them at collection time.
The pipeline at a glance
Treat collection as a pipeline with stages you can rerun, not as a one-off labelling job. Each stage has its own failure modes, and most of the bugs that hurt a preference-tuned model are visible at one specific stage if you instrument it. The stages below follow the diagram.
Sourcing prompts
The prompt pool defines what the model is being taught to be good at. Pull prompts from the traffic you actually expect: real user requests with personal data removed, support tickets, internal tasks, and a deliberate slice of hard cases such as ambiguous instructions, requests that should be refused and multi-step tasks. Purely synthetic prompts are cheap and broad, but they are written in a model's voice and under-represent the messy phrasing of real users; use them to fill gaps, not as the backbone. Synthetic data generation covers how to make those gap-fillers diverse.
Deduplicate prompts by near-duplicate matching, not only exact match, and tag each prompt with a category and difficulty so you can balance the mix and report quality per slice. Hold out a set of prompts entirely, never used for collection, as the evaluation set for the reward model or the tuned policy. If evaluation prompts leak into training pairs, held-out preference accuracy will look better than the model really is.
Generating candidates
Candidates should come from the policy you are about to train, or something close to it. Preference pairs teach the model which of two of its own plausible outputs is better; pairs built from a much weaker or much stronger model teach distinctions that never arise when your policy generates text. This is why mature projects collect in rounds: train, sample from the new policy on fresh prompts, judge, train again.
Within a round, aim for candidates that differ in ways that matter. Sampling several responses at a moderately high temperature from one policy gives natural variation. Mixing in a second checkpoint, a different system prompt or a deliberately constrained variant such as a short answer widens the spread. If both candidates are nearly identical, the judgement carries little information and raters will mark it a tie or guess. If one is obviously broken, the pair is easy and teaches only what the SFT stage already taught.
Control length explicitly. Record token counts, look at the length distribution of the candidate pool before judging, and where possible include pairs where the shorter response is the better one. Length is the most common shortcut a reward model learns, and it is far cheaper to prevent in the data than to correct in the loss.
The judging interface and guidelines
There are three common designs. Pairwise choice shows two responses and asks which is better; it is fast and maps directly onto the training loss. Ranking shows K responses and asks for an order; InstructGPT used rankings of between four and nine responses, because one ranking of K items yields K times (K - 1) / 2 pairs for the cost of reading K responses. Graded pairwise adds a strength, such as much better, better or slightly better, which can be used to drop weak preferences or to weight the loss; Llama 2 collected a four-level strength scale alongside each choice.
| Design | Strength | Weakness | Use when |
|---|---|---|---|
| Pairwise | Fast, simple to explain, directly trainable | One bit per read of two responses | Default for most projects |
| Ranking of K | Many pairs per prompt | Pairs from one ranking are correlated; fatigue grows with K | Responses are short and comparable |
| Graded pairwise | Lets you filter or weight weak preferences | Raters disagree more on strength than on direction | You want to drop near-ties |
| Separate axes | Helpfulness and safety judged independently | More work per item | Safety matters and trades off against helpfulness |
Always allow a tie or both-bad option. Forcing a choice between two equally poor answers creates a confident-looking pair that is pure noise. Randomise which response appears first and record the order. Show the full conversation context the model saw, including any system prompt, because a response can only be judged against the instructions it was given.
Guidelines are the real product of this stage. Write them as a priority order with worked examples: for instance correctness first, then following explicit instructions, then safety, then concision, then style. Include examples of ties and of tempting wrong answers, such as a confident, well-formatted response with a factual error beating a plain correct one. Version the guideline and store the version on every judgement. When raters disagree systematically on a category, the fix is usually a clearer rule, not a different rater.
Worked example: a 600-prompt round
Take a support assistant built on a small model. You sample 600 prompts from the pool across six categories, generate four responses per prompt from the current SFT checkpoint at temperature 0.9, and ask raters to rank them with ties allowed. Each complete ranking without ties yields six pairs, so the round can produce up to 3,600 pairs from 600 tasks.
from itertools import combinations
import random
def ranking_to_pairs(prompt_id, responses, ranking, max_pairs=None, seed=0):
"""responses: list of texts; ranking: indices best-first, ties as nested lists.
Example ranking [2, [0, 3], 1] means 2 > 0 = 3 > 1."""
tiers = [t if isinstance(t, list) else [t] for t in ranking]
pairs = []
for hi, lo in combinations(range(len(tiers)), 2): # hi ranked above lo
for a in tiers[hi]:
for b in tiers[lo]:
pairs.append({"prompt_id": prompt_id,
"chosen": responses[a], "rejected": responses[b],
"gap": lo - hi}) # tier distance as strength
if max_pairs and len(pairs) > max_pairs:
random.Random(seed).shuffle(pairs)
pairs = sorted(pairs, key=lambda p: -p["gap"])[:max_pairs]
return pairsDo not keep all of them blindly. Pairs from one ranking share a prompt and responses, so they are correlated, and adjacent ranks are often near-ties. A common choice is to cap pairs per prompt, prefer pairs with a larger rank gap, and make sure each prompt contributes at least one pair so the prompt mix survives. Send about 10 percent of tasks to a second rater for agreement measurement, and mix in gold tasks, where the answer is known and unambiguous, at a rate high enough that each rater sees a few dozen per round. Then compare the length of chosen and rejected responses: if chosen is longer in the large majority of pairs, investigate before training.
Quality control: gold items and agreement
Two signals catch most problems. Gold accuracy measures whether a rater gets the unambiguous cases right; it catches inattentive or misunderstanding raters. Agreement measures how often a rater matches the other raters' majority on overlapping items; it catches raters who apply a different standard and also tells you how noisy the task itself is. Pairwise agreement in preference tasks is well below 100 percent even with careful raters, because many comparisons are genuinely close. That is expected; what matters is the trend per rater and per category.
from collections import defaultdict
def rater_report(judgements, gold):
"""judgements: list of (rater, item, choice); gold: {item: correct_choice}."""
by_item = defaultdict(list)
stats = defaultdict(lambda: {"gold_n": 0, "gold_ok": 0, "agree_n": 0, "agree_ok": 0})
for rater, item, choice in judgements:
by_item[item].append((rater, choice))
if item in gold:
stats[rater]["gold_n"] += 1
stats[rater]["gold_ok"] += choice == gold[item]
for item, votes in by_item.items():
if len(votes) < 2:
continue
for rater, choice in votes: # agreement with the other raters' majority
others = [c for r, c in votes if r != rater]
majority = max(set(others), key=others.count)
if others.count(majority) * 2 <= len(others):
continue # the others are split: no majority
stats[rater]["agree_n"] += 1
stats[rater]["agree_ok"] += choice == majority
return {r: {"gold_acc": s["gold_ok"] / max(s["gold_n"], 1),
"agreement": s["agree_ok"] / max(s["agree_n"], 1),
"gold_n": s["gold_n"]} for r, s in stats.items()}Add time on task as a third signal: judgements made in a few seconds on long responses were not read. Review flagged raters with examples before removing them, since a cluster of disagreement often points to an ambiguous guideline. For chance-corrected agreement, Cohen kappa for two raters or Krippendorff alpha for many is more honest than raw percentage when one answer dominates.
AI feedback, and when it is acceptable
A strong model can act as the judge, an approach usually called RLAIF; Anthropic's Constitutional AI work used AI preference labels guided by written principles for harmlessness. AI judges are cheap, consistent and fast, but they have known biases: they often prefer the first or second position, longer answers and answers that sound like themselves. LLM-as-a-judge calibration covers these in detail.
def ai_preference(judge, prompt, a, b, rubric):
"""Ask the judge twice with the order swapped; keep only consistent verdicts."""
first = judge(prompt=prompt, response_1=a, response_2=b, rubric=rubric) # "1", "2" or "tie"
second = judge(prompt=prompt, response_1=b, response_2=a, rubric=rubric)
flipped = {"1": "2", "2": "1", "tie": "tie"}[second]
if first != flipped:
return None # position-dependent: route to a human
return {"1": "a", "2": "b", "tie": "tie"}[first]A safe hybrid is to let the AI judge label everything, keep only verdicts that survive an order swap, send the inconsistent ones and a random sample of the consistent ones to humans, and measure the judge against human labels on that sample every round. Never use the same model family as both the policy being trained and the only judge, because shared blind spots become reinforced preferences.
Release, versioning and splits
Release each round as an immutable, versioned dataset with a card that states the prompt sources, the candidate models, the guideline version, the rater pool, the agreement and gold figures, the tie rate and the chosen-versus-rejected length ratio. Split by prompt, not by pair, so that no prompt appears in both training and evaluation. Keep the raw judgements as well as the derived pairs, because you will want to re-aggregate them under a new rule. The dataset discipline in training data curation applies here unchanged.
Failure modes
- Length bias. Chosen responses are mostly longer, so the model learns verbosity. Detect with the length ratio; fix with length-balanced pairs and guidelines that reward concision.
- Stale candidates. Pairs come from an old checkpoint, so held-out preference on the new policy stalls. Fix: collect in rounds from the current policy.
- Position bias. One display slot wins far more often than half the time. Fix: randomise and record order, and audit per rater.
- Forced choices. No tie option, so near-identical pairs become noise. Fix: allow ties and drop them from training.
- Prompt leakage. Evaluation prompts appear in training pairs. Fix: split by prompt and hash-check before release.
- Guideline drift. Rules change mid-round without a version bump, so the dataset mixes standards. Fix: version guidelines and store the version per judgement.
What to do next
- Write down your schema with the metadata fields above before collecting a single judgement.
- Build a prompt pool from real traffic, deduplicate it, tag categories and hold out an evaluation slice.
- Generate candidates from your current policy and check their length distribution before judging.
- Draft a versioned guideline with a priority order, tie examples and tempting wrong answers.
- Add gold items, a 10 percent overlap for agreement, randomised order and time tracking.
- If you use an AI judge, swap orders, route inconsistencies to humans and measure it against human labels each round.
- Train a reward model on the result and check it for length bias as described in reward model training.