Three judging formats
In pointwise grading the judge sees one response and assigns a score or a pass/fail verdict against a rubric. It is simple, scales linearly and produces absolute numbers that can be tracked over time. Its weakness is scale drift: a judge asked for a 1 to 10 score tends to cluster its outputs on a few values, and small wording changes in the rubric shift the whole distribution.
In pairwise comparison the judge sees two responses to the same prompt and says which is better, or that they tie. Relative judgments are easier and more consistent than absolute ones, for models as for people, which makes pairwise the preferred format for choosing between two configurations. Its cost is that it only ranks what it compares and says nothing about whether either answer is acceptable.
In reference-guided grading the judge receives a reference answer or a list of required facts and checks the response against them. This is the most reliable format for factual tasks because it turns an open-ended quality judgment into a narrower verification task. It needs labeled references, which is why it pairs naturally with a curated golden dataset.
The biases you must design around
- Position bias. In pairwise mode, judges favor one position, often the first, independent of content. The bias is strongest when the two answers are close in quality, which is exactly when the verdict matters.
- Verbosity bias. Longer, more elaborate answers are rated higher even when the extra length adds nothing or introduces errors. Uncorrected, a judge will reward any change that makes outputs longer.
- Self-preference. A judge may rate outputs from its own model family more favorably. Avoid using the same model as both candidate and sole judge when comparing providers.
- Anchoring on surface features. Confident tone, formatting and markdown structure raise scores for answers that are no more correct.
- Limited verification ability. A judge cannot reliably grade math, code or domain facts it cannot itself work out. Give it a reference, or execute the code, instead of asking it to guess.
None of these biases are exotic. They show up in any team's first week of judge-based evaluation, and each one produces a systematically wrong conclusion rather than random noise, which is why they must be handled in the design rather than averaged away.
Swap-consistent pairwise verdicts
The standard defense against position bias is to run every comparison twice with the order swapped, and accept a win only when both orders agree. Disagreement between the two orders is recorded as a tie. The swap roughly doubles judge cost, but it converts position bias from a systematic error into a measurable inconsistency rate, which is itself a useful health metric for the judge: if more than a modest fraction of comparisons flip on swap, the rubric or the judge model is not discriminating well on that task.
JUDGE = """You are comparing two responses to the same user request.
Criteria, in priority order: {rubric}
Think step by step about each criterion, then output a final line:
VERDICT: A, VERDICT: B, or VERDICT: TIE.
[Request]
{prompt}
[Response A]
{a}
[Response B]
{b}"""
def verdict(judge, prompt, a, b, rubric):
out = judge.complete(JUDGE.format(rubric=rubric, prompt=prompt, a=a, b=b), temperature=0)
last = out.strip().splitlines()[-1]
return {"VERDICT: A": "A", "VERDICT: B": "B"}.get(last.strip(), "TIE")
def swap_consistent(judge, prompt, x, y, rubric):
first = verdict(judge, prompt, x, y, rubric) # x shown as A
second = verdict(judge, prompt, y, x, rubric) # x shown as B
if first == "A" and second == "B":
return "x", True
if first == "B" and second == "A":
return "y", True
both_tie = first == second == "TIE"
return "tie", both_tie # False = inconsistentAsking for reasoning before the verdict, and parsing only a fixed final line, makes outputs easier to audit and parse. Keep the reasoning in the stored results; when a judge verdict looks wrong, the reasoning usually shows whether the rubric was misread or the judge lacked the knowledge to decide.
Aggregating comparisons
For comparing one candidate against one baseline, the win rate with ties counted as half a win is enough, reported with a bootstrap confidence interval over prompts. When several configurations are compared, for example a model sweep, pairwise results can be combined with the Bradley-Terry model, which assigns each system a strength such that the probability that i beats j is s_i / (s_i + s_j). Chatbot Arena uses this family of models to produce its leaderboard from crowd votes. The fit is a short iterative computation.
from collections import defaultdict
def bradley_terry(results, iters=200):
"""results: list of (winner, loser); record a tie as both (a, b) and (b, a)."""
wins = defaultdict(float)
games = defaultdict(float)
players = set()
for w, l in results:
wins[w] += 1.0
games[frozenset((w, l))] += 1.0
players |= {w, l}
s = {p: 1.0 for p in players}
for _ in range(iters):
new = {}
for i in players:
denom = sum(games[frozenset((i, j))] / (s[i] + s[j])
for j in players if j != i and games[frozenset((i, j))])
new[i] = wins[i] / denom if denom else s[i]
norm = sum(new.values()) / len(new)
s = {p: v / norm for p, v in new.items()}
return sThis minorization-maximization update converges when every system has at least one win and one loss against the others; add a small pseudo-count otherwise. Treat the strengths as a ranking tool, and go back to direct swap-consistent comparisons for the final decision between the top two.
Calibrating against humans
A judge is only trustworthy on a task where it has been compared with human labels. Build a calibration set of a few hundred items drawn from real outputs, stratified so that it includes close calls, not just obvious passes and failures. Have at least two people label each item using the same rubric the judge uses, measure agreement between the humans first, then measure the judge against the adjudicated human label. Use Cohen's kappa rather than raw agreement, because raw agreement looks high whenever one class dominates. If the humans themselves agree poorly, the rubric is ambiguous, and no judge will fix that.
For pass/fail judges, the confusion matrix matters more than a single agreement number. A judge that passes almost everything has high agreement on a dataset where almost everything passes, and is useless for catching regressions. Record the judge's sensitivity, the fraction of human-labeled failures it catches, and its specificity, the fraction of human-labeled passes it accepts. Those two numbers allow the pass rate the judge measures to be corrected toward the true rate.
def corrected_pass_rate(observed_pass, sens_fail, spec_pass):
"""Rogan-Gladen style correction for an imperfect binary judge.
observed_pass: fraction the judge marked PASS on the eval set.
spec_pass: P(judge says PASS | human PASS)
sens_fail: P(judge says FAIL | human FAIL)"""
fpr = 1.0 - sens_fail # P(judge PASS | human FAIL)
denom = spec_pass - fpr
if denom <= 0:
raise ValueError("judge does not discriminate on this task")
return min(1.0, max(0.0, (observed_pass - fpr) / denom))
# judge passes 88% of outputs; it catches 70% of real failures, accepts 95% of real passes
print(corrected_pass_rate(0.88, sens_fail=0.70, spec_pass=0.95)) # about 0.89The example looks reassuring, but the same judge on a candidate whose observed pass rate falls to 0.84 implies a true rate near 0.83, and the 30 percent of failures the judge misses are exactly the ones a regression report will not show. When sensitivity is low, invest in a better rubric or reference-guided grading before investing in more eval cases.
Rubric and prompt design
Judges behave best when the rubric is concrete. Replace a single overall quality score with a few binary checks, such as whether the answer states the refund deadline, whether it recommends an action the policy forbids, and whether it answers in the user's language. Binary criteria are easier to calibrate, and the per-criterion results tell engineers what to fix. Supply a reference answer or required facts whenever they exist, instruct the judge to ignore length and formatting unless the rubric mentions them, and use deterministic decoding.
To control verbosity bias in pairwise comparisons, either include an explicit rubric line that penalizes unnecessary length, or measure the bias and correct for it. AlpacaEval's length-controlled win rate does the latter, regressing the judge's preference on the length difference and reporting the win rate predicted at zero length difference. A lighter-weight check is to report the candidate's win rate within length-matched buckets alongside the headline number.
Judges as versioned dependencies
The judge is part of the measurement system, so it must be pinned and versioned like any other dependency. Record the judge model identifier, prompt hash, rubric version and decoding parameters with every result. A hosted judge model that is silently updated can move every score in a dashboard overnight, and without the pin nobody can tell whether the product or the ruler changed. When the judge must change, re-run the calibration set, re-score the current baseline with the new judge, and start a new series instead of splicing old and new numbers together.
Cost and latency deserve the same care. Put cheap deterministic checks first, including schema validation, forbidden-string checks, citation presence and executable tests for code, and only send cases that pass them to the judge. Cache verdicts keyed by the judge configuration and the exact pair of inputs, so re-running an unchanged baseline costs nothing.
Failure modes in practice
- Judge and candidate share a blind spot. If both are wrong in the same way, the judge approves the error. Reference-guided grading and human spot checks are the only defense.
- Rubric leakage into the product. When the product prompt is tuned against the judge's rubric, scores rise faster than quality. Keep a held-out human-rated sample and check that gains transfer.
- Unparseable verdicts. A judge that occasionally returns free text instead of the fixed format creates silent ties. Count parse failures and alert on them.
- Inconsistency hiding in ties. A rising tie rate usually means rising swap inconsistency, not genuinely equal answers.
- Safety-sensitive tasks. Do not let a judge be the only gate for content that carries legal or safety risk; route it to human review.