A ranking and comparison UI shows a person two answers to the same prompt, asks which is better, and turns many such votes into a ranking of models. Public arenas made the format famous, but the same system runs inside most teams that ship LLM features: comparing a fine-tuned checkpoint against the current production model, choosing between providers, or collecting preference pairs for later training.

The UI looks trivial: two text panes and three buttons. The hard parts are elsewhere. The system must keep the comparison blind, which means the two answers must not give away their source through speed, length limits or formatting. It must serve two models at once on GPUs, which doubles inference cost per vote. It must choose which pairs to show so that votes are spent where the ranking is uncertain. And it must turn noisy, biased votes into ranks with honest intervals. This page builds that system end to end, with a worked Bradley-Terry fit whose output was produced by running the code shown. Labeling guidelines and rater quality control are covered separately in preference data collection.

The system at a glance

Blind side-by-side comparison: from prompt to ranked leaderboardRater UIprompt, two panespromptPair samplerchoose models A, BGeneration fan-outsame params, both in parallelGPU pool: model AvLLM or TGI replicaGPU pool: model BvLLM or TGI replicaParity bufferrelease both together, shuffle sidestwo anonymous answersvoteVote logappend-only eventsRankerBradley-Terry, style covariates, bootstrapLeaderboardscores and rank intervalsuncertainty feeds sampling
The sampler picks a pair, both models generate in parallel with identical parameters, the parity buffer releases both answers together on random sides, and votes feed a Bradley-Terry ranker whose uncertainty steers the next pairs.

Four guarantees

Before any code, write down what the system must guarantee, because every design choice below follows from these properties.

  • Blindness. The rater must not know which model produced which answer until after voting. Anything that correlates with the model identity, including latency, a characteristic opening phrase, a truncation length, or a refusal template, leaks it.
  • Parity. Both models get the same prompt, the same system prompt or none, the same sampling parameters and the same output budget. If one model gets a longer budget, you are ranking budgets.
  • Independence. Each vote should reflect one rater's judgement of one pair, not a rater voting on their own prompt for the fifth time in a row, and not a script.
  • Honest uncertainty. The leaderboard must show how sure it is. Two models whose intervals overlap share a rank.

Serving two models fairly on GPUs

Each vote needs two generations, so a comparison system is an inference system first. Fan the request out to both models concurrently, not one after the other; sequential generation doubles the wait and makes the second answer systematically later. On self-hosted GPUs each candidate model usually lives in its own serving pool, for example a vLLM or TGI deployment, and the fan-out service calls both with one shared parameter set.

import asyncio, random, time

PARAMS = {"temperature": 0.7, "top_p": 1.0, "max_tokens": 1024}   # identical for both sides

async def generate(client, model, prompt):
    t0 = time.monotonic()
    out = await client.complete(model=model, prompt=prompt, **PARAMS)
    return {"model": model, "text": out.text, "tokens": out.output_tokens,
            "finish": out.finish_reason, "latency_s": time.monotonic() - t0}

async def battle(clients, model_a, model_b, prompt):
    a, b = await asyncio.gather(
        generate(clients[model_a], model_a, prompt),
        generate(clients[model_b], model_b, prompt),
    )
    left, right = (a, b) if random.random() < 0.5 else (b, a)   # side is random per battle
    return {"left": left, "right": right}                        # both released together

The parity buffer is the line that returns both answers at once. Streaming each pane as tokens arrive is a better experience, but it shows the rater which model is faster, and speed is often recognisable. Two workable compromises are to start both streams only when both models have produced their first chunk, and to pace the faster stream to the slower one. Either way, record the true latency of each side in the log; it is useful data, just not something the rater should see.

Budget GPU capacity per vote, not per request. As an illustration, if answers average 500 output tokens and a replica sustains about 2,000 output tokens per second at your batch size, one answer costs a quarter of a replica-second and a vote about half a replica-second, plus prefill. Ten thousand votes a day then need roughly 5,000 replica-seconds, under two replica-hours. Measure your own throughput; the point is that comparison traffic is cheap relative to production traffic and can usually share the same replicas at low priority. Long contexts and reasoning models that produce many hidden tokens change this sum quickly.

The comparison screen

The rater sees the prompt, two panes labelled A and B, and four choices: A is better, B is better, tie, both are bad. The both-bad option matters; without it raters pick randomly between two failures and the vote becomes noise. Keep the panes the same width and render both with the same markdown renderer, so formatting differences come from the models rather than the page.

Reveal the model names only after the vote is recorded, and lock the vote at that moment. Do not allow regenerating one side, which breaks parity, and do not let a rater continue a conversation with one model before voting. If multi-turn comparisons are needed, send every turn to both models and vote once at the end. Keyboard shortcuts speed internal raters up considerably; log the time from answers shown to vote cast, because a vote cast faster than anyone could read both answers is a candidate for removal.

The vote log

Store votes as append-only events with everything needed to recompute the ranking later under a different method. Rankings are derived data; the log is the asset.

{
  "battle_id": "b-2026-10-03-000481",
  "ts": "2026-10-03T13:09:12Z",
  "rater": {"id": "r-1182", "pool": "internal-eng"},
  "prompt": {"id": "p-77310", "hash": "sha256:9c1f...", "category": "coding", "lang": "en"},
  "left":  {"model": "candidate-ft-0930", "tokens": 412, "finish": "stop", "latency_s": 3.9},
  "right": {"model": "prod-2026-09",      "tokens": 655, "finish": "stop", "latency_s": 6.1},
  "params": {"temperature": 0.7, "top_p": 1.0, "max_tokens": 1024},
  "vote": "left",
  "decision_ms": 18400,
  "style": {"left": {"md_headers": 0, "md_lists": 3, "md_bold": 1},
            "right": {"md_headers": 2, "md_lists": 7, "md_bold": 5}}
}

The prompt hash lets you detect one prompt submitted hundreds of times. The finish reason exposes truncation: if one model hits the token limit far more often, the comparison is partly about the limit. Style counts are computed at write time so that later analysis can control for them.

Choosing which pairs to show

With uniform sampling, every pair of models is compared equally often. That is simple and unbiased but wasteful: once a strong model has beaten a weak one two hundred times, the next vote between them teaches almost nothing. Active sampling spends votes where they change the ranking. A practical rule: sample a pair with probability proportional to how much its two score intervals overlap, plus a floor so every pair keeps getting some votes. Give newly added models extra traffic until their interval narrows. Keep the floor, because pure active sampling lets small biases in early votes steer where later votes go.

From votes to scores: Bradley-Terry with intervals

The Bradley-Terry model says the probability that model i beats model j is pi / (pi + pj), where each model has a positive strength. Unlike an online Elo update, a Bradley-Terry fit uses all votes at once and does not depend on their order, which is why the larger public arenas moved to it. The derivation and its link to reward models are in reward model math. The fit below uses the classic minorisation-maximisation update, counts a tie as half a win for each side, and reports scores on an Elo-like scale.

import math, random
from collections import defaultdict

def fit_bt(votes, models, iters=200):
    # votes: (a, b, outcome) with outcome 1.0 = a wins, 0.0 = b wins, 0.5 = tie
    wins, games = defaultdict(float), defaultdict(float)
    for a, b, o in votes:
        wins[a] += o
        wins[b] += 1.0 - o
        games[frozenset((a, b))] += 1.0
    p = {m: 1.0 for m in models}
    for _ in range(iters):
        new = {}
        for i in models:
            denom = sum(games.get(frozenset((i, j)), 0.0) / (p[i] + p[j])
                        for j in models if j != i)
            new[i] = (wins[i] + 0.01) / (denom + 0.02)   # tiny prior keeps unbeaten models finite
        g = math.exp(sum(math.log(v) for v in new.values()) / len(new))
        p = {m: v / g for m, v in new.items()}            # fix the scale: geometric mean 1
    return {m: 400 * math.log10(v) + 1000 for m, v in p.items()}

def bootstrap(votes, models, rounds=200, seed=7):
    rng, samples = random.Random(seed), defaultdict(list)
    for _ in range(rounds):
        resample = [votes[rng.randrange(len(votes))] for _ in votes]
        for m, s in fit_bt(resample, models).items():
            samples[m].append(s)
    return {m: (sorted(xs)[int(0.025 * rounds)], sorted(xs)[int(0.975 * rounds) - 1])
            for m, xs in samples.items()}

Worked example. Four simulated models with true strengths 1.6, 1.0, 0.9 and 0.5, 1,200 battles between uniformly chosen pairs, and 10 percent of votes recorded as ties. Running the code gives:

model-a   1078.6  95% CI [ 1058.9,  1100.6]
model-b   1016.7  95% CI [  995.6,  1035.1]
model-c    992.2  95% CI [  970.0,  1013.7]
model-d    912.5  95% CI [  892.7,   932.0]

Two lessons sit in four lines. Models b and c differ by about 25 points, but their intervals overlap heavily, so an honest leaderboard ranks them joint second. And the gap between a and b is about 62 points, smaller than the roughly 82 points the true strength ratio implies, because random ties pull every comparison toward even. Ties carry information about closeness, not quality; decide how to treat them before launch and keep the choice fixed.

Showing ranks honestly

Display rank as a range, not a single number. One common rule: a model's upper rank is one plus the number of models whose lower bound is above its upper bound. Show the vote count beside each model, and hide models with too few votes rather than listing them with a huge interval at the top. Break results down by category, such as coding, writing and other languages, because an overall rank often hides a model that is best at one and weak at another.

Style control

Raters prefer longer, more formatted answers, even when the content is no better. LMSYS showed in its August 2024 style control analysis that adding answer length and markdown counts (headers, bold, list items) as covariates in the Bradley-Terry regression moved several models noticeably in the ranking. The method is a logistic regression: the outcome is whether the left answer won, the features are a plus-one and minus-one indicator for the two models, plus differences in normalised length and markdown counts. The model coefficients are the style-controlled strengths; the style coefficients tell you how much your raters reward verbosity.

Report both the raw and the style-controlled ranking. If a fine-tuned checkpoint wins only in the raw view, it has learned to write longer answers, and a reward model or DPO run trained on these votes will learn the same habit. The training side of that chain is covered in reward model training on GPU.

Integrity and anti-gaming

Any public or incentivised comparison will be gamed. Common attacks and their defences:

  • Identity leaks. Raters ask the model its name. Strip or refuse identity questions consistently on both sides, or flag those battles and exclude them.
  • Vote stuffing. Rate-limit per account and per network, require login, and drop votes with implausibly short decision times.
  • Prompt repetition. Use the prompt hash to cap how many battles one prompt can contribute.
  • Selective reporting. If model owners can test many private variants and publish only the best, the leaderboard rewards the number of tries. Publish every variant's score or limit private testing. This concern was raised publicly about arena leaderboards in 2025.
  • Judge substitution. If you use an LLM judge to scale up, calibrate it against human votes and keep the two rankings separate; see LLM-as-a-judge calibration.

Failure modes

SymptomLikely causeFix
Faster model wins more than expectedStreaming reveals speedParity buffer or paced streams
One model truncates oftenShared max_tokens too low for itRaise the budget; report finish reasons
Ranking flips week to weekToo few votes, no intervals shownBootstrap intervals, minimum vote count
Fine-tune wins raw, loses controlledVerbosity preferenceStyle covariates; fix the training data
Spikes of votes for one modelStuffing or identity leakRate limits, prompt hash caps, exclusion rules
GPU queue delays comparisonsShares replicas with production at equal priorityLow-priority queue or a separate pool

What to do next

  1. Write the four guarantees (blindness, parity, independence, honest uncertainty) into the design doc and test each one before launch.
  2. Build the fan-out with one shared parameter set and a parity buffer; log true latency and finish reason per side.
  3. Store votes as append-only events with prompt hash and style counts.
  4. Fit Bradley-Terry with bootstrap intervals, and show rank ranges and vote counts on the leaderboard.
  5. Add style-controlled scores and report them next to raw scores.
  6. Add rate limits, decision-time filters and prompt-hash caps before opening the UI to anyone outside the team.
  7. Before training on the votes, read DPO training on GPU and audit the pairs for length bias.
Key takeaway: A comparison UI is an inference system plus a statistics system. Generate both answers in parallel with identical parameters, release them together on random sides so speed and position reveal nothing, and log every vote as an event with latency, finish reason, prompt hash and style counts. Fit Bradley-Terry with bootstrap intervals, show ranks as ranges, report style-controlled scores beside raw ones, spend votes where intervals overlap, and defend against identity leaks, stuffing and selective reporting.