Chatbot Arena began in 2023 as a research project from LMSYS Org and now runs as LMArena. Its idea is simple. A user types any prompt, two anonymous models answer side by side, and the user votes for the better answer. After millions of these battles, a statistical model turns the votes into one score per model. That score is now quoted in launch posts, procurement decks and investor memos. Few of the people quoting it can say what it measures, how wide its uncertainty is, or how a provider can move it without the model getting better.

This article explains the arena from the vote up. It covers the battle protocol, the Bradley-Terry model that replaced online Elo, bootstrap intervals and rank ranges, and style control. It ends with the documented ways the leaderboard can mislead. It also shows how to run the same method privately on your own models, quantisations and serving configurations. That is where most engineering teams get real value from it, because a private arena answers the question the public one cannot: which option your users prefer, on your prompts, on your hardware.

The battle protocol

A battle has four steps. A user writes a prompt. The platform picks two models and streams both answers with the names hidden. The user picks A, B, tie or both bad. Only then are the names shown. If a model reveals its identity during the conversation, for example by answering a question about who built it, the vote is discarded, because the point of the design is that the voter cannot favour a brand.

Three properties of this protocol matter for anything you conclude from it. The prompts are whatever visitors choose to type, so the distribution leans towards coding, casual questions, creative writing and puzzle-style tests of the model; it is not your workload. The judges are self-selected volunteers, not domain experts, and a single vote carries very little information. And the labels are relative. A vote says A was preferred to B on this prompt; it says nothing about whether either answer was correct.

The pipeline in the diagram below has one labelled input and a chain of statistical steps. Each step is a place where a design choice moves the published ranking, so it is worth knowing them one by one.

From one anonymous vote to a leaderboard row with an intervalUser promptreal, unscriptedPair samplerwhich two models?Model Ahidden identityModel Bhidden identityVoteA / B / tie / both badBattle logdedupe, drop leaks, botsBradley-Terry fitlogistic regressionStyle featureslength, markdownoptionalBootstraprefit on resamplesLeaderboardscore, CI, rank rangesampling weightsThe vote is the only label. Everything after it is statistics, and every statistical choice changes the published order.
Chatbot Arena as a data pipeline: anonymous pairwise votes, cleaned, fitted with Bradley-Terry, then bootstrapped into intervals and rank ranges.

From online Elo to Bradley-Terry

The first leaderboard used online Elo, the chess rating update. After each battle the winner gains points and the loser loses them, scaled by how surprising the result was. Online Elo has a known defect for this use: the final ratings depend on the order in which battles are processed, and recent battles weigh more than old ones. Models do not change their weights between battles, so that recency bias is noise. The arena moved to a Bradley-Terry fit around the end of 2023.

Bradley-Terry assumes each model i has a fixed strength and that the probability i beats j depends only on the difference. On the Elo-style scale the arena reports:

P(i beats j) = 1 / (1 + 10 ** ((R_j - R_i) / 400))

Fitting this to all battles at once is a logistic regression. Each battle becomes one row with +1 in the column of model A, -1 in the column of model B, and label 1 if A won. The coefficients are the strengths. The batch fit uses every vote equally and gives the same answer whatever order the battles arrived in. Ratings are only defined up to an additive constant, so one anchor is fixed, for example a reference model at 1000.

import numpy as np
from sklearn.linear_model import LogisticRegression

def fit_bt(battles, models, scale=400, base=10, init=1000):
    # battles: list of (model_a, model_b, winner) with winner in {"a", "b", "tie"}
    idx = {m: k for k, m in enumerate(models)}
    rows, labels, weights = [], [], []
    for a, b, w in battles:
        x = np.zeros(len(models))
        x[idx[a]], x[idx[b]] = 1.0, -1.0
        if w == "tie":                      # one common choice: half a win each way
            rows += [x, x]; labels += [1, 0]; weights += [0.5, 0.5]
        else:
            rows.append(x); labels.append(1 if w == "a" else 0); weights.append(1.0)
    X = np.array(rows) * np.log(base)       # makes coef_ land on the base-10 scale
    clf = LogisticRegression(fit_intercept=False, C=1e6, max_iter=1000)  # ~unregularised
    clf.fit(X, labels, sample_weight=weights)
    return {m: init + scale * clf.coef_[0][idx[m]] for m in models}

How ties are handled is a modelling choice, and "both bad" votes are another. Some pipelines drop them, others treat them as ties. Write the choice down, because it shifts models that tie often, typically the ones close in strength.

Reading the scale: a worked example

Make the scale concrete. If model A beats model B in 60 of 100 decisive battles, the maximum-likelihood gap is 400 * log10(60/40), about 70 points. Reverse the formula and a 100-point gap means the stronger model is expected to win about 64 percent of the time; 200 points means about 76 percent. Scores near the top of the board often sit 5 to 20 points apart, which is a preference of roughly 51 to 53 percent. That is a coin that is barely bent.

The model also forces transitivity. If A is 70 points above B and B is 70 above C, the fit predicts A beats C about 69 percent of the time even if A and C never met. That is what lets a sparse comparison graph produce a full ranking. It is also a strong assumption. Real preferences are not always transitive: a terse coding model can beat a chatty generalist on code prompts and lose on everything else, and one scalar per model averages that away.

How many votes does a pair need? Near a 50 percent win rate the standard error of a win fraction is sqrt(0.25 / n). With 1,000 decisive votes that is 1.6 percent, so a 95 percent interval is about plus or minus 3.1 percent. Near 50 percent one percentage point of win rate is about 7 rating points, so the interval is roughly plus or minus 22 points. To separate two models that differ by 10 points with confidence you need several thousand votes between them, or many indirect comparisons through shared opponents.

Intervals and rank ranges

A single fitted score hides its uncertainty, so the arena publishes intervals from the bootstrap. Resample the battle log with replacement, refit, repeat a few hundred times, and take percentiles of each model's score. The intervals widen for new models with few votes and for models compared mostly against one opponent.

def bootstrap_bt(battles, models, rounds=500, seed=0):
    rng = np.random.default_rng(seed)
    n = len(battles)
    samples = {m: [] for m in models}
    for _ in range(rounds):
        pick = rng.integers(0, n, n)
        fit = fit_bt([battles[i] for i in pick], models)
        for m in models:
            samples[m].append(fit[m])
    return {m: (np.percentile(v, 2.5), np.median(v), np.percentile(v, 97.5))
            for m, v in samples.items()}

def rank_upper_bound(ci):
    # rank = 1 + number of models whose lower bound beats my upper bound
    return {m: 1 + sum(ci[o][0] > ci[m][2] for o in ci if o != m) for m in ci}

The rank function matters more than the score. A rank derived this way lets several models share first place, which is the honest reading when their intervals overlap. When a press release says a model is number one, check whether it shares that rank with three others. Two cautions: percentile bootstrap intervals assume the battles are independent draws, and they ignore the sampling policy, which decides which pairs get compared and how often.

Style control

Voters prefer long, nicely formatted answers even when they are no more accurate. In 2024 the arena introduced style control. It adds features for the difference in answer length and in the count of markdown headers, bold spans and list items to the same logistic regression. The style coefficients absorb the preference for presentation, and the model coefficients estimate what is left. The arena publishes raw and style-controlled views, and the order shifts between them.

def style_features(resp_a, resp_b):
    def counts(t):
        return np.array([len(t.split()),                  # length proxy
                         len(re.findall(r"^#+ ", t, re.M)),   # headers
                         t.count("**") // 2,                   # bold spans
                         len(re.findall(r"^\s*([-*]|\d+\.) ", t, re.M))])  # list items
    a, b = counts(resp_a), counts(resp_b)
    return (a - b) / np.maximum(a + b, 1)                # normalised difference

# design row = [+1/-1 model columns] + style_features(...)
# fit as before; report only the model coefficients

Style control is a regression adjustment, not a causal fix. It removes the part of the preference that correlates with these four features in a linear way. A model that is verbose in a way the features do not count, such as long caveats or restating the question, still benefits. Look at both views. A large gap between the raw and controlled scores tells you how much of a model's lead comes from presentation.

Where the leaderboard misleads

The arena is useful and also easy to game. In 2025 the paper The Leaderboard Illusion (Singh et al., Cohere and academic co-authors) analysed about two million battles. It reported that some providers tested many private variants before release and published only the best, citing one provider with 27 private variants. Picking the maximum of many noisy scores inflates the result even when no variant is better. It also reported that proprietary models were sampled far more often than open-weight ones, with two providers estimated at about 19 and 20 percent of all arena data, and that training on arena-style data gave large gains on ArenaHard, a test set drawn from the same distribution. LMArena publicly disputed parts of it. The mechanism is still worth understanding whoever is right about the details.

  • Best-of-N selection. Ten variants with the same true strength and plus or minus 15 point noise will produce a maximum well above the true value. Treat a launch score as optimistic until it has settled with more votes.
  • Distribution mismatch. Arena prompts are not your traffic. Use the category views (coding, hard prompts, longer queries, languages) and still expect differences.
  • Serving differences. An arena entry is a model plus a system prompt, sampling settings and a serving stack. The same weights served at a different temperature or a lower-precision quantisation are effectively a different contestant.
  • Overfitting to the judge. Optimising for human preference on chat rewards confidence, length and friendliness. None of these are correctness.

Running a private arena

The method is more valuable inside your company than on a public website. Typical questions: does FP8 serving lose quality that users notice, does a smaller distilled model hold up, does a new system prompt help. Benchmarks with reference answers cover some of this. Open-ended work such as support replies or drafting needs preferences.

A private arena needs four parts. A shadow pairing service sends a sample of real requests to two candidates. A blind review queue shows the pairs to trained raters, or to an LLM judge calibrated against them, in random left-right order. A battle log stores prompt hash, candidates, order, verdict, rater and timestamp. And a nightly fit runs the Bradley-Terry and bootstrap code above.

import random, hashlib

def make_pair(request, candidates, weights):
    a, b = random.choices(candidates, weights=weights, k=2)
    while b == a:
        b = random.choices(candidates, weights=weights, k=1)[0]
    if random.random() < 0.5:          # randomise screen position
        a, b = b, a
    return {"prompt_id": hashlib.sha256(request.encode()).hexdigest()[:16],
            "left": a, "right": b}

Two GPU-side details decide whether the comparison is fair. First, pin the serving stack per candidate: engine version, quantisation, KV-cache dtype, batch limits and sampling parameters. Record them in the battle log, because a silent engine upgrade changes the contestant. Second, hide latency from raters. If one candidate streams slower because it runs on a busier pool, show both answers only when both are complete, or the vote will measure your scheduler. Budget two generations per battle, plus judge tokens if an LLM judge is used. At a few thousand battles per comparison that is usually cheap next to production traffic.

Failure modes

FailureSymptomFix
Identity leakageRaters recognise a house style or a self-descriptionStrip signatures, discard battles that mention model names
Position biasLeft answer wins more than half the ties-excluded votesRandomise order and log it; check left win rate per rater
Unpinned servingScore jumps after an infra change with no model changeVersion every serving parameter as part of the contestant
Sparse graphWide intervals, ranks that flip nightlySample pairs with overlapping intervals more often
Selective reportingBest of many variants publishedRegister variants before testing; report all of them
Stale fitNew model ranked on 200 votesHide scores until a minimum vote count and interval width

Trade-offs

Pairwise preference is cheap to collect and measures what users like, but it is not correctness. Pair it with reference tests for anything that has a right answer. Bradley-Terry gives order-free, transitive scores at the cost of hiding non-transitive, per-category behaviour; publish category boards when the workload is mixed. Style control removes some presentation bias and may remove real value: sometimes the formatted answer really is easier to use. LLM judges make private arenas cheap, but they have their own position and length biases and must be calibrated against humans before their votes count.

What to do next

  1. Read a public score as an interval and a rank range, never a point; check the style-controlled and category views.
  2. Before quoting an arena score in a decision, compare the arena's prompt mix with your traffic.
  3. Copy fit_bt and bootstrap_bt and fit them on a public battle dataset to see how intervals behave.
  4. Run a private arena for your next quantisation or model swap: shadow pairs, blind review, a nightly fit.
  5. Pin and log every serving parameter as part of the contestant identity.
  6. Calibrate any LLM judge against a few hundred human votes before trusting it; see LLM-as-a-judge calibration.
  7. Combine preferences with reference-based tests from a regression harness and the scoring patterns in evaluating LLM outputs.
Key takeaway: Chatbot Arena is a logistic regression over anonymous human preferences. A score is only meaningful with its interval, its rank range, its style-controlled twin and the prompt mix behind it. Small gaps at the top are near coin flips, best-of-many submissions inflate launch scores, and serving settings are part of the contestant. Use the public board as one signal, and run the same method privately on your own prompts to make real decisions.