Arena-Hard-Auto is an automatic benchmark built to predict how a chat model would rank on Chatbot Arena without waiting for thousands of human votes. It takes hard, real user prompts from Arena, has the candidate answer them, and asks a strong LLM judge to compare each answer against a fixed baseline model's answer. The output is a win rate against that baseline with a bootstrap interval.

For teams that train or serve models, its value is practical: it runs in hours on your own GPUs plus a judge API, it responds to changes that multiple-choice benchmarks miss, such as a new chat template, a fine-tune or an aggressive quantization, and it is open source. This article explains how the prompts were chosen, exactly how the judging and scoring code works, how to serve a candidate for a run, how to use the result as a regression gate, and where it misleads. The Bradley-Terry mathematics and human-vote arena are covered in LMSYS Chatbot Arena, in depth.

What Arena-Hard is and what its numbers mean

The first release, Arena-Hard-Auto v0.1 (2024), had 500 prompts, used gpt-4-1106-preview as judge and gpt-4-0314 as the baseline. The accompanying paper, "From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline", reports 98.6 percent correlation with Chatbot Arena rankings and three times the separability of MT-Bench, measured on a set of 20 top models from April 2024, at about $20 per evaluated model at 2024 judge prices. Those figures describe that model set and that judge; they are not a guarantee for your models.

Arena-Hard-v2.0, released in April 2025 as a preview, replaced the prompts with 500 fresh challenging queries, including open-ended software engineering and math, plus 250 creative writing queries, all sourced from Chatbot Arena. The judges became GPT-4.1 and Gemini-2.5, as the repository names them. Each set is a fixed, versioned release; the pipeline makes a refresh possible, but the benchmark does not update itself.

How the prompts were chosen

Random Arena prompts make a poor benchmark: many are greetings or trivia that every model answers well, so they compress scores together. BenchBuilder, the curation pipeline, keeps the hard ones. For v0.1 the paper describes deduplicating and removing multi-turn and non-English prompts from an initial pool of 200,000 Arena queries, clustering them into roughly 4,000 topics, and having an LLM annotator mark each prompt against seven criteria: specificity, domain knowledge, complexity, problem-solving, creativity, technical accuracy and real-world application. The quality score is the number of criteria met.

Prompts and whole clusters below a score threshold are dropped, and the final 500 were sampled as two prompts from each of 250 high-quality clusters, which keeps any single topic from dominating. The exact thresholds differ between sources: the paper states that prompts below 6 and clusters with a mean below 5 were discarded, while the repository's BenchBuilder README describes a cheap GPT-3.5-Turbo first pass followed by a GPT-4-Turbo pass that keeps prompts scoring 6 or more in clusters averaging 6 or more. If you build your own set, treat the thresholds as tuning knobs. The v2.0 curation details are not spelled out in the same way, so do not assume v0.1's numbers carry over.

The judging protocol

Each prompt is judged in two games. In the first the baseline's answer is shown as assistant A and the candidate's as B; in the second the positions are swapped. That cancels the judge's position bias on average. The standard judge prompt asks the judge to write its own answer first, then compare both answers against it for correctness, helpfulness, relevance, conciseness and missing information, and finally to output one of five verdicts: [[A>>B]], [[A>B]], [[A=B]], [[B>A]] or [[B>>A]]. The creative-writing prompt drops the own-answer step, since there is no correct answer to write.

SettingPromptsBaselineOfficial judge
v0.1500gpt-4-0314gpt-4-1106-preview
v2.0 hard prompts500o3-mini-2025-01-31gemini-2.5 with style control, or gpt-4.1
v2.0 creative writing250gemini-2.0-flash-001gpt-4.1 and gemini-2.5 ensemble

A v2.0 run with one judge therefore costs 1,500 judge calls per candidate: 750 prompts times two games. With reasoning judges and long candidate answers, each call can carry many thousands of tokens, so budget judge spend before you schedule nightly runs.

From verdicts to a score

The scoring in show_result.py is short enough to understand completely. Verdicts become a list of outcomes from the candidate's point of view: a slight win is one 1, a significant win is three 1s (the weight=3 default), a tie is 0.5, and losses are zeros on the same scheme. Both games for a prompt contribute. The plain leaderboard is the mean outcome, and the interval comes from 100 bootstrap resamples, reporting the 5th and 95th percentiles, which is a 90 percent interval even though the column is labelled CI. A minimal re-implementation:

import random

W = 3
SCORE = {"A>>B": [1]*W, "A>B": [1], "A=B": [0.5], "A<B": [0], "A<<B": [0]*W,
         "B>>A": [0]*W, "B>A": [0], "B=A": [0.5], "B<A": [1], "B<<A": [1]*W}
# judgments whose verdict matches no label are dropped before scoring, as the repo does

def outcomes(game1, game2):
    # game1: baseline is A, candidate is B -> flip; game2: candidate is A
    flip = [1 - s for s in SCORE[game1]]
    return flip + SCORE[game2]

def score(judgments, rounds=100, seed=0):
    xs = [o for g1, g2 in judgments for o in outcomes(g1, g2)]
    rng = random.Random(seed)
    boots = sorted(sum(rng.choices(xs, k=len(xs))) / len(xs) for _ in range(rounds))
    mean = sum(boots) / rounds
    return 100 * mean, 100 * boots[int(0.05 * rounds)], 100 * boots[int(0.95 * rounds) - 1]

Two consequences matter for interpretation. Strong verdicts move the score three times as much as slight ones, so a judge that is decisive inflates spread. And the bootstrap resamples outcomes, not prompts, so the interval understates prompt-level variance; for comparing two of your own checkpoints, the paired prompt-level analysis below is more honest. With --control-features markdown length, the script instead fits a Bradley-Terry model with answer length and markdown density as extra features, so a model is not rewarded merely for writing longer, more formatted answers.

Reading a leaderboard row

Take two rows from the repository's v2.0 hard-prompt leaderboard with style control and Gemini-2.5 as judge: o3-2025-04-16 at 85.9 with an interval of minus 0.8 to plus 0.9, and gpt-4.1 at 50.0 with minus 1.9 to plus 1.7. The baseline, o3-mini-2025-01-31, is fixed at 50.0 by construction, so gpt-4.1 scoring 50.0 means the judge, after style adjustment, found it about as good as the baseline on these prompts. It does not mean the two are equal on your workload, and it does not mean o3 is 72 percent better than gpt-4.1; win rates against a baseline are not a ratio scale.

The same repository shows how much the judge matters. With GPT-4.1 judging instead, gpt-4.1 scores 58.3 and gemini-2.5 drops from 79.0 to 49.1. Neither number is wrong; each is a measurement under a stated protocol. The rule that follows is simple: only compare scores produced with the same prompt set, judge, baseline and style setting, and quote all four whenever you report one.

Architecture of a run

An Arena-Hard-Auto run against a model served on your GPUsPrompt setv2.0: 500 hard + 250 creativevLLM / SGLang servercandidate, OpenAI-compatiblegen_answer.pycached per promptBaseline answersprecomputedanswersgen_judgment.pytwo games per prompt, positions swappedJudge APIgpt-4.1, gemini-2.5show_result.pyverdicts to scores, bootstrapadd_markdown_info.pylength and markdown featuresstyle controlWin rate vs baselinewith bootstrap intervalGPU cost lives in answer generation;judge cost is API tokens.
Answers come from your GPUs, judgments from an API; scoring and style features are local and cheap.

The repository is driven by three YAML files. config/api_config.yaml defines every model endpoint, including your own; config/gen_answer_config.yaml lists the models to generate answers for; config/arena-hard-v2.0.yaml sets the judge, its temperature and token limit, and the models to judge. Both generation scripts cache: they skip any prompt that already has an answer or judgment, so an interrupted run resumes cheaply, and a stale answer file silently survives a model change unless you delete it.

Serving the candidate

Serve the candidate behind an OpenAI-compatible server and register it the way the repository's own examples do:

# shell: an 8B model fits one GPU; add replicas rather than tensor parallelism
vllm serve /models/my-chat-8b --served-model-name my-chat-8b \
    --tensor-parallel-size 1 --max-model-len 32768 --port 8000

# config/api_config.yaml
my-chat-8b:
    model: my-chat-8b
    endpoints:
        - api_base: http://127.0.0.1:8000/v1
          api_key: '-'
    api_type: openai
    parallel: 128
    max_tokens: 8192
    temperature: 0.0

# then
python gen_answer.py
python gen_judgment.py
python show_result.py --judge-names gpt-4.1 --control-features markdown length

Four settings decide whether the number means anything. The chat template must be the one the model was trained with; a wrong template is the most common cause of a mysteriously low score. max_tokens must leave room for reasoning models, whose answers are otherwise truncated mid-thought and lose to the baseline. parallel should keep the server saturated; continuous batching makes 750 prompts a short job for a mid-sized model, but long reasoning traces turn generation into a KV-cache capacity problem, so watch preemptions in the server metrics. Finally, record the exact weights, engine version, precision and sampling settings with each run, because every one of them can move the score.

Worked example: gating an FP8 build

Suppose you want to ship an FP8 version of a fine-tuned 8B model and need evidence that quality did not regress. Generate answers for both the BF16 and FP8 builds, judge both against the baseline, and compare them prompt by prompt rather than comparing two headline numbers:

def paired_delta(judg_a, judg_b, rounds=2000, seed=1):
    # judg_x: dict uid -> (game1, game2) for each build against the same baseline
    uids = sorted(set(judg_a) & set(judg_b))
    per = {u: sum(outcomes(*judg_b[u])) / len(outcomes(*judg_b[u]))
              - sum(outcomes(*judg_a[u])) / len(outcomes(*judg_a[u])) for u in uids}
    rng = random.Random(seed)
    boots = sorted(sum(per[u] for u in rng.choices(uids, k=len(uids))) / len(uids)
                   for _ in range(rounds))
    return 100 * sum(per.values()) / len(uids), 100 * boots[int(0.025 * rounds)], 100 * boots[int(0.975 * rounds)]

Resampling prompts captures the variance that matters, and pairing removes prompt difficulty from the comparison. A sensible gate is that the 95 percent interval of the FP8 minus BF16 delta stays above an agreed margin, for example minus two points. If the mean is near zero but the interval is wide, rerun the judge once to separate judge noise from model change. Pair this with programmatic checks from custom benchmarks and task suites from eval frameworks, because a judged win rate does not catch a broken tool-call format.

Failure modes

The failure modes are mostly bookkeeping, which is why they bite.

  • Null judgments. If the judge output does not match the verdict regex, the pair is dropped and the script prints how many. A rising count shrinks your effective sample.
  • Judge drift. A provider updating the judge behind a stable alias changes scores. Pin dated model versions where the provider offers them and never compare runs across judges.
  • Self-preference. Judges tend to favour answers from their own model family; the paper's ensemble experiments were motivated by this. Read judge calibration and bias.
  • Contamination. The prompts are public. A model tuned on Arena-Hard prompts or judgments will score well without generalising.
  • Style gaming. Without style control, length and headers buy wins. Report the style-controlled score.
  • Distribution mismatch. Hard technical prompts may not resemble your users. A high score is evidence about Arena-like traffic, not yours.

Trade-offs

Arena-Hard sits between cheap automatic metrics and expensive human evaluation. Its strengths are sensitivity to open-ended quality and a cost low enough to run per checkpoint. Its weaknesses are dependence on a proprietary judge, a public prompt set, and a single baseline that saturates as models improve, which is why v2.0 moved to a stronger baseline. Use it to rank and gate checkpoints, not to make absolute capability claims, and keep a private set drawn from your own traffic for decisions that matter.

Key takeaway: <p>What to do next:</p><p><ol><li>Clone the repository, download the precomputed baseline answers, and reproduce one published row before scoring your own model.</li><li>Serve your candidate with its training chat template and a generous max_tokens.</li><li>Pin and record judge version, engine version, precision and sampling settings per run.</li><li>Report the style-controlled score and remember the printed interval is 90 percent.</li><li>Gate releases on paired, prompt-level deltas between builds, not on two headline numbers.</li><li>Clear cached answers whenever weights or serving settings change.</li><li>Keep a private judged set from your own traffic alongside the public one.</li></ol></p><p>Arena-Hard is a fast, honest proxy for chat quality when you control the bookkeeping. Treat its number as a measurement with an interval and a protocol, and it becomes a reliable gate.</p>