Static benchmarks decay. Once a test set has been on the internet for a year, there is a good chance some of it is in the next model's pre-training data, and a high score starts to measure memory rather than ability. LiveBench, introduced by Colin White and colleagues in the paper LiveBench: A Challenging, Contamination-Limited LLM Benchmark (arXiv 2406.19314, an ICLR 2025 spotlight), attacks the problem from two directions: questions are built from recent sources and replaced in dated releases, and every question has an objective ground-truth answer scored by code, not by another LLM.

For a team that serves its own models, LiveBench is useful in a second way: it is a hard, automatically scored, broad test you can run against a model on your own GPUs to catch regressions from quantisation, engine upgrades or decoding changes. This article explains how the benchmark is built, how its scores are computed, how to point it at a vLLM server, how to budget the run, how to compare two serving configurations properly, and the mistakes that make LiveBench numbers misleading.

What LiveBench measures, and how releases work

The README still describes 18 tasks across six categories: math, coding, reasoning, language, data analysis and instruction following. The changelog has since added an agentic coding category and several new tasks, so count tasks from the release you actually download rather than trusting any fixed number. Sources are chosen to be recent: math competitions, arXiv papers, news articles, movie synopses and new datasets. Several tasks are harder versions of older benchmarks such as Big-Bench Hard, AMPS and IFEval. Instruction-following items, for example, apply verifiable constraints to recent Guardian articles. When the paper appeared, the best models scored below 70 per cent; that figure has aged, which is the point of releases.

The project's changelog shows what a release means in practice: tasks are refreshed, hardened or replaced as models saturate or contaminate them.

ReleaseWhat changed (from the changelog)
2024-11-25New Guardian articles for instruction following, new Connections puzzles, harder zebra puzzles, answers in solution tags
2025-04-02Newer coding questions, refreshed typos and plot unscrambling, 2024 AMC questions, harder web-of-lies items
2025-04-25New coding questions on real-world library use; tablejoin and tablereformat refreshed; cta retired
2025-05-30New agentic coding category: multi-turn issue resolution in real repositories, run in Docker
2025-10-03Agentic coding switched to Mini-SWE-Agent with a 250-step limit; results rerun
2025-11-25Refreshed Connections, math competitions, instruction following and zebra puzzles; theory_of_mind replaced web_of_lies_v3
2025-12-23 and 2026-01-08New tasks: logic with navigation; integrals with game; consecutive events

Two consequences follow. First, the newest questions are not always public: the README notes that not every question in recent releases is on Hugging Face, and tells you to pass --livebench-release-option 2024-11-25 to evaluate all categories on public data. So a local run and the public leaderboard may be scoring different question sets. Second, a score is only meaningful next to its release date; comparing a 2024 number with a 2026 number compares two different tests.

Which release to use depends on the question you are asking. For a paired regression check, such as BF16 against INT4 of the same checkpoint, an older public release is fine: whatever contamination exists affects both sides equally and cancels in the difference. For an absolute capability claim about a 2025 or 2026 model, an older release may sit inside its training data, so you need a release newer than the model's data cutoff, which may not be fully public yet.

The pipeline and how scores are computed

LiveBench run against a model on your GPUsQuestion releaseHugging Face, by dategen_api_answer--parallel-requests NvLLM server on GPUsOpenAI-compatible /v1, TP across GPUsrequestsanswersmodel_answer/*.jsonlone line per questiongen_ground_truth_judgmentper-task scorer in process_resultsground_truth_judgment.jsonlquestion_id, task, category, scoreshow_livebench_resulttask means, category means, averageYour analysis: paired per-question deltas between serving configurationsquantisation, engine version, max tokens, temperature
Answers are generated against your server, scored by per-task code, aggregated into task and category means, then compared per question across configurations.

The pipeline has three scripts, which run_livebench.py runs in order. gen_api_answer.py sends each question to an OpenAI-compatible endpoint and writes model_answer/<model>.jsonl. gen_ground_truth_judgment.py applies the scorer for each task, kept under livebench/process_results, and appends to model_judgment/ground_truth_judgment.jsonl. show_livebench_result.py prints the table and writes all_tasks.csv and all_groups.csv.

The aggregation is worth knowing exactly, because it decides what a point is worth. Reading show_livebench_result.py: judgments with score -1 are dropped, scores are scaled to 0-100, each task's score is the mean over its questions, each category's score is the mean of its task means, and the headline average is the mean of the categories. Tasks therefore carry equal weight within a category regardless of question count, and a category with three tasks counts the same as one with four. Answers that end as $ERROR$ after retries, for example from a provider-side content filter, count as wrong.

Objective scoring removes judge bias and judge cost, but it does not make the benchmark bias-free. Scorers parse a specific answer format, such as text inside solution tags, and a correct answer in the wrong format scores zero. Prompt formatting, chat templates and truncation all move scores, as the next sections show.

Serving your model on GPUs for LiveBench

The README says local in-process inference is unmaintained and recommends serving the model behind an OpenAI-compatible API with vLLM, then running the normal scripts against it. On a node with two GPUs:

# 1. Serve the model (tensor parallel over 2 GPUs)
vllm serve ./checkpoints/my-model \
  --served-model-name my-model \
  --tensor-parallel-size 2 \
  --max-model-len 32768 \
  --max-num-seqs 64 \
  --gpu-memory-utilization 0.90 \
  --port 8000

# 2. From the livebench/ directory of the LiveBench repo
python run_livebench.py \
  --model my-model \
  --model-display-name my-model-bf16 \
  --api-base http://localhost:8000/v1 \
  --api-key EMPTY \
  --bench-name live_bench \
  --livebench-release-option 2024-11-25 \
  --parallel-requests 48 \
  --max-tokens 16384

Each flag here is a decision. --max-tokens defaults to 4096 unless a model config overrides it; a reasoning model that thinks for 10,000 tokens is cut off mid-thought and scored wrong, so set it from the model's real output lengths, and make --max-model-len on the server large enough for the longest prompt plus that budget. --parallel-requests is the number of questions in flight per task process; keep it at or below the server's --max-num-seqs so requests batch instead of queueing, and remember that --mode parallel runs one tmux session per category, multiplying the load. The served model name must match --model; --model-display-name sets the name the answer and judgment files are saved under, so give each configuration its own, for example rerunning against the quantised server as my-model-int4. --force-temperature pins sampling temperature when you need runs to be repeatable. How vLLM turns these limits into KV-cache memory is covered in vLLM on GPU, in depth.

The coding tasks execute model-written code during scoring and need the packages in livebench/code_runner/requirements_eval.txt. Run scoring in a throwaway container or VM, not on a shared GPU host. The agentic coding category also needs Docker and, per the README, up to 150 GB of task images; leave it out of a quick regression run unless agentic coding is what you ship.

Worked example: budgeting a run

Budget the run before booking GPUs. Count questions after python download_questions.py by counting lines in the downloaded question.jsonl files for your release and categories. Then the generation time is roughly total output tokens divided by aggregate decode throughput at your concurrency. Take an illustrative case, with numbers you should replace with your own measurements:

QuantityValueWhere it comes from
Questions1,000Line count of question files (assumed here)
Mean output tokens, non-reasoning model700Pilot run of 50 questions
Mean output tokens, reasoning model6,000Pilot run; long tail to 16,000
Aggregate decode throughput at 48 in flight2,500 tokens/sServer metrics during the pilot

The non-reasoning model needs 700,000 output tokens, about 280 seconds of decoding. The reasoning model needs 6 million, about 40 minutes, and the tail matters: the last few questions with 16,000-token answers run at low concurrency and can add a large fraction to wall time. Prefill is small next to this unless prompts are long. Scoring is CPU work and usually minutes, except agentic coding. If the pilot shows throughput collapsing as concurrency rises, the KV cache is full and requests are being preempted; lower --parallel-requests rather than letting preemption waste work. Always run the pilot on 50 questions first; it also catches chat-template and answer-format problems before they cost a full run.

While the full run executes, watch the server rather than the progress bar. vLLM exposes Prometheus metrics for running and waiting requests, KV-cache usage and preemptions. A healthy run shows the running count near your concurrency, the waiting queue near zero and no steady stream of preemptions. Record the GPU model, driver, engine version and tensor-parallel size next to the scores: the same checkpoint can produce slightly different answers on different kernels and batch shapes, and that variation belongs in the noise floor you compare against.

Comparing two serving configurations

The main reason to run LiveBench on your own hardware is a change you control: a 4-bit quantised checkpoint against BF16, a new vLLM release, a different max-tokens budget. Comparing two headline averages hides most of the signal. The same questions were answered by both configurations, so compare per question and bootstrap the difference. Because the headline weights tasks equally, resample within each task:

import json, random, collections, glob

def load(model, bench="live_bench"):
    rows = {}
    for path in glob.glob(f"data/{bench}/**/model_judgment/ground_truth_judgment.jsonl",
                          recursive=True):
        for line in open(path, encoding="utf-8"):
            j = json.loads(line)
            if j["model"] == model and j["score"] != -1:
                rows[j["question_id"]] = (j["task"], j["category"], j["score"])
    return rows

a, b = load("my-model-bf16"), load("my-model-int4")
common = sorted(set(a) & set(b))
by_task = collections.defaultdict(list)
for q in common:
    task, cat, sa = a[q]
    by_task[(cat, task)].append(b[q][2] - sa)          # per-question delta

def headline(sample):                                    # same weighting as LiveBench
    cats = collections.defaultdict(list)
    for (cat, task), ds in sample.items():
        cats[cat].append(sum(ds) / len(ds))
    return 100 * sum(sum(v) / len(v) for v in cats.values()) / len(cats)

rng = random.Random(0)
boot = sorted(headline({k: rng.choices(v, k=len(v)) for k, v in by_task.items()})
              for _ in range(2000))
print("delta", round(headline(by_task), 2), "95% CI", round(boot[50], 2), round(boot[1949], 2))

If the interval excludes zero, the change moved the score; if it straddles zero, you have not shown a difference, and should not report one. Look at the per-task deltas too: quantisation often leaves language tasks flat and costs points on math or long-chain reasoning. With sampling enabled, run each configuration two or three times and average per question first, or the noise floor swamps small effects. Custom LLM benchmarks covers paired statistics and the GPU noise floor in more depth.

Failure modes

  • Wrong release. Leaderboard numbers and your local run use different question sets unless you match --livebench-release-option. Always record it.
  • Truncated reasoning. The 4096-token default silently fails long answers. Check the share of answers that hit the limit.
  • Format, not ability. Missing solution tags or a broken chat template zero out whole tasks. Inspect a few raw answers from each task after the pilot.
  • Server errors as wrong answers. Timeouts and 5xx responses become $ERROR$ and count as incorrect. Run scripts/error_check.py and rerun with --resume --retry-failures.
  • Contamination of your own making. If you fine-tune on data scraped after a release date, you can train on LiveBench questions. Evaluate on releases newer than your data cutoff.
  • Unsandboxed code execution. Scoring runs model-written code. Isolate it.
  • Comparing across releases. A score is tied to its release; never trend it across releases as if the test were fixed.

Trade-offs

Compared with harness-based static suites run through tools like those in LLM eval frameworks, LiveBench trades stability for freshness: you get harder, less contaminated questions, but the test changes under you and the newest questions may be private. Compared with preference-based arenas such as Chatbot Arena, it measures verifiable correctness rather than what people prefer, with no judge cost or judge bias, and says little about tone, helpfulness or safety. Compared with a custom benchmark built from your own traffic, it is broad but not about your users. Most teams should use it as a general regression detector alongside a small domain suite, not as the deciding metric.

What to do next

  1. Clone the repository, install it, and run python download_questions.py to count questions per category for the release you will use.
  2. Serve your model with vLLM and run a 50-question pilot on one category to check templates, answer formats and output lengths.
  3. Set --max-tokens from the pilot's longest answers and size --parallel-requests to the server's concurrency.
  4. Run the full public release, then check for $ERROR$ answers and retry them.
  5. Store the release option, model revision, engine version, flags and GPU type with the results.
  6. Compare configurations with the per-question paired bootstrap, not headline averages.
  7. Run the scorer for coding tasks inside a disposable sandbox.
Key takeaway: LiveBench limits contamination with dated releases of recent questions and scores every answer with code against ground truth. Run it against your own vLLM server with the release option recorded, max tokens set from real output lengths and coding scorers sandboxed, and judge serving changes with a per-question paired bootstrap, never by eyeballing two headline averages.