HELM, the Holistic Evaluation of Language Models from Stanford's Center for Research on Foundation Models, is best known as a leaderboard: dozens of models scored on many scenarios with several metrics each. That reputation hides its most practical use for people who run models rather than train them. HELM is an open-source harness, installed with pip install crfm-helm, that can point at your own inference server and tell you whether a change to the serving stack changed what the model says.
That question comes up constantly. You quantize weights to FP8, upgrade the inference engine, switch the attention kernel, enable speculative decoding, change the chat template or the default sampling parameters. Each change is sold as free throughput, and each can quietly move accuracy. Latency dashboards will not show it. This article explains how HELM is built, how to wire it to a vLLM-style OpenAI-compatible endpoint, how to read its outputs, and how to turn two runs into a statistically honest pass or fail. One naming note: this page is about the evaluation framework, not the Kubernetes package manager of the same name.
How HELM is put together
HELM separates an evaluation into pieces you can recombine. A scenario is a dataset with instances, each having an input and references such as the correct answer (MMLU, GSM8K, NarrativeQA and so on). An adapter turns instances into prompts: how many in-context examples, what instructions, whether the task is multiple choice by generation or by comparing probabilities, and the decoding settings. A model deployment says where requests go. Metrics score the responses. A run entry names one combination in a compact string such as mmlu:subject=philosophy,model=openai/gpt2, and a suite is a named folder of runs you want to compare.
The holistic part is the metric set. Beyond accuracy, the original paper (Liang et al., 2022) measured calibration, robustness, fairness, bias, toxicity and efficiency, because a model that is equally accurate but badly calibrated or brittle under typos is not equally good. For a serving team, robustness matters in a specific way: HELM can apply perturbations such as typos or casing changes to inputs, and a quantized model sometimes holds its clean accuracy while losing more under perturbation.
Two properties make HELM useful as a gate. Every request and response is recorded, so a disagreement can be inspected rather than argued about. And requests are cached, so rerunning a summary costs nothing, which is a feature right up to the moment it silently reuses answers from the wrong server, discussed under failure modes.
Pointing HELM at your server
HELM reads its environment from prod_env/ by default, or from the directory given by --local-path. Two YAML files describe your model: model_metadata.yaml declares the model as an entity, and model_deployments.yaml declares how to reach it. For a vLLM server the client class is helm.clients.vllm_client.VLLMClient, which speaks the legacy text-completions API, or VLLMChatClient in the same module for chat completions. Define one deployment per serving configuration you want to compare, each with its own model name so results never mix.
# prod_env/model_metadata.yaml
models:
- name: acme/assistant-8b-bf16
display_name: Assistant 8B (bf16 baseline)
description: Production baseline served in bf16
creator_organization_name: Acme
access: open
release_date: 2026-09-01
tags: [TEXT_MODEL_TAG]
- name: acme/assistant-8b-fp8
display_name: Assistant 8B (fp8 candidate)
description: Same weights, FP8 quantized, new engine version
creator_organization_name: Acme
access: open
release_date: 2026-09-01
tags: [TEXT_MODEL_TAG]
# prod_env/model_deployments.yaml
model_deployments:
- name: vllm/assistant-8b-bf16
model_name: acme/assistant-8b-bf16
tokenizer_name: acme/assistant-8b # must match what the server tokenizes with
max_sequence_length: 8192
client_spec:
class_name: "helm.clients.vllm_client.VLLMClient"
args:
base_url: http://baseline.serving.internal:8000/v1/
vllm_model_name: assistant-8b
- name: vllm/assistant-8b-fp8
model_name: acme/assistant-8b-fp8
tokenizer_name: acme/assistant-8b
max_sequence_length: 8192
client_spec:
class_name: "helm.clients.vllm_client.VLLMClient"
args:
base_url: http://candidate.serving.internal:8000/v1/
vllm_model_name: assistant-8bThe tokenizer entry matters more than it looks. HELM uses it to count tokens and to truncate prompts that would overflow max_sequence_length. If you point it at a different tokenizer than the server's, long-context scenarios are truncated in places the server would not, and the comparison measures your configuration rather than the model. A custom tokenizer may need its own entry in a tokenizer config file; check the HELM documentation page on adding new models for your installed version.
Then run the same run entries against both deployments into one suite. Keep the scenario list small and stable: a handful of tasks that resemble your traffic beats the full leaderboard set for a gate that runs on every release.
SUITE=fp8-rollout-2026-10
for M in acme/assistant-8b-bf16 acme/assistant-8b-fp8; do
helm-run \
--run-entries "mmlu:subject=college_computer_science,model=$M" \
"gsm:model=$M" \
"narrative_qa:model=$M" \
--suite "$SUITE" \
--max-eval-instances 500 \
--num-threads 8 \
--disable-cache
done
helm-summarize --suite "$SUITE"
helm-server --suite "$SUITE" # local web UI to browse runs and raw requestsRun-entry names vary between HELM releases, so list the available scenarios in your installed version before copying these. --disable-cache is deliberate here: you want fresh answers from each server, not cached ones. --max-eval-instances samples a fixed random subset, so both models see the same instances as long as the seed and HELM version match. --num-threads (default 4) controls client concurrency; it affects wall time, not scores, unless your server behaves differently under batching, which is itself worth knowing.
Reading the output and comparing runs
Each run writes a directory under benchmark_output/runs/<suite>/ (move it with --output-path). Five files matter: run_spec.json records the scenario, adapter and metrics; scenario.json the instances; scenario_state.json every request and response; per_instance_stats.json a score per instance; and stats.json the aggregates. In current releases a stat carries a name object (metric name, split, sub-split, perturbation) plus count, sum and mean, and each per-instance record has instance_id, train_trial_index, perturbation and a list of stats. Field layouts have changed between versions, so pin the HELM version in your gate and test the parser against a real run.
Leaderboard comparisons use aggregate means. A gate should not, because two aggregate means on 500 instances each can differ by a point purely from noise. Since both deployments answered the same instances, compare them per instance. The paired difference removes the variance that comes from some questions simply being harder.
import json, random
from pathlib import Path
def per_instance(run_dir, metric="exact_match"):
rows = json.loads((Path(run_dir) / "per_instance_stats.json").read_text())
out = {}
for r in rows:
if r.get("perturbation"): # compare clean inputs; do perturbed separately
continue
for s in r["stats"]:
if s["name"]["name"] == metric and s.get("count"):
key = (r["instance_id"], r.get("train_trial_index", 0))
out[key] = s["sum"] / s["count"]
return out
def paired_gate(base_dir, cand_dir, metric="exact_match",
max_drop=0.01, n_boot=5000, seed=0):
b, c = per_instance(base_dir, metric), per_instance(cand_dir, metric)
keys = sorted(b.keys() & c.keys())
if len(keys) < 0.95 * max(len(b), len(c)):
raise RuntimeError("runs did not score the same instances")
d = [c[k] - b[k] for k in keys]
mean = sum(d) / len(d)
rng = random.Random(seed)
boots = sorted(sum(rng.choice(d) for _ in d) / len(d) for _ in range(n_boot))
lo, hi = boots[int(0.025 * n_boot)], boots[int(0.975 * n_boot)]
flips = sum(1 for x in d if x != 0)
return {"n": len(d), "delta": mean, "ci95": (lo, hi), "changed": flips,
"pass": lo > -max_drop}The rule lo > -max_drop passes the candidate only if the whole 95 percent interval sits above the tolerated drop. That is a non-inferiority test, which is the right question for an optimisation: you are not trying to prove the candidate is better, only that it is not meaningfully worse. The changed count is worth printing too. A delta of zero made of 40 wins and 40 losses is a very different model from one with no changed answers.
What the efficiency numbers are, and are not
HELM also records efficiency stats per request, including inference_runtime and, where it can estimate them, inference_denoised_runtime and inference_idealized_runtime. The paper introduced the latter two to compare models fairly across providers: denoised runtime strips out variation caused by shared API load, and idealized runtime estimates time on standardised hardware. Neither is a serving benchmark. HELM sends requests from a small thread pool, prompts follow the benchmark's length distribution rather than your traffic, and nothing controls arrival rate. The numbers describe how long the harness waited, not time to first token at your target concurrency.
Use HELM for the quality half of a rollout decision and a load generator for the performance half. The LLMPerf article covers concurrency sweeps and goodput, and the serving SLO guide covers which latency targets to hold. A change ships only when both halves pass.
Worked example: an FP8 rollout
Suppose the candidate is the same 8B model quantized to FP8 on a newer engine, promising about 1.6 times the decode throughput in your load test. You run three scenarios with 500 instances each. The figures below are illustrative of the shape such a comparison takes, not measurements of any real model.
| Scenario | bf16 mean | fp8 mean | Paired delta | 95% interval | Changed answers | Gate (max drop 1 pt) |
|---|---|---|---|---|---|---|
| MMLU college CS | 0.612 | 0.606 | -0.006 | -0.020 to +0.008 | 13 (5 up, 8 down) | fail: interval reaches -2.0 |
| GSM8K | 0.774 | 0.752 | -0.022 | -0.038 to -0.006 | 17 (3 up, 14 down) | fail |
| NarrativeQA (F1) | 0.701 | 0.699 | -0.002 | -0.008 to +0.004 | 58 | pass |
Three lessons come out of a table like this. First, the MMLU aggregate dropped by only 0.6 points, yet the gate fails, because 500 instances cannot rule out a two-point drop; the fix is more instances, not a looser threshold. Second, the GSM8K loss is real and concentrated, 14 answers lost against 3 gained: open scenario_state.json for the 17 changed answers and you will often find a pattern, such as long multi-step arithmetic where a quantization error compounds, or a chat-template difference that changes when the model stops. Third, NarrativeQA changed 58 answers with almost no net effect, which is normal for generation tasks under any numeric change and is why per-instance flips alone are not a verdict.
The decision is then an engineering one: keep FP8 for the layers that tolerate it, test a different calibration, or accept the arithmetic loss for traffic that never does arithmetic. HELM's job was to make the trade visible and attributable before it reached users.
Failure modes
- Stale cache. Without
--disable-cache, a rerun can be served from cached responses recorded against an earlier server. Use a fresh cache, or a distinct deployment name per configuration, and check request timestamps inscenario_state.json. - Template mismatch.
VLLMClientsends raw completions, so no chat template is applied; a chat-tuned model evaluated that way scores below its real behaviour. Use the chat client when your production traffic is chat. - Different sampling. HELM sets temperature from the adapter. If your server overrides or clamps sampling parameters, both deployments must do so identically.
- Truncation drift. A wrong
tokenizer_nameormax_sequence_lengthtruncates prompts differently from production. - Contamination. Public scenarios may be in the model's training data. That inflates absolute scores but cancels in a paired comparison of one model's weights, which is another reason to gate on deltas.
- Version drift. Run-entry names, instance sampling and file layouts change between HELM releases. Pin the version and rebuild the baseline when you upgrade it.
Trade-offs
HELM's strengths are breadth, a disciplined data model and full request logs. Its costs are setup effort and opinionated scenario definitions. lm-evaluation-harness is lighter and has a larger task catalogue, which makes it the common choice inside training loops, while HELM's metric breadth and perturbations suit release review. Neither replaces a custom benchmark built from your own traffic, which is the only set that measures the tasks you actually serve. A sound arrangement uses HELM scenarios as a broad tripwire, a custom set as the primary gate, and a load test, in the spirit of MLPerf Inference, for performance claims.
What to do next
- Install a pinned
crfm-helmversion in its own environment and run the quick-start example to confirm the toolchain. - Add your production model to
prod_env/with the correct tokenizer and the client that matches your traffic, completions or chat. - Pick three to six scenarios that resemble your workload and freeze them as the gate set.
- Run the baseline deployment once, with the cache disabled, and archive the run directory.
- Wire the paired gate into the rollout pipeline for every quantization, engine or template change.
- Size
--max-eval-instancesso the interval is narrower than your tolerated drop. - Read the changed answers in
scenario_state.jsonbefore accepting or rejecting. - Pair every HELM pass with a load test at target concurrency before shipping.