An LLM eval framework turns a benchmark definition into GPU work and the GPU's outputs back into a number. EleutherAI's lm-evaluation-harness (the lm_eval package), Hugging Face's lighteval, the UK AI Security Institute's Inspect and Stanford's HELM all do this, with different emphases. When a score moves by a point between two runs, the cause is usually not the model. It is how the framework built prompts, batched requests, chose a backend or computed a metric.

This page looks at eval frameworks from the GPU side, using lm-evaluation-harness as the reference implementation because its internals are open and widely copied. It covers the request types and what each costs on a GPU, how log-likelihood scoring works and where it goes wrong, batching and ordering, the HF and vLLM backends with tensor and data parallelism, a compute budget for a real suite, writing a custom task, and making scores reproducible. For choosing between frameworks, see AI evaluation frameworks.

The request-centric pipeline

Every framework runs the same pipeline. A task names a dataset split, a prompt template, a target and a metric. The framework loads documents, formats each one with optional few-shot examples into a context, and emits requests. A model backend answers the requests in batches, filters post-process the answers (for example extracting the final number from a chain-of-thought), metrics score each document, and aggregation reports a mean with a standard error. The important design decision in lm-evaluation-harness is that tasks never call the model. They only produce requests, so the framework can batch, reorder, cache and shard all of them together.

Inside an eval framework: from task config to GPU batches to a score with an error barTask configYAML: data, prompt, metricDocs + few-shotformat contextsRequestsloglikelihood / generateCache lookupskip repeated requestsReorder + batchlongest first, auto sizemissesHF backend1 GPU or DP replicasvLLM backendTP x DP, prefix cacheAPI backendremote endpointFilters + metricsacc, acc_norm, exact matchAggregatemean ± stderrResults + samplesJSON, configs, versionsPrefill-only requests dominate multiple-choice suites; decode-heavy requests dominate generative ones.
The request-centric pipeline. Tasks emit requests; the framework owns caching, ordering, batching and backend choice.

Three request types and what they cost

A backend in lm-evaluation-harness implements exactly three request types. Their GPU profiles differ enormously, and that difference decides how long a suite takes.

RequestInputReturnsGPU work
loglikelihood(context, continuation)log P(continuation | context), is-greedy flagOne forward pass; prefill only; compute-bound
loglikelihood_rollinga whole texttotal log-likelihood, for perplexityForward passes over windows of the max length
generate_until(context, generation kwargs)generated textPrefill plus token-by-token decode; memory-bandwidth-bound

A four-choice multiple-choice question becomes four loglikelihood requests that share the same context. No tokens are sampled, so the score is deterministic given the logits, and a whole batch runs as one dense prefill, which is the most GPU-efficient thing a transformer does. A generative maths task instead decodes hundreds of tokens per item, where throughput depends on batch size and KV-cache capacity, exactly as in serving.

Log-likelihood scoring, done correctly

Log-likelihood scoring sounds trivial and has one classic trap: tokenisation at the boundary. Tokenising the context and the continuation separately and concatenating them can produce token sequences the model never sees in training, because tokenisers merge a leading space into the next word. The robust approach is to tokenise the full string, then score only the positions that belong to the continuation.

import torch, torch.nn.functional as F

@torch.no_grad()
def loglikelihood(model, tok, context: str, continuation: str):
    whole = tok(context + continuation, return_tensors="pt").input_ids.to(model.device)
    n_ctx = len(tok(context).input_ids)            # tokens that belong to the context
    logits = model(whole).logits.float()           # [1, T, vocab]; upcast before softmax
    logp = F.log_softmax(logits[0, :-1], dim=-1)   # position t predicts token t+1
    targets = whole[0, 1:]
    cont = slice(n_ctx - 1, whole.shape[1] - 1)    # predictions for continuation tokens
    tok_lp = logp[cont].gather(-1, targets[cont, None]).squeeze(-1)
    greedy = bool((logp[cont].argmax(-1) == targets[cont]).all())
    return tok_lp.sum().item(), greedy

def pick(model, tok, context, choices):
    scores = [loglikelihood(model, tok, context, " " + ch) for ch in choices]
    acc = max(range(len(choices)), key=lambda i: scores[i][0])
    # acc_norm divides by the choice's length in characters, so long answers are not penalised
    norm = max(range(len(choices)), key=lambda i: scores[i][0] / len(choices[i]))
    return acc, norm

Two metrics fall out. acc picks the choice with the highest summed log-probability; acc_norm divides by the choice's length in characters, because longer answers accumulate more negative log-probability. Benchmarks such as HellaSwag are usually reported with acc_norm, and comparing one model's acc against another's acc_norm is a common way to publish a meaningless table. Note the float upcast: computing log-softmax in bfloat16 loses enough precision to flip close choices.

Ordering, batching and prefix reuse

lm-evaluation-harness sorts requests by descending total length before batching. Two benefits follow. The first batch is the largest, so an out-of-memory error happens in the first seconds rather than after an hour, and --batch_size auto can probe the largest batch that fits once, at the worst case. Batches of similar lengths also waste less compute on padding. Results are restored to the original order afterwards. auto:N re-probes N times during the run, since later batches are shorter and fit more rows, and --max_batch_size caps the probe.

Multiple-choice suites also repeat each context once per choice. A serving engine with prefix caching, such as vLLM with prefix caching enabled, computes the shared context once and reuses its KV blocks for each continuation, which can cut prefill work by close to the number of choices when contexts are long. Few-shot prompts amplify the effect, since the few-shot block is often shared across many documents too.

Backends and parallelism

The hf backend runs a Transformers model in-process. One process uses one GPU; launching with Accelerate gives data parallelism, one full model replica per GPU with the requests sharded between them. For a model too large for one GPU, the parallelize=True model argument splits layers across devices, which fits the model but leaves most GPUs idle at any moment. The vllm backend hands requests to vLLM, which brings continuous batching, paged KV cache and tensor parallelism, and is usually much faster for generative tasks.

# Single GPU, Transformers backend, automatic batch size
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-3.1-8B,dtype=bfloat16 \
  --tasks hellaswag,arc_challenge --num_fewshot 0 \
  --batch_size auto --device cuda:0 \
  --output_path results/ --log_samples

# 8 GPUs as data-parallel replicas of the HF backend
accelerate launch -m lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-3.1-8B,dtype=bfloat16 \
  --tasks mmlu --num_fewshot 5 --batch_size auto

# vLLM: 2-way tensor parallel x 4 data-parallel replicas on 8 GPUs
lm_eval --model vllm \
  --model_args pretrained=meta-llama/Llama-3.1-70B-Instruct,tensor_parallel_size=2,data_parallel_size=4,gpu_memory_utilization=0.85,dtype=auto \
  --tasks gsm8k --apply_chat_template --batch_size auto

Recent releases also accept an explicit lm-eval run subcommand alongside the older single-command form; check lm_eval --help for the version you have installed, since flags evolve. Prefer the largest data-parallel degree that fits the model: replicas scale nearly linearly, while tensor parallelism adds communication on every layer.

Worked example: a compute budget for MMLU

Estimate before you queue jobs. MMLU's test split has 14,042 questions. With four choices that is about 56,000 log-likelihood requests. A five-shot context averages several hundred tokens; take 600 as an order-of-magnitude figure. Prefill costs about 2 × parameters FLOPs per token, so for an 8B model:

requests = 14_042 * 4                     # ~56k
tokens   = requests * 600                 # ~3.4e7 prefill tokens
flops    = 2 * 8e9 * tokens               # ~5.4e17 FLOPs
peak_bf16 = 989e12                        # H100 SXM dense BF16, vendor spec
mfu       = 0.4                           # realistic for padded eval batches
seconds   = flops / (peak_bf16 * mfu)     # ~1,360 s, about 23 min on one GPU
with_prefix_reuse = seconds / 4           # shared context computed once per question

So a naive run takes around twenty minutes on one H100-class GPU, and prefix reuse or data parallelism over eight GPUs brings it to a few minutes. Generative suites behave differently: GSM8K with chain-of-thought produces a few hundred tokens per item and is limited by decode throughput, so the serving-engine backend matters far more there than for MMLU. Measure one task with --limit first and extrapolate, rather than trusting arithmetic alone.

Writing a custom task

Custom tasks are YAML files. A multiple-choice task over your own dataset looks like this:

task: internal_incident_triage
dataset_path: json
dataset_kwargs:
  data_files: {test: data/triage_test.jsonl}
test_split: test
output_type: multiple_choice
doc_to_text: "Incident: {{summary}}\nSeverity:"
doc_to_choice: ["low", "medium", "high", "critical"]
doc_to_target: "{{label}}"
metric_list:
  - metric: acc
    aggregation: mean
    higher_is_better: true

Point the harness at the directory with --include_path and run it by name. Keep templates boring: a trailing space in doc_to_text or an extra newline shifts tokenisation and therefore every log-likelihood. For anything where the item set itself is the hard part, see building custom benchmarks.

Reproducible scores

A score is only meaningful with its configuration. Things that change the number without changing the model include: batch size and padding side (different kernels and reduction orders), dtype, backend (HF and vLLM can differ slightly on the same weights), chat template on or off, few-shot sampling seed, tokenizer version, the task version, and for generative tasks every generation parameter. Pin and record all of them: the harness writes the resolved config and versions into its results JSON, --log_samples keeps every prompt and response so you can diff two runs item by item, and --use_cache stores responses in SQLite so a crashed run resumes without recomputing.

Then read the standard error. On 1,000 binary items at 70% accuracy the standard error is sqrt(0.7 × 0.3 / 1000), about 1.4 points, so a one-point gap is noise. For two models scored on the same items, a paired test on the per-item differences is far more sensitive than comparing two independent error bars, because most items are answered the same way by both models and only the disagreements carry information. Treat a results file without its config, sample log and standard error as an anecdote, and keep it out of release decisions.

Other frameworks, same pipeline

The other major frameworks make different trade-offs on the same pipeline. lighteval, from Hugging Face, covers similar academic benchmarks with several model backends, including vLLM, and is built to plug into Hugging Face training and leaderboard workflows. Inspect, from the UK AI Security Institute, is a Python framework where an evaluation is a dataset plus a solver plus a scorer; solvers can be multi-turn agents with tools running in sandboxes, which makes it a better fit for agentic and safety evaluations than for log-likelihood suites. HELM, from Stanford CRFM, evaluates many scenarios against several metrics at once, such as accuracy, calibration, robustness and efficiency, and publishes standardised comparisons across models.

From the GPU's point of view they converge: the expensive part is a model backend answering batched prefill or decode work, and the backend should be a serving engine, not a hand-written loop, whenever generation dominates. Whichever you choose, run one shared task through two frameworks once. The disagreement shows you how much of a published score is framework rather than model.

Failure modes

  • Boundary tokenisation. Scoring separately tokenised continuations biases results; tokenise the whole string.
  • Chat template mismatch. Evaluating an instruct model without its template, or a base model with one, can move scores by many points.
  • Metric mixing. Comparing acc with acc_norm, or 0-shot with 5-shot.
  • Silent truncation. Contexts longer than the model's window are cut from the left; few-shot examples disappear without an error.
  • Stale cache. A response cache keyed on the model name survives a change of weights at the same path. Clear it when weights change.
  • Contamination. Public test sets leak into pre-training data; a high score may measure memory.
  • OOM late in a run. Custom backends that do not sort longest-first fail hours in.

What to do next

  1. Install lm-evaluation-harness, run one task with --limit 100 and --log_samples, and read the actual prompts it sent.
  2. Re-run with a different batch size and dtype, and diff the per-item log-likelihoods to see your noise floor.
  3. Compare the HF and vLLM backends on one generative task for speed and score.
  4. Write the compute estimate for your full suite before scheduling it, then check it against a limited run.
  5. Port one internal test set to a YAML task and add it to CI with a pinned config.
  6. Keep learning: custom benchmarks and paired statistics, GPU reproducibility, prefix caching and vLLM on the GPU.
Key takeaway: An eval framework is a request compiler: tasks emit loglikelihood, rolling or generate requests, and the framework caches, sorts, batches and shards them onto a backend. Multiple-choice suites are prefill-bound and benefit from prefix reuse; generative suites are decode-bound and benefit from a serving engine. Tokenise at the boundary correctly, report the right metric, pin every setting, keep samples, and read the standard error before believing a difference.