Every training run ends with a number, and the number is only as trustworthy as the code that produced it. Hugging Face's evaluate library exists so that "accuracy", "F1", "BLEU" or "ROUGE-L" mean the same thing in your notebook, your training job and a paper you are comparing against. It wraps reference implementations of dozens of metrics behind one interface, handles accumulating predictions across batches and processes, and adds tools for comparing models and measuring datasets.
This article explains how the library actually works, so you can use it without being surprised: how a metric name turns into downloaded code, how incremental and distributed computation work, how to wire it into a training loop, how to get confidence intervals instead of single numbers, and where the metrics quietly disagree with each other. It also covers when not to use it. As of the check made for this article on 2026-10-03, the latest release on PyPI is 0.4.6 (September 2025), and the project README itself recommends Hugging Face's newer LightEval library for evaluating large language models. Evaluate remains a sound choice for classic supervised tasks: classification, tagging, translation, summarisation and speech recognition.
Metrics, comparisons and measurements
The library has three kinds of module, and the distinction tells you what inputs each expects. A metric scores predictions against references, such as accuracy or ROUGE. A comparison takes predictions from two models plus references and tells you whether they differ, such as McNemar's test. A measurement describes a dataset without any model, such as word length or label distribution. All three share the base class EvaluationModule and the same compute method.
Each module is a small Python script. Canonical modules live on the Hugging Face Hub under three organisations, evaluate-metric, evaluate-comparison and evaluate-measurement, and the loader tries those namespaces when you pass a bare name. Community modules are loaded by repository id, such as username/my_metric, and a local path to a script or directory works too. The signature of the loader is worth reading once:
evaluate.load(
path, # "accuracy", "user/metric", or a local path
config_name=None, # e.g. "mrpc" for the glue metric
module_type=None, # "metric" | "comparison" | "measurement"
process_id=0, num_process=1,
cache_dir=None, experiment_id=None, keep_in_memory=False,
download_config=None, download_mode=None,
revision=None, # pin a Hub revision of the metric script
**init_kwargs,
)Two properties follow from this design. First, loading a metric executes code fetched from the Hub. Treat a community metric like any third-party dependency: read the script, and pin it with revision so a later edit to the repository cannot change your numbers or your security posture. Second, a module describes its inputs with features and documents itself through description, citation and inputs_description. Printing metric.inputs_description before first use answers most questions about argument names and shapes.
How a metric accumulates and computes
There are two ways to feed a metric. Passing everything to compute(predictions=..., references=...) at once is the simplest. The incremental path calls add for one example or add_batch for a batch inside the evaluation loop, then compute() with no arguments at the end. Incremental calls write rows into an Apache Arrow table, on disk in the cache directory by default or in memory with keep_in_memory=True, so you do not hold every prediction in Python lists.
The reason for the Arrow buffer is that many metrics are not additive. You cannot average per-batch F1 or corpus BLEU and get the right answer, because they depend on counts over the whole set. So the library stores raw predictions and references and computes once over all of them. That same buffer is how distributed evaluation works: each process loads the metric with its own process_id and the shared num_process, writes its own cache file, and on compute() process 0 waits on file locks for every other process, reads all the files, and returns the result. Every other process gets None. The source documents a default wait timeout of 100 seconds.
This design has operational consequences. All processes must see the same filesystem path, so on a multi-node cluster the cache directory must be on shared storage. Two concurrent evaluation jobs writing to the same cache must use different experiment_id values or they collide. And keep_in_memory is incompatible with the distributed mode, because other processes cannot read your memory. If you already gather predictions yourself, for example with Accelerate's gather_for_metrics, the simpler path is to gather to rank 0 and call compute there with a single-process metric.
Using it in training and evaluation code
Here is the common case: a text classifier evaluated inside a custom PyTorch loop, then the same metrics wired into the Trainer. Note that F1 is loaded separately with an explicit average argument. The F1 metric wraps scikit-learn and defaults to binary averaging, which raises an error on multi-class labels, and evaluate.combine passes keyword arguments to every module rather than routing them, so per-metric options are clearest on their own module.
import evaluate
import numpy as np
import torch
accuracy = evaluate.load("accuracy")
f1 = evaluate.load("f1")
model.eval()
for batch in eval_loader:
with torch.no_grad():
logits = model(input_ids=batch["input_ids"],
attention_mask=batch["attention_mask"]).logits
preds = logits.argmax(dim=-1).cpu().numpy()
refs = batch["labels"].cpu().numpy()
accuracy.add_batch(predictions=preds, references=refs)
f1.add_batch(predictions=preds, references=refs)
results = {**accuracy.compute(), **f1.compute(average="macro")}
print(results) # {'accuracy': ..., 'f1': ...}
# The same metrics for transformers.Trainer
def compute_metrics(eval_pred):
logits, labels = eval_pred
preds = np.argmax(logits, axis=-1)
return {
**accuracy.compute(predictions=preds, references=labels),
**f1.compute(predictions=preds, references=labels, average="macro"),
}With the Trainer, the expensive part is not the metric but holding logits. For a classifier the logits are small; for token-level or generative models they are batch x sequence x vocabulary floats, which can exhaust memory before compute_metrics ever runs. Pass preprocess_logits_for_metrics to reduce logits to predicted ids on the device, and set eval_accumulation_steps to move results to the CPU periodically. Our Trainer article explains where these hooks sit in the evaluation loop.
For pipelines, the evaluator API runs inference and scoring together. evaluate.evaluator("text-classification") returns an object whose compute(model_or_pipeline=..., data=..., metric=..., label_mapping=...) runs a pipeline over a dataset and returns the metric plus total_time_in_seconds, samples_per_second and latency_in_seconds. Supported tasks include text, token and audio classification, question answering, summarisation, translation, text generation, text-to-text generation, image classification and speech recognition. label_mapping is the argument people forget: pipelines return string labels such as POSITIVE, datasets store integers, and without the mapping every prediction counts as wrong.
Text generation metrics that disagree
Generation metrics are where most reporting errors happen, because the same name hides different implementations and preprocessing choices.
- BLEU. The
bleuandsacrebleumodules both exist. SacreBLEU fixes tokenisation and reports a signature so results are comparable across papers; plain BLEU numbers depend on how you tokenised. References are a list of reference lists, one list per prediction, andsacrebleuexpects the same number of references for every prediction. Report which one you used. - ROUGE. The
rougemodule returnsrouge1,rouge2,rougeLandrougeLsum.rougeLsumtreats newlines as sentence boundaries, so summaries must be split into sentences with newlines or it silently equalsrougeL. Whether stemming is enabled changes scores by a noticeable margin; state the setting. - seqeval scores named-entity tagging at the entity level from string tags such as
B-PERandI-PER. Feeding integer ids or forgetting to drop the ignored label used for sub-word padding gives meaningless numbers. - WER for speech depends heavily on text normalisation: case, punctuation and number formatting. Normalise predictions and references with the same function before scoring.
- perplexity loads a causal language model by id and runs it, so it is an expensive computation, not a cheap formula, and its value depends on the tokenizer.
Confidence intervals and paired comparisons
A single accuracy figure hides how much it would move on a different sample. The evaluator can bootstrap: with strategy="bootstrap" and n_resamples, each metric comes back as a dictionary with score, confidence_interval and standard_error. For comparing two models on the same test set, a paired test is more powerful than eyeballing two intervals, because both models saw the same hard examples. The mcnemar comparison takes predictions1, predictions2 and references and returns stat and p.
from evaluate import evaluator
from datasets import load_dataset
import evaluate
data = load_dataset("csv", data_files={"test": "tickets_test.csv"})["test"]
task = evaluator("text-classification")
for name in ["org/ticket-clf-v1", "org/ticket-clf-v2"]:
res = task.compute(
model_or_pipeline=name, data=data, metric=evaluate.load("accuracy"),
input_column="text", label_column="label",
label_mapping={"billing": 0, "bug": 1, "account": 2, "other": 3},
strategy="bootstrap", n_resamples=1000,
)
print(name, res["accuracy"])
mcnemar = evaluate.load("mcnemar", module_type="comparison")
print(mcnemar.compute(predictions1=preds_v1, predictions2=preds_v2, references=labels))
Worked example: should version 2 ship?
A support team fine-tunes a ticket classifier and wants to know whether version 2 should replace version 1. The test set has 2,400 labelled tickets in four classes. The numbers below are illustrative, chosen to show the reasoning rather than measured from a real model.
The first run reports accuracy 0.874 for v1 and 0.881 for v2, a 0.7 point gain, and the team is ready to ship. Three checks change the conversation. The bootstrap interval for each model is roughly plus or minus 1.3 points, so the two intervals overlap almost completely. McNemar's test on the paired predictions counts the tickets where exactly one model is right: v2 fixes 61 tickets that v1 got wrong but breaks 44 that v1 got right, and the test returns a p-value of about 0.1, so the gain is not distinguishable from noise on this test set. Finally, macro F1 tells a different story from accuracy: v2 is better on the large "bug" class and worse on the small "account" class, which is the class that escalates to a human when misrouted.
The decision follows from the evidence, not the headline: keep v1 in production, collect more "account" examples, and set a ship rule in advance, for example that a new model must improve macro F1 with McNemar p below 0.05 and must not reduce any class's recall by more than two points. The team saves the result with evaluate.save("./results/", experiment="ticket-v2", **scores, **hparams), which writes a JSON file with the scores plus metadata, and records the final numbers in the model card; evaluate.push_to_hub can write them into the card's model-index metadata. Our model cards article shows what a reader of that card needs alongside the numbers.
Writing a custom metric
When no existing module fits, write one. A module subclasses evaluate.Metric, declares its inputs as datasets.Features, and implements _compute. The evaluate-cli create command scaffolds a repository with this structure, a README card and tests.
import datasets
import evaluate
class CostWeightedError(evaluate.Metric):
# Misrouting cost: some wrong labels are more expensive than others.
def _info(self):
return evaluate.MetricInfo(
description="Average misclassification cost from a cost matrix.",
citation="",
inputs_description="predictions, references: int class ids; cost: KxK list",
features=datasets.Features({
"predictions": datasets.Value("int64"),
"references": datasets.Value("int64"),
}),
)
def _compute(self, predictions, references, cost):
total = sum(cost[r][p] for p, r in zip(predictions, references))
return {"cost_weighted_error": total / max(len(references), 1)}
# metric = evaluate.load("path/to/cost_weighted_error")
# metric.compute(predictions=preds, references=labels, cost=COST)Keep custom metrics deterministic and dependency-light, give them unit tests with hand-computed answers, and version them in a Hub repository or your own codebase so the number in last quarter's report can be reproduced exactly.
Failure modes and trade-offs
| Failure | What you see | Prevention |
|---|---|---|
| Binary F1 on multi-class data | ValueError about the average setting | Pass average="macro" or "weighted" explicitly |
| Missing label_mapping | Accuracy near zero from the evaluator | Map pipeline label strings to dataset ids |
| Metric script changed upstream | Scores shift with no code change | Pin revision, or vendor the script |
| Shared cache, concurrent jobs | Lock timeouts or mixed results | Unique experiment_id per job, shared storage across nodes |
| Hub unreachable on a cluster | Load fails at the end of a long run | Load metrics at startup; pre-cache them for offline nodes |
| Averaging per-batch scores | F1 or BLEU that matches nothing | add_batch then compute once |
| Logits held for the whole eval set | Out of memory during evaluation | preprocess_logits_for_metrics, eval_accumulation_steps |
| Different preprocessing per run | WER or ROUGE jumps between reports | One normalisation function, stated in the report |
The trade-offs are clear once the internals are. Evaluate buys standard, citable implementations and painless accumulation at the price of runtime code downloads and a filesystem-based distributed protocol. For a few classification metrics in a tight inner loop, calling scikit-learn directly is simpler and has no network dependency. For LLM benchmarks, LightEval or a dedicated harness handles prompting, few-shot formatting and generation settings that Evaluate never modelled. The library's offline behaviour is controlled by its own configuration flag, HF_EVALUATE_OFFLINE; whatever you set, the robust pattern is to load every module once at job start, so a missing download fails in the first minute instead of after the last epoch. The test data matters as much as the metric, and our datasets article covers building splits that do not leak.
What to do next
- List every metric your team reports and, for each, record the exact module, its revision, and its arguments (averaging mode, stemming, tokeniser, normalisation).
- Move metric loading to the start of training and evaluation jobs so network or cache problems fail fast.
- Replace any averaging of per-batch metrics with
add_batchplus onecompute, or gather to rank 0 first. - Add bootstrap confidence intervals to your evaluation report and a McNemar test to every model-versus-model decision.
- Report per-class recall alongside accuracy for any imbalanced task, and set the ship rule before you look at results.
- Read and pin any community metric before relying on it, or vendor it into your repository.
- For LLM evaluation, evaluate LightEval or another dedicated harness rather than stretching Evaluate beyond its design.