Every training run ends with a number, and the number is only as trustworthy as the code that produced it. Hugging Face's evaluate library exists so that "accuracy", "F1", "BLEU" or "ROUGE-L" mean the same thing in your notebook, your training job and a paper you are comparing against. It wraps reference implementations of dozens of metrics behind one interface, handles accumulating predictions across batches and processes, and adds tools for comparing models and measuring datasets.

This article explains how the library actually works, so you can use it without being surprised: how a metric name turns into downloaded code, how incremental and distributed computation work, how to wire it into a training loop, how to get confidence intervals instead of single numbers, and where the metrics quietly disagree with each other. It also covers when not to use it. As of the check made for this article on 2026-10-03, the latest release on PyPI is 0.4.6 (September 2025), and the project README itself recommends Hugging Face's newer LightEval library for evaluating large language models. Evaluate remains a sound choice for classic supervised tasks: classification, tagging, translation, summarisation and speech recognition.

Advertisement

Metrics, comparisons and measurements

The library has three kinds of module, and the distinction tells you what inputs each expects. A metric scores predictions against references, such as accuracy or ROUGE. A comparison takes predictions from two models plus references and tells you whether they differ, such as McNemar's test. A measurement describes a dataset without any model, such as word length or label distribution. All three share the base class EvaluationModule and the same compute method.

Each module is a small Python script. Canonical modules live on the Hugging Face Hub under three organisations, evaluate-metric, evaluate-comparison and evaluate-measurement, and the loader tries those namespaces when you pass a bare name. Community modules are loaded by repository id, such as username/my_metric, and a local path to a script or directory works too. The signature of the loader is worth reading once:

evaluate.load(
    path,                  # "accuracy", "user/metric", or a local path
    config_name=None,      # e.g. "mrpc" for the glue metric
    module_type=None,      # "metric" | "comparison" | "measurement"
    process_id=0, num_process=1,
    cache_dir=None, experiment_id=None, keep_in_memory=False,
    download_config=None, download_mode=None,
    revision=None,         # pin a Hub revision of the metric script
    **init_kwargs,
)

Two properties follow from this design. First, loading a metric executes code fetched from the Hub. Treat a community metric like any third-party dependency: read the script, and pin it with revision so a later edit to the repository cannot change your numbers or your security posture. Second, a module describes its inputs with features and documents itself through description, citation and inputs_description. Printing metric.inputs_description before first use answers most questions about argument names and shapes.

How a metric accumulates and computes

evaluate.load(name)local path or Hub idResolve moduleevaluate-metric/NAME etc.Download scriptPython code, then importedEvaluationModulefeatures, _computeadd / add_batchrows to an Arrow cache fileOther processesone cache file eachcompute()process 0 gathers, others get NoneResult dictsave, push_to_hub, model cardimportread all filesfile locksA metric is downloaded code plus an Arrow buffer of predictions and references
The lifecycle of an Evaluate module: load resolves a name to a script on the Hub or on disk and imports it; add_batch appends predictions and references to an Arrow cache; compute on process 0 gathers every process's file and runs the metric once over all rows.

There are two ways to feed a metric. Passing everything to compute(predictions=..., references=...) at once is the simplest. The incremental path calls add for one example or add_batch for a batch inside the evaluation loop, then compute() with no arguments at the end. Incremental calls write rows into an Apache Arrow table, on disk in the cache directory by default or in memory with keep_in_memory=True, so you do not hold every prediction in Python lists.

The reason for the Arrow buffer is that many metrics are not additive. You cannot average per-batch F1 or corpus BLEU and get the right answer, because they depend on counts over the whole set. So the library stores raw predictions and references and computes once over all of them. That same buffer is how distributed evaluation works: each process loads the metric with its own process_id and the shared num_process, writes its own cache file, and on compute() process 0 waits on file locks for every other process, reads all the files, and returns the result. Every other process gets None. The source documents a default wait timeout of 100 seconds.

This design has operational consequences. All processes must see the same filesystem path, so on a multi-node cluster the cache directory must be on shared storage. Two concurrent evaluation jobs writing to the same cache must use different experiment_id values or they collide. And keep_in_memory is incompatible with the distributed mode, because other processes cannot read your memory. If you already gather predictions yourself, for example with Accelerate's gather_for_metrics, the simpler path is to gather to rank 0 and call compute there with a single-process metric.

Advertisement

Using it in training and evaluation code

Here is the common case: a text classifier evaluated inside a custom PyTorch loop, then the same metrics wired into the Trainer. Note that F1 is loaded separately with an explicit average argument. The F1 metric wraps scikit-learn and defaults to binary averaging, which raises an error on multi-class labels, and evaluate.combine passes keyword arguments to every module rather than routing them, so per-metric options are clearest on their own module.

import evaluate
import numpy as np
import torch

accuracy = evaluate.load("accuracy")
f1 = evaluate.load("f1")

model.eval()
for batch in eval_loader:
    with torch.no_grad():
        logits = model(input_ids=batch["input_ids"],
                       attention_mask=batch["attention_mask"]).logits
    preds = logits.argmax(dim=-1).cpu().numpy()
    refs = batch["labels"].cpu().numpy()
    accuracy.add_batch(predictions=preds, references=refs)
    f1.add_batch(predictions=preds, references=refs)

results = {**accuracy.compute(), **f1.compute(average="macro")}
print(results)   # {'accuracy': ..., 'f1': ...}

# The same metrics for transformers.Trainer
def compute_metrics(eval_pred):
    logits, labels = eval_pred
    preds = np.argmax(logits, axis=-1)
    return {
        **accuracy.compute(predictions=preds, references=labels),
        **f1.compute(predictions=preds, references=labels, average="macro"),
    }

With the Trainer, the expensive part is not the metric but holding logits. For a classifier the logits are small; for token-level or generative models they are batch x sequence x vocabulary floats, which can exhaust memory before compute_metrics ever runs. Pass preprocess_logits_for_metrics to reduce logits to predicted ids on the device, and set eval_accumulation_steps to move results to the CPU periodically. Our Trainer article explains where these hooks sit in the evaluation loop.

For pipelines, the evaluator API runs inference and scoring together. evaluate.evaluator("text-classification") returns an object whose compute(model_or_pipeline=..., data=..., metric=..., label_mapping=...) runs a pipeline over a dataset and returns the metric plus total_time_in_seconds, samples_per_second and latency_in_seconds. Supported tasks include text, token and audio classification, question answering, summarisation, translation, text generation, text-to-text generation, image classification and speech recognition. label_mapping is the argument people forget: pipelines return string labels such as POSITIVE, datasets store integers, and without the mapping every prediction counts as wrong.

Text generation metrics that disagree

Generation metrics are where most reporting errors happen, because the same name hides different implementations and preprocessing choices.

  • BLEU. The bleu and sacrebleu modules both exist. SacreBLEU fixes tokenisation and reports a signature so results are comparable across papers; plain BLEU numbers depend on how you tokenised. References are a list of reference lists, one list per prediction, and sacrebleu expects the same number of references for every prediction. Report which one you used.
  • ROUGE. The rouge module returns rouge1, rouge2, rougeL and rougeLsum. rougeLsum treats newlines as sentence boundaries, so summaries must be split into sentences with newlines or it silently equals rougeL. Whether stemming is enabled changes scores by a noticeable margin; state the setting.
  • seqeval scores named-entity tagging at the entity level from string tags such as B-PER and I-PER. Feeding integer ids or forgetting to drop the ignored label used for sub-word padding gives meaningless numbers.
  • WER for speech depends heavily on text normalisation: case, punctuation and number formatting. Normalise predictions and references with the same function before scoring.
  • perplexity loads a causal language model by id and runs it, so it is an expensive computation, not a cheap formula, and its value depends on the tokenizer.

Confidence intervals and paired comparisons

A single accuracy figure hides how much it would move on a different sample. The evaluator can bootstrap: with strategy="bootstrap" and n_resamples, each metric comes back as a dictionary with score, confidence_interval and standard_error. For comparing two models on the same test set, a paired test is more powerful than eyeballing two intervals, because both models saw the same hard examples. The mcnemar comparison takes predictions1, predictions2 and references and returns stat and p.

from evaluate import evaluator
from datasets import load_dataset
import evaluate

data = load_dataset("csv", data_files={"test": "tickets_test.csv"})["test"]
task = evaluator("text-classification")

for name in ["org/ticket-clf-v1", "org/ticket-clf-v2"]:
    res = task.compute(
        model_or_pipeline=name, data=data, metric=evaluate.load("accuracy"),
        input_column="text", label_column="label",
        label_mapping={"billing": 0, "bug": 1, "account": 2, "other": 3},
        strategy="bootstrap", n_resamples=1000,
    )
    print(name, res["accuracy"])

mcnemar = evaluate.load("mcnemar", module_type="comparison")
print(mcnemar.compute(predictions1=preds_v1, predictions2=preds_v2, references=labels))

Worked example: should version 2 ship?

A support team fine-tunes a ticket classifier and wants to know whether version 2 should replace version 1. The test set has 2,400 labelled tickets in four classes. The numbers below are illustrative, chosen to show the reasoning rather than measured from a real model.

The first run reports accuracy 0.874 for v1 and 0.881 for v2, a 0.7 point gain, and the team is ready to ship. Three checks change the conversation. The bootstrap interval for each model is roughly plus or minus 1.3 points, so the two intervals overlap almost completely. McNemar's test on the paired predictions counts the tickets where exactly one model is right: v2 fixes 61 tickets that v1 got wrong but breaks 44 that v1 got right, and the test returns a p-value of about 0.1, so the gain is not distinguishable from noise on this test set. Finally, macro F1 tells a different story from accuracy: v2 is better on the large "bug" class and worse on the small "account" class, which is the class that escalates to a human when misrouted.

The decision follows from the evidence, not the headline: keep v1 in production, collect more "account" examples, and set a ship rule in advance, for example that a new model must improve macro F1 with McNemar p below 0.05 and must not reduce any class's recall by more than two points. The team saves the result with evaluate.save("./results/", experiment="ticket-v2", **scores, **hparams), which writes a JSON file with the scores plus metadata, and records the final numbers in the model card; evaluate.push_to_hub can write them into the card's model-index metadata. Our model cards article shows what a reader of that card needs alongside the numbers.

Writing a custom metric

When no existing module fits, write one. A module subclasses evaluate.Metric, declares its inputs as datasets.Features, and implements _compute. The evaluate-cli create command scaffolds a repository with this structure, a README card and tests.

import datasets
import evaluate

class CostWeightedError(evaluate.Metric):
    # Misrouting cost: some wrong labels are more expensive than others.

    def _info(self):
        return evaluate.MetricInfo(
            description="Average misclassification cost from a cost matrix.",
            citation="",
            inputs_description="predictions, references: int class ids; cost: KxK list",
            features=datasets.Features({
                "predictions": datasets.Value("int64"),
                "references": datasets.Value("int64"),
            }),
        )

    def _compute(self, predictions, references, cost):
        total = sum(cost[r][p] for p, r in zip(predictions, references))
        return {"cost_weighted_error": total / max(len(references), 1)}

# metric = evaluate.load("path/to/cost_weighted_error")
# metric.compute(predictions=preds, references=labels, cost=COST)

Keep custom metrics deterministic and dependency-light, give them unit tests with hand-computed answers, and version them in a Hub repository or your own codebase so the number in last quarter's report can be reproduced exactly.

Failure modes and trade-offs

FailureWhat you seePrevention
Binary F1 on multi-class dataValueError about the average settingPass average="macro" or "weighted" explicitly
Missing label_mappingAccuracy near zero from the evaluatorMap pipeline label strings to dataset ids
Metric script changed upstreamScores shift with no code changePin revision, or vendor the script
Shared cache, concurrent jobsLock timeouts or mixed resultsUnique experiment_id per job, shared storage across nodes
Hub unreachable on a clusterLoad fails at the end of a long runLoad metrics at startup; pre-cache them for offline nodes
Averaging per-batch scoresF1 or BLEU that matches nothingadd_batch then compute once
Logits held for the whole eval setOut of memory during evaluationpreprocess_logits_for_metrics, eval_accumulation_steps
Different preprocessing per runWER or ROUGE jumps between reportsOne normalisation function, stated in the report

The trade-offs are clear once the internals are. Evaluate buys standard, citable implementations and painless accumulation at the price of runtime code downloads and a filesystem-based distributed protocol. For a few classification metrics in a tight inner loop, calling scikit-learn directly is simpler and has no network dependency. For LLM benchmarks, LightEval or a dedicated harness handles prompting, few-shot formatting and generation settings that Evaluate never modelled. The library's offline behaviour is controlled by its own configuration flag, HF_EVALUATE_OFFLINE; whatever you set, the robust pattern is to load every module once at job start, so a missing download fails in the first minute instead of after the last epoch. The test data matters as much as the metric, and our datasets article covers building splits that do not leak.

What to do next

  1. List every metric your team reports and, for each, record the exact module, its revision, and its arguments (averaging mode, stemming, tokeniser, normalisation).
  2. Move metric loading to the start of training and evaluation jobs so network or cache problems fail fast.
  3. Replace any averaging of per-batch metrics with add_batch plus one compute, or gather to rank 0 first.
  4. Add bootstrap confidence intervals to your evaluation report and a McNemar test to every model-versus-model decision.
  5. Report per-class recall alongside accuracy for any imbalanced task, and set the ship rule before you look at results.
  6. Read and pin any community metric before relying on it, or vendor it into your repository.
  7. For LLM evaluation, evaluate LightEval or another dedicated harness rather than stretching Evaluate beyond its design.
Key takeaway: Evaluate gives standard metric implementations behind one interface, but a metric is downloaded code plus an Arrow buffer, so pin revisions, load modules at startup and give concurrent jobs distinct experiment ids. Accumulate with add_batch and compute once, because most metrics are not additive. Pass task options such as F1 averaging explicitly, state preprocessing for generation metrics, and back every model decision with bootstrap intervals and a paired test. For large language model benchmarks, the project itself points to LightEval.