A small language model's model card usually arrives with a row of benchmark numbers. Most teams read that row the way they read a spec sheet: higher is better, pick the top. That fails in three ways. Some benchmarks are saturated or at chance for small models, so their differences are noise. Some are too small to resolve the gaps being compared. And none of them is your task.

This article treats benchmarks as measuring instruments. It covers what the common public benchmarks measure, which ones still discriminate between models of one to eight billion parameters, how many items a benchmark needs to detect a given difference, how to run a reproducible panel, and how to build the practical tests that make the actual decision. The evaluation system around this, meaning registries, CI gates and device lanes, is covered in SLM evaluation architecture, and scoring traps such as chat templates and answer extraction are covered in evaluation pitfalls. This page is about choosing and reading the instruments.

Advertisement

A benchmark is an instrument with a resolution

Every benchmark score is an estimate from a finite sample of items. For a model with true accuracy p on a benchmark of n independent items, the standard error is the square root of p(1 - p)/n, and a 95% interval is about 1.96 standard errors either side. That one formula explains most misreadings of leaderboards:

Benchmark size nAccuracy pStandard error95% interval, plus or minus
164 (HumanEval)0.403.8 points7.5 points
198 (GPQA Diamond)0.303.3 points6.4 points
541 (IFEval prompts)0.702.0 points3.9 points
1,319 (GSM8K test)0.601.3 points2.6 points
about 12,000 (MMLU-Pro)0.400.45 points0.9 points

So a 3-point gap on GPQA Diamond between two small models tells you nothing, while the same gap on MMLU-Pro is well outside sampling noise. The interval is also optimistic. It ignores prompt sensitivity, which can move small models by several points on its own. When two models are scored on the same items, a paired test is much sharper than comparing two independent intervals, because most items are answered the same way by both models and cancel out.

import math, random

def accuracy_ci(correct: int, n: int, z: float = 1.96):
    """Normal-approximation 95% interval for one model on one benchmark."""
    p = correct / n
    half = z * math.sqrt(p * (1 - p) / n)
    return p, p - half, p + half

def paired_diff_ci(a, b, iters=10_000, seed=0):
    """a, b: per-item 0/1 scores for two models on the SAME items, same order.
    Resampling items (not models) keeps the pairing, which is what makes the test sharp."""
    rng = random.Random(seed)
    diffs = [x - y for x, y in zip(a, b)]
    n = len(diffs)
    boot = sorted(sum(diffs[rng.randrange(n)] for _ in range(n)) / n for _ in range(iters))
    return sum(diffs) / n, boot[int(0.025 * iters)], boot[int(0.975 * iters)]

# If the interval on the difference includes 0, you have not shown that one model is better.

The catalogue: what each benchmark measures

BenchmarkWhat it measuresFormat and sizeReading it for small models
MMLUBroad knowledge across 57 subjects4-option multiple choice, 14,042 test itemsWidely contaminated and near ceiling for strong models; use as a sanity check only
MMLU-ProHarder knowledge and reasoning10 options, about 12,000 itemsChance is 10%, so less guessing; large n gives fine resolution. A good default
GPQA DiamondGraduate-level science4 options, 198 itemsMost small models sit near 25% chance; too small to separate them
GSM8KGrade-school arithmetic word problemsGenerative, 1,319 test itemsStill useful below about 3B; good small instruct models are near ceiling
MATH (Level 5)Competition mathematicsGenerative, exact-answer matchVery sensitive to answer extraction; separates reasoning-tuned models
IFEvalFollowing verifiable instructions (length, format, keywords)541 prompts, rule-checkedDirectly relevant to products; no judge model needed
BBHMulti-step reasoning23 hard BIG-Bench tasksMid-range for small models, so it discriminates; few-shot format matters
HellaSwag, ARCCommonsense completion; grade-school scienceMultiple choice; ARC-Challenge has 1,172 test itemsMostly saturated for instruct models; useful for tracking base-model pretraining
HumanEval, MBPPWriting short Python functions164 and 500 test problems, unit-test scoredSmall n; prefer larger or newer code suites for decisions
Function-calling suites (e.g. BFCL)Choosing a tool and filling argumentsMany categories; versions changeClosest public proxy for agent use; record the version you ran

Two context notes. Hugging Face's Open LLM Leaderboard, which for years was the default place to compare open models, was retired in March 2025. Its second version used IFEval, BBH, MATH Level 5, GPQA, MuSR and MMLU-Pro, and those tasks live on in lm-evaluation-harness under the leaderboard task group, so you can still reproduce that panel yourself. Second, a model card's numbers come from the vendor's own harness, prompts and settings. Treat them as claims to reproduce, not as measurements comparable with another vendor's card.

Advertisement

Floors, ceilings and the useful middle

A benchmark only discriminates where the models being compared score in its middle range. At the floor, scores near chance are dominated by guessing, so differences reflect answer-letter bias more than knowledge. GPQA for 1-3B models is the usual example. At the ceiling, the remaining errors are often mislabelled or ambiguous items, so a higher score can mean better memorisation of the test set rather than a better model.

The useful middle moves with model size and training. A practical rule: for each candidate set, discard any benchmark where every candidate is within its 95% interval of chance, or where every candidate is above about 90%. What remains are instruments that can actually rank this set of models. For one to four billion parameters in 2026, that typically leaves IFEval, MMLU-Pro, BBH, GSM8K for the smaller models, MATH for reasoning-tuned ones, and a function-calling suite if you build agents. See the SLM landscape for the model families themselves.

Formulation also moves where the middle is. Base models are often scored by log-likelihood over the answer options, while chat models are scored by generating an answer and parsing it. The same model can differ by many points between the two. Compare models only under one formulation, the one that matches how you will call them.

Running a reproducible panel

Use one harness, pin everything, and keep the per-item outputs. EleutherAI's lm-evaluation-harness is the common choice for open models. The command below scores a chat model through its chat template, with few-shot examples presented as turns, and logs every sample so you can audit extraction failures and run paired tests later:

pip install "lm_eval[hf]"       # pin the version you used in your results record

lm_eval --model hf \
  --model_args pretrained=your-org/candidate-3b-instruct,revision=COMMIT_SHA,dtype=bfloat16 \
  --tasks leaderboard_ifeval,leaderboard_mmlu_pro,gsm8k \
  --apply_chat_template --fewshot_as_multiturn \
  --batch_size auto --seed 1234 \
  --log_samples --output_path results/candidate-3b

Record the harness version, the model commit, dtype, the task list and versions the harness reports, the seed and the hardware. If you will ship a quantised model, run the panel on the quantised artifact as well. Quantisation can cost small models several points on reasoning tasks while leaving knowledge tasks nearly untouched, and you only see that by measuring both. For quick iteration, --limit runs a subset, but go back to the full sets for decisions; a 100-item subset has an interval of about plus or minus ten points.

Practical tests: the lane that decides

Public benchmarks screen for general capability. The decision belongs to practical tests built from your own task, which most small-model deployments narrow to a few skills: classifying, extracting fields into JSON, calling a tool, rewriting or summarising in a fixed style, and refusing what is out of scope.

Build the test set from real inputs: sample 300 to 1,000 items from logs or tickets, stratified so that rare but important cases appear, and label them carefully. Hold them out of any fine-tuning data. Then score each item with a rule, not a judge, wherever possible. Rule-based scoring is cheap, deterministic, and suits small models, whose failures are mostly mechanical: malformed JSON, a missing field, the wrong tool, an extra sentence. A tool-calling item can be scored in three tiers:

import json
import jsonschema

def score_tool_call(output: str, expected: dict, schema: dict) -> dict:
    """One practical test item: did the model emit a parseable, valid, correct call?"""
    try:
        call = json.loads(output)
    except json.JSONDecodeError:
        return {"parsed": 0, "valid": 0, "correct": 0}
    try:
        jsonschema.validate(call, schema)
    except jsonschema.ValidationError:
        return {"parsed": 1, "valid": 0, "correct": 0}
    correct = call.get("name") == expected["name"] and call.get("arguments") == expected["arguments"]
    return {"parsed": 1, "valid": 1, "correct": int(correct)}

The tiers tell you what to fix. Low parse rates call for constrained decoding, as in SLM function calling; valid-but-wrong calls call for better prompts, examples or fine-tuning. For subjective outputs such as summaries, use pairwise judging with a larger model, randomise the order of the two answers, and spot-check a sample by hand.

Finally, measure what the device imposes: time to first token, tokens per second, and peak memory on the target hardware with the shipped quantisation and context length. A model that wins on quality but misses the latency budget has not won.

Three lanes, three jobs: screen with public benchmarks, decide with task tests, ship after device checksCandidate models5-10 open SLMs, pinned revisionsLane 1: public panelIFEval, MMLU-Pro, GSM8K ...Shortlist2-3 models survivedrop the clear losersLane 2: task tests300-1,000 items from your trafficPaired comparisonbootstrap CI on the differenceLane 3: device checksquantised artifact, real hardwareDecisionquality, latency, memoryPublic scores answer "is this model broadly capable?"; only lanes 2 and 3 answer "will it do my job on my device?"
Public benchmarks shortlist, task tests decide, device checks confirm. Each lane answers a different question.

Worked example: picking a 3B model for ticket triage

A team needs a model that reads a support ticket and emits {"queue": ..., "priority": ..., "product": ...} on a CPU-only server. They start with six open models between 1.5B and 4B parameters. The figures below are illustrative, showing the reasoning rather than any real model's scores.

Lane 1: they run IFEval, MMLU-Pro and GSM8K with the command above. Two models trail the rest on IFEval by more than ten points and are dropped. GSM8K places all six within six points, mostly inside one another's intervals, so it does not influence the choice. One more model is dropped for an MMLU-Pro score near the bottom and a licence the legal team will not accept.

Lane 2: three models remain. The team labels 600 historical tickets and scores exact match on each field and on the whole JSON object. Model A scores 81.0% whole-object accuracy, model B 79.5%, and model C 72.3%. Unpaired intervals at n = 600 are about plus or minus 3.1 points, which would suggest A and B are tied. The paired bootstrap on the difference between A and B gives 1.5 points with an interval of roughly -0.6 to +3.6. So A is not shown to be better than B. C is clearly worse and is dropped.

Lane 3: on the target CPU at 4-bit, model B produces its JSON in about half the time of model A, because it is smaller and its tokenizer emits fewer tokens for the same output. With quality statistically tied, the team ships B. They add the 600 tickets as a CI regression gate for future fine-tunes, and plan to re-label 100 fresh tickets each quarter so the set tracks drift.

Failure modes

SymptomCauseFix
Rankings flip between runsDifferences within noise; prompt sensitivityPaired tests; several prompt variants; larger n
Card numbers not reproducibleDifferent harness, prompts, shots or formulationReproduce under your own pinned harness before comparing
Benchmark winner fails in productionBenchmark does not resemble the taskDecide on task tests; use public panels only to screen
Suspiciously high score on an old benchmarkContaminationPrefer newer or held-out sets; see the pitfalls guide
Quality drops after quantisingPanel ran on full-precision weightsEvaluate the shipped artifact
Generative scores far below likelihood scoresAnswer extraction failingRead logged samples; fix the parser or template

Trade-offs

Public benchmarks are cheap, comparable and broad, but they measure someone else's problem and age quickly as models train on them. Task tests are expensive to label and narrow, but they measure the decision you actually face. Rule-based scorers are fast and reproducible but only fit outputs with a checkable shape; model judges cover open-ended outputs at the cost of bias and variance. For classification-shaped tasks, also consider whether a fine-tuned encoder beats any generative model; small models for classification covers that choice.

What to do next

  1. List the three to five skills your product needs from the model, and write down the latency and memory budget on the target device.
  2. Pick a public panel of three or four benchmarks in the useful middle for your size class, and drop any that are at chance or saturated for every candidate.
  3. Run that panel with one pinned harness, with --log_samples, on the quantised artifacts you would ship.
  4. Build a 300 to 1,000 item task set from real inputs with rule-based scorers, and keep it out of training data.
  5. Compare finalists with a paired bootstrap on the same items; treat intervals that include zero as ties.
  6. Break ties on latency, memory and licence, then turn the task set into a regression gate for every new model or fine-tune.
Key takeaway: Benchmarks are instruments with a resolution set by their size: about plus or minus 6 points on GPQA Diamond's 198 items, under 1 point on MMLU-Pro. Use public benchmarks in the useful middle for your size class, such as IFEval, MMLU-Pro, BBH and GSM8K for smaller models, to screen candidates under one pinned harness. Then decide with a few hundred rule-scored items from your own task, compared pairwise with bootstrap intervals, and confirm latency and memory on the quantised artifact you will ship.