Small language models are rarely asked to be good at everything. A 1B to 8B parameter model is usually chosen for one job: classify tickets, extract fields, call tools, answer from a fixed knowledge base, run on a phone. That changes what evaluation means. A public leaderboard average says little about whether your fine-tuned, quantized build still extracts invoice totals correctly on the device you ship it to.
This article treats evaluation as a system with its own architecture: a versioned dataset registry, a run configuration that pins everything that changes outputs, scorers matched to task types, statistics that separate real changes from noise, a regression gate wired into CI, and a separate lane that tests the quantized artifact on target hardware. It includes runnable code for the statistical core, a worked example, and the failure modes that make SLM scores lie.
Why small models need their own evaluation architecture
Four properties of SLMs make generic evaluation misleading. First, narrowness: a small model fine-tuned for a task can lose general ability while getting better at the task, so a general benchmark can fall while the product improves, or rise while it degrades. Second, sensitivity: small models are more sensitive to prompt wording, few-shot examples and especially the chat template; the same weights with the wrong template can look broken. Third, the deployed artifact is different from the trained one: quantization to 4 or 8 bits, a different runtime and a shorter context limit all change outputs. Fourth, cheap iteration: you may train dozens of variants a week, so evaluation must be automatic and fast, and its verdicts must be trustworthy enough to block a release.
The answer is to evaluate the product task first and general capability second, to pin every input that can change an output, and to test the thing you ship, not a close relative of it.
Layer 1: the dataset registry
The most valuable evaluation set is built from your own traffic. Sample real inputs, have people label the expected outputs, and tag each example with slices that matter to the business: language, customer tier, input length bucket, document type, whether the input is adversarial. A few hundred well-labelled examples per important slice beat ten thousand scraped ones. Add a small set of hard cases that previously failed in production; each incident should leave behind at least one regression example.
Store datasets as immutable, versioned files with a content hash, and never edit a released version in place. When labels are corrected, publish a new version and re-run the baseline on it, otherwise score changes mix model changes with label changes. Keep a held-out split that no one tunes prompts against, and run an n-gram overlap check between every training dataset and every evaluation set before each training run, because contamination of a small task set is easy and silently inflates scores.
{"id": "inv-00412", "input": "Invoice #88-231 ... Total due: EUR 1,240.50",
"expected": {"invoice_id": "88-231", "total": "1240.50", "currency": "EUR"},
"slices": ["lang:en", "doc:invoice", "len:short"],
"source": "prod-sample-2026-08", "label_version": 3}
Layer 2: a run configuration that pins everything
An evaluation result is only meaningful with the full list of things that produced it. Record, and hash into the run id: the weights file hash (not a model name, because names get reused), the quantization format and tool version, the runtime and its version, the chat template text, the system prompt, few-shot examples, decoding parameters (temperature, top-p, maximum new tokens, stop sequences), any constrained decoding grammar, the dataset version and the scorer versions. Use greedy decoding for regression tests so reruns are comparable; test sampled decoding separately if the product samples.
For public capability benchmarks, the EleutherAI lm-evaluation-harness is the common runner. It loads Hugging Face models directly or talks to a local OpenAI-compatible server or vLLM, and it can log every sample so you can inspect failures rather than trusting a single number:
lm_eval --model hf \
--model_args pretrained=./checkpoints/support-3b-v7,dtype=bfloat16 \
--tasks hellaswag,arc_easy \
--num_fewshot 0 \
--batch_size 8 \
--output_path ./eval_runs/support-3b-v7 \
--log_samplesTreat these benchmarks as a canary for general damage, such as catastrophic forgetting after fine-tuning, not as the release criterion. Your task suite is the release criterion. Check your installed version's help output for chat-template options before comparing an instruction-tuned model, because scoring a chat model without its template understates it.
Layer 3: scorers matched to the task
Pick the cheapest scorer that measures what users care about. For structured tasks, which is most SLM work, programmatic scoring is both cheapest and most reliable: parse the output, validate it against the schema, then compare field by field so you can report that currency is always right but totals fail on European number formats. For tool calling, check the function name, argument validity and argument values separately, and execute calls against a sandbox when possible; the function calling article covers what those outputs look like.
| Scorer | Good for | Watch out for |
|---|---|---|
| Exact or normalised match | Classification, extraction of fixed fields, short answers | Formatting noise; normalise case, whitespace and number formats before comparing |
| Schema and program checks | JSON output, tool calls, SQL that must parse, code that must pass tests | Valid is not correct; pair with a field-level check |
| Log-likelihood choice | Multiple-choice public benchmarks | Measures ranking of options, not whether the model generates the answer |
| Model judge | Open-ended answers, summaries, tone | Position and length bias, self-preference; must be calibrated against human labels |
| Human review | Ground truth for the judge, launch decisions | Slow and expensive; use blind, paired comparisons on a sample |
Model judges are useful for open-ended outputs but are instruments that need calibration. Have humans label a few hundred outputs, run the judge on the same outputs, and measure agreement before trusting it. Randomise the order of the two answers in pairwise judging to cancel position bias, cap or normalise length, use a judge from a different model family than the candidate where you can, and version the judge prompt and model like any other dependency. Re-run calibration whenever the judge changes.
Layer 4: statistics that separate signal from noise
With a few hundred examples per slice, a one or two point difference is often noise. Because the baseline and the candidate answer the same examples, use a paired comparison: compute the per-example difference in score and bootstrap over examples. This is far more sensitive than comparing two independent accuracies, because per-example difficulty cancels out.
import random
def paired_bootstrap(base, cand, iters=10_000, seed=0):
"""base, cand: per-example scores (0/1 or 0..1) on the same examples."""
assert len(base) == len(cand) and base
rng = random.Random(seed)
diffs = [c - b for b, c in zip(base, cand)]
n = len(diffs)
mean = sum(diffs) / n
samples = []
for _ in range(iters):
s = sum(diffs[rng.randrange(n)] for _ in range(n)) / n
samples.append(s)
samples.sort()
lo, hi = samples[int(0.025 * iters)], samples[int(0.975 * iters)]
return mean, lo, hi
def verdict(mean, lo, hi, max_drop=0.02):
if lo > 0:
return "improved"
if hi < 0 and -mean > max_drop:
return "regressed"
if hi < 0:
return "minor-regression" # real drop, but within tolerance
if lo < -max_drop:
return "possible-regression" # interval allows a drop larger than tolerated
return "no-significant-change"Report the interval, not only the mean, and refuse to compute a verdict for slices below a minimum size (for example 50 examples); report them as underpowered instead. When you compare many slices at once, some will look significant by chance, so treat a single marginal slice as a prompt to look at examples rather than an automatic block.
Layer 5: regression gates in CI
Wire the suite into the training and release pipeline. Every candidate checkpoint runs the task suite against the current production baseline. The gate reads the per-slice verdicts and applies explicit rules: block if any critical slice regressed; warn if a non-critical slice shows a possible regression; require schema validity at or above a fixed floor, because an unparseable output is an outage, not a quality dip. The gate's output should include a diff of the examples whose scores changed, since a reviewer looking at twenty flipped examples learns more than from any aggregate.
Keep two budgets: a fast suite of a few hundred examples that runs on every change in minutes, and a full suite that runs nightly and before release. Store every run's raw outputs so a later question such as when did the model start dropping currencies can be answered by querying history instead of retraining old checkpoints.
The device lane: test the artifact you ship
Quantization and runtime changes alter outputs in ways that average metrics hide. Build the exact artifact that ships, for example a 4-bit GGUF file for llama.cpp or a vendor NPU package, and run the task suite on it on representative hardware. Measure three things. Quality: the same scorers and paired statistics against the full-precision build. Agreement: the fraction of examples where the quantized build produces the same output as the reference under greedy decoding; a drop in agreement concentrated in one slice, such as long inputs or numbers, pinpoints where quantization hurts. Resources: time to first token, tokens per second, peak memory and behaviour under sustained load, since phones throttle when hot and a model that is fast for ten seconds may not be fast for a minute.
Include inputs at the context limit you advertise. Quantization error and small KV cache budgets often show up only on long inputs, which are exactly the ones a short-example suite skips.
Worked example: shipping a support-ticket model
A team fine-tunes a 3B model to classify support tickets into 14 queues and extract order ids, then quantizes it to 4 bits for an on-premises CPU server. The numbers here are illustrative. The task suite has 1,200 labelled tickets in six slices. Version 7 against the production version 6 shows queue accuracy up 2.1 points overall with a paired interval of +1.2 to +3.0: a real improvement. But the German slice (160 examples) shows -4.4 with an interval of -8.1 to -0.6, a regression, and the gate blocks.
The example diff shows the cause: the new fine-tuning mix contained almost no German tickets, and the model now routes German billing tickets to general support. The team adds 400 German examples, retrains, and version 7b passes every slice. On the device lane, the 4-bit build agrees with the bf16 build on 97 percent of classifications but only 88 percent of extracted order ids, all failures on ids longer than twelve characters. Switching to a grammar-constrained id field and an 8-bit format for the output layer brings agreement to 99 percent, and the release ships with the long-id examples added to the suite permanently.
Failure modes
| Failure | Symptom | Guard |
|---|---|---|
| Template mismatch | A fine-tuned model scores far below its training loss suggests | Pin and hash the chat template in the run config |
| Contamination | Public benchmark jumps after adding a web-scraped dataset | N-gram overlap check between training data and every eval set |
| Noise read as signal | A 1-point gain reverses on rerun | Paired bootstrap intervals; minimum slice sizes |
| Judge drift | Scores move when only the judge model changed | Version the judge; re-run calibration on every judge change |
| Quantization blind spot | Server eval passes, device build fails on long inputs | Device lane on the shipped artifact with long-context slices |
| Aggregate hides slice | Overall up, one language or customer down sharply | Gate per slice, not only on the mean |
Trade-offs
Every choice here trades cost for trust. Bigger task sets narrow intervals but cost labelling time; start with a few hundred per critical slice and grow where decisions are close. Model judges scale cheaply but add a dependency that can drift; programmatic checks are preferable wherever outputs can be structured, which is one reason to design SLM tasks with structured outputs in the first place. Greedy decoding makes regressions reproducible but hides sampling-only failures. Gating per slice catches silent harm to minority users but produces more false alarms; the answer is a tiered policy, not a looser gate.
What to do next
- Sample 300 or more real inputs per critical slice, label them, and publish them as a hashed, versioned dataset with a held-out split.
- Define the run configuration record, including the chat template hash and quantization details, and refuse to report scores without it.
- Implement programmatic, field-level scorers for every structured task before adding any model judge.
- Add the paired bootstrap to your comparison report and set per-slice gate rules with a minimum slice size.
- Run a public benchmark such as hellaswag through lm-evaluation-harness after each fine-tune as a forgetting canary.
- Build the shipped quantized artifact in CI and run the suite plus latency and memory checks on target hardware, including maximum-length inputs.
- After every production incident, add the failing inputs to the regression set.