A 1B to 8B parameter model is more sensitive to how you measure it than a frontier model. Change the prompt template, the way multiple-choice answers are scored, or the regular expression that pulls out the final answer, and a small model's score can move by more than the improvement you were trying to measure. The result is familiar: a fine-tune that looked like a win on the dashboard, a model card number nobody can reproduce, or a quantized build that passed evaluation on a GPU server and fails on the phone.

This article is a field guide to those failure mechanisms. Each pitfall comes with how to reproduce it, how to detect it and how to fix it. The architecture of a full evaluation system, with dataset registries, run configs and CI gates, is covered in SLM evaluation architecture; here the focus is on the specific ways numbers go wrong, so that you can audit a score before you trust it.

Advertisement

Why small models amplify evaluation choices

Large models are robust to format: they follow almost any reasonable instruction, answer in the requested shape and ignore stray whitespace. Small models have less capacity to spare. They are more likely to answer in the wrong format, to follow a few-shot pattern too literally, to run past the token limit, or to be thrown off by a template they were not trained on. So the measurement pipeline contributes a larger share of the final number, and two teams evaluating the same checkpoint with different harness settings can report scores several points apart.

Where a small-model score goes wrong: every stage between the dataset and the numberEval setitems, labelsPrompt buildtemplate, shotsArtifactweights, runtimeDecode / scoreloglik or generateExtractparse the answerAggregatemean, intervalContaminationseen in trainingFormattemplate, BOS, shotsWrong artifactfp16 vs shipped int4Formulationcloze vs lettersTruncationmax tokens hitNoisen too smallBlue: pipeline stages. Red: the pitfall that enters at each stage. Each one can move a small model's score more than the change you are testing.
Each stage of the evaluation pipeline, with the pitfall that enters there.

Pitfall 1: the scoring formulation

Multiple-choice benchmarks can be scored two ways. In the cloze formulation, the model sees the question, and each answer text is scored by the log-probability of generating it as a continuation; the highest wins. In the multiple-choice formulation, the answers are listed with letters and the model is scored on the probability of the correct letter. The OLMES evaluation standard from AI2 documents that smaller base models often need the cloze formulation because they have not learned to map letters to answers, while larger models do better with letters. Report a small base model with letter scoring and you may be measuring whether it understands the answer format, not whether it knows the answer.

Cloze scoring has its own trap: longer answers accumulate more negative log-probability, so raw scores favour short options. Length normalisation divides by answer length, which is what lm-evaluation-harness reports as acc_norm alongside the raw acc. The two can differ substantially on the same run, so a comparison that mixes them is meaningless.

import torch

@torch.no_grad()
def continuation_logprob(model, tok, context, continuation):
    """Sum of log-probabilities of `continuation` given `context`."""
    ctx = tok(context, add_special_tokens=True).input_ids
    cont = tok(continuation, add_special_tokens=False).input_ids   # tokenise separately:
    ids = torch.tensor([ctx + cont], device=model.device)           # no merge across the boundary
    logprobs = torch.log_softmax(model(ids).logits[0, :-1].float(), dim=-1)
    targets = ids[0, 1:]
    per_token = logprobs[torch.arange(len(targets)), targets]
    return per_token[-len(cont):].sum().item()

def pick(model, tok, context, choices, normalise=False):
    scores = []
    for choice in choices:
        lp = continuation_logprob(model, tok, context, " " + choice)  # leading space: " Paris", not "Paris"
        scores.append(lp / len(choice) if normalise else lp)          # per-character normalisation
    return max(range(len(choices)), key=scores.__getitem__)

Two details in that code are regular sources of bugs. The continuation is tokenised separately and appended, so the tokenizer cannot merge the last context token with the first answer token and shift the boundary. And the leading space matters: for most BPE vocabularies, a word with a leading space is a different token from the same word without one, and scoring the wrong variant lowers every option's probability unevenly.

Advertisement

Pitfall 2: chat templates and special tokens

Instruction-tuned models were trained on text wrapped in a specific template: role markers, special tokens and a generation prompt that tells the model it is its turn. Evaluate an instruct model on raw text and you are testing it off-distribution. Evaluate a base model with a chat template it never saw and you do the same in reverse. Either can cost many points on a small model.

The subtle version is the duplicate beginning-of-sequence token. Many templates already include the BOS token, and many tokenizers add it again by default. The model then sees two BOS tokens, a sequence it never saw in training. Nothing crashes; scores just drop. Render the template once, tokenise without special tokens and assert:

def build_prompt(tok, user_text):
    msgs = [{"role": "user", "content": user_text}]
    text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
    ids = tok(text, add_special_tokens=False).input_ids    # the template already added any BOS
    if tok.bos_token_id is not None and ids.count(tok.bos_token_id) > 1:
        raise ValueError("duplicate BOS: template and tokenizer both added one")
    return ids

Pin the template by hashing the rendered text of one fixed example and storing the hash in the run record. Template files change between model revisions, and a harness that silently picks up a new one will move scores for no reason you can see. For how one family's format is laid out, see the Llama chat template.

Pitfall 3: answer extraction and truncation

Generative tasks such as maths word problems are scored by pulling a final answer out of free text. Small models are inconsistent about format: one output ends with the requested marker, the next writes 'The answer is 42.', another stops mid-sentence at the token limit. A strict regular expression then scores correct reasoning as wrong, and a loose one grabs an intermediate number. The fix is to extract in documented tiers and to count what happened:

import re

NUM = r"-?\d[\d,]*(?:\.\d+)?"

def extract_answer(text):
    """Strict first, then a documented fallback. Returns (value, how)."""
    m = re.search(r"####\s*(" + NUM + ")", text)
    if m:
        return m.group(1).replace(",", ""), "strict"
    nums = re.findall(NUM, text)
    if nums:
        return nums[-1].replace(",", ""), "fallback-last-number"
    return None, "no-answer"

def score(outputs, golds):
    stats = {"correct": 0, "strict": 0, "fallback-last-number": 0, "no-answer": 0, "truncated": 0}
    for out, gold in zip(outputs, golds):
        value, how = extract_answer(out["text"])
        stats[how] += 1
        stats["truncated"] += out["finish_reason"] == "length"
        stats["correct"] += value is not None and float(value) == float(gold)
    return stats          # report every counter, not only accuracy

If the no-answer or truncated counters are more than a few percent, the score is measuring format and budget, not ability. Raise the generation limit for reasoning tasks, check that the stop sequences do not fire inside the answer, and when format compliance is itself what you need, measure it separately or enforce it with guided decoding rather than burying it in accuracy.

Pitfall 4: prompt and few-shot sensitivity

The number of few-shot examples, their order, which examples were chosen, and the wording of the instruction all move small-model scores. A single prompt gives one draw from that distribution. Before comparing two models, run each on a handful of prompt variants, such as three instructions and two example orderings, and report the spread as well as the mean. If the spread across prompts is wider than the gap between models, the comparison is not telling you much.

Keep the prompt identical across the models you compare, including shot count and ordering, and record it in full. A common error is comparing your model's zero-shot score with a published five-shot number. Another is tuning the prompt on the test set until your model wins, which is just overfitting done by hand.

Pitfall 5: contamination

Popular benchmarks leak into web-scale pre-training data, and fine-tuning sets built by scraping or by distilling from larger models can contain test items or close paraphrases. The symptom is a benchmark score that jumps after adding a dataset while held-out tasks of the same kind do not move. A basic check is n-gram overlap between each evaluation item and the training data. Long n-grams, such as the 13-grams used in the GPT-3 contamination analysis, avoid flagging common phrases:

def ngrams(tokens, n=13):
    return {tuple(tokens[i:i + n]) for i in range(len(tokens) - n + 1)}

def build_index(training_docs, n=13):
    index = set()
    for doc in training_docs:                 # stream shards; use a Bloom filter at scale
        index |= ngrams(doc.lower().split(), n)
    return index

def flag_items(eval_items, index, n=13):
    return [i for i, item in enumerate(eval_items)
            if ngrams(item.lower().split(), n) & index]

Overlap checks miss paraphrases and translations, so pair them with a private held-out set written after the training data cutoff, in your own domain, that never leaves your storage. When a public score and the private score disagree, trust the private one. Researchers who built fresh versions of popular benchmarks have found some model families scoring noticeably lower on the new items than on the originals, which is what contamination looks like from the outside.

Pitfall 6: evaluating an artifact you do not ship

The model on the device is rarely the model that was evaluated. It may be quantized to 4 bits, served by a different runtime with its own tokenizer implementation and sampling defaults, limited to a shorter context, and run with temperature instead of greedy decoding. Each difference can change outputs. Quantization in particular tends to hurt some capabilities, such as arithmetic and long-context retrieval, more than others, so an unchanged average can hide a broken slice.

Evaluate the exact file you ship, through the runtime you ship, with the decoding settings you ship, at the context lengths your users send. Compare it with the full-precision model on the same items and look at the per-item disagreements, not just the average. The measurement side of this is covered in evaluating quantized models.

Pitfall 7: reading noise as signal

An accuracy measured on n items has a standard error of about the square root of p(1 - p) / n. On the 1,319 problems of the GSM8K test set at 50 percent accuracy, the 95 percent interval is about plus or minus 2.7 points. On a 200-item custom set at 70 percent it is about plus or minus 6.4 points. A 2-point gain on the small set is not evidence of anything. Because both models answer the same items, a paired comparison over per-item differences is far more sensitive than comparing two independent intervals; the paired bootstrap in the architecture article linked above does this. Fine-tuning seeds add a second source of variance, so a result that matters should survive a rerun with a different seed.

A worked audit

Consider a hypothetical report: a team's fine-tuned 3B model beats its base model by 4 points on a 200-item internal maths set, 66 to 70 percent. Walk the pipeline before accepting it. The interval on each score is roughly plus or minus 6 points, so the headline gap needs a paired test. The extraction counters show 9 percent of base-model outputs hit the token limit, against 2 percent for the fine-tune, which was trained on shorter solutions: part of the gain is fewer truncations, which a higher limit would remove. The base model was evaluated without its chat template while the fine-tune used one. And an n-gram check finds that 11 items overlap the fine-tuning data, which was distilled from a larger model prompted with similar problems.

After rerunning with the same template for both, a higher token limit and the overlapping items removed, the honest statement is whatever the paired interval then says, and it may well be that the difference is not distinguishable from zero. The audit took an afternoon, and it replaced a claim with a measurement.

Trade-offs

Cloze scoring is fairer to small base models but does not reflect how an assistant is used; generative evaluation reflects use but adds extraction noise. Strict extraction is reproducible but penalises harmless format drift; lenient extraction hides format failures users will see. Larger evaluation sets narrow intervals but cost labelling time. Model-based judges scale well but bring position and length biases of their own, described in LLM-as-judge calibration. Choose deliberately, write the choice into the run record, and never compare numbers made under different choices.

What to do next

  1. For every score you report, record the formulation (cloze or letters), normalisation, shot count, template hash, decoding settings and token limit.
  2. Add the duplicate-BOS assertion and the extraction counters to your harness and fail runs where no-answer or truncation exceeds a threshold you set.
  3. Run each comparison on at least three prompt variants and report the spread alongside the mean.
  4. Build an n-gram index of your fine-tuning data and check every evaluation set against it; keep a private held-out set that never leaves your storage.
  5. Evaluate the shipped quantized file through the shipped runtime, and diff its per-item results against full precision.
  6. Compute an interval or a paired test for every claimed improvement, and rerun important results with a second seed.
Key takeaway: Small models are sensitive enough that the evaluation pipeline can move their scores more than the change being tested. Pick the scoring formulation that suits the model and report normalisation, render chat templates exactly once and check for duplicate BOS tokens, count extraction failures and truncations instead of hiding them in accuracy, measure prompt spread, check for contamination against both n-gram overlap and a private held-out set, evaluate the artifact you actually ship, and put an interval on every claimed gain.