Most teams start evaluating LLMs with a notebook: a list of prompts, a loop, a few eyeballed answers. That stops working as soon as two people need to agree on whether a change helped. Then the questions become architectural. Where do test items live, and how are they versioned? What exactly identifies a run? Who scores outputs, and how is the scorer itself versioned? How do you decide that a difference is real? And how do production failures become tomorrow's test cases?

This article describes an evaluation system as a set of components with explicit contracts, following one record from dataset to release gate and back from production. Individual pieces have their own articles. A practical regression harness covers suite mechanics, and LLM-as-a-judge calibration covers model-graded scoring. Here the focus is how they fit together, with code for the schema, the runner and the statistics.

Advertisement

Three questions, one system

An evaluation system answers three different questions, and conflating them causes most of the confusion in practice. Capability: how good is this model or prompt at a task, in absolute terms? Public benchmarks and broad internal suites answer this, and they are used to choose between models. Regression: is candidate B worse than baseline A on anything we care about? This is a paired comparison on a fixed suite, and it gates releases. Production health: is the deployed system doing its job on real traffic right now? That needs online scoring of sampled traffic, since you have no gold answers.

These need different datasets, cadences and statistics. They should still share one item schema, one scorer library and one results store. When they don't, the offline judge and the production judge drift apart, and a score of 0.82 means different things on two dashboards.

The reference architecture

An LLM evaluation system: one record format from dataset to release gate to productionDataset registryversioned items, splitsRun specmodel, prompt, paramsRunnerfan-out, retries, cacheModel adaptersAPI, vLLM, localScorer pipelineexact, code, judgeResults storeper-item rowsAnalysis and gatepaired CI, thresholdsDashboards, triagediffs, slicesProduction trafficsampled, redactedOnline scorersjudge, user signalsLabelled failurespromoted into datasetspromptoutputscoresnew itemsOffline and online paths share the item schema, the scorers and the results store, so one number means one thing.
Components of an evaluation system. The offline path (top) and online path (bottom) share scorers and storage, and labelled production failures flow back into the registry.

Each component has one job and a narrow contract:

  • Dataset registry. Immutable, versioned sets of items with stable IDs and slice tags. A new version is a new name, such as support_qa@v14, never an in-place edit.
  • Run spec. Everything that determines an output: model identifier with digest, the full prompt template text, sampling parameters, dataset version and scorer versions. Its hash is the run key.
  • Runner. Fans items out to a model adapter with bounded concurrency, records every output and error, and caches outputs by run key and item ID so a rerun only pays for what changed.
  • Model adapters. One interface over hosted APIs, a vLLM server or a local checkpoint. Retries, rate limits and timeouts live here, not in the runner.
  • Scorer pipeline. Versioned functions from (item, output) to named scores, cheapest first.
  • Results store. One row per run key, item and scorer, with tags. Aggregates are computed from rows, never stored instead of them.
  • Analysis and gate. Paired comparisons, per-slice deltas and the thresholds that turn them into pass or fail.
Advertisement

The unit of record

The schema is the most important design decision, because every other component depends on it. Two rules carry most of the weight. First, an item's ID is stable across dataset versions, so you can compare a model's answer to the same question months apart. Second, a run is identified by a hash of its complete specification, including the prompt text. A prompt referred to by name will silently change under you.

import hashlib, json
from dataclasses import dataclass, field, asdict

@dataclass(frozen=True)
class EvalItem:
    item_id: str                 # stable across dataset versions
    input: dict                  # messages, tools, attachments
    reference: dict | None       # gold answer, rubric, or test cases
    tags: tuple = ()             # slice labels: "billing", "long_context", "de"

@dataclass(frozen=True)
class RunSpec:
    model: str                   # provider id or checkpoint path plus digest
    prompt_template: str         # full text, not a name
    params: dict = field(default_factory=dict)   # temperature, max_tokens, seed
    dataset: str = ""            # "support_qa@v14"
    scorers: tuple = ()          # ("exact@1", "judge_helpful@v3")

    def run_key(self) -> str:
        blob = json.dumps(asdict(self), sort_keys=True).encode()
        return hashlib.sha256(blob).hexdigest()[:16]

Store outputs verbatim, including errors and refusals, with latency and token counts. An item that failed with a timeout is a data point about the system. If the runner drops it, the candidate that times out more often looks better, because its hardest items vanish from the denominator.

The runner and its data flow

The runner is deliberately boring. It takes a spec and items, looks up cached outputs, generates the missing ones under a concurrency limit, scores everything and writes rows. Caching by run key means changing only a scorer re-scores existing outputs without paying for generation again. That matters, because generation is usually most of the cost.

import asyncio, time

async def run_eval(spec, items, adapter, scorers, store, concurrency=16):
    sem = asyncio.Semaphore(concurrency)
    key = spec.run_key()

    async def one(item):
        cached = store.get_output(key, item.item_id)
        if cached is None:
            async with sem:
                t0 = time.monotonic()
                try:
                    out = await adapter.generate(spec, item.input)   # retries live in the adapter
                    err = None
                except Exception as e:                                # record, don't drop
                    out, err = None, repr(e)
                cached = {"output": out, "error": err,
                          "latency_s": time.monotonic() - t0}
                store.put_output(key, item.item_id, cached)
        scores = {} if cached["error"] else {
            s.name: await s.score(item, cached["output"]) for s in scorers}
        store.put_scores(key, item.item_id, scores, item.tags)

    await asyncio.gather(*(one(i) for i in items))
    return key

Two choices in this sketch are deliberate. Scores are computed after caching, so scorer changes never invalidate outputs. Errors are stored as data and scored explicitly by the analysis layer, normally as failures. For sampled decoding, either fix a seed where the provider honours one or generate several samples per item and treat the item's score as the mean. Otherwise run-to-run noise hides real differences.

Scorers as a tiered pipeline

Order scorers by cost and trust. Deterministic checks, such as exact match, regex, JSON schema validity or refusal detection, are free and unambiguous, so run them on everything. Programmatic checks execute something: unit tests for generated code, SQL against a fixture database, a tool call replayed against a mock. Model-graded scorers handle open-ended quality with a rubric, and they are themselves models, so they need a version, a calibration set with human labels and a measured agreement rate. Human review is the most expensive tier. Spend it on disagreements between scorers and on a small random audit sample.

Treat each scorer as code with tests. A judge prompt edit is a new scorer version. Re-score the baseline with it before comparing, or you are measuring the judge change rather than the model change. Keep judge and candidate from the same model family apart where you can, since self-preference bias is well documented.

Deciding whether a difference is real

Two runs on the same items give paired data. Every item was answered by both systems, so the right quantity is the per-item difference, not two independent accuracies. Pairing removes the variance that comes from items being easy or hard, and it often narrows the interval considerably.

import numpy as np

def paired_bootstrap(base, cand, n_boot=10_000, seed=0):
    """base, cand: per-item scores aligned by item_id. Resample ITEMS, keep pairs together."""
    d = np.asarray(cand, float) - np.asarray(base, float)
    rng = np.random.default_rng(seed)
    idx = rng.integers(0, len(d), size=(n_boot, len(d)))
    means = d[idx].mean(axis=1)
    lo, hi = np.percentile(means, [2.5, 97.5])
    return d.mean(), lo, hi

Worked example. A suite has 400 items. The baseline scores 0.780 and the candidate 0.795. Per item, 34 flip from wrong to right, 28 flip from right to wrong, and 338 are unchanged. The mean difference is (34 - 28)/400 = 0.015. The per-item differences take values -1, 0 and +1, with variance 62/400 - 0.015^2, about 0.155, so the standard deviation is about 0.394 and the standard error is 0.394/20, about 0.0197. The 95 percent interval is roughly 0.015 plus or minus 0.039, or -0.024 to +0.054. The bootstrap gives almost the same answer. It includes zero, so this suite cannot tell whether the candidate is better.

Treating the runs as independent would give a standard error of about 0.029, half again as wide, so pairing helps. To detect a 1.5-point gain with the same flip rate, you need the standard error below about 0.0077, which means roughly 2,600 items. The practical lessons are to report intervals, look at the flips as well as the net score, and size suites for the effect you need to detect. For gates, test for non-inferiority. Here the lower bound is -2.4 points, so these 400 items can support "not worse by more than 2.5 points"; a 1-point margin, like a 1.5-point gain, needs a suite in the thousands.

From offline to online

Offline suites are fixed and labelled; production is neither. The architecture bridges them in stages. Shadow runs the candidate on mirrored live traffic without showing its output to anyone, and scores both systems with the online scorers. Canary serves a small share of users and compares guardrail metrics such as refusal rate, error rate, latency and judge scores. A/B testing measures what offline evaluation cannot: whether users complete tasks, retry, escalate or churn.

In steady state, sample production traffic, redact it according to your data policy, and score it with the same scorer versions used offline. Low scores, user complaints and escalations go to a triage queue. Once a human labels a failure, promote it into the registry as a new item in the next dataset version. That loop is what keeps a suite representative; without it, the suite measures last year's product. For agentic systems the same flow applies to whole trajectories, as described in agent evaluation at scale.

Dataset lifecycle and contamination

Split every dataset into a development set that prompt engineers may look at and a held-out set that only the gate reads. Otherwise the suite gets tuned to, one prompt tweak at a time. Rotate held-out items periodically, and retire items that every candidate answers correctly, since they add cost without signal.

Public benchmarks are likely to appear in pre-training data. Treat them as a capability signal and never as your only gate. For internal data, keep test items out of anything used for fine-tuning, and check with n-gram overlap against training sets before each release. Perplexity-based metrics have a place for base models, as explained in perplexity and evaluation metrics. For instruction-tuned systems, however, task-level scores on your own data matter more. AI evaluation frameworks covers public benchmark hygiene in more detail.

Operating the system

  • Cadence tiers. A smoke suite of 50 to 100 items runs on every prompt or code change in minutes. The full regression suite runs nightly and before every release. Capability suites run when you consider changing the model.
  • Cost control. Output caching, cheap scorers first and judge sampling on large suites keep cost roughly linear in what changed. Track cost per run as a first-class metric.
  • Pinning. Hosted model aliases move. Record the exact model version the provider returns, and alert when it changes under a fixed alias.
  • Slices. Report per-tag deltas next to the headline number. A flat average can hide a 10-point drop on one language or customer segment.
  • Ownership. Every dataset and scorer has an owner and a changelog, and gate thresholds change by review, not by whoever is blocked.

Failure modes

  • Unversioned prompts or judges. Scores move and nobody can say why. Hash everything that affects an output or a score.
  • Dropped errors. Timeouts and malformed outputs vanish from the denominator and flatter the worse system.
  • Single-sample noise. One sampled generation per item at temperature above zero makes identical systems look different.
  • Judge drift. A provider update to the judge model shifts every score. Re-run the judge's calibration set on a schedule.
  • Suite overfitting. Prompts tuned against the held-out set stop generalizing. Enforce the dev/held-out split in tooling.
  • Online-offline mismatch. Different scorers or preprocessing online and offline make production numbers incomparable with gate numbers.

Trade-offs

ChoiceGainCost
Model-graded vs deterministic scoringcovers open-ended qualityjudge bias, cost, need for calibration
Large suite vs small suitedetects smaller effectsslower gates, more labelling
Pairwise preference vs absolute scoresmore sensitive, fewer tiesno absolute level; N-way comparisons grow
Build vs adopt a frameworkexact fit to your record formatmaintenance; adopting gives faster start
Online scoring of all traffic vs samplingcatches rare failurescost and privacy exposure

What to do next

  1. Write down the three questions you need answered (capability, regression, production health) and the owner of each.
  2. Define the item and run-spec schemas, with stable item IDs and run keys that hash the full prompt text and parameters.
  3. Build or adopt a runner that caches outputs, records errors as data and separates generation from scoring.
  4. Order scorers from deterministic to human, version every one, and calibrate each judge against human labels.
  5. Gate releases on paired intervals and per-slice non-inferiority, and size the suite for the effect you need to detect.
  6. Sample and score production traffic with the same scorers, and promote labelled failures into the next dataset version every sprint.
Key takeaway: An LLM evaluation system works when it is built as a pipeline of versioned parts with narrow contracts: a registry of immutable datasets, a run spec whose hash identifies every output, a runner that caches and records errors, tiered and versioned scorers, and a results store of per-item rows. Decide on paired intervals, not raw averages, and size suites for the effect you need to detect. Use the same scorers online and offline, and feed labelled production failures back into the datasets, so the suite keeps measuring the product you actually ship.