What a regression harness is for

Public benchmarks measure general capability; a regression harness measures your product's contract. Its cases come from your traffic, your bug reports and your policies, and its scorers encode what your users consider a correct answer. That distinction drives every design decision. Coverage matters more than size, stability of the case set matters more than novelty, and the most valuable output is a diff of which cases changed outcome, not a single aggregate number.

Three kinds of cases usually belong in the set. Capability cases sample the real task distribution, stratified by intent or domain so each slice has enough examples to measure. Regression cases are frozen reproductions of past incidents, each of which must keep passing. Policy cases cover safety, refusal, privacy and formatting rules that are non-negotiable. Treating these separately in the gate, rather than averaging them together, is what makes the harness useful.

Advertisement

Case format and versioning

A case is data, not code. Store it in a reviewable text format with a stable id, the input messages or template variables, any fixture it depends on such as retrieved documents or tool responses, the scorers that apply, and tags for slicing. Hash the canonical serialization of each case and record the hash with every result, so a silently edited case can never be compared against results from its previous version.

# cases/refunds.yaml
- id: refund-policy-017
  tags: [billing, policy, must_pass]
  input:
    messages:
      - role: user
        content: "I was charged twice for order 8812. Can I get one refunded?"
  fixtures:
    tools:
      lookup_order: {order_id: "8812", charges: 2, status: "delivered"}
  expect:
    - scorer: tool_called
      args: {name: issue_refund, order_id: "8812"}
    - scorer: json_schema
      args: {schema: schemas/support_reply.json}
    - scorer: judge
      args: {rubric: rubrics/refund_tone.md, min_score: 4}

Fixtures matter as much as inputs. If a case depends on live retrieval or a live tool, its outcome changes whenever the index or the downstream service changes, and the harness starts reporting regressions that have nothing to do with the change under test. Freeze tool responses and retrieved passages for most cases, and keep a separate, clearly labelled end-to-end suite that exercises live dependencies.

Advertisement

Pinning the run

A result is only comparable if everything that produced it is recorded: model identifier and version, provider or engine version, prompt template version, decoding parameters, system prompt hash, tool schema hash and harness version. Write these into a run manifest and refuse to compare runs whose manifests differ in anything other than the variable under test.

Temperature zero does not make LLM output deterministic. Batch composition, kernel selection and floating-point reduction order can change which token wins a near tie, and hosted APIs add their own variation. Measure this directly: run the baseline configuration several times and record which cases flip between runs with no change at all. That flake rate is the noise floor, and any harness that ignores it will produce regressions on every pull request until people learn to ignore the gate.

A response cache keyed by model, rendered prompt and parameters makes the harness cheap enough to run on every change. When only a scorer or the gate logic changes, cached responses are re-scored without new model calls; when a prompt changes, only affected cases are regenerated. Include a switch that bypasses the cache for scheduled full runs, so cached answers never hide drift in a hosted model.

Scorers, cheapest first

Order scorers by cost and reliability. Deterministic checks come first: exact or normalized match for short answers, regular expressions for required phrases or forbidden content, JSON Schema validation for structured output, and assertions on which tools were called with which arguments. Executable checks come next: run generated code against tests, run generated SQL against a fixture database, or check a computed number within a tolerance. These are fast, stable and explainable.

Model-graded scorers handle what rules cannot, such as faithfulness to a source, helpfulness or tone. Treat the judge as a component with its own version and its own tests. Give it a narrow rubric with anchored score levels, ask for a short justification before the score, and validate it against a few hundred human labels before trusting it in a gate. Pin the judge model; upgrading it changes scores for the baseline and candidate alike, but it invalidates comparisons with historical runs.

SCORERS = {}

def scorer(name):
    def register(fn):
        SCORERS[name] = fn
        return fn
    return register

@scorer("json_schema")
def json_schema(output, args, ctx):
    import json, jsonschema
    try:
        jsonschema.validate(json.loads(output.text), ctx.load(args["schema"]))
        return 1.0, "valid"
    except Exception as exc:          # parse or validation error
        return 0.0, f"invalid: {exc}"[:200]

@scorer("tool_called")
def tool_called(output, args, ctx):
    want = {k: v for k, v in args.items() if k != "name"}
    for call in output.tool_calls:
        if call.name == args["name"] and all(call.args.get(k) == v for k, v in want.items()):
            return 1.0, "called"
    return 0.0, f"missing {args['name']}"

Every scorer returns a score and a short reason. The reason is what makes a failing report actionable: a reviewer should be able to read why a case failed without re-running anything.

Comparing runs with statistics that respect noise

Because baseline and candidate are run on the same cases, the right comparison is paired. For pass or fail scores, count the cases that flipped in each direction; the discordant pairs carry all the information, and McNemar's test on them is more sensitive than comparing two independent pass rates. For graded scores, bootstrap the mean per-case difference by resampling cases, and report the confidence interval, not just the point estimate.

Set expectations about sample size before building the gate. With 200 cases and a pass rate around 80 percent, the standard error of a single pass rate is about 2.8 points, so its 95 percent interval is roughly plus or minus 5.5 points. Pairing narrows the interval for the difference considerably when most cases do not change, but a slice with 30 cases still cannot reliably detect a two-point regression. Either grow the slices that matter or accept that the gate can only catch large changes in them.

import random

def paired_bootstrap(base, cand, iters=5000, seed=7):
    """base/cand: dict case_id -> score in [0, 1]; returns mean diff and 95% CI."""
    ids = sorted(base.keys() & cand.keys())
    diffs = [cand[i] - base[i] for i in ids]
    rng = random.Random(seed)
    means = sorted(
        sum(diffs[rng.randrange(len(diffs))] for _ in diffs) / len(diffs)
        for _ in range(iters)
    )
    return sum(diffs) / len(diffs), (means[int(0.025 * iters)], means[int(0.975 * iters)])

def flips(base, cand, threshold=0.5):
    fixed = [i for i in base if base[i] < threshold <= cand.get(i, 0)]
    broken = [i for i in base if cand.get(i, 0) < threshold <= base[i]]
    return fixed, broken
Case registryversioned, content-hashedRunnerpinned model + paramsResponse cachekey: model, prompt, paramscaseslookupScorersexact / schema / exec / judgeoutputsResults storeper-case, per-run rowsComparatorpaired diff + intervalsbaseline + candidateCI gateblock / warn / passReportslices, flips, examples
Regression harness data flow: versioned cases run against a pinned configuration through a response cache, layered scorers write per-case results, and a paired comparator against the baseline feeds the CI gate and the report.

Designing the gate

A useful gate has tiers. Must-pass cases, including every frozen incident and every policy case, block the merge if any of them fails, subject to the measured flake list. Critical slices block if the upper bound of the confidence interval on their difference falls below the tolerated regression, for example minus two points. Everything else produces a warning and a report. A single threshold on the global average is the classic mistake: a large gain on an easy, well-populated slice can hide a serious regression on a small, important one.

A worked example shows why. Suppose a prompt change is evaluated on 600 cases: 450 general questions, 100 billing questions and 50 policy cases. The candidate fixes 18 general cases and breaks 3. But 8 of the 100 billing cases flip from pass to fail with none fixed, and one frozen incident case in the policy set now fails. Net, 18 fixes against 12 breaks, the global pass rate still rises by about 1.0 point, and a global threshold would approve it. The tiered gate blocks the change on the must-pass failure alone, and the billing slice would have blocked it anyway: a paired bootstrap puts its upper bound near minus three points. The author sees twelve broken cases, led by the nine in blocking tiers, instead of a misleadingly positive average.

The report should lead with flips, not averages. List each case that went from pass to fail with its input, both outputs and the scorer's reason, then the cases that were fixed, then slice-level intervals. Reviewers make better decisions from ten concrete broken examples than from a table of means, and those examples are also the fastest route to a fix.

Multi-turn and agent cases

Single-turn question and answer cases are the easy part. Agents take several steps, call tools, and reach an end state, and a regression can hide in any of them. Score the trajectory as well as the final answer: which tools were called, in what order, with what arguments, how many steps it took, and whether any forbidden action was attempted. A candidate that reaches the right answer in eleven tool calls instead of four is a regression in cost and latency even if the answer scorer is satisfied.

Replay is the key technique. Record tool responses from a reference run and serve them from fixtures keyed by tool name and normalized arguments, so the agent sees a deterministic world. When the candidate makes a call the fixtures do not cover, fail the case with a clear reason rather than falling through to a live service; an unexpected call is itself a behavioral change worth reviewing. For conversations, fix the user turns in the case rather than simulating the user with another model, unless the suite is explicitly a simulation suite with its own noise measurements. Simulated users add a second source of variance that makes paired comparison much weaker.

Keeping the suite honest

  • Label errors. A fraction of expected answers in any hand-built set are wrong. Review every case that fails on both baseline and candidate; many are bad labels, not model failures.
  • Saturation. Cases that every configuration passes no longer discriminate. Keep them as a smoke test, but add harder cases from recent traffic so the suite keeps measuring something.
  • Overfitting to the suite. Prompt changes tuned until the harness is green can overfit it. Hold out a rotating subset that authors cannot see and check that gains transfer.
  • Judge drift. A hosted judge model that changes underneath you shifts every graded score. Pin versions and re-validate against human labels on any change.
  • Contamination. Cases published in documentation or examples can leak into training data for future models, inflating scores. Keep the regression set private.
  • Fixture rot. Frozen tool responses drift away from what production services now return. Refresh fixtures deliberately, as a reviewed change with its own baseline run.