Public leaderboards answer a question nobody on your team is asking: how does a model do on somebody else's tasks, with somebody else's prompts, scored by somebody else's rules. The question you actually face is narrower and more expensive to get wrong. Can we move from this model to that one, from BF16 weights to FP8, from one serving engine to another, or from four GPUs to two, without our users noticing a drop in quality? A custom benchmark is the instrument that answers it.

This article treats a custom benchmark as a measurement device with three parts: a frozen, versioned item set built from your own traffic; a pinned runner that executes every candidate on the same items on GPUs; and a statistics layer that compares candidates item by item and knows how much noise the hardware itself introduces. It covers how many items you need, why temperature zero is not deterministic on GPUs, and ends with a checklist.

What a custom benchmark is for

A benchmark exists to make a decision repeatable, so write the decision down first: swap model A for B, quantise, upgrade the engine. Each decision has a guard rail, for example no more than a two-point drop in task success on any major task type, and a reward, such as 35 percent lower cost per request.

That framing settles several design choices at once. The items must represent the traffic the decision affects, not the traffic that is easy to score. The metric must be per item so that two candidates can be compared on the same inputs. And the size of the set follows from the smallest difference the guard rail needs you to detect, not from a round number.

It is not a load test: throughput and latency under concurrency are measured separately. It measures the quality of outputs, item by item, for a fixed set of candidates.

The architecture

The diagram shows the moving parts. Items come from two sources: a sample of production requests, scrubbed and de-duplicated, and items written by people who know where the system fails. They are merged into a numbered, hashed version of the set that never changes once published. Each candidate is a fully pinned configuration: model revision, tokenizer, chat template, quantisation, engine version, parallelism and sampling parameters. The runner produces one output row per item per candidate. Scorers turn outputs into numbers, and the statistics layer compares candidates on matching item IDs.

A custom benchmark is a frozen item set, a pinned runner and paired statisticsProduction logssampled, scrubbedExpert-written itemsedge cases, policyItem set v7stratified, hashedCandidate ABF16, engine pinnedCandidate BFP8 weightsCandidate Cother modelOutputsone row per itemScorersprogram first, judge lastper-item scoresPaired statisticsbootstrap, noise floorDecision reportquality, cost, latencyEverything left of the candidates is versioned once; everything right of them is recomputed for every run.The same item IDs flow through every candidate, which is what makes paired comparison possible.
Items are versioned once and shared by every candidate; scoring and statistics are recomputed on each run.

Because outputs are stored and scoring is separate from generation, a scorer fix or a new metric is a rescore, not another GPU run.

Building the item set

Start from logs. Sample requests across a few weeks so that weekly cycles and recent product changes are both represented. Remove personal data, then de-duplicate: near-identical requests inflate the apparent sample size without adding information. Normalise whitespace and numbers, hash, and keep one item per hash.

Then stratify. Label each item with a task type, such as extraction, summarisation, tool call or free-form answer, and a difficulty band. Report results per stratum as well as overall. A quantised model that loses six points on long multi-step tool calls but nothing elsewhere can look fine in a pooled average if tool calls are 10 percent of traffic. If a stratum matters to the decision, give it enough items to be measured on its own.

Add a hand-written slice of 50 to 200 rare, expensive cases, such as policy edges and malformed inputs, each with a reference answer or a clear pass condition.

Finally, freeze and version. Store each item with a stable ID, its stratum labels, a reference or check, and its source, and publish it with a content hash. Keep a development split for prompt tuning and a held-out split that only the release process reads. Refresh items as a new version, never by editing the old one.

Scoring: programs first, judges last

Prefer scorers that are programs. If the output must be JSON, validate it against the schema. If it must call a tool, parse the call and compare the arguments with the expected ones. If it is a number, extract it and compare with a tolerance. If it is code, run the tests in a sandbox. Programmatic scorers are cheap, deterministic and easy to argue about.

Use reference-based scoring for short factual answers: normalise case and punctuation, then compare, or check that every required fact appears. Use a model as judge only for qualities programs cannot check, such as whether a summary is faithful. Then give the judge an explicit rubric, pin its model and prompt, calibrate it against a few hundred human labels, and never use a candidate as its own judge.

Every scorer returns a value between 0 and 1 for one item, plus a reason string. The registry pattern below keeps scoring separate from generation and makes it easy to rescore stored outputs.

import json, re
from jsonschema import validate, ValidationError

SCORERS = {}

def scorer(name):
    def reg(fn):
        SCORERS[name] = fn
        return fn
    return reg

@scorer("json_schema")
def json_schema(item, output):
    try:
        validate(json.loads(output), item["schema"])
        return 1.0, "valid"
    except (json.JSONDecodeError, ValidationError) as e:
        return 0.0, f"invalid: {str(e)[:80]}"

@scorer("numeric")
def numeric(item, output):
    m = re.search(r"-?\d+(?:\.\d+)?", output.replace(",", ""))
    if not m:
        return 0.0, "no number"
    ok = abs(float(m.group()) - item["answer"]) <= item.get("tol", 1e-6)
    return float(ok), m.group()

@scorer("required_facts")
def required_facts(item, output):
    text = output.lower()
    hits = [f for f in item["facts"] if f.lower() in text]
    return len(hits) / len(item["facts"]), f"{len(hits)}/{len(item['facts'])}"

def score_run(items, outputs):
    rows = []
    for it in items:
        s, why = SCORERS[it["scorer"]](it, outputs[it["id"]])
        rows.append({"id": it["id"], "stratum": it["stratum"], "score": s, "why": why})
    return rows

Paired statistics and how many items you need

Because every candidate answers the same items, compare them item by item. For item i, the paired difference is d_i = score_B(i) - score_A(i). The mean of these differences is the estimated change, and its standard error is the standard deviation of d divided by the square root of the number of items. Pairing removes the variation between easy and hard items, so it needs far fewer items than comparing two averages.

Work it through for pass or fail scores. If candidates A and B disagree on 10 percent of items, then d is non-zero on 10 percent of them and its standard deviation is about the square root of 0.1, or 0.32. With 500 items the standard error is 0.32 divided by 22.4, about 1.4 points, and the 95 percent interval is roughly plus or minus 2.8 points. That cannot confirm a guard rail of no more than a two-point drop. To detect a two-point difference with 80 percent power at that disagreement rate you need about (1.96 + 0.84)^2 * 0.1 / 0.02^2, which is roughly 1,960 items. Halve the disagreement rate and you halve the items. This is why the set size comes from the decision.

A bootstrap over items gives the same interval without distributional assumptions and works for any score, including fractional ones. Resample item IDs with replacement, recompute the mean difference, repeat a few thousand times, and read the 2.5th and 97.5th percentiles. Do it per stratum as well.

import numpy as np

def paired_bootstrap(a, b, n_boot=5000, seed=0):
    # a, b: dicts item_id -> score for two candidates on the same item set
    ids = sorted(set(a) & set(b))
    d = np.array([b[i] - a[i] for i in ids])
    rng = np.random.default_rng(seed)
    idx = rng.integers(0, len(d), size=(n_boot, len(d)))
    means = d[idx].mean(axis=1)
    lo, hi = np.percentile(means, [2.5, 97.5])
    return d.mean(), lo, hi, len(ids)

delta, lo, hi, n = paired_bootstrap(scores["bf16"], scores["fp8"])
print(f"fp8 - bf16: {delta:+.3f}  95% CI [{lo:+.3f}, {hi:+.3f}]  n={n}")
# Guard rail "no worse than -0.02": pass only if lo > -0.02.

Running candidates on GPUs reproducibly

It is tempting to assume that greedy decoding at temperature zero gives the same text every time. On GPUs it often does not, and the reason matters for how you read results. Floating-point addition is not associative, so the order in which a kernel reduces numbers changes the low bits of the result. Many inference kernels choose their reduction strategy based on the batch size and sequence lengths in flight. A request processed alongside 3 others can therefore produce slightly different logits from the same request processed alongside 40. When two tokens are nearly tied, that flips the argmax, and from that token on the outputs diverge. Thinking Machines Lab described this failure of batch invariance in detail in a September 2025 post on nondeterminism in LLM inference.

A different tensor-parallel degree, quantised kernels or an engine upgrade have the same effect, so part of any measured difference is noise from the hardware path, not a change in model quality.

So measure the noise floor. Run the baseline configuration twice, in separate processes, and compare the two runs with the same paired bootstrap. The disagreement rate and interval you get are what nothing at all looks like. A candidate's difference only means something if it clearly exceeds that floor. For sampled outputs, draw several samples per item with fixed seeds and score the mean, which turns a noisy single draw into a stable per-item rate.

Pin everything that can be pinned and record it with every run: model revision hash, tokenizer and chat template, engine version, CUDA driver, GPU type, parallelism, quantisation, and sampling parameters. The runner below uses vLLM's offline LLM class and writes a manifest next to the outputs.

import hashlib, json, platform
import vllm
from vllm import LLM, SamplingParams

def run_candidate(cfg, items, out_path):
    llm = LLM(model=cfg["model"], revision=cfg["revision"],
              tensor_parallel_size=cfg["tp"], quantization=cfg.get("quantization"),
              seed=cfg["seed"], max_model_len=cfg["max_model_len"])
    params = SamplingParams(temperature=cfg["temperature"], max_tokens=cfg["max_tokens"],
                            n=cfg["samples"], seed=cfg["seed"])
    convs = [[{"role": "system", "content": cfg["system_prompt"]},
              {"role": "user", "content": it["input"]}] for it in items]
    results = llm.chat(convs, params)   # the engine batches and orders work itself
    rows = [{"id": it["id"], "outputs": [o.text for o in r.outputs],
             "finish": [o.finish_reason for o in r.outputs]}   # "length" means truncated
            for it, r in zip(items, results)]
    manifest = {"config": cfg, "vllm": vllm.__version__, "host": platform.node(),
                "items_sha256": hashlib.sha256(json.dumps([i["id"] for i in items]).encode()).hexdigest()}
    with open(out_path, "w") as f:
        json.dump({"manifest": manifest, "rows": rows}, f)

Offline batch generation packs many prompts into each forward pass. Cap max_tokens at what the task needs; one runaway generation can dominate run time.

Worked example: BF16 versus FP8 versus 4-bit

Suppose a team serves an 8-billion-parameter model in BF16 and wants to know whether FP8 weights, or a 4-bit weight-only format, would let them halve the number of GPUs. Their benchmark has 2,000 held-out items in four strata. The table shows the shape of the result such a team might see. The numbers are illustrative, chosen to show how to read the output, not measurements of any particular model.

ComparisonOverall delta95% intervalWorst stratumReading
BF16 run 2 minus BF16 run 1+0.1-0.4 to +0.6tool calls, -0.3noise floor
FP8 minus BF16-0.3-0.9 to +0.3extraction, -0.8inside the guard rail
4-bit minus BF16-1.2-1.8 to -0.6tool calls, -4.9fails on one stratum

The noise-floor row says that two identical configurations differ by up to about half a point, so anything inside that band is not evidence. FP8's interval and its worst-stratum estimate both sit inside the two-point guard rail, so it passes. The 4-bit candidate looks acceptable overall, but its tool-call stratum fails badly, which a pooled number would hide. Read the failing outputs next: often one systematic error, such as a malformed argument type, explains them.

Put cost next to quality in the same report: GPU-hours and tokens per run, converted to cost per thousand items.

Failure modes

The failures below are the common ones, roughly in order of how often they cause wrong decisions.

  • Leakage into the held-out split. Prompts or few-shot examples are tuned while looking at held-out items, and the benchmark stops measuring generalisation. Enforce access to the split in the pipeline, not by convention.
  • Ignoring the noise floor. A half-point difference between engines is reported as a regression when two identical runs differ by as much.
  • Pooled averages hiding a stratum. Always print per-stratum intervals and gate on the worst one that matters.
  • Judge drift. The judge model is upgraded silently by a provider and scores shift for every candidate. Pin judge versions and re-run calibration when they change.
  • Truncation. A low max_tokens or max_model_len cuts off long answers for one candidate with a longer style, and it is scored as wrong. Record finish reasons and report truncation rates.

Trade-offs

More items buy smaller intervals but cost GPU time and labelling; let the guard rail set the count. Programmatic scorers are exact but only cover checkable tasks; judges cover the rest at the cost of bias. Several samples per item stabilise rates but multiply cost. A private benchmark cannot be compared with anyone else's, so keep a small public one alongside it as a sanity check.

What to do next

  1. Write down the next model or serving decision your team faces, with its guard rail and expected saving.
  2. Sample 2,000 recent requests, scrub and de-duplicate them, label task type, and add 100 expert-written edge cases.
  3. Freeze version 1 with stable IDs and a content hash, and split it into development and held-out parts.
  4. Implement programmatic scorers for every stratum that allows them; calibrate a judge only for the rest.
  5. Run the current production configuration twice to measure the noise floor.
  6. Run the candidate, compute paired bootstrap intervals overall and per stratum, and add cost per thousand items.
  7. Read at least twenty failing items from the worst stratum before you decide.
  8. Store outputs and manifests so the next scorer fix is a rescore, not a rerun.

To go further, read building golden datasets, calibrating LLM judges, an evaluation harness for regression testing, performance regression analysis and what public benchmarks measure for small models.

Key takeaway: A custom benchmark is a decision instrument: a frozen, stratified item set from your own traffic, a pinned runner, programmatic scorers wherever possible, and paired statistics on matching items. Size it from the smallest difference your guard rail must detect, measure the GPU noise floor by running the baseline twice, gate on the worst stratum that matters, and report quality next to cost.