A performance regression in LLM serving is a change that makes the same workload slower or more expensive on the same hardware: first tokens arrive later, tokens stream more slowly, or a GPU sustains fewer requests per second. They are routine. Every upgrade of the driver, CUDA, PyTorch, an attention kernel library or the serving engine changes code paths, and engines regularly change default flags between releases. Most regressions are a few percent and invisible to anyone not measuring; some are 30 percent and only show up as a capacity shortfall weeks later.
This article is about speed, not answer quality; quality regressions need an evaluation harness, which is covered elsewhere on the site. Here you will learn why each latency metric maps to a different part of the GPU's work, how to control the noise that makes small regressions undetectable, how to compare two runs statistically instead of eyeballing averages, and how to localise a confirmed regression by diffing inputs, bisecting and comparing kernel profiles. Benchmark commands use the vllm bench serve CLI as documented on 2026-10-03; the method applies to any engine.
What the metrics measure
Generation has two phases with different bottlenecks, and the metrics follow them. Prefill processes the whole prompt in parallel: large matrix multiplications and attention over many tokens, usually limited by arithmetic throughput. It determines time to first token (TTFT), together with any time the request spends queued. Decode produces one token per step for each running sequence: every step rereads the weights and the KV cache, so at serving batch sizes it is usually limited by memory bandwidth. It determines time per output token (TPOT) and the inter-token latency (ITL) a streaming user feels.
End-to-end latency is roughly TTFT plus TPOT times the output length minus one, so a change in output length moves it without anything getting slower. Throughput in tokens per second is what capacity planning uses, and goodput, requests completed within latency targets, is what users experience at load. Classifying a regression by which of these moved is the fastest first cut, because it tells you which phase, and therefore which kernels and scheduler paths, to suspect.
Where regressions come from
| Layer | Typical change | Usual symptom |
|---|---|---|
| Driver and CUDA | New driver, CUDA runtime or cuBLAS version picks different GEMM algorithms | Throughput and TTFT shift across all shapes |
| Kernel libraries | Attention or fused-MoE kernel update; a fallback path when a shape is unsupported | One phase slows, often only for some sequence lengths |
| Serving engine | Changed scheduler defaults, chunked prefill, prefix caching, CUDA graph capture sizes | TTFT or tail latency moves at load, not at batch 1 |
| Model artefacts | New tokenizer, chat template or system prompt | Prompts get longer in tokens; TTFT rises with no code change |
| Numerics | Different quantisation or KV cache dtype | TPOT changes, memory headroom changes batch size |
| Parallelism | Tensor-parallel degree, NCCL version, topology | Decode slows from exposed communication |
| Hardware state | Clock throttling, power cap, a degraded link, a noisy neighbour | Run-to-run variance, one node slower than others |
The model-artefact row catches teams most often. A template change that adds 300 tokens to every prompt is a real regression for users, but no kernel got slower, and a benchmark using random token prompts will never see it. Benchmark at least once with recorded production-like prompts, not only synthetic ones.
Controlling noise
A benchmark can only detect regressions larger than its run-to-run noise, so measure the noise first. Run the same build twice, an A/A comparison, and the spread you see is your floor. Then shrink it:
- Pin everything that is not under test: the container digest, the Python lockfile, the model revision, engine flags and the dataset seed.
- Stabilise the GPU. Persistence mode (
nvidia-smi -pm 1) avoids reinitialisation between runs, and locking graphics clocks (nvidia-smi -lgc <min>,<max>, reset with-rgc) removes boost-clock variance where your GPU and policy allow it. Recordnvidia-smi --query-gpu=clocks.sm,temperature.gpu,power.draw --format=csv -l 1during every run so throttling is visible after the fact. - Warm up. The first requests pay for CUDA graph capture, autotuning and allocator growth; discard a warm-up phase.
- Fix the output length with
--ignore-eosso both runs generate the same number of tokens. - Run the load generator on a separate machine or at least separate cores, and confirm it is not saturated.
- Interleave runs, A B A B, rather than A A B B, so slow drift such as thermal state or a neighbour's job lands on both sides.
A fixed request rate measures latency at a chosen load; --request-rate inf, the default, measures behaviour at saturation, where queueing dominates TTFT. Run both, and keep the load identical between baseline and candidate.
Collecting comparable runs
With vLLM the client is vllm bench serve. The flags below are from the CLI reference; --save-detailed adds per-request arrays to the JSON, which the statistics need.
vllm bench serve --backend vllm --model "$MODEL" \
--dataset-name random --random-input-len 1024 --random-output-len 256 \
--ignore-eos --num-prompts 1000 --request-rate 8 --seed 0 \
--percentile-metrics ttft,tpot,itl,e2el --metric-percentiles 50,90,99 \
--save-result --save-detailed --result-filename "runs/${LABEL}_${N}.json"In the source checked, the summary keys include median_ttft_ms, p99_tpot_ms, output_throughput and request_throughput, and the detailed arrays include ttfts, latencies, output_lens and errors, with times in seconds. Key names have changed across releases, so assert they exist before you compute anything.
Comparing runs statistically
Comparing two averages is how false alarms and missed regressions both happen: latency distributions are skewed, and a mean moves when a handful of outliers do. Compare percentiles with a bootstrap confidence interval instead. Resample each run's requests with replacement many times, recompute the percentile difference each time, and read off the middle 95 percent. If the interval excludes zero, the change is real; whether it matters is a separate question you answer with a threshold.
import json, random
def load(path):
r = json.load(open(path))
ok = [i for i, e in enumerate(r["errors"]) if not e]
ttft = [r["ttfts"][i] * 1000 for i in ok]
tpot = [(r["latencies"][i] - r["ttfts"][i]) * 1000 / (r["output_lens"][i] - 1)
for i in ok if r["output_lens"][i] > 1]
return ttft, tpot, len(r["errors"]) - len(ok)
def pct(xs, q):
s = sorted(xs); k = (len(s) - 1) * q / 100; f = int(k); c = min(f + 1, len(s) - 1)
return s[f] + (s[c] - s[f]) * (k - f)
def boot_ci(a, b, q, n=2000, seed=0):
rng = random.Random(seed); d = []
for _ in range(n):
d.append(pct(rng.choices(b, k=len(b)), q) - pct(rng.choices(a, k=len(a)), q))
d.sort()
return d[int(0.025 * n)], d[int(0.975 * n) - 1]
def verdict(a, b, q, floor_ms):
lo, hi = boot_ci(a, b, q); delta = pct(b, q) - pct(a, q)
if lo > 0:
return "REGRESSION" if delta > floor_ms else "real, below floor"
return "IMPROVEMENT" if hi < 0 else "no detectable change"Count errors separately and fail the comparison if they rise: a request that times out leaves the latency sample, so a broken build can look faster.
Worked example: what 1,000 requests can see
To show what the method can and cannot see, the snippet was run on synthetic data, not a real GPU: 1,000 baseline TTFTs drawn from a log-normal distribution with a median near 180 ms, a second baseline drawn the same way (the A/A run), and a candidate in which 8 percent of requests pay an extra 120 ms, the shape of a scheduling change that occasionally delays prefill. Python's random was seeded with 7, with a 5 ms floor:
| Comparison | Baseline | Candidate | 95% CI of difference | Verdict |
|---|---|---|---|---|
| A/A, p50 TTFT | 181.9 ms | 179.9 ms | -8.0 to +5.1 ms | no detectable change |
| A/A, p99 TTFT | 391.6 ms | 390.4 ms | -55.1 to +40.0 ms | no detectable change |
| Candidate, p50 TTFT | 181.9 ms | 190.2 ms | +2.0 to +15.3 ms | regression |
| Candidate, p99 TTFT | 391.6 ms | 412.2 ms | -35.9 to +51.6 ms | no detectable change |
| Candidate, share over 300 ms | 6.9% | 10.9% | +1.4 to +6.5 points | regression |
Two lessons fall out. The A/A rows show that with 1,000 requests the p99 cannot resolve anything smaller than roughly 50 ms either way; a p99 alarm on a run this size is mostly noise. And the injected tail regression is invisible at p99 but clear in the share of requests over a fixed threshold, because a proportion uses every request while a percentile is estimated from the few around it. Track threshold shares tied to your service-level objective, and use many more requests, or repeated runs, when the tail is what you care about.
Localising a confirmed regression
Once a regression is confirmed, diff the inputs before touching a profiler: compare lockfiles, container digests, engine version and the effective engine configuration printed at startup, and the token counts of your recorded prompts. A changed default flag explains a surprising share of regressions and takes minutes to test by setting the old value explicitly.
If the inputs do not explain it and the change range is a series of commits, let the benchmark bisect. git bisect run treats exit code 0 as good, 1 to 127 as bad, and 125 as untestable:
#!/usr/bin/env bash
# bisect_ttft.sh: exits 0 if p50 TTFT is within 5% of the baseline, 1 if slower, 125 if the build fails
pip install -e . >/dev/null 2>&1 || exit 125
./start_server.sh && ./bench.sh "runs/bisect_$(git rev-parse --short HEAD).json" || exit 125
python check_p50.py --baseline runs/base.json --candidate "runs/bisect_$(git rev-parse --short HEAD).json" --max-ratio 1.05
# usage: git bisect start BAD_SHA GOOD_SHA && git bisect run ./bisect_ttft.shBisection needs a threshold well above the noise floor, otherwise one unlucky run sends it down the wrong half. For the profile diff, capture both builds under the same load with Nsight Systems and compare per-kernel totals with nsys stats --report cuda_gpu_kern_sum (older releases named the report differently). A new kernel name, a kernel whose total time grew, or GPU idle gaps between kernels point respectively at a changed code path, a slower implementation, or CPU-side launch and scheduling overhead.
Failure modes
- Measuring the client. A saturated load generator caps throughput and inflates latency identically for both builds, hiding the difference.
- Unequal work. Different output lengths, a warm prefix cache in one run only, or a tokenizer change make the two runs do different work.
- Too many metrics. Checking thirty metrics independently at 95 percent confidence produces one or two false alarms per comparison on average. Pick a few primary metrics, and require a rerun before paging anyone.
- Saturation-only testing. Infinite-rate runs measure queueing; most production traffic runs below saturation, where different code paths dominate.
- One noisy node. A throttled or degraded GPU produces a convincing regression. Check the recorded clocks and run the A/A test on the same machine.
- Silent errors. Failed requests drop out of latency samples; always compare error counts.
Running it continuously
Make this a pipeline. Run a fixed benchmark matrix nightly and on every dependency bump: two or three prompt-length mixes, a fixed request rate near production load and one saturation run. Store every JSON with the build metadata, so you can plot a metric's history and spot slow drifts that no single comparison flags. Gate merges on the confirmed-regression verdict rather than on a raw percentage, and keep the noise floor itself on a chart, because a floor that grows means the harness has degraded.
The trade-off is GPU time against sensitivity. Halving the detectable effect size needs roughly four times as many requests, so decide which size of regression is worth catching, a few percent on throughput, a larger margin on tail latency, and size the runs to that, instead of running everything everywhere.
What to do next
- Run an A/A comparison on your current build to measure the noise floor for each primary metric.
- Pin the container, lockfile, model revision and engine flags, and record GPU clocks during runs.
- Save detailed per-request results and compare percentiles and threshold shares with bootstrap intervals.
- Add a benchmark with recorded production-like prompts alongside the synthetic one.
- Write a bisect script with exit codes 0, 1 and 125 and a threshold above the noise floor.
- Keep Nsight Systems kernel summaries for the baseline so a profile diff is one command away.
- Schedule the matrix nightly and on every dependency bump, and chart history plus the noise floor.
Related reading: inference latency from first principles, reading PyTorch profiler traces, Nsight Systems and Nsight Compute, LLM serving stacks and, for quality rather than speed, an evaluation harness for regression testing.