Most changes to an LLM serving stack are bets on a trade-off. FP8 weights promise more throughput for a small quality risk. Speculative decoding promises lower latency at the cost of extra GPU work. A new engine version, a larger batch limit, a different GPU type or a distilled model each move latency, cost and quality in different directions at once. Offline benchmarks narrow the options, but they run on curated prompts at fixed concurrency; real traffic has long conversations, bursty arrival and users who give up. An online A/B test is how you find out what the change does to the product.
This article is about A/B testing serving changes on GPUs: what to randomise, why sharing a GPU pool between arms silently biases the result, which statistics fit tail latency, and how to turn quality and cost into one decision. Gradual rollouts with automated aborts are a different tool, covered in LLM canary releases; a canary asks "is it safe?", an A/B test asks "is it better, and by how much?". Prompt-level experiments, with the delta method and sample-ratio checks derived in full, are covered in prompt A/B testing.
Start with the trade-off you are testing
Start by writing down the hypothesis as a trade-off with numbers: "FP8 weights on the 70B model will cut GPU-seconds per successful conversation by at least 25 percent, with task success no more than 0.5 points lower and p95 time to first token no worse." That sentence names the primary metric, the guardrails and the minimum effect that would justify shipping. Without it, a test with twenty metrics will always find something significant.
Serving changes fall into three groups. Bit-changing changes, such as quantisation, a different model or a new tokenizer, can change outputs and need quality metrics. Bit-preserving changes, such as speculative decoding with exact verification, or a scheduler tweak, should produce the same distribution of outputs; test mainly latency and cost, but still watch quality, because numeric differences in kernels change greedy outputs more often than people expect. Capacity changes, such as batch limits or GPU type, affect queueing and are the most sensitive to the interference problem below.
Randomise conversations, not requests
Randomise by user or by conversation, never by request. A conversation that switches model mid-thread produces incoherent replies and contaminates both arms, and prefix caches are per replica, so request-level splitting also destroys cache hit rates in a way production would never see. Assign deterministically so every service makes the same choice without coordination:
import hashlib
def assign(unit_id: str, experiment: str, salt: str, treatment_share: float) -> str:
"""Stable arm for a user or conversation. New salt per experiment avoids carry-over between tests."""
h = hashlib.sha256(f"{experiment}:{salt}:{unit_id}".encode()).digest()
bucket = int.from_bytes(h[:8], "big") % 10_000
return "treatment" if bucket < treatment_share * 10_000 else "control"
# Router: look up arm, send to that arm's pool, log one exposure row per request
arm = assign(conv.user_id, "fp8-70b-2026-10", salt="k3f9", treatment_share=0.5)
pool = POOLS[arm] # separate replica sets, see next section
log_exposure(unit=conv.user_id, arm=arm, request_id=req.id, ts=now())Log exposure at the router when the request is actually served, not when the user is assigned; users assigned but never served are not in the experiment. Then run the sample ratio mismatch check before reading any result: if a 50/50 split shows 50.6/49.4 over a million users, something, usually timeouts or retries in one arm, is dropping units non-randomly, and every metric is suspect.
The exposure record
Every analysis in this article depends on one table, so design it before launch. Write one row per served request, at the router, and join engine-side numbers to it by request ID. Without the engine fields you cannot match load between arms; without the unit and arm you cannot run the sample ratio check or a user-level bootstrap; without the outcome fields you can only measure speed, which is rarely the whole decision.
| Field group | Examples | Used for |
|---|---|---|
| Assignment | unit id, experiment, arm, salt version | SRM check, clustering, carry-over audits |
| Request shape | prompt tokens, output tokens, cached prefix tokens | comparing like with like |
| Latency | queue time, time to first token, inter-token latency, total | percentile comparisons |
| Load at admission | running requests, queue length, KV-cache usage on that replica | matched load bands |
| Cost | GPU-seconds attributed to the request, replica GPU type | cost per success |
| Outcome | finish reason, retry or rephrase within five minutes, feedback | task success |
Keep raw prompts and responses out of this table; store them separately under your retention policy, sampled for judge scoring, so the experiment log can be kept long enough to compare against the next test without becoming a privacy liability. Attributing GPU-seconds to individual requests in a batched engine is approximate; a simple rule, splitting each forward pass's time across the requests in it by tokens processed, is consistent between arms, and consistency matters more than precision here.
Shared GPUs bias the result
This is the GPU-specific trap. Continuous batching engines pack concurrent requests into the same forward pass, so a request's latency depends on everything else in the batch. Put both arms on one shared pool, for example two model versions behind the same scheduler or a 5 percent treatment on a few dedicated replicas, and the arms experience different load. A 5 percent arm on replicas sized for 5 percent of traffic sees noisier, smaller batches; a 5 percent arm on replicas sized at a round number often sees far less load per GPU than control. Either way, latency and cost per token measured during the test do not describe what happens at 100 percent.
Three rules fix most of it. First, give each arm its own replica pool, sized in proportion to its traffic share, so tokens per second per GPU, batch occupancy and KV-cache utilisation match between arms. Second, prefer a 50/50 split for capacity and performance experiments; it makes proportional sizing trivial and maximises statistical power. Third, record load per arm, such as running requests, queue length and KV-cache usage from the engine's own metrics, and compare latency only within matched load bands. If the treatment is faster only because it was less loaded, the load comparison shows it. The serving tiers involved are described in LLM serving architecture.
Metrics and the statistics that fit them
Use three layers of metrics with different speeds. Guardrails move within minutes: error rate, out-of-memory restarts, timeouts, p99 time to first token and the sample ratio. Breaching one stops the test; this is the canary's job and should be automated. Primary metrics move over days: task success (the user did not retry, rephrase or abandon), explicit feedback rates, conversation length, and an LLM judge's pairwise preference on a sample of logged responses. Cost metrics tie it together: GPU-seconds per request and, better, per successful conversation, because a cheaper arm that causes more retries is not cheaper.
For latency, compare percentiles, not means. Token latencies are heavy-tailed, and a change that improves the median while fattening the tail is common for speculative decoding (verification rejections) and for aggressive batching. Percentiles of pooled requests also ignore that requests from one user are correlated, so compute them with a bootstrap that resamples users, not requests:
import numpy as np
def p95_diff_ci(ttft_by_user_a, ttft_by_user_b, iters=2000, seed=0):
"""Cluster bootstrap: resample users, pool their requests, compare p95 TTFT."""
rng = np.random.default_rng(seed)
ua, ub = list(ttft_by_user_a.values()), list(ttft_by_user_b.values())
diffs = []
for _ in range(iters):
a = np.concatenate([ua[i] for i in rng.integers(0, len(ua), len(ua))])
b = np.concatenate([ub[i] for i in rng.integers(0, len(ub), len(ub))])
diffs.append(np.percentile(b, 95) - np.percentile(a, 95))
lo, hi = np.percentile(diffs, [2.5, 97.5])
return float(np.median(diffs)), (float(lo), float(hi))Ratio metrics such as tokens per second or GPU-seconds per success need the delta method or the same cluster bootstrap; a naive t-test on per-request values understates the variance because requests within a user are correlated. CUPED, using each user's pre-experiment value of the same metric as a covariate, often cuts variance substantially for engagement metrics and lets the test finish sooner.
Sample size and when to stop
Decide the sample size before launch. For a binary task-success metric with baseline 0.80, detecting a 0.5 point difference at 5 percent significance and 80 percent power takes roughly 2 x (1.96 + 0.84)^2 x 0.8 x 0.2 / 0.005^2, about 100,000 users per arm. A 2-point difference needs about 6,300 per arm. That arithmetic is why quality-risk tests on small products are often infeasible online, and why offline evaluation has to carry more of the weight for bit-changing changes. Latency effects are usually large relative to their variance and show up within a day.
Do not stop the test the first time the p-value dips below 0.05. Repeated peeking inflates false positives dramatically. Either fix the duration in advance, at least one full weekly cycle, because traffic mix differs between weekdays and weekends, or use a sequential method designed for continuous monitoring. Guardrails are the exception: stopping for harm is always allowed.
Worked example: FP8 weights for a 70B chat model
A team serves a 70B chat model in BF16 on eight-GPU nodes and wants to move to FP8 weights. Offline, FP8 matched BF16 within noise on their evaluation suite and raised throughput at their target latency by about 40 percent. The online test splits users 50/50 into two pools of eight replicas each, so per-arm load is identical; the FP8 pool therefore runs with headroom, and the team records batch occupancy to account for it.
After two weeks and 240,000 users per arm, the sample ratio is 50.02/49.98. p95 time to first token is 6 percent lower in FP8 with a bootstrap interval that excludes zero, but within matched load bands the difference shrinks to 2 percent, so most of the gain came from headroom, not speed. Task success is 81.5 percent against 81.6 percent, a difference whose 95 percent interval spans minus 0.33 to plus 0.13 points: not significant, and the whole interval sits inside the pre-registered 0.5-point margin. Pairwise judge preference on 3,000 sampled responses is 49 percent for FP8. GPU-seconds per successful conversation, computed from the throughput the FP8 pool reaches when loaded like production, falls by 27 percent. Following the pre-registered rule, they ship FP8 for the chat product and keep BF16 for a code assistant whose own test showed a significant drop in tests-passing rate.
Failure modes
- Treatment looks faster than it will be: the arm was under-loaded. Size pools in proportion and compare within load bands.
- Sample ratio mismatch: one arm drops units through timeouts, crashes or retries. Fix and restart; do not analyse.
- Cache effects masquerade as model effects: request-level randomisation broke prefix-cache locality. Randomise by conversation.
- Novelty or carry-over: users react to change itself, or a previous test's assignment correlates with this one. Use a fresh salt and look at the effect over time.
- Significant by peeking: the test was stopped on a lucky day. Fix duration or use sequential tests.
- Judge disagrees with users: the LLM judge prefers longer answers. Calibrate it against human labels and control for length.
- Cost per token improved, cost per outcome got worse: more retries or longer conversations. Make the cost metric per success.
Trade-offs
Separate pools cost GPUs: a 50/50 test on a large model briefly needs two full serving fleets' worth of flexibility, and a smaller treatment share saves capacity at the price of power and the load-matching problem. Long tests reduce false positives and catch weekly patterns, but keep a worse arm live longer. Interleaved or pairwise evaluation, showing two responses and asking which is better, is far more sensitive than A/B for quality, but cannot measure latency or cost. A practical ladder is offline evaluation, then shadow traffic to confirm stability and cost under real prompts, then a canary for safety, then the A/B test for the decision. For engine architectures that split prefill and decode, such as disaggregated serving, run the experiment at the level of complete pools, because the two tiers interfere with each other as much as arms do.
What to do next
- Write the hypothesis as a trade-off with a primary metric, guardrails and a minimum worthwhile effect.
- Randomise by user or conversation with a salted hash, and log exposure where requests are served.
- Give each arm its own replica pool sized in proportion to its traffic, and record per-arm load.
- Automate guardrails and run the sample ratio check before reading any metric.
- Compare latency percentiles with a user-level bootstrap, within matched load bands.
- Compute cost as GPU-seconds per successful conversation, not per token.
- Fix sample size and duration in advance, covering at least one weekly cycle.
- Record the decision and the numbers so the next experiment starts from a known baseline.