An offline eval tells you whether a new prompt does better on the examples you chose. A production A/B test tells you whether it does better for the users you actually have, on the metrics the business cares about, at a cost and latency you can afford. Prompt changes are cheap to write and easy to ship, which is exactly why they need the same experimental discipline as any other product change.
LLM features also bring traps of their own: non-deterministic outputs, judge-scored metrics, prompt caching that favours the incumbent prompt, models that change underneath you, and conversations that span many requests. This article covers the online half of prompt evaluation: how to assign traffic, what to log, which metrics to compute and how, the statistics that keep you honest, and the confounders specific to prompts. It assumes the candidate prompt has already passed an offline eval gate.
Where the online test sits
A prompt change moves through four stages: an offline eval gate (prompt evaluation architecture), an online A/B test, a ramp to full traffic, and a release recorded in the registry (prompt versioning and the prompt registry). The general machinery of experimentation platforms is covered in A/B experimentation platform architecture, and model-level tests in ML A/B testing. This page covers what changes when the treatment is a prompt.
Randomize by user or conversation, not by request
If each request flips its own coin, a multi-turn conversation can switch prompts halfway through. The second turn's system prompt may contradict the first turn's style, and the outcome you measure, such as whether the conversation resolved, belongs to neither arm. Users who see inconsistent behaviour across sessions also change how they behave. Assign at the level where you measure the outcome: the conversation for single-session tasks, the user for anything with return visits. A deterministic hash of the unit ID with a per-experiment salt keeps assignment stable without a lookup table and independent across experiments.
import hashlib, json, time
def bucket(unit_id: str, salt: str, buckets: int = 10_000) -> int:
h = hashlib.sha256(f"{salt}:{unit_id}".encode()).digest()
return int.from_bytes(h[:8], "big") % buckets
EXPERIMENT = {
"salt": "support-concise-2026-10", # new salt per experiment
"allocation": [("control", 0, 500), ("concise", 500, 1000)], # 5% per arm
}
def assign(user_id: str):
b = bucket(user_id, EXPERIMENT["salt"])
for arm, lo, hi in EXPERIMENT["allocation"]:
if lo <= b < hi:
return arm
return None # not in the experiment: production prompt, not analysed
def log_call(out, user, conv, arm, prompt_sha, model, usage, latency_ms):
out.write(json.dumps({
"ts": time.time(), "user": user, "conv": conv, "arm": arm,
"prompt_sha": prompt_sha, "model": model,
"in_tok": usage["input"], "out_tok": usage["output"],
"cached_tok": usage.get("cached", 0), "latency_ms": latency_ms,
}) + "\n")Log exposure at the moment the variant actually affects a response, not at assignment time. Users who are assigned but never reach the feature only dilute the effect. Log the resolved prompt content hash and model ID with every call too, so the analysis joins on what actually ran rather than on what the config said.
Choose metrics before you launch
| Tier | Example | Unit | Role |
|---|---|---|---|
| Primary | Conversation resolved without escalation | Conversation | One metric; decides the ship |
| Guardrail | Cost per conversation, p95 latency, refusal rate, safety flags, parse failures | Conversation or request | Must not regress past a set threshold |
| Diagnostic | Output tokens, turns per conversation, judge sub-scores | Various | Explains results; never decides |
User signals such as ratings, escalations, copied answers and repeated questions are sparse and skew toward unhappy users, but they are real outcomes. Judge scores are dense but are one model's opinion. Use both, and treat the judge as a measurement instrument that has to be validated.
Online judge scoring without fooling yourself
- Sample evenly: score a fixed random fraction of conversations from both arms at the same rate, never only the flagged ones.
- Blind the judge: it sees the conversation and the rubric, never the arm name or the prompt text.
- Pin it: fix the judge model version and judge prompt for the whole experiment. Changing the judge mid-test is a change of measurement that hits the two arms at different points of the ramp.
- Calibrate it: have humans label a few hundred sampled conversations and measure their agreement with the judge before trusting it. Length bias is the classic failure: judges tend to prefer longer answers, and answer length is exactly what a "be concise" prompt changes.
- Compare pairwise where you can: replay the same inputs through both prompts offline and ask the judge which answer is better. This is more sensitive than absolute 1-to-10 scores.
How many conversations you need
For a binary primary metric, the sample size depends on the baseline rate, the minimum detectable effect (MDE), the significance level and the power. With a 62% resolution rate, a two-sided alpha of 0.05 and 80% power, the standard two-proportion approximation gives:
| MDE (absolute) | Conversations per arm |
|---|---|
| 1 point | 36,791 |
| 2 points | 9,147 |
| 3 points | 4,042 |
from math import sqrt, ceil
from statistics import NormalDist
def n_per_arm(p0, mde, alpha=0.05, power=0.8):
z = NormalDist().inv_cdf
p1 = p0 + mde
pbar = (p0 + p1) / 2
num = (z(1 - alpha / 2) * sqrt(2 * pbar * (1 - pbar))
+ z(power) * sqrt(p0 * (1 - p0) + p1 * (1 - p1))) ** 2
return ceil(num / mde ** 2)
for mde in (0.01, 0.02, 0.03):
print(mde, n_per_arm(0.62, mde)) # 36791, 9147, 4042Halving the detectable effect roughly quadruples the sample. Set the MDE from the business decision: if a 1-point gain would not change what you do, do not wait for 37,000 conversations per arm. These counts also assume independent units. If you randomize by user but measure per conversation, one user's conversations are correlated and the naive standard error is too small. Aggregate to the randomization unit, or use the delta method for ratio metrics.
Ratio metrics and the delta method
Most LLM guardrails are ratios computed over units: cost per conversation, escalations per message, tokens per resolved case. A per-message t-test treats every message as independent and overstates your confidence. The delta method gives the variance of a ratio of per-unit sums correctly:
import numpy as np
def ratio_delta(num, den):
"""num, den: arrays of per-user sums, e.g. cost and conversations."""
n = len(num)
mx, my = num.mean(), den.mean()
vx, vy = num.var(ddof=1), den.var(ddof=1)
cxy = np.cov(num, den, ddof=1)[0, 1]
var = (vx / my**2 - 2 * mx * cxy / my**3 + mx**2 * vy / my**4) / n
return mx / my, var
def compare(ctrl, trt):
r0, v0 = ratio_delta(*ctrl)
r1, v1 = ratio_delta(*trt)
diff, se = r1 - r0, np.sqrt(v0 + v1)
return diff, (diff - 1.96 * se, diff + 1.96 * se)Pass in arrays with one row per user, for example total cost and conversation count per user in each arm. The interval on the difference is what you compare against the guardrail threshold.
Check sample ratio mismatch first
Before reading any metric, check that the arms received the traffic you allocated. A 50/50 split that arrives as 50,912 versus 49,088 looks close, but a chi-square test gives 33.27, with p of about 8e-9. Something is dropping or misrouting traffic in one arm. Common LLM-specific causes: the new prompt times out more often and those calls never log; a parsing error in the treatment path throws before the exposure log is written; or a response cache keyed without the arm serves control answers to treatment users. Results from an experiment with sample ratio mismatch cannot be trusted in either direction. Find the cause and restart.
Confounders that are specific to prompts
- Prompt caching favours the incumbent. Where the provider caches prompt prefixes, the control prompt is already warm, so for its first hours a new prompt shows higher latency and input cost. Log cached tokens separately, and agree a warm-up exclusion window for cost and latency before launch.
- Longer prompts cost on every call. A system prompt that grows by 800 tokens adds input cost to every turn of every conversation. Set the guardrail on cost per conversation, not cost per call.
- Model drift. If both arms use a floating model alias and the provider updates it mid-test, both arms change. If only one arm pins its model, the test mixes a prompt change with a model change. Pin the model ID in both arms and record it on every log line.
- Non-determinism. Sampling adds variance to each response. It does not bias the comparison, but it raises the sample size you need. Keep temperature the same in both arms unless temperature is what you are testing.
- Novelty and learning. Users adapt to a new style. Run for at least one full weekly cycle, and look at the effect by days since first exposure.
- Shared state. Memory features, retrieved conversation history or cached answers can carry one arm's outputs into the other arm's context.
Peeking, ramps and stopping
Checking a fixed-horizon test every day and stopping at the first significant result pushes the false-positive rate well above 5%. Either fix the sample size in advance and analyse once, or use a sequential method built for continuous monitoring, such as alpha-spending boundaries or always-valid confidence sequences. Guardrails are different. Monitor them continuously with conservative thresholds, because stopping early for harm costs little.
A typical ramp runs 1% per arm for a day to catch crashes, parse failures and sample ratio mismatch, then 5 to 10% per arm until the planned sample is reached. You ship by moving the production pointer in the registry. If you want to measure long-term effects, keep a small holdout on the old prompt for a few weeks.
Worked example: a concise support prompt
A support assistant's control prompt produces thorough answers. The candidate, "concise", asks for short answers followed by an offer to go deeper. Offline evals showed equal correctness and 35% fewer output tokens. The online plan, written before launch, was:
- Primary metric: resolution without escalation, with a 62% baseline and a 2-point MDE, so at least 9,150 conversations per arm, a floor before allowing for clustering.
- Non-inferiority margin: 2 points.
- Randomization: by conversation, at 5% per arm, since most users here ask one question.
- Guardrails: cost per conversation must not rise; p95 latency must not rise by more than 10%; escalation rate must not rise by more than 1 point.
- Warm-up: cost and latency are excluded for the first 24 hours, because of prompt caching.
The day-1 sample ratio check passed. Early on, latency for the concise arm looked 12% worse, and the logs showed near-zero cached tokens in its first hours, which is the case the warm-up rule covers. After 11 days both arms had passed 9,150 conversations. Resolution was 63.1% against 62.0%. The 95% interval on the difference ran from about -0.3 to +2.5 points, so the gain is not established. However, the lower bound sits well inside the 2-point non-inferiority margin agreed in advance. Cost per conversation fell 18% with a tight interval, and judge helpfulness was unchanged after length-bias calibration. The decision was to ship for cost, recording in the registry that the effect on resolution was inconclusive.
Failure modes
- Assigning per request, so conversations mix prompts.
- Logging exposure at assignment, which dilutes every effect.
- Using a judge that is not blind or not pinned, or never checked against human labels.
- Running per-message tests on user-randomized data, which makes intervals too narrow.
- Stopping at the first significant peek.
- Shipping on a diagnostic metric ("answers got shorter") when the primary metric did not move.
Trade-offs
| Choice | You gain | You pay |
|---|---|---|
| User rather than conversation as the unit | Consistent experience; captures return effects | Fewer units, so longer tests |
| Judge metrics | Dense, fast signal | Instrument bias; needs calibration |
| Small MDE | Detects small gains | Long tests and more exposure to a worse arm |
| Sequential testing | Valid early stopping | Wider intervals at a fixed sample size |
| Holdout after shipping | Measures long-term effects | Some users keep the old prompt |
What to do next
- Write a one-page plan before launch: unit, primary metric, MDE, guardrails with thresholds, warm-up window and analysis method.
- Implement hashed assignment with a fresh salt per experiment, and log the prompt hash, model ID and cached tokens on every call.
- Validate your judge against a few hundred human labels, and check it for length bias.
- Compute the sample size with the code above and estimate the test's duration from real traffic.
- Automate a sample ratio check on day 1 and every day after.
- Analyse ratio metrics per unit with the delta method, then record the decision and its evidence in the prompt registry.