An offline eval tells you whether a new prompt does better on the examples you chose. A production A/B test tells you whether it does better for the users you actually have, on the metrics the business cares about, at a cost and latency you can afford. Prompt changes are cheap to write and easy to ship, which is exactly why they need the same experimental discipline as any other product change.

LLM features also bring traps of their own: non-deterministic outputs, judge-scored metrics, prompt caching that favours the incumbent prompt, models that change underneath you, and conversations that span many requests. This article covers the online half of prompt evaluation: how to assign traffic, what to log, which metrics to compute and how, the statistics that keep you honest, and the confounders specific to prompts. It assumes the candidate prompt has already passed an offline eval gate.

Advertisement

Where the online test sits

A prompt change moves through four stages: an offline eval gate (prompt evaluation architecture), an online A/B test, a ramp to full traffic, and a release recorded in the registry (prompt versioning and the prompt registry). The general machinery of experimentation platforms is covered in A/B experimentation platform architecture, and model-level tests in ML A/B testing. This page covers what changes when the treatment is a prompt.

Prompt A/B test: assign per user, log what actually ran, analyse per unitRequestuser, conversationAssignmenthash(salt, user)Prompt registryarm to prompt hashModelpinned model IDResponseto the userExposure and call logarm, prompt hash, model, tokens, latencylog every callJudge samplerblind, pinned judgeUser signalsescalations, ratingsPer-unit metric tableone row per user and armAnalysisSRM, delta method, guardrailsDecision recorded in the registry
The serving path resolves a user's arm to a prompt hash and a pinned model, and logs both. Analysis reads per-unit tables built from those logs, judge scores and user signals.

Randomize by user or conversation, not by request

If each request flips its own coin, a multi-turn conversation can switch prompts halfway through. The second turn's system prompt may contradict the first turn's style, and the outcome you measure, such as whether the conversation resolved, belongs to neither arm. Users who see inconsistent behaviour across sessions also change how they behave. Assign at the level where you measure the outcome: the conversation for single-session tasks, the user for anything with return visits. A deterministic hash of the unit ID with a per-experiment salt keeps assignment stable without a lookup table and independent across experiments.

import hashlib, json, time

def bucket(unit_id: str, salt: str, buckets: int = 10_000) -> int:
    h = hashlib.sha256(f"{salt}:{unit_id}".encode()).digest()
    return int.from_bytes(h[:8], "big") % buckets

EXPERIMENT = {
    "salt": "support-concise-2026-10",          # new salt per experiment
    "allocation": [("control", 0, 500), ("concise", 500, 1000)],  # 5% per arm
}

def assign(user_id: str):
    b = bucket(user_id, EXPERIMENT["salt"])
    for arm, lo, hi in EXPERIMENT["allocation"]:
        if lo <= b < hi:
            return arm
    return None            # not in the experiment: production prompt, not analysed

def log_call(out, user, conv, arm, prompt_sha, model, usage, latency_ms):
    out.write(json.dumps({
        "ts": time.time(), "user": user, "conv": conv, "arm": arm,
        "prompt_sha": prompt_sha, "model": model,
        "in_tok": usage["input"], "out_tok": usage["output"],
        "cached_tok": usage.get("cached", 0), "latency_ms": latency_ms,
    }) + "\n")

Log exposure at the moment the variant actually affects a response, not at assignment time. Users who are assigned but never reach the feature only dilute the effect. Log the resolved prompt content hash and model ID with every call too, so the analysis joins on what actually ran rather than on what the config said.

Advertisement

Choose metrics before you launch

TierExampleUnitRole
PrimaryConversation resolved without escalationConversationOne metric; decides the ship
GuardrailCost per conversation, p95 latency, refusal rate, safety flags, parse failuresConversation or requestMust not regress past a set threshold
DiagnosticOutput tokens, turns per conversation, judge sub-scoresVariousExplains results; never decides

User signals such as ratings, escalations, copied answers and repeated questions are sparse and skew toward unhappy users, but they are real outcomes. Judge scores are dense but are one model's opinion. Use both, and treat the judge as a measurement instrument that has to be validated.

Online judge scoring without fooling yourself

  • Sample evenly: score a fixed random fraction of conversations from both arms at the same rate, never only the flagged ones.
  • Blind the judge: it sees the conversation and the rubric, never the arm name or the prompt text.
  • Pin it: fix the judge model version and judge prompt for the whole experiment. Changing the judge mid-test is a change of measurement that hits the two arms at different points of the ramp.
  • Calibrate it: have humans label a few hundred sampled conversations and measure their agreement with the judge before trusting it. Length bias is the classic failure: judges tend to prefer longer answers, and answer length is exactly what a "be concise" prompt changes.
  • Compare pairwise where you can: replay the same inputs through both prompts offline and ask the judge which answer is better. This is more sensitive than absolute 1-to-10 scores.

How many conversations you need

For a binary primary metric, the sample size depends on the baseline rate, the minimum detectable effect (MDE), the significance level and the power. With a 62% resolution rate, a two-sided alpha of 0.05 and 80% power, the standard two-proportion approximation gives:

MDE (absolute)Conversations per arm
1 point36,791
2 points9,147
3 points4,042
from math import sqrt, ceil
from statistics import NormalDist

def n_per_arm(p0, mde, alpha=0.05, power=0.8):
    z = NormalDist().inv_cdf
    p1 = p0 + mde
    pbar = (p0 + p1) / 2
    num = (z(1 - alpha / 2) * sqrt(2 * pbar * (1 - pbar))
           + z(power) * sqrt(p0 * (1 - p0) + p1 * (1 - p1))) ** 2
    return ceil(num / mde ** 2)

for mde in (0.01, 0.02, 0.03):
    print(mde, n_per_arm(0.62, mde))     # 36791, 9147, 4042

Halving the detectable effect roughly quadruples the sample. Set the MDE from the business decision: if a 1-point gain would not change what you do, do not wait for 37,000 conversations per arm. These counts also assume independent units. If you randomize by user but measure per conversation, one user's conversations are correlated and the naive standard error is too small. Aggregate to the randomization unit, or use the delta method for ratio metrics.

Ratio metrics and the delta method

Most LLM guardrails are ratios computed over units: cost per conversation, escalations per message, tokens per resolved case. A per-message t-test treats every message as independent and overstates your confidence. The delta method gives the variance of a ratio of per-unit sums correctly:

import numpy as np

def ratio_delta(num, den):
    """num, den: arrays of per-user sums, e.g. cost and conversations."""
    n = len(num)
    mx, my = num.mean(), den.mean()
    vx, vy = num.var(ddof=1), den.var(ddof=1)
    cxy = np.cov(num, den, ddof=1)[0, 1]
    var = (vx / my**2 - 2 * mx * cxy / my**3 + mx**2 * vy / my**4) / n
    return mx / my, var

def compare(ctrl, trt):
    r0, v0 = ratio_delta(*ctrl)
    r1, v1 = ratio_delta(*trt)
    diff, se = r1 - r0, np.sqrt(v0 + v1)
    return diff, (diff - 1.96 * se, diff + 1.96 * se)

Pass in arrays with one row per user, for example total cost and conversation count per user in each arm. The interval on the difference is what you compare against the guardrail threshold.

Check sample ratio mismatch first

Before reading any metric, check that the arms received the traffic you allocated. A 50/50 split that arrives as 50,912 versus 49,088 looks close, but a chi-square test gives 33.27, with p of about 8e-9. Something is dropping or misrouting traffic in one arm. Common LLM-specific causes: the new prompt times out more often and those calls never log; a parsing error in the treatment path throws before the exposure log is written; or a response cache keyed without the arm serves control answers to treatment users. Results from an experiment with sample ratio mismatch cannot be trusted in either direction. Find the cause and restart.

Confounders that are specific to prompts

  • Prompt caching favours the incumbent. Where the provider caches prompt prefixes, the control prompt is already warm, so for its first hours a new prompt shows higher latency and input cost. Log cached tokens separately, and agree a warm-up exclusion window for cost and latency before launch.
  • Longer prompts cost on every call. A system prompt that grows by 800 tokens adds input cost to every turn of every conversation. Set the guardrail on cost per conversation, not cost per call.
  • Model drift. If both arms use a floating model alias and the provider updates it mid-test, both arms change. If only one arm pins its model, the test mixes a prompt change with a model change. Pin the model ID in both arms and record it on every log line.
  • Non-determinism. Sampling adds variance to each response. It does not bias the comparison, but it raises the sample size you need. Keep temperature the same in both arms unless temperature is what you are testing.
  • Novelty and learning. Users adapt to a new style. Run for at least one full weekly cycle, and look at the effect by days since first exposure.
  • Shared state. Memory features, retrieved conversation history or cached answers can carry one arm's outputs into the other arm's context.

Peeking, ramps and stopping

Checking a fixed-horizon test every day and stopping at the first significant result pushes the false-positive rate well above 5%. Either fix the sample size in advance and analyse once, or use a sequential method built for continuous monitoring, such as alpha-spending boundaries or always-valid confidence sequences. Guardrails are different. Monitor them continuously with conservative thresholds, because stopping early for harm costs little.

A typical ramp runs 1% per arm for a day to catch crashes, parse failures and sample ratio mismatch, then 5 to 10% per arm until the planned sample is reached. You ship by moving the production pointer in the registry. If you want to measure long-term effects, keep a small holdout on the old prompt for a few weeks.

Worked example: a concise support prompt

A support assistant's control prompt produces thorough answers. The candidate, "concise", asks for short answers followed by an offer to go deeper. Offline evals showed equal correctness and 35% fewer output tokens. The online plan, written before launch, was:

  • Primary metric: resolution without escalation, with a 62% baseline and a 2-point MDE, so at least 9,150 conversations per arm, a floor before allowing for clustering.
  • Non-inferiority margin: 2 points.
  • Randomization: by conversation, at 5% per arm, since most users here ask one question.
  • Guardrails: cost per conversation must not rise; p95 latency must not rise by more than 10%; escalation rate must not rise by more than 1 point.
  • Warm-up: cost and latency are excluded for the first 24 hours, because of prompt caching.

The day-1 sample ratio check passed. Early on, latency for the concise arm looked 12% worse, and the logs showed near-zero cached tokens in its first hours, which is the case the warm-up rule covers. After 11 days both arms had passed 9,150 conversations. Resolution was 63.1% against 62.0%. The 95% interval on the difference ran from about -0.3 to +2.5 points, so the gain is not established. However, the lower bound sits well inside the 2-point non-inferiority margin agreed in advance. Cost per conversation fell 18% with a tight interval, and judge helpfulness was unchanged after length-bias calibration. The decision was to ship for cost, recording in the registry that the effect on resolution was inconclusive.

Failure modes

  • Assigning per request, so conversations mix prompts.
  • Logging exposure at assignment, which dilutes every effect.
  • Using a judge that is not blind or not pinned, or never checked against human labels.
  • Running per-message tests on user-randomized data, which makes intervals too narrow.
  • Stopping at the first significant peek.
  • Shipping on a diagnostic metric ("answers got shorter") when the primary metric did not move.

Trade-offs

ChoiceYou gainYou pay
User rather than conversation as the unitConsistent experience; captures return effectsFewer units, so longer tests
Judge metricsDense, fast signalInstrument bias; needs calibration
Small MDEDetects small gainsLong tests and more exposure to a worse arm
Sequential testingValid early stoppingWider intervals at a fixed sample size
Holdout after shippingMeasures long-term effectsSome users keep the old prompt

What to do next

  1. Write a one-page plan before launch: unit, primary metric, MDE, guardrails with thresholds, warm-up window and analysis method.
  2. Implement hashed assignment with a fresh salt per experiment, and log the prompt hash, model ID and cached tokens on every call.
  3. Validate your judge against a few hundred human labels, and check it for length bias.
  4. Compute the sample size with the code above and estimate the test's duration from real traffic.
  5. Automate a sample ratio check on day 1 and every day after.
  6. Analyse ratio metrics per unit with the delta method, then record the decision and its evidence in the prompt registry.
Key takeaway: An online prompt test answers what offline evals cannot: whether real users do better at an acceptable cost. Randomize by user or conversation with a salted hash, log exposure, prompt hash and model ID on every call, and fix one primary metric, guardrails and an MDE before launch. Validate and pin the judge, size the test from the baseline rate, use the delta method for per-unit ratios, check sample ratio mismatch first, and account for prompt caching, model drift and novelty. Do not stop on a lucky peek; ship by moving the registry pointer and record what was and was not shown.