Teams change agents constantly: a reworded instruction, a new tool, a different model, a planner in front of a specialist. Offline evaluation says whether a change passes the test set. It cannot say whether real users resolve more problems, abandon fewer conversations or cost less to serve, because real traffic is messier than any dataset. An A/B test answers that by randomly splitting users between the current agent and the candidate and comparing outcomes.
Agents make A/B testing harder than it is for web pages. A conversation spans many turns and model calls, outcomes arrive late, and variants can share memory and caches. This article designs an experiment for an ADK Java agent from the hypothesis down: user-level assignment persisted in session state, exposure and telemetry logged through callbacks, sizing, a sample-ratio check and a ratio-metric analysis, with a worked example and the threats specific to agents.
A/B test or canary?
A canary and an A/B test both run two configurations side by side, but they ask different questions. A canary asks whether the new version is safe to roll out: small traffic, watching for regressions, rolling back fast. That is covered in canary deployments for ADK Java agents. An A/B test asks which version is better on a metric you chose in advance: a fixed split, a fixed duration and an analysis that holds up to scrutiny. Run the canary first so the experiment does not expose half your users to a broken build; then run the A/B test to decide whether the change is worth keeping. Comparing models offline before either step is covered in comparing Gemini variants.
Design the experiment before the code
Write the design down before any code, and keep it unchanged while the test runs.
- Hypothesis. "The v2 triage instruction raises the share of support sessions resolved without human handoff."
- Primary metric. One number that decides the test: resolved sessions divided by sessions, per arm. Define resolved precisely, for example no handoff and no new ticket within 72 hours.
- Guardrails. Metrics that must not get worse: cost per session, p95 turn latency, tool error rate, safety flags, user-reported dissatisfaction.
- Unit of randomisation. The user, not the turn or the session. A user who sees both variants contaminates both, and their sessions are correlated anyway.
- Minimum detectable effect and duration. The smallest change worth shipping, which sets the sample size, and at least one full weekly cycle.
Architecture
Assignment happens once per user, before the session starts. The arm is stored in session state, both arms run in the same service against the same session store, callbacks write exposure and per-call telemetry, and outcome events that arrive later (a handoff, a reopened ticket) are joined by user and session in the analysis job.
Assigning users and running both variants
Assignment must be deterministic, uniform and independent across experiments. Hashing the user id with a per-experiment salt through SHA-256 gives all three: the same user always gets the same arm, buckets are evenly spread, and a new salt reshuffles users so one test's arms do not line up with the next. The arm is also written into state under the user: prefix, which ADK uses for state shared across a user's sessions; whether that persists across restarts depends on the session service you configure.
import com.google.adk.agents.LlmAgent;
import com.google.adk.artifacts.InMemoryArtifactService;
import com.google.adk.runner.Runner;
import com.google.adk.sessions.BaseSessionService;
import com.google.adk.sessions.Session;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import java.util.Map;
final class Experiment {
static final String KEY = "user:exp.triage_prompt_v2"; // user-scoped state
final String salt; // new salt per experiment, never reused
final int treatmentBp; // 5000 = 50%, fixed for the whole test
Experiment(String salt, int treatmentBp) { this.salt = salt; this.treatmentBp = treatmentBp; }
String arm(String userId) throws Exception {
byte[] d = MessageDigest.getInstance("SHA-256")
.digest((salt + ":" + userId).getBytes(StandardCharsets.UTF_8));
long v = 0;
for (int i = 0; i < 8; i++) v = (v << 8) | (d[i] & 0xff);
return Long.remainderUnsigned(v, 10_000) < treatmentBp ? "treatment" : "control";
}
}
Runner runnerFor(LlmAgent agent, BaseSessionService sessions) {
return Runner.builder().agent(agent).appName("support")
.sessionService(sessions) // shared by both arms
.artifactService(new InMemoryArtifactService())
.build();
}
Session startSession(Experiment exp, BaseSessionService sessions, String userId)
throws Exception {
String arm = exp.arm(userId); // deterministic, so stable
return sessions.createSession("support", userId,
Map.of(Experiment.KEY, arm), null).blockingGet();
}
// One runner per arm; agents come from buildAgent(...) in the next snippet.
Map<String, Runner> runners = Map.of(
"control", runnerFor(controlAgent, sessions),
"treatment", runnerFor(treatmentAgent, sessions));Two runners share one session service and one app name, so a user's history is the same object whichever arm serves them. Route each turn to runners.get(arm) and call runAsync(userId, sessionId, message) as usual. Keep everything except the change under test identical: same model id, tools, agent name and generation settings.
Logging exposure and telemetry with callbacks
An experiment is only as good as its log. Two ADK callbacks cover it without touching agent logic. beforeAgentCallbackSync runs before the agent does any work, so it records exposure even if the model call later fails. afterModelCallbackSync sees every model response and records tokens and the served model version. Both return Optional.empty(), which tells ADK to leave the run unchanged.
LlmAgent buildAgent(String arm, String modelId, String instruction,
List<BaseTool> tools, ExperimentLog log) {
return LlmAgent.builder()
.name("triage") // same name in both arms
.model(modelId) // from config, same in both arms
.instruction(instruction) // the only difference
.tools(tools)
.beforeAgentCallbackSync(ctx -> { // exposure: before any model call
log.exposure(ctx.userId(), ctx.sessionId(), ctx.invocationId(), arm,
String.valueOf(ctx.state().get(Experiment.KEY)));
return Optional.empty(); // never alter the run
})
.afterModelCallbackSync((ctx, resp) -> { // per-call telemetry
int in = resp.usageMetadata().flatMap(u -> u.promptTokenCount()).orElse(0);
int out = resp.usageMetadata().flatMap(u -> u.candidatesTokenCount()).orElse(0);
log.modelCall(ctx.userId(), ctx.invocationId(), arm,
resp.modelVersion().orElse("?"), in, out);
return Optional.empty(); // keep the response unchanged
})
.build();
}Logging exposure before the model call matters. If exposure were logged only after a successful response, users whose calls errored or timed out would vanish from the arm that caused the errors, and the comparison would quietly favour the worse variant. Analyse every assigned, exposed user (intent to treat), whatever happened next. More on callback composition is in ADK Java callbacks.
Sizing in users, not turns
For a proportion near p and a minimum detectable difference delta at 5 percent significance and 80 percent power, the classic per-arm size is n = 2 (1.96 + 0.84)^2 p(1 - p) / delta^2. With a baseline resolution rate of 0.61 and a 2-point effect, that is about 9,340 independent units per arm.
Sessions from one user are not independent: some users have harder problems every time. Inflate by the design effect 1 + (m - 1) rho, where m is sessions per user and rho the within-user correlation, or size directly in users from a pilot. Variance reduction can shrink the required sample. CUPED subtracts theta x (x - mean x) from each user's metric, where x is the same metric from before the test and theta = cov(y, x) / var(x); returning users with history gain the most, new users gain nothing. Then fix the duration in advance. Checking the p-value every morning and stopping when it dips below 0.05 inflates false positives badly; if you must look early, use a sequential method with adjusted thresholds, as the canary article does for rollback.
Analysis: SRM first, then the metric
Analysis runs in two steps. First, the sample ratio mismatch (SRM) check: if a 50/50 split produced visibly unequal counts, something in assignment, routing or logging is broken, and no metric from the test can be trusted until it is explained. Second, the metric. Resolution rate is a ratio of two per-user sums, so its variance must account for users with many sessions. The delta method does that without per-session independence assumptions.
import math
from statistics import fmean
def srm_p(n_a, n_b, share_a=0.5):
"""Chi-square (1 df) test that assignment counts match the planned split."""
tot = n_a + n_b
e_a, e_b = tot * share_a, tot * (1 - share_a)
chi2 = (n_a - e_a) ** 2 / e_a + (n_b - e_b) ** 2 / e_b
return math.erfc(math.sqrt(chi2 / 2))
def ratio_metric(users):
"""users: list of (resolved_sessions, sessions). Rate and delta-method variance."""
k = len(users)
y = [u[0] for u in users]; n = [u[1] for u in users]
my, mn = fmean(y), fmean(n)
r = my / mn
vy = sum((a - my) ** 2 for a in y) / (k - 1)
vn = sum((b - mn) ** 2 for b in n) / (k - 1)
cyn = sum((a - my) * (b - mn) for a, b in zip(y, n)) / (k - 1)
return r, (vy - 2 * r * cyn + r * r * vn) / (k * mn * mn)
def compare(control, treatment):
ra, va = ratio_metric(control)
rb, vb = ratio_metric(treatment)
diff, se = rb - ra, math.sqrt(va + vb)
p = math.erfc(abs(diff / se) / math.sqrt(2))
return dict(control=ra, treatment=rb, diff=diff,
ci=(diff - 1.96 * se, diff + 1.96 * se), p=p)
Worked example: a triage instruction change
The worked example uses data simulated for illustration, analysed with the code above. A support agent tests the v2 triage instruction against v1 at 50/50 for two weeks. The first readout has 10,012 users in control and 9,580 in treatment. srm_p(10012, 9580) is about 0.002, so the team stops and investigates instead of reading the metric. The cause: exposure was logged after the model response, and v2's longer prompts produced more timeouts, so failed treatment users were never counted. Moving exposure to beforeAgentCallbackSync recovers them: 10,012 and 9,988, with an SRM p-value of about 0.87.
| Control (v1) | Treatment (v2) | |
|---|---|---|
| Users | 10,012 | 9,988 |
| Sessions | 19,592 | 19,726 |
| Resolution rate | 60.3% | 63.0% |
The difference is +2.7 points with a 95 percent interval of +1.7 to +3.7 points. The delta-method standard error is about 0.0052, against 0.0049 from a naive per-session binomial formula; on real traffic with stronger per-user correlation the gap is larger, and the naive interval is the one that misleads. Guardrails decide the rest: if cost per session rose by more than the agreed limit, the win is not free. Wiring the same metrics into dashboards is covered in ADK Java metrics dashboards.
Threats to validity specific to agents
- Interference through shared state. If both arms write to the same long-term memory or semantic cache, the treatment's outputs can leak into control's context. Partition memory and caches by arm for the test.
- Silent model drift. A model alias can change under both arms mid-test. Record
modelVersion()per call and check it is identical across arms. - Novelty and learning effects. Users react to a changed agent at first and settle later. Plot the effect by day and do not stop in the first week.
- Judge bias. If quality is scored by an LLM judge, blind it to the arm and validate it against human labels on a sample.
- Carryover. Users who were in one arm of the last test carry its effects; fresh salts help, and a cooling-off period helps more.
- Late outcomes. Resolution is only known once the 72-hour window closes, so users from the last three days of the test are still open; cut the analysis at the last fully observed day for both arms.
Failure modes
- Per-turn randomisation. Users see both agents in one conversation; the comparison measures confusion.
- Metric chosen after the fact. With ten candidate metrics, one will look significant by chance.
- Ignoring SRM. The most common cause of wrong conclusions, and the cheapest to detect.
- Changing the split mid-test. Ramping from 10 to 50 percent mixes cohorts with different start dates; analyse each period separately or keep the split fixed.
- Shipping on the primary metric alone. A resolution gain that doubles cost or latency is a different decision.
What to do next
- Write a one-page design: hypothesis, primary metric, guardrails, unit, effect size and duration.
- Add the hashed assignment and the
user:state key to session creation. - Add the two callbacks and confirm exposure counts arrive for both arms in a staging run.
- Run an A/A test (both arms identical) for a week; the SRM check should pass and the difference should be indistinguishable from zero.
- Size the real test from the A/A data and fix its duration.
- Analyse with SRM first, then the delta method, then guardrails; write up the decision.