A good A/B test of an agent is designed before any code runs: the metric, the unit of randomisation, the sample size and the analysis. That half is covered in A/B Testing Agent Variants in ADK Java, including hashing users to arms, exposure logging, SRM checks and the delta method. This article is about the other half: operating the experiment while real users are in it.
Running is where agent experiments go wrong in ways web experiments rarely do. A conversation started on one variant cannot safely continue on the other. Ramping up allocation changes who is in each arm. Two variants share a model quota and a memory store, so one can damage the other. And a treatment that occasionally issues a refund twice has to be stopped in hours, not at the end of a two-week test, without the repeated checking turning noise into false alarms. ADK details are from the google-adk 1.11.0 jar.
The experiment at run time
Three pieces of state drive everything: the experiment config, which only operators and the monitor change; the arm stored in each session; and the exposure and outcome logs tagged with the arm. The rest of this article is about keeping those three consistent while traffic flows.
Variants as specs, one runner per arm
Describe each arm as data, not as a code branch. A variant spec names the model, the instruction version and the tool set, and the experiment config names the two specs, the salt, the treatment threshold and a kill flag. Build one agent and one Runner per arm at startup, and give both the same app name and the same session, artifact and memory services. The runners differ only in the agent they execute.
record VariantSpec(String id, String model, String instructionVersion, List<String> tools) {}
Runner runnerFor(VariantSpec v) {
LlmAgent agent = LlmAgent.builder()
.name("support") // same name in both arms
.model(v.model())
.instruction(instructions.load(v.instructionVersion()))
.tools(toolRegistry.resolve(v.tools()))
.build();
return new Runner(agent, APP, artifacts, sessions, memory, List.of(exposurePlugin));
}
Map<String, Runner> arms = Map.of(
"control", runnerFor(cfg.control()),
"treatment", runnerFor(cfg.treatment()));Keeping the agent name identical matters: event authors and any analysis keyed on them then compare like with like. Treat a spec as immutable once the experiment starts. Editing the treatment instruction mid-test silently creates a third variant whose results are pooled with the second; if you need a change, end the test and start a new one with a new salt. Before any of this goes live, compare the two specs offline on a fixed dataset, as in dataset-driven evaluation; an online test should confirm a promising change, not discover a broken one.
Pin the arm to the session
Assignment hashes the user; execution must follow the session. A session's history contains the instructions, tool calls and tool results of the arm that produced it. Hand that history to the other arm and it may see calls to tools it does not have, or continue a plan written under different rules. So the arm is decided once, when the session is created, written into session state, and read on every later turn.
static final String ARM_KEY = "exp.reply_style_v3.arm";
Runner resolve(String userId, String sessionId) {
Session s = sessions.getSession(APP, userId, sessionId, Optional.empty()).blockingGet();
String arm = (String) s.state().get(ARM_KEY);
if (arm == null) arm = "control"; // sessions older than the test
if (cfg.killed() && arm.equals("treatment")) {
return drainPolicy.route(s); // finish on treatment, or hand off to a NEW control session
}
return arms.get(arm);
}
Session newSession(String userId) {
String arm = (!cfg.killed() && bucket(cfg.salt(), userId) < cfg.thresholdBp())
? "treatment" : "control";
return sessions.createSession(APP, userId, Map.of(ARM_KEY, arm), null).blockingGet();
}The pinning rule has one consequence you must decide in advance: what happens to treatment sessions when the experiment is stopped. If the treatment is merely worse, letting open sessions finish on it is cleanest. If it is harmful, the drain policy hands each one to a new control session seeded with a short summary, never the treatment's raw history, and records the link between the two sessions for analysis. Write the policy down; the moment you need it is a bad moment to design it.
Ramping without Simpson's paradox
Nobody should send half of production to an untested agent on day one. Ramps go 1%, 5%, 20%, 50%, each stage held long enough to check guardrails. Two rules keep a ramp analysable.
Make it monotone. With bucket(salt, user) < thresholdBp on a 0 to 9,999 bucket, raising the threshold only moves users from control to treatment, never back. A user in treatment at 5% is still in treatment at 20%. Crossover does not vanish: the users who move keep their old control sessions, so they own sessions in both arms. But it becomes one-directional and countable, while a scheme that reshuffles at each stage moves users both ways and contaminates both arms.
Never pool across stages. The allocation changes over time, and so does the baseline. Suppose week one runs at 5% treatment and week two at 50%, and week two has harder traffic:
| Control users | Control resolved | Treatment users | Treatment resolved | |
|---|---|---|---|---|
| Week 1 (5%) | 95,000 | 60% | 5,000 | 62% |
| Week 2 (50%) | 50,000 | 50% | 50,000 | 51% |
| Pooled | 145,000 | 56.6% | 55,000 | 52.0% |
Treatment wins in both weeks and loses by 4.6 points pooled, because most of its users arrived in the harder week. This is Simpson's paradox, and ramps manufacture it. Analyse each stage separately and combine with stage weights, or, more simply, make the decision on the final stable stage only and use the earlier stages for guardrails. If one user can appear in several stages, attribute them to the stage and arm of their first exposure, or exclude the movers from the comparison.
Interference between arms
Randomisation assumes one user's arm does not affect another user's outcome. Agents break that assumption through shared infrastructure:
| Shared resource | How arms interfere | Mitigation |
|---|---|---|
| Model quota and rate limits | a token-hungry treatment causes 429s for control, making control look worse | separate quota pools or headroom; track throttling per arm |
| Caches (responses, tool results) | one arm's cached answers are served to the other | include the variant id in cache keys |
Memory service and user: state | treatment writes memories that control later reads | namespace memory and state keys by experiment arm, or freeze writes |
| Tool backends | treatment's extra load slows a shared API for both | load-test the treatment's call rate first |
| Human reviewers and judges | graders who can see the arm drift toward it | blind arm labels in review queues and judge prompts |
The quota row is the most common and the most misleading, because it makes the treatment look relatively better while harming the product overall. A 50/50 split hides it best; a small ramp stage with a large control arm hides it worst, because control absorbs most of the throttling.
Guardrails that stop the test
The primary metric is analysed once, at the planned end. Guardrails are watched continuously and can stop the test early. For agents, the important guardrails are rare harmful events: a write tool repeated in one invocation, a policy block, an escalation to a human, a complaint. Their counts are small, so use an exact test that works with small counts.
Conditional binomial test. Let s be the treatment's share of exposures in the stage, measured from exposure logs, not assumed from the threshold. If both arms have the same per-exposure event rate, then given k events in total, the number in treatment follows Binomial(k, s). The p-value for harm is the upper tail. No variance estimate and no normal approximation are involved.
Multiplicity. Checking G guardrails at L planned looks gives G × L chances of a false alarm. The simplest valid correction is Bonferroni over both: test each guardrail at each look at 0.05 / (G × L). With 2 guardrails checked every 6 hours over 25 looks, that is 0.05 / 50 = 0.001. It is conservative; group-sequential boundaries and always-valid methods such as mSPRT spend the error more efficiently, but the Bonferroni version is easy to audit, and for harm signals conservatism in the false-alarm direction is affordable because true harm tends to produce lopsided counts quickly.
// runs at each planned look; looks are scheduled, never triggered by a bad-looking number
boolean shouldStop(long kTotal, long kTreatment, double treatmentShare, double alphaPerCheck) {
double p = 0;
for (long x = kTreatment; x <= kTotal; x++) {
p += binomialPmf(x, kTotal, treatmentShare); // exact upper tail
}
return p < alphaPerCheck;
}
void look() {
for (Guardrail g : cfg.guardrails()) {
Counts n = outcomes.counts(g, cfg.currentStage());
if (shouldStop(n.total(), n.treatment(), exposures.treatmentShare(), 0.05 / (G * L))) {
cfg.kill("guardrail " + g.name()); // threshold to 0; new sessions get control
alerts.page("ExperimentAutoStopped", g.name());
return;
}
}
}Continuous guardrails, such as cost per session, are better handled with a pre-agreed non-inferiority bound checked at the same looks with the same alpha split, using the interval method from the design article. The stop itself is a config change: killed = true sets the effective threshold to zero, new sessions get control, and the drain policy handles open ones.
Worked example: a refund tool stopped at 10%
A team tests a treatment that lets the billing agent call issue_refund directly instead of drafting a refund request. Offline evaluation looks good. The ramp starts at 1% for a day, then 10%. Exposure logs at the 10% stage show a treatment share of 0.10. The two guardrails are repeated write tools and human escalations, checked every 6 hours, 25 looks planned: alpha per check is 0.001.
At the third look of the 10% stage there have been 9 repeated-write events in total. If 4 were in treatment, P(X ≥ 4 | 9, 0.1) is 0.0083, above 0.001: keep watching, though a person should read those four traces. If 5 were in treatment, the tail is 0.00089, which crosses. In the actual run 7 of the 9 were in treatment, a tail of 3.0 × 10⁻⁶. The monitor sets the kill flag, pages the experiment owner, and new sessions go to control.
About 120 treatment sessions are open. Because the harm is a repeated money movement, the drain policy moves them to control with a summary context. The traces show the treatment's retry instruction re-issuing refunds after tool timeouts: a fix to make the tool idempotent, not a reason to abandon the idea. The next experiment gets a new salt, a new spec id, and starts again at 1%.
Failure modes
| Failure | Effect | Prevention |
|---|---|---|
| Arm chosen per request | sessions switch arms mid-conversation | pin the arm in session state at creation |
| Reshuffling ramp | users bounce between arms | monotone threshold on a fixed salted bucket; attribute movers to first exposure |
| Pooling across ramp stages | Simpson reversal | decide on the stable stage or weight by stage |
| Spec edited mid-test | two treatments pooled as one | immutable specs; new test, new salt |
| Peeking at the primary metric | inflated false positives | only guardrails stop early, at planned looks |
| Assumed exposure share | biased guardrail tests | measure the share from exposure logs |
| Shared quota or cache | control harmed by treatment | variant-keyed caches, per-arm throttle metrics |
The auto-stop is itself a production control, so it needs the same care as the containment levers in the incident response playbook: test that setting the flag really routes new sessions to control, and alert if the monitor has not completed a look on schedule.
Ending the experiment
An experiment ends with a decision, and the decision has mechanics. To ship, raise the threshold to 10,000 so all new sessions get treatment, then make the treatment spec the new control and retire the experiment config. Remove the arm key only after open sessions have aged out, since resolve falls back to control when the key is missing. To abandon, set the kill flag and let the drain policy run. Either way, retire the salt so it is never reused, keep the exposure and outcome logs with the spec ids, and record the result where the next team will look. For changes that are rollouts rather than questions, use a canary instead, as in canary deploys for agents.
What to do next
- Write the variant specs and experiment config as versioned data; build one Runner per arm on shared services.
- Pin the arm in session state at creation and write the drain policy for both worse and harmful outcomes.
- Use a monotone threshold ramp and plan the stages and their durations up front.
- Pick two or three rare-event guardrails, schedule the looks, and compute alpha per check.
- Measure exposure share from logs and implement the exact binomial tail.
- Key caches and memory by arm, and add per-arm throttling metrics.
- Rehearse the kill flag in staging and alert on missed looks.
- Decide on the final stable stage only, and retire the salt when done.