A good A/B test of an agent is designed before any code runs: the metric, the unit of randomisation, the sample size and the analysis. That half is covered in A/B Testing Agent Variants in ADK Java, including hashing users to arms, exposure logging, SRM checks and the delta method. This article is about the other half: operating the experiment while real users are in it.

Running is where agent experiments go wrong in ways web experiments rarely do. A conversation started on one variant cannot safely continue on the other. Ramping up allocation changes who is in each arm. Two variants share a model quota and a memory store, so one can damage the other. And a treatment that occasionally issues a refund twice has to be stopped in hours, not at the end of a two-week test, without the repeated checking turning noise into false alarms. ADK details are from the google-adk 1.11.0 jar.

The experiment at run time

An experiment at run timeRequestuserId, sessionIdArm resolverpinned or assignedRunner: controlspec ARunner: treatmentspec BShared session storearm in session stateExposure + outcomestagged with armGuardrail monitorplanned looksExperiment configthreshold, killed, saltstop: threshold = 0read each turnAssignment changes only through config; the arm of an existing session never changes.
The arm resolver reads config for new sessions and session state for existing ones; the guardrail monitor can only stop the test, never extend it.

Three pieces of state drive everything: the experiment config, which only operators and the monitor change; the arm stored in each session; and the exposure and outcome logs tagged with the arm. The rest of this article is about keeping those three consistent while traffic flows.

Variants as specs, one runner per arm

Describe each arm as data, not as a code branch. A variant spec names the model, the instruction version and the tool set, and the experiment config names the two specs, the salt, the treatment threshold and a kill flag. Build one agent and one Runner per arm at startup, and give both the same app name and the same session, artifact and memory services. The runners differ only in the agent they execute.

record VariantSpec(String id, String model, String instructionVersion, List<String> tools) {}

Runner runnerFor(VariantSpec v) {
  LlmAgent agent = LlmAgent.builder()
      .name("support")                                 // same name in both arms
      .model(v.model())
      .instruction(instructions.load(v.instructionVersion()))
      .tools(toolRegistry.resolve(v.tools()))
      .build();
  return new Runner(agent, APP, artifacts, sessions, memory, List.of(exposurePlugin));
}

Map<String, Runner> arms = Map.of(
    "control",   runnerFor(cfg.control()),
    "treatment", runnerFor(cfg.treatment()));

Keeping the agent name identical matters: event authors and any analysis keyed on them then compare like with like. Treat a spec as immutable once the experiment starts. Editing the treatment instruction mid-test silently creates a third variant whose results are pooled with the second; if you need a change, end the test and start a new one with a new salt. Before any of this goes live, compare the two specs offline on a fixed dataset, as in dataset-driven evaluation; an online test should confirm a promising change, not discover a broken one.

Pin the arm to the session

Assignment hashes the user; execution must follow the session. A session's history contains the instructions, tool calls and tool results of the arm that produced it. Hand that history to the other arm and it may see calls to tools it does not have, or continue a plan written under different rules. So the arm is decided once, when the session is created, written into session state, and read on every later turn.

static final String ARM_KEY = "exp.reply_style_v3.arm";

Runner resolve(String userId, String sessionId) {
  Session s = sessions.getSession(APP, userId, sessionId, Optional.empty()).blockingGet();
  String arm = (String) s.state().get(ARM_KEY);
  if (arm == null) arm = "control";                    // sessions older than the test
  if (cfg.killed() && arm.equals("treatment")) {
    return drainPolicy.route(s);      // finish on treatment, or hand off to a NEW control session
  }
  return arms.get(arm);
}

Session newSession(String userId) {
  String arm = (!cfg.killed() && bucket(cfg.salt(), userId) < cfg.thresholdBp())
      ? "treatment" : "control";
  return sessions.createSession(APP, userId, Map.of(ARM_KEY, arm), null).blockingGet();
}

The pinning rule has one consequence you must decide in advance: what happens to treatment sessions when the experiment is stopped. If the treatment is merely worse, letting open sessions finish on it is cleanest. If it is harmful, the drain policy hands each one to a new control session seeded with a short summary, never the treatment's raw history, and records the link between the two sessions for analysis. Write the policy down; the moment you need it is a bad moment to design it.

Ramping without Simpson&#x27;s paradox

Nobody should send half of production to an untested agent on day one. Ramps go 1%, 5%, 20%, 50%, each stage held long enough to check guardrails. Two rules keep a ramp analysable.

Make it monotone. With bucket(salt, user) < thresholdBp on a 0 to 9,999 bucket, raising the threshold only moves users from control to treatment, never back. A user in treatment at 5% is still in treatment at 20%. Crossover does not vanish: the users who move keep their old control sessions, so they own sessions in both arms. But it becomes one-directional and countable, while a scheme that reshuffles at each stage moves users both ways and contaminates both arms.

Never pool across stages. The allocation changes over time, and so does the baseline. Suppose week one runs at 5% treatment and week two at 50%, and week two has harder traffic:

Control usersControl resolvedTreatment usersTreatment resolved
Week 1 (5%)95,00060%5,00062%
Week 2 (50%)50,00050%50,00051%
Pooled145,00056.6%55,00052.0%

Treatment wins in both weeks and loses by 4.6 points pooled, because most of its users arrived in the harder week. This is Simpson's paradox, and ramps manufacture it. Analyse each stage separately and combine with stage weights, or, more simply, make the decision on the final stable stage only and use the earlier stages for guardrails. If one user can appear in several stages, attribute them to the stage and arm of their first exposure, or exclude the movers from the comparison.

Interference between arms

Randomisation assumes one user's arm does not affect another user's outcome. Agents break that assumption through shared infrastructure:

Shared resourceHow arms interfereMitigation
Model quota and rate limitsa token-hungry treatment causes 429s for control, making control look worseseparate quota pools or headroom; track throttling per arm
Caches (responses, tool results)one arm's cached answers are served to the otherinclude the variant id in cache keys
Memory service and user: statetreatment writes memories that control later readsnamespace memory and state keys by experiment arm, or freeze writes
Tool backendstreatment's extra load slows a shared API for bothload-test the treatment's call rate first
Human reviewers and judgesgraders who can see the arm drift toward itblind arm labels in review queues and judge prompts

The quota row is the most common and the most misleading, because it makes the treatment look relatively better while harming the product overall. A 50/50 split hides it best; a small ramp stage with a large control arm hides it worst, because control absorbs most of the throttling.

Guardrails that stop the test

The primary metric is analysed once, at the planned end. Guardrails are watched continuously and can stop the test early. For agents, the important guardrails are rare harmful events: a write tool repeated in one invocation, a policy block, an escalation to a human, a complaint. Their counts are small, so use an exact test that works with small counts.

Conditional binomial test. Let s be the treatment's share of exposures in the stage, measured from exposure logs, not assumed from the threshold. If both arms have the same per-exposure event rate, then given k events in total, the number in treatment follows Binomial(k, s). The p-value for harm is the upper tail. No variance estimate and no normal approximation are involved.

Multiplicity. Checking G guardrails at L planned looks gives G × L chances of a false alarm. The simplest valid correction is Bonferroni over both: test each guardrail at each look at 0.05 / (G × L). With 2 guardrails checked every 6 hours over 25 looks, that is 0.05 / 50 = 0.001. It is conservative; group-sequential boundaries and always-valid methods such as mSPRT spend the error more efficiently, but the Bonferroni version is easy to audit, and for harm signals conservatism in the false-alarm direction is affordable because true harm tends to produce lopsided counts quickly.

// runs at each planned look; looks are scheduled, never triggered by a bad-looking number
boolean shouldStop(long kTotal, long kTreatment, double treatmentShare, double alphaPerCheck) {
  double p = 0;
  for (long x = kTreatment; x <= kTotal; x++) {
    p += binomialPmf(x, kTotal, treatmentShare);        // exact upper tail
  }
  return p < alphaPerCheck;
}

void look() {
  for (Guardrail g : cfg.guardrails()) {
    Counts n = outcomes.counts(g, cfg.currentStage());
    if (shouldStop(n.total(), n.treatment(), exposures.treatmentShare(), 0.05 / (G * L))) {
      cfg.kill("guardrail " + g.name());                // threshold to 0; new sessions get control
      alerts.page("ExperimentAutoStopped", g.name());
      return;
    }
  }
}

Continuous guardrails, such as cost per session, are better handled with a pre-agreed non-inferiority bound checked at the same looks with the same alpha split, using the interval method from the design article. The stop itself is a config change: killed = true sets the effective threshold to zero, new sessions get control, and the drain policy handles open ones.

Worked example: a refund tool stopped at 10%

A team tests a treatment that lets the billing agent call issue_refund directly instead of drafting a refund request. Offline evaluation looks good. The ramp starts at 1% for a day, then 10%. Exposure logs at the 10% stage show a treatment share of 0.10. The two guardrails are repeated write tools and human escalations, checked every 6 hours, 25 looks planned: alpha per check is 0.001.

At the third look of the 10% stage there have been 9 repeated-write events in total. If 4 were in treatment, P(X ≥ 4 | 9, 0.1) is 0.0083, above 0.001: keep watching, though a person should read those four traces. If 5 were in treatment, the tail is 0.00089, which crosses. In the actual run 7 of the 9 were in treatment, a tail of 3.0 × 10⁻⁶. The monitor sets the kill flag, pages the experiment owner, and new sessions go to control.

About 120 treatment sessions are open. Because the harm is a repeated money movement, the drain policy moves them to control with a summary context. The traces show the treatment's retry instruction re-issuing refunds after tool timeouts: a fix to make the tool idempotent, not a reason to abandon the idea. The next experiment gets a new salt, a new spec id, and starts again at 1%.

Failure modes

FailureEffectPrevention
Arm chosen per requestsessions switch arms mid-conversationpin the arm in session state at creation
Reshuffling rampusers bounce between armsmonotone threshold on a fixed salted bucket; attribute movers to first exposure
Pooling across ramp stagesSimpson reversaldecide on the stable stage or weight by stage
Spec edited mid-testtwo treatments pooled as oneimmutable specs; new test, new salt
Peeking at the primary metricinflated false positivesonly guardrails stop early, at planned looks
Assumed exposure sharebiased guardrail testsmeasure the share from exposure logs
Shared quota or cachecontrol harmed by treatmentvariant-keyed caches, per-arm throttle metrics

The auto-stop is itself a production control, so it needs the same care as the containment levers in the incident response playbook: test that setting the flag really routes new sessions to control, and alert if the monitor has not completed a look on schedule.

Ending the experiment

An experiment ends with a decision, and the decision has mechanics. To ship, raise the threshold to 10,000 so all new sessions get treatment, then make the treatment spec the new control and retire the experiment config. Remove the arm key only after open sessions have aged out, since resolve falls back to control when the key is missing. To abandon, set the kill flag and let the drain policy run. Either way, retire the salt so it is never reused, keep the exposure and outcome logs with the spec ids, and record the result where the next team will look. For changes that are rollouts rather than questions, use a canary instead, as in canary deploys for agents.

What to do next

  1. Write the variant specs and experiment config as versioned data; build one Runner per arm on shared services.
  2. Pin the arm in session state at creation and write the drain policy for both worse and harmful outcomes.
  3. Use a monotone threshold ramp and plan the stages and their durations up front.
  4. Pick two or three rare-event guardrails, schedule the looks, and compute alpha per check.
  5. Measure exposure share from logs and implement the exact binomial tail.
  6. Key caches and memory by arm, and add per-arm throttling metrics.
  7. Rehearse the kill flag in staging and alert on missed looks.
  8. Decide on the final stable stage only, and retire the salt when done.
Key takeaway: Running an agent A/B test means keeping assignment, sessions and logs consistent while traffic flows. Describe arms as immutable specs with one Runner each on shared services, pin the arm in session state, ramp with a monotone threshold and never pool across stages, and isolate shared quota, caches and memory. Let only rare-harm guardrails stop the test early, with an exact binomial test at planned looks and an alpha split over guardrails and looks.