A regression test answers one question: does behaviour that used to be correct still work after this change? For ordinary code the answer is a deterministic pass or fail. For an agent, the same change can pass on one run and fail on the next, because the model samples, because the provider updated the model behind an alias, or because a tool description edit nudged the model toward a different plan. Teams respond in one of two bad ways: they stop trusting the suite and ignore red builds, or they run every case live on every commit and pay for it in hours and tokens.

This article treats the regression suite as an engineered artifact. It covers where cases should come from, three layers with different costs and guarantees, which change triggers which layer, how many repeated trials a live verdict needs, and how to quarantine and triage without hiding real failures. Companion pages cover neighbouring problems: recorded cassettes and the CI baseline gate and detecting behaviour drift in production. ADK signatures were read from the google-adk 1.11.0 jar with javap.

Where regression cases come from

A regression suite is only as good as its cases, and the best source is your own history. Every incident, escalation or bug report that involved the agent should end with a case that would have caught it: the input, the relevant session state, the expected tool trajectory or a set of must-call and must-not-call rules, and a response check. Give each case an ID that links to the incident, so that when the case fails a year later someone knows why it exists. Add contract cases for behaviour you promised (the agent never quotes a price without calling the pricing tool) and a small set of golden conversations for the main journeys. How to write the checks themselves is covered in writing eval test cases.

Size matters less than coverage of failure history. Two hundred cases, each tied to a real failure or promise, beat two thousand synthetic paraphrases. Tag each case with the intents, tools and procedures it exercises; the tags drive change-impact selection later.

Layer 1: a scripted model

The first layer removes the model. A scripted BaseLlm returns pre-written responses, so the test exercises everything around the model deterministically: tool argument parsing, callbacks, state writes, confirmation handling, the instruction that was built, and the backends your tools call. In ADK Java a model is any subclass of BaseLlm passed to LlmAgent.Builder.model(BaseLlm); it must implement generateContent(LlmRequest, boolean) and connect(LlmRequest).

/** A model that replays a fixed script: deterministic, free, and fast. */
final class ScriptedLlm extends BaseLlm {
  private final Deque<LlmResponse> script;
  final List<LlmRequest> seen = new ArrayList<>();

  ScriptedLlm(List<LlmResponse> turns) {
    super("scripted");
    this.script = new ArrayDeque<>(turns);
  }

  @Override
  public Flowable<LlmResponse> generateContent(LlmRequest request, boolean stream) {
    seen.add(request);
    LlmResponse next = script.poll();
    return next == null
        ? Flowable.error(new AssertionError("model called more times than scripted"))
        : Flowable.just(next);
  }

  @Override
  public BaseLlmConnection connect(LlmRequest request) {
    throw new UnsupportedOperationException("live mode not scripted");
  }
}

static LlmResponse call(String tool, Map<String, Object> args) {
  return LlmResponse.builder()
      .content(Content.fromParts(Part.fromFunctionCall(tool, args))).build();
}

@Test
void refundOverLimitCreatesApprovalBeforeReplying() {
  ScriptedLlm model = new ScriptedLlm(List.of(
      call("lookupOrder", Map.of("orderId", "A-17")),
      call("createApproval", Map.of("orderId", "A-17")),
      LlmResponse.builder().content(Content.fromParts(Part.fromText("Approval requested."))).build()));
  LlmAgent agent = SupportAgent.builder().model(model).build();   // real tools, fake backends
  List<Event> events = run(agent, "Refund order A-17 please");   // InMemoryRunner + blockingIterable
  assertThat(backends.approvals()).containsExactly("A-17");
  assertThat(String.join("\n", model.seen.get(0).getSystemInstructions()))
      .contains("refund.over_limit v3");
}

These tests answer "if the model does the right thing, does our code do the right thing?" and they catch most regressions caused by code changes: a renamed state key, a callback that now blocks a valid call, a tool that serialises a date differently. They also let you assert on what the model was sent, which is where instruction regressions often hide. They say nothing about whether the real model will choose those calls; that is the job of the other two layers. Unit-level patterns for tools and callbacks are in the agent unit testing article.

Layers 2 and 3, and which change runs which

Which change runs which layer of the regression suiteTool or callback codeInstruction / procedureTool schema or descriptionModel version or ADK upgradeL1 scripted modelevery commit, secondsL2 recorded replayevery PR, minutesL3 live, k trials per caseon trigger, tens of minutesVerdict per casepass / regress / flakyQuarantine + triage queueCheap deterministic layers run always; expensive statistical runs only when the model's input or the model changes.
Code changes run the scripted layer; changes to the model's input or the model itself also run recorded replay and live repeated trials for the affected cases.

The second layer replays recorded model responses keyed by the request, so a change that leaves model inputs identical replays deterministically, and a change that alters them shows up as a cache miss that must be re-recorded. It is cheap enough for every pull request. The third layer calls the real model, several times per case, and is the only layer that can detect a regression caused by the model's own behaviour. Run it when the model's input or the model changes, as the diagram shows, and on a schedule to catch silent provider updates.

ChangeL1 scriptedL2 recordedL3 live, k trials
Tool implementation, callback, state handlingYesYesNo
Instruction or procedure textYesRe-record missesYes, tagged cases
Tool name, description or schemaYesRe-record missesYes, cases using the tool
Model version, generation config, ADK upgradeYesRe-record allYes, full suite
Nothing (nightly)NoNoYes, rotating sample

How many trials a live verdict needs

A live case is a coin with an unknown bias. If a case passes with probability 0.9 per run, it fails one run in ten with no regression at all, and the chance that it passes five runs in a row is 0.9 to the fifth, about 0.59. Single runs therefore produce a stream of false alarms, and the fix is repeated trials with a decision rule that accounts for noise. Run each case k times on the baseline and k times on the candidate, and compare pass counts with a one-sided Fisher exact test, which is valid at these small counts.

Baseline passesCandidate passesOne-sided pVerdict at p < 0.05
10/107/100.105Not significant; run more trials
20/2017/200.115Not significant
20/2014/200.010Regression

Two lessons fall out of those numbers. Ten trials cannot reliably tell a case that always passes from one that passes 70 percent of the time, so do not let one red L3 run block a release. And even 20 out of 20 passes only bounds the true pass rate above about 0.84 at 95 percent confidence (the Wilson lower bound; for 10 out of 10 it is about 0.72). Use a sequential plan: start with 5 trials per side, stop early when both sides are 5 out of 5, and extend to 20 only for cases that drop. With many cases, also control for multiple comparisons; out of 200 independent tests at p < 0.05 you expect about ten false alarms, so either lower the threshold or require a failing case to reproduce on a second run. For a case the business cannot tolerate failing even occasionally, such as never charging without confirmation, measure the stronger property that all k trials pass, and treat any single failure as a defect.

The verdict function and versioned baselines

The decision rule fits in a small function. passes runs the case through a fresh InMemoryRunner k times and scores each run; Fisher is a few lines of binomial-coefficient arithmetic. Log every verdict with the model ID, instruction hash and tool-schema hash, so a later reader can reproduce the comparison.

record Verdict(String caseId, int basePass, int candPass, int k, String outcome) {}

Verdict judge(Case c, Agent base, Agent cand) {
  int k = 5;
  int b = passes(c, base, k), n = passes(c, cand, k);
  if (b == k && n == k) return new Verdict(c.id(), b, n, k, "PASS");
  k = 20;                                               // extend only when needed
  b += passes(c, base, k - 5); n += passes(c, cand, k - 5);
  double p = Fisher.oneSidedLess(n, k, b, k);           // P(candidate this low | same rate)
  if (c.critical() && n < k) return new Verdict(c.id(), b, n, k, "REGRESS");
  if (p < 0.01) return new Verdict(c.id(), b, n, k, "REGRESS");
  if (b < k * 0.8) return new Verdict(c.id(), b, n, k, "FLAKY_BASELINE");
  return new Verdict(c.id(), b, n, k, "PASS");
}

Baselines are versioned, not overwritten. Store pass counts per case for each released configuration, and compare a candidate against the configuration it replaces. When an intended change alters behaviour, such as a new confirmation step, update the expected trajectory in the same pull request with a reviewer's approval; a baseline that moves silently is how a suite turns into a record of whatever the agent currently does.

Quarantine and triage

Some cases are flaky on the baseline itself: they pass 60 percent of the time no matter what you change. Quarantine them automatically when the baseline rate drops below a threshold, keep running them, and report their rates separately so the gate stays meaningful. Quarantine is a queue with an owner and a deadline, not a graveyard; a case quarantined for a month is either rewritten with a looser check (for example, must-call rules instead of an exact trajectory) or deleted with a note.

When a case regresses, triage in a fixed order. First, did the model's input change? Diff the recorded request against the baseline's; most regressions are an instruction or tool description edit. Second, did the model change? Check the model ID and provider release notes. Third, is the check wrong? A correct new behaviour that fails an over-strict assertion is a test bug. Only then debug the agent. Write the conclusion on the case, so the next person who sees it red has history to read.

Worked example: a tool description edit

A team changes the description of checkReturnWindow to mention store credit. L1 passes: no code changed. L2 reports 31 cache misses, all in cases tagged with that tool, and the change-impact rule schedules those 31 for L3. At 5 trials per side, 27 cases are 5/5 on both and stop. Four extend to 20. Three end within noise. One, INC-2291-refund-damaged-item, goes from 20/20 to 13/20, a one-sided p of about 0.004: with the new description the model now offers store credit for damaged items, which the incident that created the case had ruled out.

Triage takes ten minutes because the case links to its incident and the request diff shows the description edit. The fix is one sentence in the description, the case returns to 20/20, and the release proceeds. Total cost: about 430 live agent runs (31 cases at 5 trials per side, plus 15 more per side for four cases) instead of 8,000 for the full 200-case suite at 20 trials per side.

Failure modes

  • Single-run gating: noise fails builds and the team learns to ignore red. Use repeats and a decision rule.
  • Baselines that drift: re-recording everything whenever something fails turns the suite into a snapshot of current behaviour. Require review for expected-value changes.
  • Model alias moves underneath you: an alias like "latest" changes without a commit. Pin model versions in tests and run a nightly live sample.
  • Cases without provenance: nobody knows whether a failing case still matters. Link every case to an incident or promise.
  • Over-strict trajectories: exact call sequences fail on harmless reordering. Prefer must-call, must-not-call and argument rules where order does not matter.
  • Shared state between trials: reusing a session or backend fixture makes trial two depend on trial one. Use a fresh runner, session and fake backend per trial.

Trade-offs

More trials buy confidence linearly in cost and only with a square-root return in precision, so spend them where changes land, not uniformly. Scripted tests are cheap and precise but can pass while the real model fails; live tests are realistic but slow and noisy; the recorded layer sits between and goes stale when inputs change. Statistical thresholds trade missed regressions for false alarms, and the right point differs per case, which is why critical cases use all-trials-pass while ordinary ones use a significance test. The broader framework these layers plug into is described in the ADK Java evaluation framework article.

What to do next

  1. Turn your last ten agent incidents into cases with IDs that link back to them, and tag each case with intents and tools.
  2. Write a ScriptedLlm and an L1 test for each journey; run them on every commit.
  3. Add recorded replay for pull requests and route cache misses to the live layer.
  4. Implement the repeated-trial verdict with early stopping and a Fisher test, plus an all-pass rule for critical cases.
  5. Version baselines per released configuration and require review for expected-value changes.
  6. Create the quarantine queue with an owner, and a nightly live sample against pinned models.
Key takeaway: Build the regression suite from incidents and promises, not paraphrases. Test the code around the model deterministically with a scripted BaseLlm on every commit, replay recordings on pull requests, and run live repeated trials only for cases whose model input changed. Decide live verdicts with early stopping and a Fisher test, require all trials to pass for critical cases, version baselines with review, and give flaky cases an owner instead of a blind eye.