An agent regression is a change in what the agent does, not only in whether a test passes. It calls a tool it used to skip, asks a clarifying question it used to answer directly, takes six model turns where it took three, or starts refusing a class of requests. These changes come from places a unit test does not watch: an edited instruction, a reworded tool description, a new model version behind the same name, a new tool that competes with an old one, or a shift in what users ask. Many of them leave the final answer looking fine on a small test set while cost, latency or error rates move in production.

This article is about detecting those changes in an ADK Java agent. It builds a behavioural fingerprint from the event stream the runner already emits, diffs fingerprints between two versions of the agent on the same inputs, tests whether the difference is real, and watches production with a CUSUM alarm that catches gradual drift. Building eval sets and scoring trajectories is covered in the ADK Java evaluation framework article; recorded model cassettes and the CI gate are in Integrating Evals Into CI. This page assumes you have some cases and focuses on noticing when behaviour moves.

What counts as a behaviour regression

Start with a taxonomy, because each kind needs a different signal.

KindExampleSignal
TrajectoryCalls searchOrders before getOrderStatus where it used to go directTool sequence edit distance
OutcomeFinal answer wrong or missingCase pass rate, answered rate
EfficiencyThree model turns become sixTurns and tool calls per run, tokens
ReliabilityMalformed arguments, tool errorsTool error rate per run
PolicyNew refusals or escalationsRefusal and handoff rates

The common thread is that every signal can be computed from one run's events, so the same extractor serves offline comparison and production monitoring. If the two lanes compute metrics differently, an offline result tells you nothing about production.

Two lanes, one fingerprint: offline comparison and online driftOffline: before a change shipsGolden + sampled casesinputs with expectationsReplay baseline + candidatesame inputs, sandboxed toolsDiff fingerprintstrajectory edits, flipsOnline: after it shipsProduction runsevents + version stampPer-run fingerprinttools, errors, outcomeCUSUM per metricper route, per versionFingerprint extractorone code path for bothTriage queueflipped cases, alarmsblock or acceptalarmAttribute: version diff, bisect
Both lanes share the fingerprint extractor. Offline comparison blocks or accepts a change before it ships; online CUSUM catches drift that no pre-release test sees, such as a model update behind a stable name or a shift in traffic.

A fingerprint from the event stream

ADK Java's runner returns a stream of Event objects for each invocation. Events carry function calls the model requested, function responses the tools returned, and final responses. A fingerprint is a small record derived from them:

import com.google.adk.events.Event;
import com.google.genai.types.FunctionCall;
import com.google.genai.types.FunctionResponse;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;

public record Fingerprint(
    List<String> tools,      // tool names in call order
    int toolCallTurns,       // events that requested at least one tool
    int toolErrors,          // responses whose payload carries an "error" key
    boolean answered,        // a final response with non-blank text
    int answerChars) {

  public static Fingerprint of(List<Event> events) {
    List<String> tools = new ArrayList<>();
    int turns = 0, errors = 0, chars = 0;
    boolean answered = false;
    for (Event e : events) {
      if (!e.functionCalls().isEmpty()) turns++;
      for (FunctionCall call : e.functionCalls()) {
        tools.add(call.name().orElse("?"));
      }
      for (FunctionResponse r : e.functionResponses()) {
        Map<String, Object> body = r.response().orElse(Map.of());
        if (body.containsKey("error")) errors++;
      }
      if (e.finalResponse()) {
        String text = e.stringifyContent();
        if (!text.isBlank()) { answered = true; chars += text.length(); }
      }
    }
    return new Fingerprint(List.copyOf(tools), turns, errors, answered, chars);
  }
}

The error convention is yours to define. The code above assumes tools return a map with an error key on failure, a common ADK Java pattern; whatever you choose, make every tool follow it, or tool errors become invisible. Collect events with runner.runAsync(userId, sessionId, content, RunConfig.builder().build()) followed by toList().blockingGet() offline, or attach the extractor to your existing event logging in production. Stamp every fingerprint with the agent version, a hash of the instruction text, the model identifier and a hash of the tool schemas, so any later difference can be attributed to a concrete change.

Diffing trajectories between versions

Comparing two versions starts with replaying the same inputs through both. Use a fixed golden set plus a sample of recent production inputs, and run tools against a sandbox or with side-effecting tools stubbed: a replay must never send a second refund. Then compare trajectories case by case with an edit distance over tool names:

static int editDistance(List<String> a, List<String> b) {
  int[][] d = new int[a.size() + 1][b.size() + 1];
  for (int i = 0; i <= a.size(); i++) d[i][0] = i;
  for (int j = 0; j <= b.size(); j++) d[0][j] = j;
  for (int i = 1; i <= a.size(); i++) {
    for (int j = 1; j <= b.size(); j++) {
      int sub = a.get(i - 1).equals(b.get(j - 1)) ? 0 : 1;
      d[i][j] = Math.min(Math.min(d[i - 1][j] + 1, d[i][j - 1] + 1),
                         d[i - 1][j - 1] + sub);
    }
  }
  return d[a.size()][b.size()];
}

enum Change { SAME, EXTRA_CALL, MISSING_CALL, DIFFERENT_TOOL, REORDERED }

static Change classify(List<String> base, List<String> cand) {
  if (base.equals(cand)) return Change.SAME;
  List<String> sortedBase = new java.util.ArrayList<>(base);
  List<String> sortedCand = new java.util.ArrayList<>(cand);
  java.util.Collections.sort(sortedBase);
  java.util.Collections.sort(sortedCand);
  if (sortedBase.equals(sortedCand)) return Change.REORDERED;  // same calls, same counts
  if (cand.size() > base.size() && editDistance(base, cand) == cand.size() - base.size())
    return Change.EXTRA_CALL;
  if (base.size() > cand.size() && editDistance(base, cand) == base.size() - cand.size())
    return Change.MISSING_CALL;
  return Change.DIFFERENT_TOOL;
}

Histograms of these classes are far more useful in review than one similarity score. "31 cases gained an extra searchOrders call" points straight at a tool description; "12 cases lost the verifyIdentity call" is a policy problem that must block the release whatever the pass rate says. Some changes are not regressions: a candidate that reaches the same answer in fewer calls is an improvement, so judge trajectory diffs together with outcome, never alone. For models that sample, run each case several times per version and compare distributions of fingerprints rather than single runs.

Worked example: are these flips real?

Outcome comparison should be paired, because the same inputs went through both versions. Count cases that flipped. Suppose 200 cases ran on both versions: 14 passed on the baseline and failed on the candidate, 5 did the reverse, and the rest agreed. The pass rate fell by 9 cases, 4.5 points. Only the 19 discordant cases carry information, and if the change made no difference each would be equally likely to go either way. The exact McNemar test asks how surprising 5 or fewer favourable flips out of 19 would be under that coin-flip assumption:

// P(X <= 5) for X ~ Binomial(19, 0.5)
// = (C(19,0) + C(19,1) + ... + C(19,5)) / 2^19
// = (1 + 19 + 171 + 969 + 3876 + 11628) / 524288
// = 16664 / 524288 = 0.0318 one-sided, 0.0636 two-sided
static double mcnemarOneSided(int favourable, int discordant) {
  double p = 0, term = Math.pow(0.5, discordant);   // C(n,0) / 2^n
  for (int k = 0; k <= favourable; k++) {
    p += term;
    term = term * (discordant - k) / (k + 1);        // C(n,k+1) from C(n,k)
  }
  return p;
}

A one-sided p of 0.032 is fairly strong evidence that the candidate is worse; the two-sided 0.064 misses a conventional 0.05 line. Do not let that line decide. The 14 regressed cases are concrete inputs you can read today, and reading them usually settles the question faster than more statistics. The test matters most in the other direction: it stops you chasing a 2-case drop that is noise. Confidence intervals and stratified runs are covered in the dataset-driven evaluation article.

Catching drift in production with CUSUM

Offline comparison cannot see changes that arrive without a release: a provider updating the model behind a stable name, a downstream API that starts timing out, or users asking new questions. For those, watch production rates continuously. A threshold on a daily average is slow and noisy; a cumulative sum (CUSUM) accumulates small, persistent excesses and alarms when they add up.

For a per-run 0/1 signal such as "this run had a tool error", choose a baseline rate p0 from a quiet period and the rate p1 you care to detect. A common reference value is the midpoint k = (p0 + p1) / 2. Each run adds its value minus k to a running sum floored at zero, and the alarm fires when the sum crosses a threshold h:

final class Cusum {
  private final double k, h;
  private double s;
  Cusum(double p0, double p1, double h) { this.k = (p0 + p1) / 2; this.h = h; }

  /** Feed one run; returns true when the alarm fires. */
  boolean update(boolean event) {
    s = Math.max(0, s + (event ? 1.0 : 0.0) - k);
    if (s > h) { s = 0; return true; }
    return false;
  }
}

// One detector per (metric, route, agent version), fed from the fingerprint stream.
Cusum toolErrors = new Cusum(0.02, 0.05, 4.0);
if (toolErrors.update(fp.toolErrors() > 0)) alerts.raise("tool-error drift", fp);

Work through the numbers. With p0 = 0.02 and p1 = 0.05, k = 0.035. An error-free run subtracts 0.035 and an erroring run adds 0.965. At the baseline rate the sum drifts down by 0.015 per run on average and stays near zero; after a shift to 5% it drifts up by 0.015 per run, so with h = 4 the alarm fires roughly 4 / 0.015, about 270 runs, after the shift. A larger h means fewer false alarms and slower detection. Choose h by replaying a month of known-good fingerprints through the detector and counting alarms, not by intuition.

Two practices keep CUSUM honest. Run a detector per route or intent: a traffic mix that shifts towards a hard intent raises the global error rate without any regression, and per-route detectors tell the two apart. And reset baselines deliberately when a release intentionally changes behaviour, recording why. Feed alarms into the same tracing you already have; the ADK Java observability article covers exporting runs as traces.

Failure modes

  • Lanes compute different things. Offline and online fingerprints from separate code drift apart. Share one extractor.
  • Missing version stamps. An alarm with no agent version, prompt hash or model id attached cannot be attributed. Make the stamp mandatory at logging time.
  • Replays with side effects. A comparison run that calls a real payments tool is an incident. Sandbox or stub every write.
  • Single-sample comparisons. One run per case per version mistakes sampling noise for change; repeat cases and compare distributions.
  • Score-only gates. A pass rate that holds while a safety-relevant tool call disappears is a regression. Make specific trajectory changes blocking in their own right.
  • Alarm fatigue. Detectors tuned by guesswork fire daily and get muted. Calibrate h on historical data and review the false-alarm count monthly.

Trade-offs

MethodCatchesMissesCost
Golden-set replayKnown behaviours, before releaseUnseen inputs, post-release driftLow
Production sample replayRealistic distributionChanges outside the sample windowMedium; needs sandboxing
Paired flip analysisReal outcome changes, with evidenceSmall effects on small setsLow
CUSUM on productionGradual drift, silent model updatesProblems with no metricLow once wired
Shadow trafficLive behaviour without user impactEffects of stateful toolsHigh: double inference

What to do next

  1. Implement the fingerprint record and log one per production run, with version stamps.
  2. Replay your golden set through the current and previous agent versions and produce the trajectory change histogram.
  3. Add flip counts and the exact McNemar p-value to the comparison report, and list every regressed case by input.
  4. Mark the trajectory changes that must block a release regardless of score.
  5. Wire one CUSUM detector on tool error rate per route, and calibrate h on a month of history.
  6. Write a runbook entry: on alarm, diff the version stamps of runs before and after, then bisect.
Key takeaway: Derive one behavioural fingerprint per run from the ADK event stream and use it everywhere: diff trajectories and count flipped cases between versions on the same inputs, test flips with an exact McNemar test but read the regressed cases, and watch production with per-route CUSUM detectors calibrated on history. Stamp every run with its versions so any alarm points at a change.