An agent regression is a change in what the agent does, not only in whether a test passes. It calls a tool it used to skip, asks a clarifying question it used to answer directly, takes six model turns where it took three, or starts refusing a class of requests. These changes come from places a unit test does not watch: an edited instruction, a reworded tool description, a new model version behind the same name, a new tool that competes with an old one, or a shift in what users ask. Many of them leave the final answer looking fine on a small test set while cost, latency or error rates move in production.
This article is about detecting those changes in an ADK Java agent. It builds a behavioural fingerprint from the event stream the runner already emits, diffs fingerprints between two versions of the agent on the same inputs, tests whether the difference is real, and watches production with a CUSUM alarm that catches gradual drift. Building eval sets and scoring trajectories is covered in the ADK Java evaluation framework article; recorded model cassettes and the CI gate are in Integrating Evals Into CI. This page assumes you have some cases and focuses on noticing when behaviour moves.
What counts as a behaviour regression
Start with a taxonomy, because each kind needs a different signal.
| Kind | Example | Signal |
|---|---|---|
| Trajectory | Calls searchOrders before getOrderStatus where it used to go direct | Tool sequence edit distance |
| Outcome | Final answer wrong or missing | Case pass rate, answered rate |
| Efficiency | Three model turns become six | Turns and tool calls per run, tokens |
| Reliability | Malformed arguments, tool errors | Tool error rate per run |
| Policy | New refusals or escalations | Refusal and handoff rates |
The common thread is that every signal can be computed from one run's events, so the same extractor serves offline comparison and production monitoring. If the two lanes compute metrics differently, an offline result tells you nothing about production.
A fingerprint from the event stream
ADK Java's runner returns a stream of Event objects for each invocation. Events carry function calls the model requested, function responses the tools returned, and final responses. A fingerprint is a small record derived from them:
import com.google.adk.events.Event;
import com.google.genai.types.FunctionCall;
import com.google.genai.types.FunctionResponse;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
public record Fingerprint(
List<String> tools, // tool names in call order
int toolCallTurns, // events that requested at least one tool
int toolErrors, // responses whose payload carries an "error" key
boolean answered, // a final response with non-blank text
int answerChars) {
public static Fingerprint of(List<Event> events) {
List<String> tools = new ArrayList<>();
int turns = 0, errors = 0, chars = 0;
boolean answered = false;
for (Event e : events) {
if (!e.functionCalls().isEmpty()) turns++;
for (FunctionCall call : e.functionCalls()) {
tools.add(call.name().orElse("?"));
}
for (FunctionResponse r : e.functionResponses()) {
Map<String, Object> body = r.response().orElse(Map.of());
if (body.containsKey("error")) errors++;
}
if (e.finalResponse()) {
String text = e.stringifyContent();
if (!text.isBlank()) { answered = true; chars += text.length(); }
}
}
return new Fingerprint(List.copyOf(tools), turns, errors, answered, chars);
}
}The error convention is yours to define. The code above assumes tools return a map with an error key on failure, a common ADK Java pattern; whatever you choose, make every tool follow it, or tool errors become invisible. Collect events with runner.runAsync(userId, sessionId, content, RunConfig.builder().build()) followed by toList().blockingGet() offline, or attach the extractor to your existing event logging in production. Stamp every fingerprint with the agent version, a hash of the instruction text, the model identifier and a hash of the tool schemas, so any later difference can be attributed to a concrete change.
Diffing trajectories between versions
Comparing two versions starts with replaying the same inputs through both. Use a fixed golden set plus a sample of recent production inputs, and run tools against a sandbox or with side-effecting tools stubbed: a replay must never send a second refund. Then compare trajectories case by case with an edit distance over tool names:
static int editDistance(List<String> a, List<String> b) {
int[][] d = new int[a.size() + 1][b.size() + 1];
for (int i = 0; i <= a.size(); i++) d[i][0] = i;
for (int j = 0; j <= b.size(); j++) d[0][j] = j;
for (int i = 1; i <= a.size(); i++) {
for (int j = 1; j <= b.size(); j++) {
int sub = a.get(i - 1).equals(b.get(j - 1)) ? 0 : 1;
d[i][j] = Math.min(Math.min(d[i - 1][j] + 1, d[i][j - 1] + 1),
d[i - 1][j - 1] + sub);
}
}
return d[a.size()][b.size()];
}
enum Change { SAME, EXTRA_CALL, MISSING_CALL, DIFFERENT_TOOL, REORDERED }
static Change classify(List<String> base, List<String> cand) {
if (base.equals(cand)) return Change.SAME;
List<String> sortedBase = new java.util.ArrayList<>(base);
List<String> sortedCand = new java.util.ArrayList<>(cand);
java.util.Collections.sort(sortedBase);
java.util.Collections.sort(sortedCand);
if (sortedBase.equals(sortedCand)) return Change.REORDERED; // same calls, same counts
if (cand.size() > base.size() && editDistance(base, cand) == cand.size() - base.size())
return Change.EXTRA_CALL;
if (base.size() > cand.size() && editDistance(base, cand) == base.size() - cand.size())
return Change.MISSING_CALL;
return Change.DIFFERENT_TOOL;
}Histograms of these classes are far more useful in review than one similarity score. "31 cases gained an extra searchOrders call" points straight at a tool description; "12 cases lost the verifyIdentity call" is a policy problem that must block the release whatever the pass rate says. Some changes are not regressions: a candidate that reaches the same answer in fewer calls is an improvement, so judge trajectory diffs together with outcome, never alone. For models that sample, run each case several times per version and compare distributions of fingerprints rather than single runs.
Worked example: are these flips real?
Outcome comparison should be paired, because the same inputs went through both versions. Count cases that flipped. Suppose 200 cases ran on both versions: 14 passed on the baseline and failed on the candidate, 5 did the reverse, and the rest agreed. The pass rate fell by 9 cases, 4.5 points. Only the 19 discordant cases carry information, and if the change made no difference each would be equally likely to go either way. The exact McNemar test asks how surprising 5 or fewer favourable flips out of 19 would be under that coin-flip assumption:
// P(X <= 5) for X ~ Binomial(19, 0.5)
// = (C(19,0) + C(19,1) + ... + C(19,5)) / 2^19
// = (1 + 19 + 171 + 969 + 3876 + 11628) / 524288
// = 16664 / 524288 = 0.0318 one-sided, 0.0636 two-sided
static double mcnemarOneSided(int favourable, int discordant) {
double p = 0, term = Math.pow(0.5, discordant); // C(n,0) / 2^n
for (int k = 0; k <= favourable; k++) {
p += term;
term = term * (discordant - k) / (k + 1); // C(n,k+1) from C(n,k)
}
return p;
}A one-sided p of 0.032 is fairly strong evidence that the candidate is worse; the two-sided 0.064 misses a conventional 0.05 line. Do not let that line decide. The 14 regressed cases are concrete inputs you can read today, and reading them usually settles the question faster than more statistics. The test matters most in the other direction: it stops you chasing a 2-case drop that is noise. Confidence intervals and stratified runs are covered in the dataset-driven evaluation article.
Catching drift in production with CUSUM
Offline comparison cannot see changes that arrive without a release: a provider updating the model behind a stable name, a downstream API that starts timing out, or users asking new questions. For those, watch production rates continuously. A threshold on a daily average is slow and noisy; a cumulative sum (CUSUM) accumulates small, persistent excesses and alarms when they add up.
For a per-run 0/1 signal such as "this run had a tool error", choose a baseline rate p0 from a quiet period and the rate p1 you care to detect. A common reference value is the midpoint k = (p0 + p1) / 2. Each run adds its value minus k to a running sum floored at zero, and the alarm fires when the sum crosses a threshold h:
final class Cusum {
private final double k, h;
private double s;
Cusum(double p0, double p1, double h) { this.k = (p0 + p1) / 2; this.h = h; }
/** Feed one run; returns true when the alarm fires. */
boolean update(boolean event) {
s = Math.max(0, s + (event ? 1.0 : 0.0) - k);
if (s > h) { s = 0; return true; }
return false;
}
}
// One detector per (metric, route, agent version), fed from the fingerprint stream.
Cusum toolErrors = new Cusum(0.02, 0.05, 4.0);
if (toolErrors.update(fp.toolErrors() > 0)) alerts.raise("tool-error drift", fp);Work through the numbers. With p0 = 0.02 and p1 = 0.05, k = 0.035. An error-free run subtracts 0.035 and an erroring run adds 0.965. At the baseline rate the sum drifts down by 0.015 per run on average and stays near zero; after a shift to 5% it drifts up by 0.015 per run, so with h = 4 the alarm fires roughly 4 / 0.015, about 270 runs, after the shift. A larger h means fewer false alarms and slower detection. Choose h by replaying a month of known-good fingerprints through the detector and counting alarms, not by intuition.
Two practices keep CUSUM honest. Run a detector per route or intent: a traffic mix that shifts towards a hard intent raises the global error rate without any regression, and per-route detectors tell the two apart. And reset baselines deliberately when a release intentionally changes behaviour, recording why. Feed alarms into the same tracing you already have; the ADK Java observability article covers exporting runs as traces.
Failure modes
- Lanes compute different things. Offline and online fingerprints from separate code drift apart. Share one extractor.
- Missing version stamps. An alarm with no agent version, prompt hash or model id attached cannot be attributed. Make the stamp mandatory at logging time.
- Replays with side effects. A comparison run that calls a real payments tool is an incident. Sandbox or stub every write.
- Single-sample comparisons. One run per case per version mistakes sampling noise for change; repeat cases and compare distributions.
- Score-only gates. A pass rate that holds while a safety-relevant tool call disappears is a regression. Make specific trajectory changes blocking in their own right.
- Alarm fatigue. Detectors tuned by guesswork fire daily and get muted. Calibrate h on historical data and review the false-alarm count monthly.
Trade-offs
| Method | Catches | Misses | Cost |
|---|---|---|---|
| Golden-set replay | Known behaviours, before release | Unseen inputs, post-release drift | Low |
| Production sample replay | Realistic distribution | Changes outside the sample window | Medium; needs sandboxing |
| Paired flip analysis | Real outcome changes, with evidence | Small effects on small sets | Low |
| CUSUM on production | Gradual drift, silent model updates | Problems with no metric | Low once wired |
| Shadow traffic | Live behaviour without user impact | Effects of stateful tools | High: double inference |
What to do next
- Implement the fingerprint record and log one per production run, with version stamps.
- Replay your golden set through the current and previous agent versions and produce the trajectory change histogram.
- Add flip counts and the exact McNemar p-value to the comparison report, and list every regressed case by input.
- Mark the trajectory changes that must block a release regardless of score.
- Wire one CUSUM detector on tool error rate per route, and calibrate h on a month of history.
- Write a runbook entry: on alarm, diff the version stamps of runs before and after, then bisect.