An agent that passed its offline eval set on Monday can be failing real users by Wednesday, without any code change. The traffic mix shifted, a tool's backend began returning a new error shape, or the model endpoint behind the same name was updated. Offline evaluation answers "is this release good on the cases we thought of?" Continuous evaluation answers a different question: "how good is the release on the requests it is actually getting, right now, and is that worse than before?"
This article builds that loop for an ADK Java agent. We capture each turn as it happens, without slowing it down. Cheap deterministic checks run on every turn. A rubric judge runs on a stratified sample, with weights so the numbers still describe the whole population. Scores are aggregated into windows per release and per slice, with honest confidence intervals, and compared against a baseline release. Failures flow back into the offline eval sets. The worked example shows how a pooled pass rate can raise a false alarm and hide the real one, and how per-slice weighted numbers find it.
It assumes you already have an offline harness and a judge. If not, start with the ADK Java evaluation framework and the LLM-as-judge scorer.
What continuous evaluation is made of
Continuous evaluation has four parts, and each one can fail separately.
Capture turns a live turn into a record you can score later: the user text, the final answer, the tool calls, latency, errors, and the labels you will slice by (release, route or intent, tenant tier). Scoring attaches verdicts to records. There is no reference answer in production, so production scorers must be reference-free: rubric judgements such as "answers the question asked" and "does not claim an action the tools did not perform", structural checks, and behavioural signals from users. Aggregation turns verdicts into rates with intervals, per window, release and slice. Action means alerting, rollback (the canary controller in ADK Java canary deployment can consume these scores), and harvesting failing turns into eval cases.
The principle that shapes the design: evaluation must never change what the user experiences. It runs after the response is returned, it can be dropped under load, and its failure must never fail the request.
Capturing turns without slowing them
Wrap the runner instead of changing the agent. The wrapper below consumes the event stream from runAsync the same way the canary and replay harnesses on this site do. It collects function-call names and the final response text, then builds a TurnRecord in a finally block, so failed turns are captured too. Failed turns are often the most informative ones.
public record TurnRecord(String turnId, String sessionId, String release, String slice,
String userText, String finalText, List<String> toolCalls,
long latencyMs, String error, Instant at) {}
public final class EvaluatedRunner {
private final Runner runner; // the agent's normal ADK runner
private final String release; // e.g. "support-agent@r42", from the release manifest
private final Sampler sampler;
private final EvalQueue queue;
private final Redactor redactor; // strips emails, card numbers, tokens
public String handleTurn(String userId, String sessionId, String slice, String userText) {
long start = System.nanoTime();
List<String> tools = new ArrayList<>();
StringBuilder answer = new StringBuilder();
String error = null;
try {
Content msg = Content.fromParts(Part.fromText(userText));
for (Event e : runner.runAsync(userId, sessionId, msg).blockingIterable()) {
for (FunctionCall fc : e.functionCalls()) tools.add(fc.name().orElse(""));
if (e.finalResponse()) answer.append(e.stringifyContent());
}
return answer.toString();
} catch (RuntimeException ex) {
error = ex.getClass().getSimpleName();
throw ex; // the user sees the real failure
} finally {
try {
TurnRecord r = new TurnRecord(UUID.randomUUID().toString(), sessionId, release, slice,
redactor.apply(userText), redactor.apply(answer.toString()), List.copyOf(tools),
(System.nanoTime() - start) / 1_000_000, error, Instant.now());
queue.offer(r, sampler.decide(r)); // non-blocking; never throws to the caller
} catch (RuntimeException ignored) {
queue.captureFailures().increment();
}
}
}
}Three details matter. First, redaction happens before the record leaves the process, so the evaluation store, the judge's prompt and the eval-set harvest never see raw personal data. Second, slice is assigned by the caller from information it already has, such as the route, a classifier's intent label or the tenant tier. You cannot add slices to old data later, so decide them now. Third, the whole capture path sits inside its own try block. A bug in the redactor increments a counter; it does not turn a good answer into a 500.
Census checks and a weighted judge sample
Running a judge on every production turn is usually unaffordable and always unnecessary. At 200,000 turns a day, a 2% sample is 4,000 judge calls a day, which is plenty to detect a few points of regression within a day on the large slices. So split the scorers by cost.
Census checks run on every turn because they cost microseconds. Did the turn error? Did a tool call fail? Is structured output valid against its schema? Did the answer claim to have done something ("I have cancelled your order") with no matching tool call in toolCalls? Is latency over the SLO? These catch most outright breakage, and with no sampling there is no sampling error.
Sampled judging runs the rubric judge on a fraction of sessions. The fraction differs per slice: oversample slices that are rare, risky or just changed, and undersample the bulk. Every sampled verdict carries a weight of one over its sampling rate, so a pooled estimate can be rebuilt to match the real traffic mix. Sampling by a salted hash of the session ID makes the decision deterministic and keeps whole conversations together, which a multi-turn rubric needs.
public record Decision(boolean judge, double weight) {}
public final class Sampler {
private final Map<String, Double> judgeRate; // per slice: billing 0.04, howto 0.01
private final String salt; // rotate per evaluation period
public Decision decide(TurnRecord r) {
double rate = judgeRate.getOrDefault(r.slice(), 0.01);
// Hash the SESSION, not the turn: a sampled conversation is judged end to end.
// Use a real 64-bit hash (Murmur3, xxHash) in production; hashCode() keeps this short.
long bucket = Math.floorMod((salt + ":" + r.sessionId()).hashCode(), 10_000);
boolean in = bucket < rate * 10_000;
return new Decision(in, in ? 1.0 / rate : 0.0);
}
}
// Worker loop, N threads, fed by an ArrayBlockingQueue whose offer() drops when full.
void work() throws InterruptedException {
while (running) {
Job j = queue.poll(1, TimeUnit.SECONDS);
if (j == null) continue;
for (Check check : checks) // census: every turn, weight 1
store.put(j.record(), check.name(), check.apply(j.record()), 1.0);
if (j.decision().judge()) {
Verdict v = judge.score(j.record()); // PASS, FAIL or ERROR
store.put(j.record(), "judge:" + judge.version(), v, j.decision().weight());
}
}
}The judge returns three outcomes, not two. A judge timeout or an unparsable verdict is ERROR, tracked as its own rate and left out of the pass-rate denominator. Counting it as a failure lets a judge outage look like a quality regression. Version the judge (model, rubric text, temperature) and store the version with every verdict. A rubric edit is a measurement change, and windows scored by different judge versions must never be compared directly.
Windows, intervals and baselines
Aggregation keys are (release, slice, scorer, window). For each key, keep the weighted and unweighted counts of judged and passed turns. Report a pass rate with a 95% Wilson interval, which stays sensible near 0% and 100% and at small counts. The textbook normal-approximation interval does neither. To compare a candidate release with the baseline, use a two-proportion z statistic per slice. The same arithmetic appears in dataset-driven evaluation; the difference here is that the data keeps arriving.
Two rules prevent most false alarms. Never evaluate a window below a minimum number of judged turns per slice, so that early on the answer is "not enough data" rather than OK or REGRESSION. And because you will look at the numbers every hour, raise the alarm threshold to account for repeated looks: z above 2.5 to 3 rather than 1.96. Better still, use a sequential test designed for continuous monitoring. Compare releases on the same calendar window wherever traffic allows (canary versus baseline). Comparing this week's release with last week's mixes the release effect with the traffic effect.
Worked example: the alarm that was wrong and the regression it hid
A support agent's traffic is 25% billing questions and 75% how-to questions. Baseline release r41 has 2,400 judged turns from a proportional sample: 2,208 pass, 92.0% (Wilson 95% interval 90.9% to 93.0%). Billing is 534 of 600 (89.0%) and how-to is 1,674 of 1,800 (93.0%).
Release r42 changes the billing tool's prompt, so its sampler oversamples billing. After day two r42 has 600 judged turns, split 300 billing and 300 how-to. The pooled raw rate is 531 of 600, 88.5%, and the pooled z against r41 is 2.72. An alert keyed on that number fires. But the sample is half billing while traffic is a quarter billing, and billing is the harder slice. Pooling an oversampled hard slice drags the average down even if nothing changed.
Per slice: how-to went from 93.0% to 95.0% (285 of 300, interval 91.9% to 97.0%), so it is not worse. Billing went from 89.0% to 82.0% (246 of 300, interval 77.3% to 85.9%, against 86.2% to 91.3% for r41), with z = 2.91. Reweighted to the true mix, r42 scores 0.25 x 82.0% + 0.75 x 95.0% = 91.75%, almost the same as r41's 92.0%. The correct reading is that r42 has a real, localised billing regression that a traffic-weighted average nearly hides, plus a pooled number that would have raised the alarm for the wrong reason. Act per slice: roll back the billing prompt change and keep the how-to improvement.
Note too what day one looked like: 300 judged turns, pooled 88.7%, z = 1.97. With a naive 1.96 threshold and hourly looks, that is a coin-flip alert. The minimum-sample and raised-threshold rules exist for exactly this.
Closing the loop: harvesting failures
Continuous evaluation earns its cost when production failures become regression tests. A nightly job selects judge FAILs, census failures, thumbs-down turns and escalations. It clusters near-duplicates (embedding similarity or simple text shingles) and queues one representative per cluster for human review. A reviewer confirms the failure, writes the expected behaviour or rubric criteria, and commits it to the versioned eval set used by the CI eval gate. The next release cannot regress on that case without failing the build.
Do not auto-append raw judge FAILs to the eval set. The judge has a false-failure rate, and unreviewed cases teach the team to ignore the suite. Keep provenance on each harvested case (source turn ID, release, date) so it can be traced, and expire cases whose underlying feature has been removed.
Operational guidance
- Export the queue's depth, drop count and capture-failure count as metrics alongside your agent's other telemetry. A silent evaluator looks exactly like a healthy agent.
- Run a daily calibration sample: 50 to 100 judged turns also labelled by a human. Track judge-human agreement per slice. When agreement drops, the judge is measuring something else now, and the dashboards are untrustworthy until fixed.
- Budget judge cost explicitly: sample rate x traffic x tokens per judge call. Put the rates in configuration so incident response can raise a slice to 100% for an hour without a deploy.
- Keep the judge's model and data region within the same data-handling rules as the agent. Production transcripts are user data even after redaction.
- Tag every record with the full release identity: agent code version, instruction version, model ID and tool versions. "r42" is useless if two of those changed under it.
Failure modes
| Failure mode | What you see | Fix |
|---|---|---|
| Evaluator blocks the request path | p99 latency rises with judge latency | Capture in finally, bounded queue, offer() never put() |
| Judge outage counted as FAIL | Pass rate collapses across every slice at once | Three outcomes; ERROR excluded and alerted separately |
| Pooled rate over a stratified sample | Alarm fires, or a regression is masked | Weight by 1 / rate; decide per slice |
| Rubric edited mid-comparison | Step change on the edit date | Version the judge; compare only same-version windows |
| Turn-level sampling | Multi-turn rubrics see half a conversation | Hash the session ID |
| Unredacted capture | Personal data in the eval store and the judge's logs | Redact before the record leaves the process |
| Hourly looks at a 1.96 threshold | Frequent false regressions | Minimum samples, higher threshold or a sequential test |
| Queue drops under peak load | Peak hours under-represented | Track drops per slice; scale workers or reweight |
Trade-offs
Sampling rate trades cost against detection speed. Halving the minimum detectable regression needs roughly four times the judged turns, so pick rates per slice from the regression size you care about, not one global rate. Judges trade breadth against trust: deterministic checks are exact but narrow, and judges are broad but noisy and need calibration. Most teams get the best return from many cheap census checks plus a small, well-calibrated judge sample. User signals are free and plentiful but biased: unhappy users click more, and silence is not success. Use them as triggers for judging, not as the quality metric.
What to do next
- Wrap your runner so every turn produces a redacted
TurnRecordwith release and slice labels, captured in afinallyblock behind a bounded queue. - Write five census checks from your last ten incidents, such as tool errors, schema failures and claimed-but-not-performed actions, and run them on 100% of turns.
- Pick per-slice judge rates from the regression size you need to see within a day, and store the weight with every verdict.
- Build the per-slice dashboard with Wilson intervals and a minimum-sample rule before you build any pooled number.
- Version the judge and run a daily human calibration sample.
- Start the nightly harvest: cluster failures, have a person review them, and commit the confirmed ones to the CI eval set.