A handful of hand-picked eval cases tells you whether an agent can do a task. It cannot tell you whether yesterday's prompt change made the agent better, because a dozen cases cannot separate a real regression from the noise of a model that answers differently on every call. Dataset-driven evaluation treats the eval set as a versioned dataset of hundreds or thousands of rows, each tagged with the slice of traffic it represents. It runs the agent against that dataset under controlled conditions and compares runs statistically, row by row.
This article builds that pipeline for ADK Java. ADK Java v1.11.0 has no evaluator of its own, so the pieces here sit on top of the harness from the evaluation framework overview: its eval-set loader, its replay() on a fresh InMemoryRunner and its trajectory and response scorers. That page scores one case. This one is about everything that changes when the dataset becomes the unit: versioning, slicing, sampling, running at scale, confidence intervals and paired comparison.
The pipeline at a glance
Every stage of the pipeline writes output keyed by two identifiers: the content hash of the dataset and the version of the agent (git commit, model id, prompt hash). Without both, a result cannot be reproduced and two results cannot be compared. The figure shows the flow.
Most teams that say their evals are flaky have a problem in the last two stages: they compare pass rates from different row sets, or treat a two-point move on 50 cases as a signal.
The dataset as a versioned artifact
Store the dataset as JSONL, one row per line, in the same repository as the agent. Each row wraps one eval case in the Python ADK eval-set shape (so the existing loader parses it) and adds the metadata the pipeline needs:
{"id": "refund-0142", "slice": "refunds", "tags": ["multi_turn", "tool_error"],
"split": "dev", "source": "prod-2026-09", "added": "2026-09-18",
"case": {"eval_id": "refund-0142", "conversation": [ ... ], "session_input": { ... }}}slice is the single category used for reporting, and every row has exactly one. tags are free-form and overlap. split is dev for rows anyone may look at while tuning, or holdout for rows only CI reads. source records where the row came from, so a bad batch of mined production rows can be found and removed together. The loader computes a hash over the canonical bytes and refuses duplicate ids:
public record Row(String id, String slice, List<String> tags, String split,
String source, EvalCase evalCase) {}
public record Dataset(String name, String sha256, List<Row> rows) {
static Dataset load(Path jsonl) throws IOException {
List<String> lines = Files.readAllLines(jsonl, StandardCharsets.UTF_8).stream()
.map(String::strip).filter(l -> !l.isEmpty()).sorted().toList(); // order-independent hash
MessageDigest sha;
try { sha = MessageDigest.getInstance("SHA-256"); }
catch (NoSuchAlgorithmException e) { throw new IllegalStateException(e); }
List<Row> rows = new ArrayList<>();
Set<String> seen = new HashSet<>();
for (String line : lines) {
sha.update(line.getBytes(StandardCharsets.UTF_8));
sha.update((byte) '\n');
Row r = RowParser.parse(line); // wraps the overview's case parser
if (!seen.add(r.id())) throw new IllegalArgumentException("duplicate id " + r.id());
rows.add(r);
}
return new Dataset(jsonl.getFileName().toString(), HexFormat.of().formatHex(sha.digest()), rows);
}
}Sorting before hashing means reordering the file does not change the hash, but editing any row does. That is the point: if someone rewrites the expected answer of a failing row, the dataset is a new version, and its results are not comparable with the old version's. Changes to the dataset go through review like code, and the commit message says why a row was added, changed or removed.
Choosing which rows run
Running 2,000 rows with three model calls each, repeated three times, is 18,000 calls: too slow for every pull request. Run a stratified sample on pull requests and the full dataset nightly and before release. Stratify by slice so that small slices are not drowned out. Make the sample deterministic so that the same seed picks the same rows for baseline and candidate:
static List<Row> stratified(List<Row> rows, int perSlice, long seed) {
Map<String, List<Row>> bySlice = new TreeMap<>();
for (Row r : rows) bySlice.computeIfAbsent(r.slice(), k -> new ArrayList<>()).add(r);
List<Row> out = new ArrayList<>();
bySlice.forEach((slice, rs) -> {
List<Row> copy = new ArrayList<>(rs);
copy.sort(Comparator.comparing(Row::id)); // input order must not matter
Collections.shuffle(copy, new Random(seed * 31 + slice.hashCode()));
out.addAll(copy.subList(0, Math.min(perSlice, copy.size())));
});
return out;
}A pull-request run with perSlice = 40 over twelve slices is 480 rows, small enough to finish in a few minutes and large enough that each slice has a usable interval. Keep the seed fixed for a week, then rotate it, so that prompts are not tuned against one fixed sample.
Running rows at scale
The runner turns rows into outcomes. Three rules matter more than speed. First, each row runs on a fresh InMemoryRunner with its own session, exactly as the single-case harness does, so no state leaks between rows. Second, concurrency is bounded by the model quota, not by CPU: virtual threads are cheap, rate-limited model calls are not (the runner below needs Java 21). Third, an outcome is one of three verdicts. PASS and FAIL come from scorers. ERROR means the run never produced something to score, for example a timeout or a 429.
enum Verdict { PASS, FAIL, ERROR }
public record Outcome(String rowId, String slice, int repeat, Verdict verdict,
String detail, long latencyMs) {}
List<Outcome> runAll(List<Row> rows, int repeats, int maxInFlight) throws InterruptedException {
Semaphore permits = new Semaphore(maxInFlight);
List<Future<Outcome>> futures = new ArrayList<>();
try (ExecutorService pool = Executors.newVirtualThreadPerTaskExecutor()) {
for (Row r : rows) {
for (int k = 0; k < repeats; k++) {
final int rep = k;
futures.add(pool.submit(() -> {
permits.acquire();
try { return runWithRetry(r, rep); } finally { permits.release(); }
}));
}
}
} // close() waits for all tasks
List<Outcome> out = new ArrayList<>();
for (Future<Outcome> f : futures) {
try { out.add(f.get()); }
catch (ExecutionException e) { throw new IllegalStateException(e.getCause()); }
}
return out;
}
Outcome runWithRetry(Row r, int rep) throws InterruptedException {
for (int attempt = 0; ; attempt++) {
List<Invocation> inv = replay(agentFactory.get(), r.evalCase()); // from the overview
Optional<String> err = inv.stream().map(Invocation::error).filter(Objects::nonNull).findFirst();
if (err.isPresent() && isTransient(err.get()) && attempt < 3) {
Thread.sleep((1L << attempt) * 2_000 + ThreadLocalRandom.current().nextLong(1_000));
continue;
}
if (err.isPresent()) return new Outcome(r.id(), r.slice(), rep, Verdict.ERROR, err.get(), 0);
return score(r, rep, inv); // trajectory + response (+ judge) -> PASS or FAIL
}
}Repeats deal with non-determinism: one row can pass on one call and fail on the next. Run each row k times (three is common) and record every repeat; a row that passes one or two of three is flaky, which is information. Write every outcome as one JSONL line together with the dataset hash, the agent version, the model id and the repeat index. That results file is the input to everything that follows, so a report can be recomputed without calling the model again.
From outcomes to evidence
A pass rate without an interval is not a result. With 60 rows in a slice and 47 passes, the pass rate is 78%. The 95% interval, however, runs from about 66% to 87%. The Wilson score interval behaves well near 0% and 100% and for small n, which is where agent slices live, so use it instead of the textbook normal approximation:
/** Wilson score interval for k successes in n trials; z = 1.96 for 95%. */
static double[] wilson(int k, int n, double z) {
if (n == 0) return new double[] {0, 1};
double p = (double) k / n, z2 = z * z;
double denom = 1 + z2 / n;
double centre = (p + z2 / (2.0 * n)) / denom;
double half = z * Math.sqrt(p * (1 - p) / n + z2 / (4.0 * n * n)) / denom;
return new double[] {Math.max(0, centre - half), Math.min(1, centre + half)};
}Comparing two runs needs a different tool. Two overlapping intervals do not prove that nothing changed, because the baseline and the candidate ran on the same rows, and that pairing contains most of the information. Collapse each row to pass or fail per run (for example, pass on at least two of three repeats). Then count the rows that flipped. Let b be the rows that passed on the baseline and failed on the candidate, and c the rows that went the other way. Rows that did not change say nothing about the difference. Under the null hypothesis that the change did nothing, each flip is equally likely to go either way, which gives the exact McNemar test:
/** Two-sided exact McNemar p-value from discordant counts b (broke) and c (fixed). */
static double mcnemarExact(int b, int c) {
int n = b + c, m = Math.min(b, c);
if (n == 0) return 1.0;
// Underflows to p = 0 beyond about 1,000 discordant rows; use log-space terms there.
double tail = 0, term = Math.pow(0.5, n); // C(n,0) / 2^n
for (int k = 0; k <= m; k++) {
tail += term;
term = term * (n - k) / (k + 1); // C(n,k+1) / 2^n
}
return Math.min(1.0, 2 * tail);
}Run the test overall and for each slice. With twelve slices you are making twelve tests, so treat a slice p-value near 0.05 as something to investigate rather than an automatic block, or apply a Holm correction before gating.
Worked example: a regression that cancels out
A support agent has a 400-row holdout split, with 60 rows in the refunds slice and 340 in the rest. A pull request rewrites the system prompt to be more concise. The baseline passes 352 rows (88%, Wilson 84.4% to 90.8%). The candidate passes 356 rows (89%, 85.6% to 91.7%). Read on its own, that is a small win.
The paired view says otherwise:
| Scope | Baseline | Candidate | Broke (b) | Fixed (c) | Exact McNemar p |
|---|---|---|---|---|---|
| All rows (400) | 352 | 356 | 10 | 14 | 0.54 |
refunds (60) | 54 | 47 | 8 | 1 | 0.039 |
| Other slices (340) | 298 | 309 | 2 | 13 | 0.0074 |
Overall, 24 rows flipped, split 10 to 14: noise. By slice, two real effects cancel. The shorter prompt helped general questions (13 fixed, 2 broken) and damaged refunds (8 broken, 1 fixed). The flip list shows why. In seven of the eight broken refund rows the agent no longer calls lookup_order before quoting the refund window, because the sentence that told it to was removed. The fix is to restore that instruction and keep the rest. A single pass rate would have shipped the regression.
Keeping the dataset healthy
A dataset decays: expected answers go stale, mined rows carry labelling mistakes, and tuning overfits the prompt to them. Habits that keep it healthy:
- Audit the failures, not just the rate. Every week, review a sample of failing rows. Some will be label errors, where the agent was right. Fix them as a dataset change with a new hash.
- Keep the holdout untouched. Engineers tune on
dev.holdoutrows appear only in CI reports as aggregate numbers and flip ids. Refresh a share of the holdout each quarter with new production rows. - Watch slice coverage. Compare the slice distribution of the dataset with recent traffic. A slice that is 15% of traffic and 2% of rows is under-tested. See writing eval test cases for mining and labelling rows.
Failure modes
- Comparing across dataset versions. Someone fixes three labels, the pass rate jumps, and the gain is credited to the prompt change in the same PR. Refuse comparisons whose dataset hashes differ, and re-run the baseline on the new version.
- Errors counted as failures. A quota spike turns 40 rows into timeouts and the candidate "regresses". Report
ERRORseparately, and fail the run as invalid when the error rate exceeds a threshold such as 2%. - Judge drift. A model-judged scorer is itself a model. Pin its model version and prompt, include the judge's version in the results key, and calibrate it against human labels. See the LLM-as-judge scorer.
- Shared state between rows. Reusing one runner or session service across rows to save time lets earlier rows change later ones, and results then depend on scheduling order.
- Sample peeking. A prompt is iterated until the PR sample passes. Rotate the seed, and gate releases on the full holdout rather than on the sample.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Full run on every PR | No sampling error | Slow and expensive; developers stop waiting for it |
| Stratified PR sample + nightly full | Fast feedback with slice coverage | Small slices still have wide intervals |
| Repeats k = 3 | Measures flakiness, stabilises row scores | Three times the cost |
| Recorded model calls (cassettes) | Deterministic, free re-runs | Tests the agent's code, not the model's behaviour |
| Model-judged scoring | Scores open-ended answers | Judge cost, bias and drift to manage |
| Hard gate on slice p-values | Catches cancelling regressions | Multiple-testing false alarms without correction |
Use cassettes, as in eval CI integration, to check tool wiring; use live dataset runs to measure behaviour changes.
What to do next
- Convert your existing eval cases into a JSONL dataset with
id,slice,splitandsourceon every row, and commit it next to the agent. - Add the hashing loader and make every results file carry the dataset hash and agent version.
- Split off a holdout set of at least 20% of rows, and make CI the only reader.
- Wire the stratified sampler into pull-request CI, and run the full dataset nightly with three repeats.
- Report every slice with a Wilson interval, and list the flipped row ids for every comparison.
- Gate releases on the exact McNemar test overall and per slice, using a Holm correction or a review step for marginal slices.
- Schedule a weekly failure audit and a quarterly holdout refresh, and put both on someone's calendar.