An evaluation dataset is your agent's specification written as examples. When an evaluation run says the agent got worse, you act on it: you revert a prompt, block a release, or spend a day debugging. That only makes sense if the rows are right. In practice most bad eval signals start in the data, not in the scorer. The usual causes are a label two people would disagree on, a holdout row that is a near-copy of a dev row, synthetic variants that inherited a label they no longer deserve, and an expected trajectory that still names a tool renamed last month.
Other pages on this site cover the neighbouring problems. Dataset-driven evaluation pipelines covers the versioned JSONL file, the dev/holdout split, sampling and the statistics. Writing eval test cases covers case anatomy and mining production traces. The evaluation framework overview covers the replay harness. This page is about making each row trustworthy. It covers a labelling protocol, measuring agreement, synthetic expansion, leakage checks, migrating rows when tools change, and privacy for rows taken from production. The core google-adk 1.11.0 jar has no evaluation classes, so the code below is plain Java that you own and run next to the harness.
Where dataset quality is lost
Think of each row as passing through a small factory, shown below. Rows come from three sources, are scrubbed, labelled independently by two people, adjudicated, checked for leakage, and only then land in a dataset version. Each stage exists because a specific defect gets past the others:
| Defect | What you see | The check that catches it |
|---|---|---|
| Ambiguous label | A slice flaps between releases with no code change | Double labelling and kappa |
| Wrong label | The agent 'fails' rows where it was right | Adjudication and failure audits |
| Inherited synthetic label | Synthetic slice pass rate far below the real one | Relabel information-changing variants |
| Leakage | Holdout scores track dev too closely | Near-duplicate and prompt-overlap scan |
| Schema drift | A whole slice drops to zero after a refactor | Tool-name validation against the agent |
| Privacy debt | Customer data in a git repository | Scrubbing before drafting, retention dates |
A labelling protocol
A label is a decision about what the agent should do, so define its parts before anyone labels. For an agent that uses tools, four fields work well. The outcome class is one of a few values: resolve, ask a clarifying question, refuse, or escalate to a human. The required tools are the calls that must happen. The forbidden tools are calls that must not happen, such as issue_refund on an ineligible order. The answer facts are the short statements a correct reply must contain, which a rubric or judge checks later.
{"id": "refund-0217",
"label": {
"outcome": "clarify",
"required_tools": ["get_order"],
"forbidden_tools": ["issue_refund"],
"answer_facts": ["asks which item in the order is damaged"],
"guideline_version": "refunds-g7",
"labellers": ["lab-14", "lab-22"],
"adjudicated": true}}Three rules keep labels honest. First, labellers see the conversation and the tool fixtures but not the agent's output. Seeing a plausible answer anchors people toward accepting it, and then the label just records the current agent's behaviour. Second, each new slice is labelled by two people independently, and after that a fixed share of new rows (20% is a reasonable start) keeps getting a second labeller so you keep measuring. Third, a third person resolves disagreements, and every disagreement is first treated as a gap in the written guideline. Fix the guideline, bump its version, and record the version on the row, so you can find every row labelled under the old rule.
Measuring agreement with Cohen's kappa
Raw agreement flatters you when one class dominates. If 80% of refund rows are 'resolve', two labellers who both always answer 'resolve' agree 80% of the time and have learned nothing. Cohen's kappa subtracts the agreement you would expect by chance from each labeller's own class frequencies: kappa = (po - pe) / (1 - pe), where po is the observed agreement and pe is the sum over classes of the product of the two labellers' marginal rates.
/** Cohen's kappa for two labellers over the same rows, outcome class only. */
static double cohensKappa(List<String> a, List<String> b) {
if (a.size() != b.size() || a.isEmpty()) throw new IllegalArgumentException("paired labels required");
int n = a.size(), agree = 0;
Map<String, Integer> ca = new HashMap<>(), cb = new HashMap<>();
for (int i = 0; i < n; i++) {
if (a.get(i).equals(b.get(i))) agree++;
ca.merge(a.get(i), 1, Integer::sum);
cb.merge(b.get(i), 1, Integer::sum);
}
double po = (double) agree / n, pe = 0;
for (var e : ca.entrySet()) pe += (e.getValue() / (double) n) * (cb.getOrDefault(e.getKey(), 0) / (double) n);
return pe == 1.0 ? 1.0 : (po - pe) / (1 - pe);
}Compute kappa per slice, not over the whole dataset. A strong global number can hide one slice where people are guessing. As a working rule, below about 0.6 the guideline for that slice is ambiguous, and the slice should not gate a release until it is rewritten. Between 0.6 and 0.8, review the disagreements. Above 0.8 the labels are dependable. These cut-offs are conventions, not laws, so use them to choose which slices to inspect. Tool sets are multi-label, so compare them separately, as per-tool agreement on 'required: yes/no'. A slice can agree on outcomes and still disagree on whether get_order is mandatory.
Synthetic variants without fooling yourself
Real rows are scarce for rare intents, so teams generate variants with a model. Useful transforms are paraphrase, typos and informal register, irrelevant extra detail, two intents in one message, and adversarial content, such as an instruction planted in a tool result. The trap is that the variant inherits the seed's label automatically. Some transforms change the information in the message. A paraphrase that drops the order number turns a 'resolve' row into a 'clarify' row. A combined intent can add a required tool. Mark each transform as information-preserving or not, and send every variant from an information-changing transform back to labelling.
enum Transform {
PARAPHRASE(true), TYPOS(true), NOISE_DETAIL(true),
DROP_IDENTIFIER(false), MULTI_INTENT(false), INJECTED_TOOL_TEXT(false);
final boolean preservesLabel;
Transform(boolean preservesLabel) { this.preservesLabel = preservesLabel; }
}
record Variant(String seedId, Transform t, String userText) {
String id() { return seedId + "~" + t.name().toLowerCase(Locale.ROOT); }
boolean needsRelabel() { return !t.preservesLabel; }
}Three more rules apply. Tag every variant with source: synthetic and report synthetic and real pass rates separately. A gap between them tells you about the generator as much as about the agent. Cap synthetic rows per slice, for example at 30%, so that generated phrasing does not dominate the signal. And if the generator is the same model family as the agent, expect its paraphrases to be easier than real users' writing. Treat synthetic rows as extra coverage, never as a replacement for mined ones.
Leakage and near-duplicates
Leakage makes the holdout look better than it is. Three kinds of leak show up in agent datasets. The first is near-duplicates across the split: two rows mined from the same incident, one in dev and one in holdout. The second is rows copied into the agent's instruction or few-shot examples while someone fixed a failure. The third is rows used as fine-tuning data. Exact-match checks miss all three, because the copies differ in names and dates. Word shingles with Jaccard similarity catch them cheaply:
static Set<String> shingles(String text, int k) {
String[] w = text.toLowerCase(Locale.ROOT).replaceAll("[^a-z0-9 ]", " ").trim().split("\\s+");
Set<String> out = new HashSet<>();
for (int i = 0; i + k <= w.length; i++) out.add(String.join(" ", Arrays.asList(w).subList(i, i + k)));
return out;
}
static double jaccard(Set<String> x, Set<String> y) {
if (x.isEmpty() && y.isEmpty()) return 1.0;
Set<String> inter = new HashSet<>(x); inter.retainAll(y);
return inter.size() / (double) (x.size() + y.size() - inter.size());
}
/** Fraction of the row's shingles that appear in the prompt text: catches copied examples. */
static double containment(Set<String> row, Set<String> prompt) {
if (row.isEmpty()) return 0;
return row.stream().filter(prompt::contains).count() / (double) row.size();
}Join each row's user turns into one string, take shingles with k = 3, and compare every holdout row with every dev row. For a few thousand rows a plain double loop takes seconds. Past roughly 50,000 rows, switch to MinHash with locality-sensitive hashing. Flag pairs above about 0.8 Jaccard for a person to review rather than deleting them automatically, because short template-like messages ('where is my order') are legitimately similar. Run the containment check against the agent's instruction and example text on every pull request that edits the prompt. A row whose shingles are more than half inside the prompt has become a test of memory.
Migrating rows when tools change
Tool refactors break datasets quietly. Rename lookup_order to get_order and every row that requires the old name fails. The failure looks like a regression in the agent, and someone may 'fix' it by loosening matchers. Treat tool changes as dataset migrations. Write the rename (and any argument renames) into a small map in the repository, apply it with a script that produces the next dataset version, and add a test that fails when the dataset names a tool the agent does not declare:
@Test
void datasetOnlyNamesDeclaredTools() {
Set<String> declared = rootAgent.tools().blockingGet().stream()
.map(BaseTool::name).collect(Collectors.toSet());
List<String> unknown = Dataset.load(Path.of("eval/refunds.jsonl")).rows().stream()
.flatMap(r -> Stream.concat(r.label().requiredTools().stream(), r.label().forbiddenTools().stream())
.filter(t -> !declared.contains(t)).map(t -> r.id() + ":" + t))
.toList();
assertTrue(unknown.isEmpty(), "rows name undeclared tools: " + unknown);
}In ADK Java 1.11.0 LlmAgent.tools() returns a Single<List<BaseTool>>, hence the blockingGet() in a test. If your agent uses toolsets that resolve per context, list the names from the same configuration the agent is built from. A forbidden tool that no longer exists is just as stale as a required one. The rule is vacuously true, so the row looks safe while testing nothing.
Privacy for rows mined from production
Rows mined from production contain customer data, and a dataset in git lasts for years. Scrub before drafting, not afterwards. Replace identifiers with keyed pseudonyms, so the same customer gets the same token across all of a conversation's turns and the row stays coherent:
static String pseudonym(String kind, String value, byte[] key) throws GeneralSecurityException {
Mac mac = Mac.getInstance("HmacSHA256");
mac.init(new SecretKeySpec(key, "HmacSHA256"));
byte[] h = mac.doFinal((kind + ":" + value).getBytes(StandardCharsets.UTF_8));
return kind + "_" + HexFormat.of().formatHex(h, 0, 6); // e.g. order_3fa91c0b27d4
}Keep the key out of the repository, so nobody can reverse a token by hashing guesses. Replace live tool responses with fixtures that use the same pseudonyms. Store a hashed source session id and a retain_until date on every mined row. Then a deletion request becomes a query, and expired rows can be dropped in a dataset change like any other.
Worked example: an unstable slice that was a labelling problem
Take a refunds slice of 120 rows whose pass rate has bounced between 68% and 79% over four releases, with no relevant code changes. A second labeller relabels 60 rows blind. Observed agreement on the outcome class is 0.75, and expected agreement from the two labellers' class rates is 0.48, so kappa is (0.75 - 0.48) / 0.52 = 0.52. That is in the ambiguous band. Of the 15 disagreements, 11 involve partial refunds on orders older than 30 days, a case the guideline never mentions. One labeller expected escalation and the other expected a clarifying question.
The team adds a rule (older than 30 days and partially shipped: escalate), bumps the guideline to g8, and relabels the 31 rows the rule touches, which changes the expected outcome on 14 of them. A repeat sample gives kappa 0.84. The agent's pass rate on the slice is now 83%, and it holds steady across the next two releases. Nothing in the agent changed. The earlier 'regressions' were the model choosing between two answers the data itself had not decided between. The same review runs the leakage scan and finds nine holdout rows with Jaccard above 0.85 against dev rows, all mined from one outage. They move to dev, and the holdout gets nine fresh rows from a later week.
Failure modes
- Labels copied from agent output. Draft rows built from observed behaviour lock in its bugs. Labellers must decide the expected behaviour without seeing it.
- One global kappa. A good average hides the one slice where labels are coin flips. Compute it per slice.
- Synthetic flood. Generated rows outnumber real ones, and the headline pass rate measures how the agent handles the generator's style.
- Silent migration. A tool rename turns required-tool rules into failures and forbidden-tool rules into checks that can never fail.
- Deleting near-duplicates blindly. Short, common intents are supposed to look alike. Review flagged pairs and move them between splits; do not purge them.
- Scrubbing after commit. Git history keeps the original. Scrub in the mining job, before a row is ever written into the repository.
Trade-offs
Double labelling roughly doubles labelling cost. Spend it where it buys the most: on new slices, slices with low kappa, and rows that gate releases, and use single labels with spot checks elsewhere. Synthetic rows are cheap coverage with a known bias, so take the coverage and report the bias. Strict leakage thresholds protect the holdout but shrink it, so refresh it from later traffic rather than relaxing the threshold.
What to do next
- Write a labelling guideline with the four label fields, give it a version, and put that version on every row.
- Have a second labeller relabel 50-60 rows of your noisiest slice blind, and compute kappa with the function above.
- Rewrite the guideline where disagreements cluster, then relabel only the affected rows as a new dataset version.
- Add the transform enum to your synthetic generator, and send every information-changing variant back to labelling.
- Run the shingle scan across dev and holdout, and the containment check against your prompt, in CI.
- Add the declared-tools test, and keep a rename map for every tool refactor.
- Move pseudonymization into the mining job and add
retain_untilto mined rows. Then feed the cleaned dataset into your CI eval gate and continuous evaluation.