An agent eval compares what the agent did with what it should have done. The second half, the ground truth, is the part teams spend least time on and trust most. When an eval fails, the first question should be whether the agent is wrong or the label is, and most suites cannot answer it, because nobody recorded who wrote the label, under which policy, or when it stops being true.

This page treats ground truth as managed data with its own lifecycle. The dataset-driven evaluation article covers storing cases as a versioned artifact and choosing which rows run, and writing eval test cases covers what a good case contains. Here the subject is the label itself: what it should record, how two people agree on it, how it expires, and how label errors distort the numbers you report. The code is plain Java with no framework dependency, so it sits beside whichever harness replays your cases. The eval framework overview tracks what ADK Java itself provides, and you should check it against the release you run.

Four kinds of truth for an agent

For a chat model, ground truth is often one reference answer. An agent needs more, because it acts. Separate four kinds of truth, since each is labelled, checked and broken differently:

  • Trajectory truth: the tools that must be called, the tools that must not be called, and argument constraints that matter (the refund amount, not the request id). Required and forbidden lists age better than an exact sequence, because harmless reorderings stop counting as failures.
  • Outcome truth: a set of acceptable final answers, or the facts the answer must state, not one golden string. Two correct phrasings should both pass.
  • Rubric truth: criteria for open-ended answers, which a person or a judge model applies. The rubric is the label; the judge only applies it.
  • State truth: what should be different in the world afterwards: the order row marked refunded, no duplicate ticket created. This is the most reliable truth, because it checks the effect rather than the narration.

Every kind depends on context the label writer saw: the policy document, the tool data, and the date. A label without that context is an opinion with no expiry date.

A truth record and its validator

Store truth beside the cases, not inside them. A truth file is JSONL keyed by case id and revision, so the case (the user's turns and seeded state) can stay stable while its label evolves. Each record carries the expectation and its provenance:

enum Status { DRAFT, SILVER, GOLD, DISPUTED, RETIRED }

/** One label for one eval case: what "correct" means, who said so, and until when. */
record TruthRecord(String caseId, int revision, Status status,
                   List<String> requiredTools, List<String> forbiddenTools,
                   List<String> acceptableAnswers, String rubric,
                   String policyVersion, LocalDate asOf, LocalDate expires,
                   List<String> labelers, String adjudicator, String rationale) {}

/** Problems that must block a run, checked before any agent is called. */
static List<String> validate(TruthRecord t, String currentPolicy, LocalDate today) {
    List<String> p = new ArrayList<>();
    if (t.status() != Status.GOLD && t.status() != Status.SILVER)
        p.add(t.caseId() + ": status " + t.status() + " is not scorable");
    if (t.status() == Status.GOLD && (t.labelers().size() < 2 || t.adjudicator() == null))
        p.add(t.caseId() + ": gold needs two labelers and an adjudicator");
    if (!t.policyVersion().equals(currentPolicy))
        p.add(t.caseId() + ": labelled under policy " + t.policyVersion()
              + ", current is " + currentPolicy);
    if (t.expires() != null && !today.isBefore(t.expires()))
        p.add(t.caseId() + ": expired on " + t.expires());
    if (t.acceptableAnswers().isEmpty() && t.rubric() == null)
        p.add(t.caseId() + ": no acceptable answer and no rubric");
    if (!Collections.disjoint(t.requiredTools(), t.forbiddenTools()))
        p.add(t.caseId() + ": a tool is both required and forbidden");
    return p;
}

Run validate over the whole truth file before the first model call. For a gold record for refund-0142, labelled under refunds-v7 on 2026-09-18 with a six-month expiry, it returns an empty list today. Run it again with the policy at refunds-v8 on 2027-04-01 and it returns two problems: the policy moved, and the label expired on 2027-03-18. Those cases are excluded and queued for relabelling, not silently scored against a stale answer.

Three rules keep the file trustworthy. A (caseId, revision) pair is immutable, so a past run can always be re-scored against exactly the truth it used. Changes arrive as new revisions in reviewed commits, with an owner group required on the truth directory. And the rationale field is mandatory, because the next person to dispute the label needs the reason, not just the verdict.

Join truth to cases at load time. The case files keep the Python ADK eval-set shape the overview describes: eval_cases, each with an eval_id and a conversation of turns carrying final_response and intermediate_data.tool_uses. Load the truth JSONL, keep the latest revision per eval_id, and only then filter by status, so a DISPUTED revision hides the older GOLD one instead of falling back to it:

Map<String, TruthRecord> latest = records.stream().collect(Collectors.toMap(
        TruthRecord::caseId, t -> t, (x, y) -> x.revision() >= y.revision() ? x : y));
Set<String> scorable = latest.values().stream()
        .filter(t -> t.status() == Status.GOLD || t.status() == Status.SILVER)
        .map(TruthRecord::caseId).collect(Collectors.toSet());
List<String> excluded = caseIds.stream().filter(id -> !scorable.contains(id)).toList();

Scoring then reads the observed run, not the case file. The required and forbidden tools are checked against the tool names in the observed tool_uses, and the acceptable answers against the observed final response. The expectations written in the case file stay as readable examples, but the truth record wins when they differ, and a lint step should flag the difference. Report the excluded list with every run, so a suite that quietly lost half its cases to expiry is visible.

The life of one label: every transition is a reviewed commitDRAFTmined case, no labelSILVERone labeler or LLM draftGOLDtwo labelers + adjudicatorRETIREDkept for historylabelagreeobsoleteDISPUTEDexcluded from scoresaudit finds errorrelabeltriggerspolicy change, expiry, tool changescorer readsGOLD for gates, SILVER for trendsA new revision never edits an old one: (caseId, revision) is immutable, so past runs stay reproducible
Label lifecycle. Only GOLD labels gate releases; SILVER labels feed trend dashboards; DISPUTED labels are excluded until relabelled.

Agreement and adjudication

A label is only as good as the agreement behind it. For gold status, two people label each case independently, then an adjudicator resolves differences and writes the rationale. Measure the independent agreement, because it tells you whether the task is well defined. Raw percent agreement overstates it when one answer dominates, so use Cohen's kappa, which subtracts the agreement two labelers would reach by chance given how often each says each label:

static double kappa(List<String> a, List<String> b) {
    if (a.size() != b.size() || a.isEmpty()) throw new IllegalArgumentException("paired labels required");
    int n = a.size(), agree = 0;
    Map<String, Integer> ca = new HashMap<>(), cb = new HashMap<>();
    for (int i = 0; i < n; i++) {
        if (a.get(i).equals(b.get(i))) agree++;
        ca.merge(a.get(i), 1, Integer::sum);
        cb.merge(b.get(i), 1, Integer::sum);
    }
    double po = (double) agree / n, pe = 0;
    for (var e : ca.entrySet())
        pe += (e.getValue() / (double) n) * (cb.getOrDefault(e.getKey(), 0) / (double) n);
    return pe == 1.0 ? 1.0 : (po - pe) / (1 - pe);
}

Worked example: two labelers judge 20 agent transcripts as pass or fail. They agree on 17, which is 85%. Labeler A says pass 12 times and labeler B 13 times, so chance agreement is 0.60 x 0.65 + 0.40 x 0.35 = 0.53. Kappa is (0.85 - 0.53) / (1 - 0.53) = 0.681, which kappa returns for these labels. That is respectable but not strong; the three disagreements are where your rubric is ambiguous. Read them, sharpen the rubric, and relabel a fresh sample. A common rule of thumb treats kappa above about 0.8 as strong agreement, but treat any such cut-off as a convention, not a law.

Track kappa per slice, not globally. A refund slice at 0.9 and a tone slice at 0.4 average to a comfortable number that hides the fact that the tone labels are close to noise. Low-agreement slices should gate nothing until the rubric improves.

Policy, world and tool changes

Agent truth expires for three reasons, and each needs a field in the record.

  • Policy changes. The refund window moves from 30 to 14 days, and every refund label written under the old policy may now be wrong. Tag labels with policyVersion. On a policy change, query every record with the old version and open a relabelling campaign, rather than discovering them as failures one at a time.
  • World changes. 'The order is inside the window' was true on the asOf date and false a month later. Either freeze the world by replaying recorded tool responses, which is the cassette approach in eval CI integration, or give the label an expires date. Prefer frozen fixtures for anything that gates a release.
  • Tool changes. A tool is renamed or split in two, and trajectory labels naming the old tool fail every case. Keep a rename map applied at load time and bump the revision of every affected record in one commit, so the diff is reviewable.

Label noise and how to find it

Label errors do not just add noise; they bias the measurement in a predictable direction. Suppose a fraction e of labels is wrong, and the agent's true success rate is a. On open-ended tasks a wrong agent answer almost never matches a wrong label, so the measured pass rate is close to a(1 - e). Two consequences follow:

  • A ceiling. With 5% bad labels no agent can score above 95%. A candidate stuck at 94% may be finished, and the remaining work is in the labels.
  • Compressed differences. With e = 0.05, an agent at a = 0.90 measures 0.855 and one at 0.93 measures about 0.884. A real 3-point gain shows as 2.85 points, and with a few hundred cases it can fall inside run-to-run variance.

Find bad labels by mining disagreement. Each week, collect the cases where two strong candidate agents both fail but agree with each other, or where an LLM judge passes an answer the exact-match check fails. Send those to a human audit. This queue should hold far more label errors than a random sample, so audit time goes where the errors are; measure its overturn rate to confirm. Record each audit outcome: an overturned label becomes DISPUTED, then a new revision.

Model-drafted labels have a place. A model can draft expected answers for mined cases quickly, and those are SILVER: good enough for trend dashboards, never for a release gate. Two rules avoid circularity. Never let the model under evaluation draft its own labels, and never let the judge model that scores a case also have drafted its label, or both will share the same blind spot.

Failure modes

Failure modes seen in real suites:

  • Editing a label in place. Last month's pass rate can no longer be reproduced, and nobody can tell whether a jump came from the agent or the truth. Revisions only.
  • Truth leaking into prompts. Someone pastes expected answers into few-shot examples or into a retrieval index the agent can search. Keep holdout truth in a directory the agent's build cannot read, and plant canary strings to detect leaks.
  • One golden string. Correct paraphrases fail, engineers tune the prompt toward one phrasing, and the suite rewards memorisation.
  • Unowned labels. Labels written by a contractor who has left, under a policy nobody recorded. Require labelers and rationale before GOLD.
  • Comparing across truth versions. Fix three labels and the pass rate rises in the same PR as a prompt change. Put the truth file's hash in every results key and refuse comparisons that cross hashes.
  • Scoring disputed cases. A label under audit keeps failing the build. Exclude DISPUTED records automatically and report how many were excluded.

Trade-offs

ChoiceGainCost
Double labelling plus adjudicationMeasured agreement, defensible gatesAbout twice the labelling time
Acceptable-answer setsParaphrases pass; less prompt overfittingSets must be maintained as answers evolve
Rubrics applied by a judgeCovers open-ended answersJudge drift; needs calibration against human labels
Frozen tool fixturesLabels stay true indefinitelyFixtures drift from production behaviour
Expiry datesLive data stays honestSteady relabelling workload
SILVER model draftsFast coverage of new casesCannot gate releases; risk of shared blind spots

What to do next

  1. Split truth out of your case files into a JSONL file keyed by (caseId, revision), and add status, policy version, as-of, expiry, labelers, adjudicator and rationale.
  2. Run the validator before every eval run, and fail the run if a gating case has an expired or out-of-policy label.
  3. Double-label a sample of 50 cases per slice, compute kappa, and fix the rubric for any slice that comes out low.
  4. Replace single golden strings with acceptable-answer sets, and replace exact trajectories with required and forbidden tool lists.
  5. Start a weekly disagreement-mining audit, and track how many labels it overturns.
  6. Add the truth file's hash to every results key, and refuse comparisons across hashes.
Key takeaway: Treat labels as managed data: keep them beside cases, never edit them in place, and record who wrote each one, under which policy and until when. Double-label and measure kappa before trusting a gate. Remember that a 5% label error rate caps every score at 95% and shrinks real gains, and mine disagreements to find the bad labels.