Automated evaluation scales, but something has to define what good looks like, and for most agent behaviour that something is people. An LLM judge is only as trustworthy as its agreement with careful human raters; a regression suite is only as useful as the labels on its cases. Human evaluation is how you produce those labels, and it is easy to do badly: vague rubrics, a single rater, unblinded comparisons, and no measure of whether the raters agree with each other.

ADK Java does not ship a human evaluation module; the 1.11.0 core library has no evaluation classes at all. You build the program from the session store, a rating tool and a few statistics. This article covers the rubric, exporting cases from ADK Java sessions, task design, inter-rater agreement with a worked Cohen's kappa, a pairwise comparison with a confidence interval, calibrating an LLM judge, and running the program. Sampling production traffic for automated judging is covered in continuous evaluation.

When human evaluation is worth it

Human rating is slow and expensive, so spend it where nothing else will do:

  • Calibrating judges. Before an LLM judge gates releases, measure its agreement with humans on a few hundred items.
  • Subjective qualities. Tone, helpfulness and whether an answer would actually resolve a customer's problem resist automatic metrics.
  • High-stakes domains. Medical, legal and financial answers need domain experts, not generalist raters or a judge.
  • Big changes. A new model, a rewritten instruction or a new tool set deserves a blind comparison against the current version.
  • Building regression sets. Adjudicated human labels become the expected outcomes in dataset-driven evaluation.

Do not use humans for what code can check: whether the right tool was called with the right arguments, whether output parses, whether a forbidden action happened. Those belong in deterministic cases, described in writing eval test cases.

Designing the rubric

A human evaluation loop for an ADK Java agentSession storegetSession(...)Case extractorredact, render, hashTask builderblind, shuffle, goldRater toolrubric, 2-3 ratersLabels tablerater, item, scoresAnalysiskappa, win rate, CILLM judge calibrationjudge vs human kappaRegression setadjudicated casesRater calibration: gold items, disagreement review, rubric editsHumans set the standard; the judge and the regression set carry it into every build.
Cases flow from the session store through redaction and blinding to raters; their labels calibrate the judge and become the regression set.

A rubric turns an opinion into a measurement. Each criterion needs a name, a question a rater can answer from the transcript alone, and anchored levels with an example of each. Prefer a few binary or three-level criteria over one 1-to-10 score: raters agree far more on "did the answer contain a factual error" than on "rate the quality".

CriterionQuestionLevels
Task resolvedWould this response let the user complete what they asked?yes / partly / no
Factual accuracyDoes the response state anything false about the account, policy or world?no errors / minor / material
Tool useDid the agent take an action the user did not ask for or would not want?no / yes
ToneIs the tone appropriate for a support conversation?yes / no
SafetyDoes the response violate the content policy?no / yes (flag for review)

Write the rubric with examples drawn from real transcripts, including borderline ones, and pilot it: three raters label the same 30 items, you read every disagreement together, and you rewrite the questions that caused them. Version the rubric like code. Labels collected under rubric v2 are not directly comparable with v1.

Exporting cases from ADK Java sessions

Raters need a readable transcript: the user messages, the agent's answers, and which tools were called with what result, because judging tool use without seeing it is guesswork. The session store already holds all of that as events. This extractor turns one session into a case record, redacting as it goes.

record Turn(String role, String text, List<String> toolCalls, List<String> toolResults) {}
record Case(String caseId, String appName, String sourceSession, String agentVersion,
            List<Turn> turns) {}

final class CaseExtractor {
  private final BaseSessionService sessions;
  private final Redactor redactor;               // same detector as production redaction

  Case extract(String app, String user, String sessionId, String agentVersion) {
    Session s = sessions.getSession(app, user, sessionId, Optional.empty())
        .blockingGet();                           // batch job, blocking is fine here
    if (s == null) throw new IllegalArgumentException("no session " + sessionId);
    List<Turn> turns = new ArrayList<>();
    for (Event e : s.events()) {
      if (e.partial().orElse(false)) continue;    // keep only complete events
      String text = e.content().map(Texts::visibleText).orElse("");
      List<String> calls = e.functionCalls().stream()
          .map(fc -> fc.name().orElse("?") + "(" + redactor.redact(String.valueOf(fc.args().orElse(Map.of()))) + ")")
          .toList();
      List<String> results = e.functionResponses().stream()
          .map(fr -> fr.name().orElse("?") + " -> " + redactor.redact(String.valueOf(fr.response().orElse(Map.of()))))
          .toList();
      if (text.isBlank() && calls.isEmpty() && results.isEmpty()) continue;
      turns.add(new Turn(e.author(), redactor.redact(text), calls, results));
    }
    String caseId = Hashing.sha256().hashString(app + "/" + sessionId, UTF_8).toString().substring(0, 16);
    return new Case(caseId, app, sessionId, agentVersion, turns);
  }
}

Three details matter. Stamp each case with the agent version (instruction hash, model, tool set) so results attach to a build. Redact before the case leaves your environment: raters, especially external ones, should not see customer data, and the same detector you use in production keeps behaviour consistent; see PII redaction. And choose sessions by stratified sampling over intent and outcome, not the most recent N, or the evaluation overrepresents whatever was busy last week.

Task design: absolute and pairwise

Two task formats cover most needs. Absolute rating shows one transcript and asks the rubric questions; it tells you how good a version is. Pairwise comparison shows two responses to the same conversation and asks which is better on each criterion, with a tie option; it is more sensitive to small differences because people compare more consistently than they score.

For pairwise comparisons, generate both responses from the same conversation prefix by replaying the user turns against each agent version, then:

  • Blind the raters: no version names, no model names, identical formatting.
  • Randomise position per item, because raters favour whichever side they read first or last.
  • Overlap: give every item to at least two raters, or a fixed 20% of items to three, so agreement can be measured.
  • Seed gold items with known answers at around 5% of the queue; a rater who misses several is retrained before their labels count.

Do the raters agree? Cohen&#x27;s kappa, worked

Raw agreement overstates reliability, because two raters who both answer "yes" most of the time agree often by chance. Cohen's kappa corrects for that: kappa = (po - pe) / (1 - pe), where po is observed agreement and pe is the agreement expected if each rater labelled at random with their own base rates.

Worked example. Two raters label the same 100 transcripts as resolved or not. Both say resolved on 60, both say not resolved on 20, rater A says resolved where B says not on 12, and the reverse on 8. Observed agreement is 80 of 100, so po = 0.80. Rater A says resolved 72 times and B 68 times, so chance agreement is 0.72 x 0.68 + 0.28 x 0.32 = 0.4896 + 0.0896 = 0.5792. Kappa is (0.80 - 0.5792) / (1 - 0.5792) = 0.2208 / 0.4208, about 0.52.

So 80% raw agreement is only moderate reliability. The conventional Landis and Koch labels call 0.41 to 0.60 moderate and 0.61 to 0.80 substantial; they are rules of thumb, not laws. A kappa near 0.5 on a criterion means the rubric question is ambiguous, and the fix is to read the 20 disagreements and rewrite the question, not to add raters. For more than two raters or missing labels, use Fleiss' kappa or Krippendorff's alpha.

static double cohensKappa(int[][] m) {          // m[i][j]: rater A said i, rater B said j
  int k = m.length; double n = 0, agree = 0;
  double[] rowSum = new double[k], colSum = new double[k];
  for (int i = 0; i < k; i++)
    for (int j = 0; j < k; j++) {
      n += m[i][j]; rowSum[i] += m[i][j]; colSum[j] += m[i][j];
      if (i == j) agree += m[i][j];
    }
  double po = agree / n, pe = 0;
  for (int i = 0; i < k; i++) pe += (rowSum[i] / n) * (colSum[i] / n);
  return (po - pe) / (1 - pe);
}
// cohensKappa(new int[][] {{60, 12}, {8, 20}}) == 0.5247...

Worked example: is the new instruction better?

A team compares a new instruction (B) against the current one (A) on 200 conversations, pairwise and blind. B wins 92, A wins 70, and 38 are ties. Dropping ties, which is one common convention (report the tie rate alongside), B's win rate is 92 of 162, about 0.568.

Is that real? A Wilson score interval at 95% for 92 successes in 162 trials runs from about 0.491 to 0.642. The interval includes 0.5, so the data do not show B is better, even though B won 22 more comparisons. To detect a true 57% win rate reliably you would need several hundred non-tied comparisons. The honest report is: "B is not worse, and may be better; we are extending the study to 400 items", not "B wins 57%".

static double[] wilson(int wins, int n, double z) {
  double p = (double) wins / n, z2 = z * z;
  double centre = (p + z2 / (2 * n)) / (1 + z2 / n);
  double half = z * Math.sqrt(p * (1 - p) / n + z2 / (4.0 * n * n)) / (1 + z2 / n);
  return new double[] {centre - half, centre + half};
}
// wilson(92, 162, 1.96) -> [0.491, 0.642]

Look at the per-criterion results too. A new version can win overall while losing on factual accuracy, which a single preference question would hide.

Calibrating an LLM judge with human labels

Once humans have labelled a few hundred items, use them to calibrate the LLM judge that will do the daily work. Run the judge on the same items with the same rubric text, then compute judge-to-human kappa per criterion, using the adjudicated human label (majority or reviewed) as the reference. Compare it with human-to-human kappa: a judge cannot be expected to agree with humans more than humans agree with each other, but it should come close.

Read every disagreement. Patterns are usually systematic: the judge rewards length, misses errors in tool arguments it cannot see, or applies a stricter tone standard than people. Fix the judge prompt or give it the tool calls, re-run on the same items, and keep a held-out set the prompt was never tuned on so you do not overfit. Re-check calibration whenever the judge model, the rubric or the agent changes substantially. Adjudicated items graduate into the regression set that CI evaluation runs on every build.

Running the program

Budget first. At about two minutes per absolute rating, 300 items with three raters each is 1,800 minutes, or 30 rater-hours, before adjudication. Pairwise items take longer. Most teams run a small standing panel of trained raters for routine work and bring in domain experts for specialised criteria.

Track rater health as data: agreement with gold items, agreement with peers, speed (very fast raters are often not reading) and drift over time. Keep rater identities in the labels table so a bad batch can be removed. Give raters an "I can't tell" option rather than forcing a guess, and count it, since a high rate means missing context. If raters will see harmful content, limit exposure and make opting out easy.

Failure modes

  • Single rater. No agreement measure means no idea whether labels are signal. Overlap at least a fifth of items.
  • Unblinded comparisons. Raters who know which version is new favour it. Hide versions and randomise sides.
  • Convenience sampling. Rating whatever was easy to export skews toward short, successful sessions. Stratify by intent and outcome.
  • Missing context. Raters judging an answer without the tool results cannot tell a hallucination from a correct lookup. Include tool calls and results.
  • Leaking data. Unredacted transcripts in a rating vendor's tool are a privacy incident. Redact at extraction.
  • Overreading small samples. A 57% win rate on 162 items is not a finding. Report intervals.

What to do next

  1. Write a rubric of four or five binary or three-level criteria with anchored examples, and pilot it with three raters on 30 items.
  2. Build a case extractor over BaseSessionService.getSession that includes tool calls, redacts and stamps the agent version.
  3. Sample 200 to 300 sessions stratified by intent and outcome.
  4. Run blind, position-randomised tasks with 20% overlap and 5% gold items.
  5. Compute kappa per criterion and rewrite any question below about 0.6.
  6. Report pairwise results with Wilson intervals and per-criterion breakdowns.
  7. Calibrate your LLM judge against the adjudicated labels and move those cases into the regression set.
Key takeaway: Human evaluation defines what good means for your agent, and everything automated inherits its quality. ADK Java gives you the session store, not an evaluation module, so build a small program: an anchored rubric, redacted cases with tool calls, blind and position-randomised tasks, overlap and gold items, kappa per criterion, and intervals on every comparison. Then use the adjudicated labels to calibrate the LLM judge and seed the regression set.