Automated evaluation scales, but something has to define what good looks like, and for most agent behaviour that something is people. An LLM judge is only as trustworthy as its agreement with careful human raters; a regression suite is only as useful as the labels on its cases. Human evaluation is how you produce those labels, and it is easy to do badly: vague rubrics, a single rater, unblinded comparisons, and no measure of whether the raters agree with each other.
ADK Java does not ship a human evaluation module; the 1.11.0 core library has no evaluation classes at all. You build the program from the session store, a rating tool and a few statistics. This article covers the rubric, exporting cases from ADK Java sessions, task design, inter-rater agreement with a worked Cohen's kappa, a pairwise comparison with a confidence interval, calibrating an LLM judge, and running the program. Sampling production traffic for automated judging is covered in continuous evaluation.
When human evaluation is worth it
Human rating is slow and expensive, so spend it where nothing else will do:
- Calibrating judges. Before an LLM judge gates releases, measure its agreement with humans on a few hundred items.
- Subjective qualities. Tone, helpfulness and whether an answer would actually resolve a customer's problem resist automatic metrics.
- High-stakes domains. Medical, legal and financial answers need domain experts, not generalist raters or a judge.
- Big changes. A new model, a rewritten instruction or a new tool set deserves a blind comparison against the current version.
- Building regression sets. Adjudicated human labels become the expected outcomes in dataset-driven evaluation.
Do not use humans for what code can check: whether the right tool was called with the right arguments, whether output parses, whether a forbidden action happened. Those belong in deterministic cases, described in writing eval test cases.
Designing the rubric
A rubric turns an opinion into a measurement. Each criterion needs a name, a question a rater can answer from the transcript alone, and anchored levels with an example of each. Prefer a few binary or three-level criteria over one 1-to-10 score: raters agree far more on "did the answer contain a factual error" than on "rate the quality".
| Criterion | Question | Levels |
|---|---|---|
| Task resolved | Would this response let the user complete what they asked? | yes / partly / no |
| Factual accuracy | Does the response state anything false about the account, policy or world? | no errors / minor / material |
| Tool use | Did the agent take an action the user did not ask for or would not want? | no / yes |
| Tone | Is the tone appropriate for a support conversation? | yes / no |
| Safety | Does the response violate the content policy? | no / yes (flag for review) |
Write the rubric with examples drawn from real transcripts, including borderline ones, and pilot it: three raters label the same 30 items, you read every disagreement together, and you rewrite the questions that caused them. Version the rubric like code. Labels collected under rubric v2 are not directly comparable with v1.
Exporting cases from ADK Java sessions
Raters need a readable transcript: the user messages, the agent's answers, and which tools were called with what result, because judging tool use without seeing it is guesswork. The session store already holds all of that as events. This extractor turns one session into a case record, redacting as it goes.
record Turn(String role, String text, List<String> toolCalls, List<String> toolResults) {}
record Case(String caseId, String appName, String sourceSession, String agentVersion,
List<Turn> turns) {}
final class CaseExtractor {
private final BaseSessionService sessions;
private final Redactor redactor; // same detector as production redaction
Case extract(String app, String user, String sessionId, String agentVersion) {
Session s = sessions.getSession(app, user, sessionId, Optional.empty())
.blockingGet(); // batch job, blocking is fine here
if (s == null) throw new IllegalArgumentException("no session " + sessionId);
List<Turn> turns = new ArrayList<>();
for (Event e : s.events()) {
if (e.partial().orElse(false)) continue; // keep only complete events
String text = e.content().map(Texts::visibleText).orElse("");
List<String> calls = e.functionCalls().stream()
.map(fc -> fc.name().orElse("?") + "(" + redactor.redact(String.valueOf(fc.args().orElse(Map.of()))) + ")")
.toList();
List<String> results = e.functionResponses().stream()
.map(fr -> fr.name().orElse("?") + " -> " + redactor.redact(String.valueOf(fr.response().orElse(Map.of()))))
.toList();
if (text.isBlank() && calls.isEmpty() && results.isEmpty()) continue;
turns.add(new Turn(e.author(), redactor.redact(text), calls, results));
}
String caseId = Hashing.sha256().hashString(app + "/" + sessionId, UTF_8).toString().substring(0, 16);
return new Case(caseId, app, sessionId, agentVersion, turns);
}
}Three details matter. Stamp each case with the agent version (instruction hash, model, tool set) so results attach to a build. Redact before the case leaves your environment: raters, especially external ones, should not see customer data, and the same detector you use in production keeps behaviour consistent; see PII redaction. And choose sessions by stratified sampling over intent and outcome, not the most recent N, or the evaluation overrepresents whatever was busy last week.
Task design: absolute and pairwise
Two task formats cover most needs. Absolute rating shows one transcript and asks the rubric questions; it tells you how good a version is. Pairwise comparison shows two responses to the same conversation and asks which is better on each criterion, with a tie option; it is more sensitive to small differences because people compare more consistently than they score.
For pairwise comparisons, generate both responses from the same conversation prefix by replaying the user turns against each agent version, then:
- Blind the raters: no version names, no model names, identical formatting.
- Randomise position per item, because raters favour whichever side they read first or last.
- Overlap: give every item to at least two raters, or a fixed 20% of items to three, so agreement can be measured.
- Seed gold items with known answers at around 5% of the queue; a rater who misses several is retrained before their labels count.
Do the raters agree? Cohen's kappa, worked
Raw agreement overstates reliability, because two raters who both answer "yes" most of the time agree often by chance. Cohen's kappa corrects for that: kappa = (po - pe) / (1 - pe), where po is observed agreement and pe is the agreement expected if each rater labelled at random with their own base rates.
Worked example. Two raters label the same 100 transcripts as resolved or not. Both say resolved on 60, both say not resolved on 20, rater A says resolved where B says not on 12, and the reverse on 8. Observed agreement is 80 of 100, so po = 0.80. Rater A says resolved 72 times and B 68 times, so chance agreement is 0.72 x 0.68 + 0.28 x 0.32 = 0.4896 + 0.0896 = 0.5792. Kappa is (0.80 - 0.5792) / (1 - 0.5792) = 0.2208 / 0.4208, about 0.52.
So 80% raw agreement is only moderate reliability. The conventional Landis and Koch labels call 0.41 to 0.60 moderate and 0.61 to 0.80 substantial; they are rules of thumb, not laws. A kappa near 0.5 on a criterion means the rubric question is ambiguous, and the fix is to read the 20 disagreements and rewrite the question, not to add raters. For more than two raters or missing labels, use Fleiss' kappa or Krippendorff's alpha.
static double cohensKappa(int[][] m) { // m[i][j]: rater A said i, rater B said j
int k = m.length; double n = 0, agree = 0;
double[] rowSum = new double[k], colSum = new double[k];
for (int i = 0; i < k; i++)
for (int j = 0; j < k; j++) {
n += m[i][j]; rowSum[i] += m[i][j]; colSum[j] += m[i][j];
if (i == j) agree += m[i][j];
}
double po = agree / n, pe = 0;
for (int i = 0; i < k; i++) pe += (rowSum[i] / n) * (colSum[i] / n);
return (po - pe) / (1 - pe);
}
// cohensKappa(new int[][] {{60, 12}, {8, 20}}) == 0.5247...
Worked example: is the new instruction better?
A team compares a new instruction (B) against the current one (A) on 200 conversations, pairwise and blind. B wins 92, A wins 70, and 38 are ties. Dropping ties, which is one common convention (report the tie rate alongside), B's win rate is 92 of 162, about 0.568.
Is that real? A Wilson score interval at 95% for 92 successes in 162 trials runs from about 0.491 to 0.642. The interval includes 0.5, so the data do not show B is better, even though B won 22 more comparisons. To detect a true 57% win rate reliably you would need several hundred non-tied comparisons. The honest report is: "B is not worse, and may be better; we are extending the study to 400 items", not "B wins 57%".
static double[] wilson(int wins, int n, double z) {
double p = (double) wins / n, z2 = z * z;
double centre = (p + z2 / (2 * n)) / (1 + z2 / n);
double half = z * Math.sqrt(p * (1 - p) / n + z2 / (4.0 * n * n)) / (1 + z2 / n);
return new double[] {centre - half, centre + half};
}
// wilson(92, 162, 1.96) -> [0.491, 0.642]Look at the per-criterion results too. A new version can win overall while losing on factual accuracy, which a single preference question would hide.
Calibrating an LLM judge with human labels
Once humans have labelled a few hundred items, use them to calibrate the LLM judge that will do the daily work. Run the judge on the same items with the same rubric text, then compute judge-to-human kappa per criterion, using the adjudicated human label (majority or reviewed) as the reference. Compare it with human-to-human kappa: a judge cannot be expected to agree with humans more than humans agree with each other, but it should come close.
Read every disagreement. Patterns are usually systematic: the judge rewards length, misses errors in tool arguments it cannot see, or applies a stricter tone standard than people. Fix the judge prompt or give it the tool calls, re-run on the same items, and keep a held-out set the prompt was never tuned on so you do not overfit. Re-check calibration whenever the judge model, the rubric or the agent changes substantially. Adjudicated items graduate into the regression set that CI evaluation runs on every build.
Running the program
Budget first. At about two minutes per absolute rating, 300 items with three raters each is 1,800 minutes, or 30 rater-hours, before adjudication. Pairwise items take longer. Most teams run a small standing panel of trained raters for routine work and bring in domain experts for specialised criteria.
Track rater health as data: agreement with gold items, agreement with peers, speed (very fast raters are often not reading) and drift over time. Keep rater identities in the labels table so a bad batch can be removed. Give raters an "I can't tell" option rather than forcing a guess, and count it, since a high rate means missing context. If raters will see harmful content, limit exposure and make opting out easy.
Failure modes
- Single rater. No agreement measure means no idea whether labels are signal. Overlap at least a fifth of items.
- Unblinded comparisons. Raters who know which version is new favour it. Hide versions and randomise sides.
- Convenience sampling. Rating whatever was easy to export skews toward short, successful sessions. Stratify by intent and outcome.
- Missing context. Raters judging an answer without the tool results cannot tell a hallucination from a correct lookup. Include tool calls and results.
- Leaking data. Unredacted transcripts in a rating vendor's tool are a privacy incident. Redact at extraction.
- Overreading small samples. A 57% win rate on 162 items is not a finding. Report intervals.
What to do next
- Write a rubric of four or five binary or three-level criteria with anchored examples, and pilot it with three raters on 30 items.
- Build a case extractor over
BaseSessionService.getSessionthat includes tool calls, redacts and stamps the agent version. - Sample 200 to 300 sessions stratified by intent and outcome.
- Run blind, position-randomised tasks with 20% overlap and 5% gold items.
- Compute kappa per criterion and rewrite any question below about 0.6.
- Report pairwise results with Wilson intervals and per-criterion breakdowns.
- Calibrate your LLM judge against the adjudicated labels and move those cases into the regression set.