Every agent evaluation ends with a comparison: did the agent's answer match what we expected? The two obvious ways to answer are at opposite ends. Exact match says yes only when the strings are identical, so it rejects correct answers that are phrased differently. Semantic similarity says yes when the meanings are close, so it can accept an answer that differs in the one word that mattered. Neither is right in general. The skill is picking the strategy per field, normalizing before comparing, and calibrating any threshold against cases a person has labelled.
This article builds that skill for ADK Java evaluations. It covers what a scorer's two kinds of error cost, a normalization ladder that makes exact match usable, per-field scoring for structured output, where token overlap and embedding similarity fail, with computed numbers, how to calibrate a threshold, and a cascade scorer that combines all of them. The harness that replays cases and the judge are covered in the evaluation framework and the LLM-as-judge scorer.
What a scorer can get wrong
A scorer turns an answer into pass or fail, so it can be wrong in two ways. A false fail rejects a correct answer: the team loses trust in the suite, starts ignoring red builds, and eventually raises thresholds until real regressions slip through. A false pass accepts a wrong answer: a regression ships with a green build. Exact match produces many false fails and almost no false passes. Similarity scorers trade some of each depending on the threshold. The question for each field is which error is cheaper. For a refund amount, a false pass is a customer told the wrong number; prefer strict. For a one-sentence explanation, a false fail on rephrasing is noise; prefer tolerant.
That decision belongs to the case author, not to the scorer, so make it explicit in the eval case: each expected field names its strategy. The guide to writing eval test cases shows how to record that in a case file.
Exact match and the normalization ladder
Exact match becomes far more useful once you normalize both sides with the same pipeline. Each rung removes one class of difference that does not change meaning, and you stop at the rung where the remaining differences start to matter.
| Rung | Transformation | Safe for |
|---|---|---|
| 1 | Unicode NFKC, trim, collapse whitespace | almost everything |
| 2 | Case fold | prose, labels; not case-sensitive ids |
| 3 | Strip trailing and surrounding punctuation | short answers, labels |
| 4 | Canonicalize numbers and dates | amounts, dates, quantities |
| 5 | Sort or set-compare list items | unordered lists only |
static String normalize(String s, int rung) {
String out = Normalizer.normalize(s, Normalizer.Form.NFKC).strip().replaceAll("\\s+", " ");
if (rung >= 2) out = out.toLowerCase(Locale.ROOT);
if (rung >= 3) out = out.replaceAll("^[\\p{Punct}\\s]+|[\\p{Punct}\\s]+$", "");
if (rung >= 4) out = Canonical.numbersAndDates(out); // "12 March 2026" -> "2026-03-12"
return out;
}
static boolean exactAt(String expected, String actual, int rung) {
return normalize(expected, rung).equals(normalize(actual, rung));
}With the reference Order A-1042 ships on 12 March 2026., the candidate order a-1042 ships on 12 March 2026 fails raw comparison but matches at rung 3. The candidate Order A-1042 ships on 2026-03-12. needs rung 4, and only if your canonicalizer parses both date forms. Rung 4 is the one to build carefully: parse with an explicit locale, because 03/04/2026 means different days in different countries, and keep the canonicalizer's tests next to the scorer's.
The other lever is to accept more than one reference. Many questions have a small number of correct answers rather than one: a yes/no question may be answered as yes, correct or that is right. Store a list of acceptable references per case and pass when the candidate matches any of them at the chosen rung. Grow the list from reviewed false fails: when a person confirms a rejected answer was correct, add its normalized form. Over a few weeks most of the false fails on short answers disappear without ever loosening the comparison, and every addition is a reviewed, auditable decision rather than a threshold nudge. Cap the list and review it, though; a reference list that keeps growing usually means the case is really asking for a free-text answer and belongs to a softer strategy.
Structured output: score fields, not strings
Agents increasingly return structured output, or a final answer you can parse into fields. Score the fields, not the serialized string. A JSON object with the same content but different key order or spacing is identical, and comparing strings calls it different. More importantly, per-field scoring lets each field use the strategy that fits it.
| Field kind | Strategy | Reason |
|---|---|---|
| Ids, codes, enums | exact, rung 1 | one character changes the referent |
| Amounts, counts | numeric equality with a stated tolerance | 0.1 + 0.2 and currency rounding |
| Dates and times | parse, compare instants or days | many valid spellings |
| Sets of items | set equality or Jaccard | order is not meaning |
| Free-text explanation | semantic or judge | phrasing varies legitimately |
Combine field results with an explicit rule: every hard field must pass, and soft fields contribute a weighted score. A failing hard field ends the case as a fail before any similarity is computed, which also saves embedding calls.
Token overlap and where it fails
Token overlap, such as the unigram F1 on the framework page, sits between exact and semantic. It is cheap and deterministic, and it is blind to the words that carry most of the meaning in agent answers. Three pairs scored with lowercase alphanumeric tokens show the problem:
| Reference | Candidate | Unigram F1 | Correct? |
|---|---|---|---|
| The refund was issued to the original card. | Refund issued to your original card. | 0.714 | yes |
| The refund was issued to the original card. | The refund was not issued to the original card. | 0.941 | no |
| Your order ships on 12 March. | Your order ships on 21 March. | 0.833 | no |
The correct paraphrase scores lowest. The negated answer scores highest, because adding one word costs little precision and no recall. The swapped date scores 0.833 because one token of six differs. No threshold separates these three correctly: anything that passes the first passes the other two. Token overlap is a reasonable signal for drift across many cases, and a poor gate for any single answer where negation or numbers decide correctness.
Semantic similarity with embeddings
Embedding similarity maps each text to a vector and compares the vectors, usually by cosine. It handles paraphrase much better than token overlap. It does not reliably fix negation or numbers, because a sentence and its negation, or two dates in the same template, often sit close together in embedding space. Treat that as a property to measure on your own data, not as a fixed fact about any one model.
Keep the embedding client behind a small interface so the scorer can be tested without network calls and so you can swap models. Wire it to your embedding provider; on Google Cloud that is a Gemini embedding model through the GenAI SDK.
public interface Embedder {
float[] embed(String text); // wire to your embedding client
}
public final class CosineScorer {
private final Embedder embedder;
private final Map<String, float[]> cache = new ConcurrentHashMap<>(); // references repeat
public CosineScorer(Embedder embedder) { this.embedder = embedder; }
public double score(String expected, String actual) {
float[] a = cache.computeIfAbsent(expected, embedder::embed);
float[] b = embedder.embed(actual);
double dot = 0, na = 0, nb = 0;
for (int i = 0; i < a.length; i++) {
dot += a[i] * b[i]; na += a[i] * a[i]; nb += b[i] * b[i];
}
return dot / (Math.sqrt(na) * Math.sqrt(nb));
}
}Pin the embedding model's name and version in the eval report. A cosine of 0.85 under one model means nothing under another, so a model upgrade invalidates every threshold and must be followed by recalibration.
Calibrating a threshold on labelled pairs
A threshold is a claim that answers above it are acceptable. Test the claim. Take pairs from real agent output, have a person label each as acceptable or not, score them, and read precision and recall at candidate thresholds. The table uses illustrative cosine scores for 20 labelled pairs, 12 acceptable and 8 not; the precision and recall are computed from those scores.
| Threshold | True passes | False passes | False fails | Precision | Recall |
|---|---|---|---|---|---|
| 0.75 | 12 | 6 | 0 | 0.667 | 1.000 |
| 0.80 | 10 | 5 | 2 | 0.667 | 0.833 |
| 0.85 | 7 | 3 | 5 | 0.700 | 0.583 |
| 0.88 | 5 | 1 | 7 | 0.833 | 0.417 |
Read it as a decision, not a search for the best number. To accept every good answer you also accept six of eight bad ones. To get precision above 0.8 you reject seven good answers. When the curves look like this, the fix is not a cleverer threshold. Look at the bad pairs with high scores: if they are negations and wrong numbers, as they usually are, add field rules that catch them before similarity is consulted, and send the remaining gray zone to a judge. Keep the labelled pairs as a versioned set, as ground-truth management describes, and rerun the calibration when the model, the agent or the data changes.
A scoring cascade
The cascade in the diagram puts these pieces in order of cost. Exact and normalized checks are free and decisive when they pass. Field rules are free and decisive when they fail. Only what survives reaches the embedding call, and only the band between two calibrated thresholds reaches the judge.
public Verdict score(ExpectedAnswer exp, String actual) {
if (exactAt(exp.text(), actual, exp.rung())) return Verdict.pass("normalized-exact");
for (FieldRule rule : exp.hardFields()) { // numbers, ids, dates, negation
if (!rule.holds(actual)) return Verdict.fail("field:" + rule.name());
}
double cos = cosine.score(exp.text(), actual);
if (cos >= exp.passAbove()) return Verdict.pass("semantic", cos);
if (cos < exp.failBelow()) return Verdict.fail("semantic", cos);
return judge.decide(exp, actual).withScore(cos); // gray zone only
}Every verdict records which stage decided it. That turns the report into a diagnostic: a jump in field:amount failures after a prompt change is a very different story from a drift in semantic scores.
Failure modes
- Thresholds copied from elsewhere. A threshold from a paper or another framework means nothing for your model and data. Calibrate.
- Normalizing away meaning. Case-folding an id or sorting an ordered list turns a real difference into a pass. Choose the rung per field.
- Similarity on numbers. Dates, amounts and counts compared by overlap or embeddings. Parse them and compare exactly.
- Silent model change. The embedding model is upgraded and every score shifts. Pin and report the model; recalibrate on change.
- Averages hiding failures. A suite average of 0.86 hides five critical cases at zero. Gate on per-case verdicts, then aggregate as dataset-driven evaluation does.
- Reference rot. The expected answer was right last quarter. Expire and re-verify references that encode facts about the world.
Trade-offs
Strict scoring is cheap, reproducible and trusted, and it needs many references or careful normalization to avoid false fails. Semantic scoring tolerates phrasing and adds a model dependency, a network call per answer, and a threshold that has to be maintained. Judges handle nuance at the highest cost and with their own biases. The cascade gets most of the benefit of each by spending the expensive stages only where the cheap ones cannot decide.
What to do next
- For every expected field in your cases, write down whether a false pass or a false fail is worse, and pick the strategy from the field table.
- Implement the normalization ladder with a rung per field and unit-test the date and number canonicalizer.
- Add hard field rules for amounts, ids, dates and negation before any similarity scoring.
- Label 50 to 100 real answer pairs, compute precision and recall at several thresholds, and set pass and fail bands.
- Put the cascade in your harness and record the deciding stage on every verdict.
- Pin the embedding model and recalibrate whenever it, the agent or the data changes.