Every team that ships an LLM feature reaches the same moment: someone asks whether the outputs are good, and nobody can answer with a number they trust. The usual reaction is to adopt a ready-made metric suite, watch a dashboard of helpfulness and coherence scores, and discover months later that the scores never moved when users were unhappy. The problem is not a lack of metrics. It is that evaluation was designed before anyone looked at what the model actually got wrong, and that the evaluators themselves were never checked.
This article builds output evaluation from first principles: what a single verdict should judge, how error analysis turns raw traces into criteria worth measuring, how to pick the cheapest evaluator that can decide each criterion, how to prove an evaluator agrees with experts, and how to handle sampling noise and small test sets. A worked example follows a support-ticket summariser through the whole loop, and the article ends with a checklist.
What a verdict should judge
An evaluation is a function from an example to a verdict. The example is more than the output text: it is the input, any retrieved context or tool results the model saw, the output itself, and sometimes a reference answer. The verdict should answer one question about one property. 'Is this summary good?' is not a question any evaluator can answer reliably; 'does the summary state the action the customer asked for?' is. Each property you care about becomes its own criterion with its own evaluator, and the quality of a release becomes a small table of criterion pass rates rather than one blended score.
Prefer binary verdicts to 1-to-5 scales. Annotators and model judges disagree about what separates a 3 from a 4, and averaging ordinal scores hides whether the cases that mattered improved. A pass or fail with a written reason is easier to agree on, audit and act on. When you need gradation, split it into several binary criteria (mentions the order number, gives the refund amount) instead of asking for a number.
Start from error analysis, not a metric list
The most common mistake is choosing metrics before reading outputs. Generic criteria are cheap to adopt and rarely track the failures your users actually hit. Instead, collect around 100 real traces, from production logs if you have them or from a synthetic set that covers your main user intents if you do not, and read them.
For each trace, write a short free-text note on the first thing that went wrong, if anything. This is open coding: you are noticing, not filling in a form. After a pass through the set, group the notes into categories, which is axial coding: 'invented an order number', 'ignored the stated date range', 'answered in the wrong language', 'refused a legitimate request'. Count each category. That table is your first real evaluation result, and it tells you where evaluators are worth building. A category with two occurrences and an obvious prompt fix needs the fix, not an evaluator. A category that keeps recurring and is costly when it does deserves an automated check.
The evaluator ladder
Once a criterion is defined, choose the cheapest evaluator that can decide it reliably. The options form a ladder from deterministic and cheap to judgement-based and expensive.
| Evaluator | Decides | Cost | Typical weakness |
|---|---|---|---|
| Exact or normalised match | Labels, extracted IDs, short factual answers | Microseconds | Fails on valid paraphrase; normalise case, whitespace and number formats first |
| Structural validation | JSON schema, required fields, enum values, length limits | Microseconds | Says nothing about whether the content is right |
| Executable checks | Generated code, SQL, tool calls: run tests or execute against a fixture | Seconds | Needs a sandbox; weak tests pass wrong answers |
| Content assertions | Required facts, banned phrases, citation format | Microseconds | Brittle when over-specified |
| Reference overlap (BLEU, ROUGE) | Closeness to a reference text | Milliseconds | Rewards shared words, not correctness |
| Embedding similarity | Topical closeness | Milliseconds | Nearly blind to negation and swapped numbers |
| LLM judge with a rubric | Semantic or subjective criteria | One model call | Biased and noisy until validated |
| Human review | Anything, including new failure types | Minutes per item | Slow, costly, and humans disagree too |
Climb only as high as you must. Many real failures (wrong format, missing field, an identifier not in the source, code that does not compile) are decided at the bottom of the ladder, cheaply enough to check every production output. Be wary of the overlap and similarity rungs: 'the refund was approved' and 'the refund was not approved' share almost every token and sit close together in embedding space, so neither metric is a verdict on correctness.
A scorer stack in code
A practical pattern runs deterministic checks first and short-circuits on hard failures, so a model judge only sees outputs that are structurally valid. Each scorer returns a verdict and a reason, and the reasons are what you read in the next round of error analysis.
import json
from dataclasses import dataclass
@dataclass
class Verdict:
criterion: str
passed: bool
reason: str
def check_schema(out):
try:
data = json.loads(out)
except json.JSONDecodeError as e:
return Verdict("valid_json", False, f"parse error: {e}"), None
missing = {"summary", "action", "order_ids"} - data.keys()
if missing:
return Verdict("valid_json", False, f"missing {sorted(missing)}"), None
return Verdict("valid_json", True, "ok"), data
def check_grounded_ids(data, source_text):
invented = [i for i in data["order_ids"] if i not in source_text]
return Verdict("ids_grounded", not invented,
f"not in source: {invented}" if invented else "ok")
def evaluate(example, output, judge=None):
first, data = check_schema(output)
verdicts = [first]
if data is None:
return verdicts # nothing else is meaningful
verdicts.append(check_grounded_ids(data, example["ticket"]))
if judge is not None: # semantic, expensive, last
verdicts.append(judge("action_correct", example, data))
return verdictsThe grounded-ID check deserves attention. Inventing identifiers, dates and amounts is one of the most damaging forms of hallucination in business workflows, and a substring check against the source catches the plain cases without any model call.
Validating an evaluator against experts
Any evaluator involving judgement must itself be evaluated before its numbers mean anything. Have a domain expert label 100 to 200 outputs pass or fail on the criterion, run the evaluator on the same sample, and compare.
Report the comparison as a confusion matrix, not one agreement percentage. Suppose the expert fails 40 of 200 outputs. The judge flags 34 of those 40 and also flags 12 of the 160 the expert passed. Raw agreement is (34 + 148) / 200 = 91 percent, which sounds excellent but is inflated because most outputs pass. The judge's true positive rate on failures is 34 / 40 = 0.85 and its true negative rate is 148 / 160 = 0.925.
Cohen's kappa corrects agreement for chance. The judge fails 23 percent of outputs and the expert 20 percent, so chance agreement is 0.23 x 0.20 + 0.77 x 0.80 = 0.662. Kappa is (0.91 - 0.662) / (1 - 0.662), about 0.73. Values above roughly 0.6 are conventionally read as substantial agreement; below 0.4 the evaluator is not measuring what you think it is.
Known error rates also let you correct the headline number. If the judge reports a 23 percent failure rate on new data, the estimated true rate is (0.23 + 0.925 - 1) / (0.85 + 0.925 - 1) = 0.155 / 0.775 = 20 percent. Keep the labelled set and re-run it whenever the judge prompt or model changes, because judges drift like any other dependency. Judge-specific biases and pairwise formats are covered in LLM-as-judge calibration.
Sampling: one output is not the behaviour
The same input yields different outputs at non-zero temperature, and many serving stacks are not bit-for-bit deterministic even at zero. One sample per case measures a draw, not the behaviour, so for variance-sensitive criteria draw n samples per case.
pass@k, from the Codex paper (Chen et al., 2021), is the probability that at least one of k samples passes; it suits workflows where a user or verifier picks the best attempt. pass^k, introduced with the tau-bench agent benchmark, is the probability that all k attempts pass; it measures reliability, which a customer-facing agent needs. Both have unbiased estimators from n samples with c passes:
from math import comb
def pass_at_k(n, c, k):
# P(at least one of k samples passes), from n samples with c passing
if n - c < k:
return 1.0
return 1.0 - comb(n - c, k) / comb(n, k)
def pass_hat_k(n, c, k):
# P(all k samples pass)
return comb(c, k) / comb(n, k)
print(pass_at_k(10, 3, 1), pass_at_k(10, 3, 5)) # 0.3 0.9167
print(pass_hat_k(10, 8, 3)) # 0.4667The last line is the sobering one: a case that passes 8 times in 10 passes all of three independent attempts less than half the time. If users retry, or an agent chains three dependent calls, per-call accuracy overstates what people experience.
How many cases you need
A pass rate from a few dozen cases is mostly noise, so put a confidence interval on every rate you report. For a proportion, the Wilson score interval behaves well even near 0 or 100 percent. With 172 passes out of 200 cases (86 percent), the 95 percent Wilson interval runs from about 80.5 to 90.1 percent.
from math import sqrt
def wilson(passes, n, z=1.96):
p = passes / n
denom = 1 + z * z / n
centre = (p + z * z / (2 * n)) / denom
half = z * sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / denom
return centre - half, centre + half
print(wilson(172, 200)) # about (0.805, 0.901)A move from 86 to 88 percent on that set is indistinguishable from noise. Detecting it needs many more cases, or a paired comparison on the same cases, which removes between-case variance; the evaluation harness article covers paired comparisons and release gates. Halving an interval's width takes roughly four times the cases, so size the set from the precision you need, and stratify by user intent so that rare but important intents get enough cases for their own rate.
Worked example: a support-ticket summariser
A team generates a JSON summary of each support ticket for human agents: a short summary, the customer's requested action and any order IDs mentioned. They read 120 production traces and code the failures. Eleven outputs contained an order ID not present in the ticket, nine misstated the requested action (usually refund versus replacement), six exceeded the length limit, four were invalid JSON after a long ticket was truncated, and two had tone issues nobody cared about.
They build four evaluators. Valid JSON and length are code. Grounded IDs is a substring check. Action correctness needs semantics, so it gets a model judge with a binary rubric that lists the five allowed actions and gives two examples of each common confusion. An expert labels 150 outputs for action correctness. The first judge prompt reaches a kappa of 0.52, mostly by failing outputs that named the right action in different words. Adding one line saying synonyms count, with an example, raises kappa to 0.78 on a fresh 100-item sample, and they accept it.
They then change the prompt (quote order IDs verbatim; choose the action from a fixed list) and re-run 300 cases. Invented IDs fall from about 9 percent to 1 percent, far outside the noise. Action errors move from 7.5 to 6 percent, inside the interval, so they report no detectable change. Code checks now run on every request; the judge runs on a 5 percent daily sample and is re-validated monthly.
Failure modes
- Metrics before reading. A generic suite stays flat while users complain, and the team concludes evaluation does not work. Read traces first.
- Unvalidated judges. A judge prompt written in an afternoon becomes the release gate. Without labels nobody knows its error rates, and a judge that passes nearly everything makes any model look great.
- Test-set leakage. Evaluation cases end up as few-shot examples in the prompt and scores jump for reasons that will not transfer. Keep a held-out set nobody tunes against.
- Stale datasets. The set reflects last quarter's traffic. Sample fresh production traces regularly and add cases for each new failure category.
- Shared blind spots. A judge from the same model family as the generator tends to favour its own style. Prefer a different family for judging, and lean on code checks where possible.
Trade-offs
Deterministic checks run everywhere and never drift, but they only capture what you can specify precisely. Model judges reach semantic criteria at moderate cost, but bring bias, noise and upkeep: a labelled set, periodic re-validation and version pinning. Human review is the ground truth you calibrate against and the only reliable way to discover new kinds of failure, but it does not scale to every output. A healthy setup uses all three in proportion: code checks on all traffic, judges on samples and test sets, and a few hours of expert reading per release.
Breadth also trades against precision: twenty criteria with fifty cases each tell you less than five with two hundred each. Evaluation cost is part of the quality, cost and latency frontier you choose models on.
What to do next
- Pull 100 recent traces and write a one-line note on the first failure in each.
- Group the notes into categories, count them, and fix the trivial ones in the prompt.
- For each remaining category, write one binary criterion and pick the cheapest evaluator on the ladder that decides it.
- Implement the code checks first and run them on all production outputs.
- For each judge, have an expert label 150 outputs, compute TPR, TNR and kappa, and iterate the rubric until kappa clears your bar.
- Sample several outputs per case for variance-sensitive criteria, and report pass^k where reliability matters.
- Attach a Wilson interval to every rate before it drives a decision; for RAG systems add the metrics in RAG evaluation.
- Schedule a monthly trace-reading session and add a case for every new failure you find.