Human review of LLM output is usually described as one thing, "a human checks it". In practice it is several different patterns. They sit at different points on the request path, cost different amounts of GPU time and latency, and measure different things. A pre-delivery gate stops bad output from reaching users but makes you buffer the whole generation. A sampled post-delivery audit adds no latency but only tells you the defect rate after the fact. Shadow mode lets you evaluate a model on real traffic before it is trusted with any of it.

This article is a catalogue of those patterns: what each one is for, what it does to the serving system, and how to measure whether it works. It covers confidence intervals for audit samples, stratified estimates, reviewer agreement and gold items. The companion piece LLM human-in-the-loop, in depth covers confidence signals, routing and sizing the review team; this one assumes you have a router and asks which review topology to put behind it.

The pattern catalogue

Review patterns placed on the request path: where the human sits and what it costsRequestprompt + contextGPU generationprefill + decodeJudge pre-screenbatched, short outputriskyPre-delivery gatehold, approve, editcleanDelivered to userstreamed only if no pre-delivery check appliesPost-delivery auditstratified sampleDual review + gold itemsagreement, reviewer accuracyMetricsdefect rate, kappareject: regenerate (prefix cache may save the prompt prefill)Shadow mode (pre-launch)model output never delivered; humans do the task; compareOnly the gate adds user-visible latency; everything below the delivery line is offline cost.
Where each pattern attaches. The pre-delivery gate is the only one on the user's critical path; audits, dual review and shadow mode run off it.
PatternPurposeUser latencyGPU costWhat it measures
Pre-delivery gateStop harmful outputQueue wait + review timeRegeneration on rejectNothing by itself; blocks
Edit-then-releaseFix drafts before sendingAs gateLow; edits replace regenerationEdit distance, edit rate
Post-delivery auditEstimate defect rateNoneNone beyond loggingDefect rate with interval
Judge pre-screenCut human volumeFull generation plus judge call if in path; none if run on a sent streamOne extra prefill-heavy callJudge precision and recall
Dual review + adjudicationReliable labelsNone (offline)NoneInter-rater agreement
Gold itemsCheck the reviewersNoneNoneReviewer accuracy
Shadow modeEvaluate before launchNone; output unusedFull generation, discardedAgreement with humans

Pre-delivery gate and edit-then-release

A pre-delivery gate holds a response until a human approves it. It is the right choice when one bad output is expensive and irreversible: a message to a regulator, a medical instruction, an action with financial consequences. The serving system changes in three ways.

No streaming. Token streaming exists to cut perceived latency, but you cannot stream text a human has not seen. The application must buffer the full response, so time to first visible token becomes generation time plus queue wait plus review time. That is seconds of GPU time and often minutes of human time. Design the product for it with an explicit "pending review" state. Do not stream to a hidden buffer and hope the user waits.

Free the GPU early. Once generation finishes, its KV cache can be released; the held response is just text in a database. Do not keep a sequence resident on the GPU while a human reads it. An engine slot held for minutes is capacity taken from everyone else.

Reject means regenerate. A rejection either sends a reviewer-written answer or triggers a new generation, often with the reviewer's note appended to the prompt. Because the original prompt is unchanged up to the note, an engine with prefix caching, such as vLLM's automatic prefix caching, can reuse the prompt prefix and pay only for the new suffix and the decode, but only if those blocks are still cached. After minutes of review on a busy server they have usually been evicted, so budget for a full prefill. Keep the note at the end of the prompt so the shared prefix stays as long as possible.

Edit-then-release is a gate in which the reviewer fixes the draft rather than rejecting it. It avoids regeneration entirely and produces the most valuable training signal you can collect: a model output paired with a human correction. Store both, plus a character-level edit distance, so you can track how much reviewers change and which kinds of request need the most editing.

Judge pre-screens and the recall trap

Gating every response does not scale, so most systems put a cheaper check in front: a classifier or LLM judge decides which responses need a human. Judge calls are dominated by prefill, since they read the full response and emit a short verdict. They batch well and can share a GPU pool with guard models; LLM guardrails covers the placement and latency budget for exactly this kind of call.

A screen in the request path has the same streaming cost as a gate: the judge reads the whole response, so everything it screens must be buffered until it finishes. The alternative is to stream immediately, screen in parallel and retract or replace the message on a fail. Users may briefly see withdrawn output, which suits mild quality issues, not serious harms.

The screen's job is recall on bad outputs. A response it misses never reaches a reviewer. Measure that recall on the audit sample described next, not on the items the judge flagged, because the flagged set by construction contains only what the judge already caught. Many teams report the judge's precision on its flagged set, see high numbers and never learn what it misses.

Post-delivery audit: intervals and stratification

A post-delivery audit samples responses that have already been delivered and has humans grade them. It adds no latency, and it is the only pattern that produces an unbiased estimate of the defect rate users actually experience. That makes it mandatory alongside any gate or screen. The statistics are simple but often done wrong.

Report an interval, not a point. With 6 defects in 400 sampled responses the point estimate is 1.5 percent, but the 95 percent Wilson interval runs from 0.69 to 3.23 percent. The true rate could be twice the point estimate. Use the Wilson interval rather than the textbook normal approximation, which misbehaves at the low rates and small counts typical of audits. With zero defects in 300 samples the upper bound is 1.26 percent, close to the "rule of three" approximation 3/300 = 1 percent. Zero defects does not mean a zero rate.

Size the sample from the precision you need. To estimate a rate near 2 percent within plus or minus 1 point at 95 percent confidence you need about 1.962 × 0.02 × 0.98 / 0.012, which is 753 graded responses. Halving the margin quadruples the sample, which is why audits tend to settle on a weekly cadence: daily samples are too small to say anything.

import math, random

def wilson(defects, n, z=1.96):
    if n == 0:
        return (0.0, 1.0)
    ph = defects / n
    den = 1 + z * z / n
    centre = (ph + z * z / (2 * n)) / den
    half = z * math.sqrt(ph * (1 - ph) / n + z * z / (4 * n * n)) / den
    return max(0.0, centre - half), min(1.0, centre + half)

def stratified_sample(log, per_stratum, key, seed):
    """log: delivered responses; key: function giving each one's stratum."""
    rng = random.Random(seed)            # fixed seed: the sample is reproducible
    strata = {}
    for item in log:
        strata.setdefault(key(item), []).append(item)
    sample = {s: rng.sample(items, min(per_stratum, len(items)))
              for s, items in strata.items()}
    sizes = {s: len(items) for s, items in strata.items()}
    return sample, sizes

def stratified_rate(sizes, graded):
    """graded[s] = (defects, n) per stratum; weights are population shares."""
    total = sum(sizes.values())
    est = var = 0.0
    for s, (x, n) in graded.items():
        w, ps = sizes[s] / total, x / n
        est += w * ps
        var += w * w * ps * (1 - ps) / n
    half = 1.96 * math.sqrt(var)
    return est, max(0.0, est - half), est + half

Stratify by risk so the rare, dangerous slice gets enough samples, then weight the results back. Suppose a day has 5,000 high-risk and 45,000 low-risk responses. You grade 200 of each and find 8 and 2 defects. Pooling the 400 naively gives 2.5 percent, which is badly inflated because high-risk items are 10 percent of traffic but half the sample. Weighting by population share gives 0.1 × 4% + 0.9 × 1% = 1.3 percent, with a 95 percent interval of roughly 0.03 to 2.57 percent. Report both the overall rate and the high-risk stratum's own 4 percent; they answer different questions. For tracking these rates over time with change detection, see AI risk monitoring.

Checking the reviewers: agreement and gold items

Review results are only as reliable as the reviewers, and reviewers disagree more than anyone expects. Two patterns measure this.

Dual review with adjudication. Route a slice of items, typically 5 to 10 percent, to two reviewers independently, and send disagreements to a senior adjudicator. Measure agreement with Cohen's kappa, which corrects for agreement you would get by chance. Raw agreement flatters reviewers when most items pass. In one worked case, two reviewers grade 200 items: both pass 150, both fail 20, and they split on 30. Raw agreement is 85 percent, which sounds fine. But reviewer A passes 84 percent of items and B passes 81 percent, so chance alone predicts 71 percent agreement, and kappa is (0.85 - 0.71) / (1 - 0.71), about 0.48. That is moderate agreement at best, and it usually means the rubric is ambiguous rather than the reviewers being careless. Rewrite the rubric around the disagreeing items and measure again.

def cohen_kappa(both_pass, a_pass_b_fail, a_fail_b_pass, both_fail):
    n = both_pass + a_pass_b_fail + a_fail_b_pass + both_fail
    observed = (both_pass + both_fail) / n
    pa = (both_pass + a_pass_b_fail) / n          # A's pass rate
    pb = (both_pass + a_fail_b_pass) / n          # B's pass rate
    chance = pa * pb + (1 - pa) * (1 - pb)
    return (observed - chance) / (1 - chance)

print(round(cohen_kappa(150, 18, 12, 20), 2))     # 0.48

Gold items. Mix items with known, adjudicated answers into each reviewer's queue without marking them, at around 2 to 5 percent of volume. Accuracy on gold items is a per-reviewer quality measure that does not depend on another reviewer. Reviewers below threshold get retraining, and their recent decisions get re-reviewed. Rotate gold items regularly; once reviewers recognise them they stop measuring anything.

Shadow mode before launch

Shadow mode runs the model on live traffic while humans keep doing the task, and the model's output is logged but never delivered. It is the safest way to evaluate a new model or a new automation before any user sees it, and it produces paired data: the human's decision and the model's on the same input.

The GPU cost is the full generation for every shadowed request, with no product value until launch. Two things keep it affordable. Shadow traffic has no latency target, so run it at low priority or in batch on spare capacity, using the offline throughput techniques from LLM data labeling. And shadow a sample, not everything: a few thousand stratified requests are usually enough to estimate agreement with the interval methods above.

Exit shadow mode on a pre-registered criterion, written down before you look at the data, for example "model-human agreement on the high-risk stratum, measured on 750 items, has a lower Wilson bound above 95 percent". Then move to a gate, relax to a screen plus audit as evidence accumulates, and keep the audit permanently.

Failure modes

  • Rubber-stamp approval. Reviewers approve nearly everything because nearly everything is fine. Watch approval rate and median review time per reviewer; gold items with known defects reveal it directly.
  • Measuring the screen on its own output. Judge precision looks great and recall is never measured. Only an independent audit sample measures misses.
  • Pooled audit rates. Unweighted stratified samples overstate or understate the defect rate depending on how you oversampled.
  • Holding GPU state during review. A sequence kept resident while a human reads it starves the batch. Persist text and free the slot.
  • Gate queue as an outage. When reviewers are unavailable the gate becomes a full stop. Decide the timeout behaviour in advance: fail closed with an apology, or fall back to a safer canned answer. Never silently release.
  • Unversioned rubrics. Defect rates jump when the rubric changes, not the model. Stamp every grade with the rubric version.

Trade-offs

The central trade-off is latency against assurance. A gate gives the strongest assurance and the worst latency; an audit gives none of the protection but all of the measurement. Most production systems combine them: a gate only for the riskiest class, a screen in front of it, and an audit over everything. The audit's job is to tell you whether the gate's routing is drawing the line in the right place. For actions rather than text, where a wrong approval has direct effects, the approval design in human-in-the-loop for high-risk actions applies on top of everything here.

Review spend also competes with model spend: a larger model that halves the defect rate may cost less than the review hours it saves. Only audit-quality defect rates for both options can tell you.

What to do next

  1. Draw your request path and mark where each review pattern attaches, and which ones add user-visible latency.
  2. Start a stratified post-delivery audit with a fixed seed and Wilson intervals, sized for the margin you need.
  3. If you run a judge pre-screen, measure its recall on the audit sample, not its precision on its own flags.
  4. For any gate, disable streaming for gated classes, persist held responses outside the GPU, and put reviewer notes at the end of the regeneration prompt.
  5. Route 5 to 10 percent of items to dual review, compute kappa, and fix the rubric if it is below about 0.6.
  6. Seed 2 to 5 percent gold items into every reviewer's queue and rotate them monthly.
  7. Run new models in shadow mode on a stratified sample, with an exit criterion written before the data comes in.
Key takeaway: Human review is a set of patterns, not one. A pre-delivery gate blocks harm at the cost of streaming and regeneration. A judge pre-screen cuts volume, but only an independent audit can measure what it misses. A stratified post-delivery audit with Wilson intervals is the one unbiased measure of what users experience. Dual review, kappa and gold items check the reviewers, and shadow mode evaluates a model on real traffic before it is trusted. Combine them, free GPU state while humans read, and keep the audit permanently.