A language model that is right 85% of the time is hard to ship if the other 15% of answers look just as confident. A verifier is a separate component whose only job is to judge a candidate answer before anyone relies on it. It changes the question from "can the model always produce a correct answer?" to "can we tell correct answers from wrong ones?". For many tasks the second is much easier: checking a SQL query by running it, or checking a proof step, is simpler than writing it.

This article treats the verifier as architecture rather than a prompt trick. It covers the generator/verifier split, the kinds of verifier and when each is trustworthy, how to choose an accept threshold from data, how to spend extra compute with best-of-N and reject-and-retry, what happens when every attempt fails, and how to reason about cost. Chain-of-Verification covers the related technique of breaking one answer into separately checked claims; here the unit is the whole candidate.

Advertisement

Why separate checking from generating

Verifier architecture: generate, score, gate on a calibrated threshold, retry with feedback, fall backrequesttask + contextgenerator1 or N candidatescandidatesverifier stackcheap checks first, thenmodel grader / reward modelscoregatescore >= t ?yesacceptreturn + logno: feedbackretry budget left?k attempts, token capyesnofallbackstronger model, abstain, humanThe threshold t is chosen on a labelled dev set for a target precision, and rechecked whenever the generator or verifier changes.Every decision (scores, attempts, outcome) is logged so the verifier itself can be evaluated.
Candidates flow through a stack of checks to a calibrated gate. Rejections go back with feedback until the budget runs out, then a fallback decides.

Asking a model in the same prompt to "double-check your answer" rarely helps much, because the same weights and the same context that produced the error are grading it. A separate verifier breaks that correlation in three ways. It can use a different signal (execution, a schema, a retrieval result), a different model, or a different view of the input (only the answer and the source, without the generator's reasoning). Each of these makes its errors less correlated with the generator's, and uncorrelated errors are what make filtering work.

The split also gives you an engineering seam. The generator can be tuned for coverage and creativity, the verifier for precision, and each can be evaluated, versioned and replaced on its own. Research on math word problems and on reward models for reasoning has repeatedly found that sampling several solutions and letting a trained verifier choose beats trusting a single sample. That is the architectural idea this page builds on.

Kinds of verifier

VerifierSignalStrengthWeakness
Programmaticrun code, parse schema, execute SQL, unit tests, regex, arithmeticexact, cheap, not fooled by fluent proseonly checks what you can write rules for
Reference-basedcompare with retrieved sources or a known answercatches unsupported factsonly as good as retrieval
Model grader (LLM judge)a prompted model scores the candidate against a rubricflexible, handles open-ended tasksbiases toward length and style; must be calibrated
Outcome reward modeltrained model scores the final answerstrong ranking signal for a fixed taskneeds labelled data; drifts with the generator
Process reward modeltrained model scores each reasoning steplocalises the first bad step, gives better feedbackneeds step-level labels; more expensive

Stack them from cheapest and most certain to most expensive. A JSON answer that does not parse should never reach an LLM judge. A good verifier stack is ordered like a firewall: deterministic rejects first, then reference checks, then model-based scoring only on what survives.

Advertisement

Designing a model grader

Model graders are the most common verifier and the easiest to get wrong. Four rules keep them honest. Grade against a rubric, not a vibe: list the specific properties that make an answer acceptable and ask for a verdict on each. Hide the generator's reasoning unless the rubric is about the reasoning, because persuasive reasoning biases the judge. Ask for a structured output with a score and short reasons, so the reasons can be fed back to the generator. Use a different model or at least a different prompt family from the generator.

GRADER_PROMPT = """You are checking an answer. Do not solve the task yourself first.
Task: {task}
Source material: {sources}
Candidate answer: {answer}

For each criterion, answer pass or fail with one sentence of evidence:
1. Every factual claim is supported by the source material.
2. The answer addresses every part of the task.
3. The output follows the required format: {format_spec}

Return JSON: {{"criteria": [{{"id": 1, "verdict": "pass|fail", "evidence": "..."}}, ...],
               "p_correct": <probability between 0 and 1 that the answer is fully acceptable>}}"""

The p_correct field is a raw score, not a probability you can trust. Treat it as a number to be calibrated in the next step. Programmatic checks and reward models also produce scores; the same calibration applies to all of them.

Choosing the threshold from data

A verifier is a classifier, and its threshold sets the trade-off between accepting wrong answers (false accepts) and throwing away right ones (false rejects). Pick it from data, not intuition. Collect a dev set of a few hundred real candidates, label each as correct or not, score them with the verifier, and choose the lowest threshold whose accepted set meets your precision target:

def pick_threshold(scores, labels, target_precision=0.97, min_accept=0.3):
    """Lowest threshold whose accepted set reaches target precision."""
    pairs = sorted(zip(scores, labels), reverse=True)
    best, correct = None, 0
    for i, (s, ok) in enumerate(pairs, start=1):
        correct += ok
        precision, accept_rate = correct / i, i / len(pairs)
        if precision >= target_precision:
            best = (s, precision, accept_rate)
    if best is None or best[2] < min_accept:
        raise ValueError("verifier cannot reach the target at a useful accept rate")
    return best          # (threshold, precision, accept_rate)

Report the precision with a confidence interval, because a few hundred labels leave real uncertainty, and use a held-out split to confirm the chosen threshold. Recalibrate whenever the generator, the verifier, the prompt or the input mix changes: a new generator produces different errors, and a threshold fitted to the old one can quietly let the new errors through. If you need the score to mean a probability, for example to combine it with other signals, fit a calibration map such as isotonic regression on the same labelled data. The broader practice of building these labelled sets is covered in prompt evaluation.

Spending compute: best-of-N and reject-and-retry

Once you can score candidates, you can buy accuracy with more generation. There are two basic patterns, and they suit different latency budgets.

  • Best-of-N. Sample N candidates in parallel, score all of them, and return the best one if it clears the threshold. Latency is one generation plus one verification round; cost is N of each. It works best when the verifier ranks well and the generator is diverse enough at non-zero temperature.
  • Reject-and-retry. Generate one candidate, verify it, and if it fails, generate again with the verifier's reasons added to the prompt. Cost is proportional to the attempts actually used, and the feedback gives each new attempt information that independent sampling lacks. Latency grows with every retry.
def answer(task, gen, verifiers, threshold, max_attempts=3, fallback=None):
    feedback, log = [], []
    for attempt in range(1, max_attempts + 1):
        cand = gen(task, feedback=feedback)
        verdict = None
        for v in verifiers:                          # cheapest first
            verdict = v(task, cand)
            if verdict.hard_fail:                     # e.g. unparsable, test failed
                break
        log.append((attempt, verdict.score, verdict.reasons))
        if not verdict.hard_fail and verdict.score >= threshold:
            return Result(cand, status="accepted", attempts=attempt, log=log)
        feedback.append(verdict.reasons)             # concrete, specific, short
    if fallback:
        return fallback(task, log)                   # stronger model, abstain or human queue
    return Result(None, status="abstained", attempts=max_attempts, log=log)

Feedback must be specific. "Criterion 1 failed: the revenue figure 4.2M does not appear in the source" helps; "please try again" mostly produces a reworded copy of the same error. Cap attempts at a small number, two to four in most systems, because the chance of success usually falls on each retry: the inputs that fail twice tend to be genuinely hard or ambiguous.

Worked example: text-to-SQL with a two-stage verifier

A support tool turns questions into SQL over an orders database. The stack has two stages. First, a programmatic check parses the query, rejects anything other than a single SELECT, runs it with EXPLAIN and then with a row limit against a read replica, and fails on an error or an empty result where rows are expected. Second, a model grader sees the question, the schema, the SQL and the first twenty result rows, and scores whether the result answers the question.

Suppose (illustrative numbers) the generator's first attempt is acceptable 78% of the time. The execution check rejects a third of the bad ones outright with exact error messages, which make excellent retry feedback. On a labelled set of 400 questions, a grader threshold of 0.8 gives 97% precision on accepted answers at a 75% first-pass accept rate. With up to three attempts and a fallback that says "I could not answer this reliably" and shows the closest query, the tool answers a little over nine questions in ten (1 - 0.25 x 0.5 x 0.5, assuming half of retries pass), and the answers it gives are right about 97% of the time. The remaining questions go to an analyst queue, where their labels feed the next calibration round.

The expected cost per request is easy to write down. With first-attempt pass rate a, retry pass rate r and generation plus verification cost g per attempt, three attempts cost g times (1 + (1 - a) + (1 - a)(1 - r)). With a = 0.75 and r = 0.5 (again illustrative) that is about 1.38g: a 38% cost increase to move from unverified answers to measured precision, before the fallback's cost.

Fallbacks and abstention

Every verifier loop needs a defined end. The options, in rising cost: abstain and say so plainly; return a hedged partial answer with the failing checks shown; escalate to a stronger model, the same idea as a routing cascade in LLM model routing; or send to a human. Choose per product. A customer-facing assistant should usually abstain or escalate; a batch pipeline can queue for review.

Abstention is a feature, not a failure, and it must be measured as one. Track the abstention rate next to precision. A verifier that reaches 99% precision by abstaining on half the traffic may be worse for users than one at 95% that answers nearly everything.

Failure modes

  • Correlated errors. The same model generates and grades, and both share the same misconception. Use a different signal or model, and check agreement on the labelled set.
  • Reward hacking. Under best-of-N with large N, the generator's most-selected outputs are those that exploit the verifier's blind spots, such as long, confident, citation-heavy answers. Keep N moderate and audit what the top-scored answers have in common.
  • Stale thresholds. A threshold fitted before a model upgrade silently changes precision. Pin model versions and recalibrate on every change.
  • Retry loops that never converge. Vague feedback produces rephrasings. Make feedback concrete and cap attempts.
  • Verifier cost overtaking generation. An expensive grader on every candidate of a large N. Put cheap programmatic checks first and grade only survivors.
  • Checks that do not match the harm. A verifier that checks format while the real risk is a wrong fact. Write the rubric from actual failures; hallucination guardrails lists the grounding checks worth having.

What to do next

  1. Collect 200 to 500 real outputs from your system and label each as acceptable or not; this set is the foundation for everything else.
  2. Write every deterministic check you can (parsing, schema, execution, tests) and measure how many bad outputs they catch alone.
  3. Add a rubric-based model grader, using a different model or prompt family from the generator, that returns structured verdicts and reasons.
  4. Pick the threshold with the precision-target procedure above, confirm it on a held-out split, and record the model versions it was fitted for.
  5. Wrap generation in a reject-and-retry loop with at most three attempts, specific feedback and a defined fallback, and log every score and decision.
  6. Dashboard precision on a sampled audit, accept rate, abstention rate, attempts per request and cost per request, and recalibrate whenever any component changes.
Key takeaway: A verifier separates checking from generating, which works because judging an answer is often easier than producing it and because a different signal or model makes errors less correlated. Stack verifiers from cheap and exact (parsing, execution, tests) to flexible and expensive (rubric graders, reward models). Treat every score as uncalibrated: choose the accept threshold on a labelled set for a target precision and refit it whenever anything changes. Spend extra compute through best-of-N or reject-and-retry with specific feedback, cap the attempts, give the loop a defined fallback, and measure precision, abstention and cost together.