Language models state false facts with the same fluency as true ones. The errors cluster in predictable places: lists of entities, dates and numbers, and the middle of long answers where one fabricated detail sits among correct ones. Chain-of-Verification, or CoVe, is a prompting architecture that attacks exactly those errors by making the model check its own draft, one claim at a time, before answering.
The method comes from Dhuliawala and colleagues at Meta AI, in the 2023 paper Chain-of-Verification Reduces Hallucination in Large Language Models (arXiv 2309.11495). This article explains why it works, turns it into a pipeline you can run and operate, marks clearly where common production additions go beyond the paper, and covers the cost, failure modes and evaluation you need before shipping it.
Why checking one fact at a time works
The observation behind CoVe is that a model often answers a short, focused question correctly even when it gets the same fact wrong inside a longer generation. Asked to list politicians born in a city, a model may include someone born elsewhere; asked directly where that person was born, it frequently gives the right answer. Long generation mixes many facts into one sampling pass, and an early mistake conditions everything after it. A short question isolates one fact and gives the model a fresh chance at it.
CoVe exploits that gap. It does not add new knowledge; in the paper the same model drafts, plans and verifies, with no retrieval. It restructures inference so that knowledge the model has but did not use in the draft gets a chance to correct the draft. That also marks its limit: facts the model does not know cannot be fixed by asking it again.
The four stages
The paper defines four steps. Baseline response: answer the question normally. Plan verifications: given the question and draft, write a list of verification questions that test the draft's factual claims. Execute verifications: answer each question. Final verified response: produce a revised answer that takes the verification results into account.
In practice each stage is one or more model calls with its own prompt, and the interesting design decisions are about what each call is allowed to see. The planner must see the draft, because it extracts claims from it. The final step must see both the draft and the evidence. The verifiers are the subtle part.
Joint, two-step, factored and factor+revise
The paper compares several ways to run the planning and execution steps, and the comparison is the most useful result for anyone building the pipeline.
| Variant | How verification runs | What goes wrong |
|---|---|---|
| Joint | One prompt plans the questions and answers them in the same generation | Answers are conditioned on the draft sitting in context and tend to repeat its errors |
| Two-step | Planning is one call; all questions are answered in a second call without the draft | Answers can still influence each other inside that one generation |
| Factored | Each question is answered in its own call, with nothing else in context | More calls; no explicit cross-check against the draft |
| Factor+revise | Factored answers, then an extra step compares each answer with the original claim | Most calls, but inconsistencies are flagged explicitly |
The paper reports that factored verification beats joint verification, and that factor+revise does best on long-form generation. The mechanism is the same throughout: every piece of the draft that a verifier can see is a chance for the model to copy a hallucination instead of recalling the fact. The paper also found that open verification questions worked better than yes/no questions, which invite the model to agree with whatever the question asserts. Design the pipeline around those two rules: isolate verifiers, and ask open questions.
A reference implementation
The code below implements factor+revise. The llm argument is any function that takes a prompt string and returns text, so it works with whichever model client you use. A separate verify_llm is optional, for reasons explained in the section on correlated errors.
import json
from concurrent.futures import ThreadPoolExecutor
MAX_QUESTIONS = 8
PLAN = """Below is a question and a draft answer. List the factual claims in the draft
and, for each, one open-ended question whose answer would confirm or refute it.
Do not ask yes/no questions. Return JSON: [{"claim_id": 1, "claim": "...", "question": "..."}]
Question: {question}
Draft: {draft}"""
VERIFY = """Answer the question concisely and factually. If you are not sure, say "unknown".
Question: {q}"""
JUDGE = """Claim: {claim}
Independent answer: {answer}
Does the independent answer SUPPORT, CONTRADICT, or say NOTHING about the claim?
Reply with exactly one word."""
FINAL = """Original question: {question}
Draft answer: {draft}
Verification results (claim, independent answer, verdict):
{ledger}
Write the final answer. Keep supported claims. Remove or correct contradicted claims.
For claims with verdict NOTHING, either drop them or mark them as uncertain.
Do not add any new factual claim that is not in the draft or the verification answers."""
def cove(question, llm, verify_llm=None):
verify_llm = verify_llm or llm
draft = llm(question) # 1. baseline
plan = json.loads(llm(PLAN.format(question=question, draft=draft)))[:MAX_QUESTIONS] # 2. plan
def check(item): # 3. factored execution
answer = verify_llm(VERIFY.format(q=item["question"])) # draft NOT in context
verdict = llm(JUDGE.format(claim=item["claim"], answer=answer)).strip().upper()
return {**item, "answer": answer, "verdict": verdict}
with ThreadPoolExecutor(max_workers=8) as pool:
ledger = list(pool.map(check, plan))
lines = "\n".join(f'- [{r["verdict"]}] {r["claim"]} | {r["answer"]}' for r in ledger)
final = llm(FINAL.format(question=question, draft=draft, ledger=lines)) # 4. revise
return final, ledgerFour decisions are built into it. Verification questions are capped, because a long answer can yield dozens of claims and cost grows linearly. The planner returns structured JSON with a claim identifier per question, so every verdict maps back to a specific sentence; wrap that parse in a retry with a schema check, as described in the structured output article. Verifiers run in parallel, so latency grows with the slowest question rather than the number of questions. And the final prompt forbids new claims, because a revision step that adds fresh details has reintroduced exactly the unverified content CoVe is meant to remove.
The claim ledger as the core data structure
The ledger, a list of records with claim, question, independent answer and verdict, is what turns CoVe from a prompt trick into something you can operate. Store it with every response. It lets you show users which statements were checked, compute how often drafts are corrected, find questions where the verifier said unknown, and build evaluation sets from real traffic. A response whose ledger contains many contradictions is a signal about the question type, not just that one answer.
Verdicts should drive behaviour explicitly rather than being left to the final prompt's judgment. A simple policy works well: supported claims stay; contradicted claims are removed or replaced by the verifier's answer; claims with no evidence either way are removed from list answers and softened in prose. When more than a threshold fraction of claims is contradicted, the draft is probably unreliable as a whole, and regenerating from the verified facts beats patching it.
Extension: tool-backed verification
The paper's verifiers use only the model's own knowledge. Many production systems go further, and it is worth being clear that this is an extension beyond the published method. The verification question is a natural search query, so a verifier can retrieve documents and answer from them, query a database for numbers, or run a calculator or code for arithmetic. That converts CoVe from self-consistency checking into grounded checking, which can catch errors the model could never catch on its own.
Tools change the threat model. Retrieved text can contain instructions, so treat it as untrusted data and apply the defences from the hallucination guardrails article. Retrieval quality now bounds verification quality: a verifier that finds nothing should return unknown, not guess. And verdicts should record their source, so a reviewer can tell a claim checked against your documentation from one checked against the model's memory.
Worked example
Take the question: list three programming languages created at Bell Labs. Suppose the baseline answer is C, C++ and Java. The planner extracts three claims and writes open questions: where was C created, where was C++ created, where was Java created. Each question runs in its own call with no draft in context. The verifiers answer Bell Labs for C (Dennis Ritchie), Bell Labs for C++ (Bjarne Stroustrup), and Sun Microsystems for Java.
The judge marks the first two SUPPORT and the third CONTRADICT. The final step keeps C and C++, removes Java, and, because the question asked for three, needs a replacement. Under the no-new-claims rule it cannot simply invent one, so there are two correct designs: return two verified items and say so, or run a second, smaller CoVe round on a candidate such as AWK, which was also created at Bell Labs, before adding it. The ledger for this response records one contradiction, which is exactly the kind of error CoVe targets: a plausible neighbour inserted into a list.
Cost, latency and when to skip it
Factor+revise with N questions costs 1 baseline call, 1 planning call, N verification calls, N judge calls and 1 final call. With N = 8 that is 19 calls for one answer, though verification and judge prompts are short. Latency is roughly baseline plus planning plus the slowest verification-and-judge pair plus final, if verifications run in parallel. Budget both before promising CoVe to a product team.
Most traffic does not need it. Route by risk: apply CoVe to answers containing lists, named entities, dates or figures, and to domains where errors are costly, and skip it for opinion, formatting and creative requests. A cheap classifier or a rule on the planner's claim count can make that decision. Where a single short fact is asked, sampling several answers and checking agreement, as in the self-consistency article, may be cheaper.
Failure modes
- Correlated errors. The same model answers its own verification questions with the same knowledge gaps; a confidently wrong fact is often confirmed. Use a different model or tools for verification on high-stakes paths.
- Leading questions. A planner that writes 'Was Java created at Bell Labs?' invites agreement. Enforce open questions in the plan prompt and reject yes/no questions in code.
- Over-deletion. A wrong verifier answer removes a correct claim. Track how often corrections were themselves wrong in your evaluation set.
- Revision drift. The final step adds new unverified details. Forbid it in the prompt and diff final claims against the ledger.
- Uncovered claims. The planner misses a claim, which then passes unchecked. Cap questions but log what was left out.
- Parse failures. Malformed planner JSON silently skips verification. Fail closed and retry.
Evaluating a CoVe deployment
Measure it against the baseline on the same questions. For list answers, precision is the natural metric: the fraction of listed items that are correct. For long-form answers, split the output into atomic claims and score the fraction supported by a trusted source, which is how the paper evaluated biographies. Track correction rate, wrong-correction rate, unknown rate, added latency and cost per answer. The evaluation article covers building the golden set and running these comparisons in CI.
What to do next
- Collect 50 real questions from your traffic that produced lists, names, dates or numbers, and label the correct answers.
- Run the reference pipeline above against them with your model, and store every ledger.
- Compare precision against the plain baseline, and read every case where verification removed a correct claim.
- Add a routing rule so only high-risk answers pay for verification.
- Try a different model, or retrieval, for the verification step, and measure whether correlated errors drop.
- Expose the ledger in logs or UI so reviewers can see which claims were checked and how.