Self-consistency is the simplest test-time compute technique that reliably works: ask a model the same reasoning question several times with sampling turned on, pull the final answer out of each response and return the most common one. Correct reasoning paths tend to converge on the same answer while mistakes scatter, so the plurality is right more often than any single path. Wang and colleagues introduced it in 2022 and reported large gains over single chain-of-thought on arithmetic and commonsense benchmarks.
The idea fits in a sentence, and the self-consistency overview explains why voting works. This article is about making it work in a real system: extracting and normalising answers so votes are counted correctly, choosing the number of samples, stopping early, extending it to free-form outputs, using the vote margin as a confidence score and proving the gain on your own data before paying for it.
The pipeline
A self-consistency call has five stages: a prompt that elicits reasoning and a fixed answer format, N sampled completions, extraction of the final answer from each, normalisation so equivalent answers compare equal, and a vote. In production add a sixth: a decision on the result based on how strongly the samples agreed.
Self-consistency only applies where answers can be compared: a number, a choice, a label, a short entity, a SQL result set or the output of running code. If the output is an essay, plain voting has nothing to count, and you need the universal variant described below.
Extraction and normalisation decide everything
Voting fails silently when extraction is sloppy. If one sample ends with “The answer is 12”, another with “12 apples” and a third with “$12.00”, a naive string vote sees three different answers and the true consensus is lost. Worse, a regex that grabs the first number in a response will vote for an intermediate result.
Fix it at both ends. In the prompt, require a terminal line in a fixed format, such as Final answer: <value>, and give one example of it. In code, take the last matching line, normalise by answer type and treat an unparseable response as an abstention, never as a vote:
import re
from fractions import Fraction
ANSWER = re.compile(r"(?im)^\s*final answer\s*:\s*(.+?)\s*$")
def extract(completion):
"""Return the last 'Final answer:' line, or None (an abstention, not a vote)."""
hits = ANSWER.findall(completion)
return hits[-1] if hits else None
def normalise(ans, kind):
if ans is None:
return None
s = ans.strip().rstrip(".").replace(",", "").replace("$", "")
if kind == "number":
lead = re.match(r"-?[\d.]+(?:/\d+)?", s) # "12 apples" -> "12"
s = lead.group(0) if lead else s
try:
return str(Fraction(s).limit_denominator(10_000)) # "0.5", "1/2" -> "1/2"
except (ValueError, ZeroDivisionError):
return None
if kind == "choice":
m = re.match(r"\(?([A-E])\)?\b", s.upper())
return m.group(1) if m else None
return " ".join(s.lower().split())Normalising numbers through Fraction makes 0.5, 1/2 and .50 one vote. For multiple choice, map to a letter. For code, compare by behaviour: run each candidate against a few generated inputs and vote on the output signature, because two different programs can both be right. For SQL, vote on the result set. Log the raw answers that failed to parse; a rising abstention rate is the first sign a prompt or model change broke your format.
A parallel implementation with early stopping
Samples are independent, so issue them concurrently. The function below takes any async generate(prompt, temperature) that wraps your provider, samples in batches and stops as soon as the leader's margin exceeds the number of samples left, at which point no outcome of the remaining calls could change the winner:
import asyncio
from collections import Counter
async def self_consistent(generate, prompt, kind, n=10, temperature=0.7, batch=5,
stop_margin=True):
"""generate(prompt, temperature) -> completion text; any provider."""
votes, abstain, used = Counter(), 0, 0
while used < n:
k = min(batch, n - used)
outs = await asyncio.gather(*(generate(prompt, temperature) for _ in range(k)))
used += k
for out in outs:
a = normalise(extract(out), kind)
if a is None:
abstain += 1
else:
votes[a] += 1
if stop_margin and votes:
ranked = votes.most_common(2)
lead = ranked[0][1] - (ranked[1][1] if len(ranked) > 1 else 0)
if lead > n - used: # remaining samples cannot change the winner
break
if not votes:
return None, 0.0, used
answer, count = votes.most_common(1)[0]
return answer, count / used, used # agreement counts abstentions against youTrace it with N of 10 and batches of 5. The first batch returns 42, 42, 42, 40 and one response with no final-answer line. Votes are 42 three times and 40 once, with one abstention; the margin is 2 and 5 samples remain, so sampling continues. The second batch returns 42, 42, 42, 38 and 42. Now 42 has 7 votes, the margin is 6 and no samples remain, so the loop ends with 42 at an agreement of 7 out of 10. Had the first batch been five clean 42s, the margin of 5 would only equal the 5 remaining samples, which still allows a tie, so the rule runs the second batch; with an odd budget of 9, the same five agreeing samples exceed the 4 remaining and the call stops at just over half its budget, which is what happens on most easy questions.
Agreement is computed over all samples used, including abstentions, so a run where half the responses were unparseable cannot report high confidence. If your provider can return several completions for one request, use that instead of separate calls, since the prompt is processed once. Either way, cache the shared prompt prefix where your provider supports prompt caching, because every sample repeats it.
Choosing N and temperature
Gains are steep for the first few samples and flatten quickly. A useful mental model: if each sample is right with probability 0.6 and wrong answers are spread across several values, a plurality of 5 is right noticeably more often than one sample, a plurality of 15 more still, and beyond that each extra sample buys little. If wrong answers concentrate on one value, typically a tempting trap, voting helps much less and can even lock in the error, because the trap wins the vote. That is why measured curves differ by task and you must measure your own.
Temperature controls diversity. Too low and the samples are near copies, so the vote adds nothing; too high and reasoning quality drops for every sample. Values around 0.5 to 0.8, with the provider's default top-p, are a reasonable starting range; sweep them on your evaluation set. Some models and reasoning modes fix or restrict sampling parameters, so check your provider's documentation rather than assuming a temperature setting is honoured.
Cost scales linearly in output tokens. With a 600-token reasoning chain and N of 10, one question costs 6,000 output tokens instead of 600. Early stopping cuts this sharply on easy questions, where the first five samples usually agree, and spends the full budget only on hard ones. Adaptive-Consistency (Aggarwal and colleagues, 2023) and Early-Stopping Self-Consistency (Li and colleagues, 2024) formalise this with statistical stopping rules and report large sample savings at similar accuracy; the margin rule above is the simplest exact version.
Weighted and universal variants
Plain plurality treats every sample equally. Three refinements are worth knowing.
| Variant | How it works | Use when |
|---|---|---|
| Weighted by likelihood | Each vote weighted by the sample's normalised log-probability | Your API returns token log-probs; gains are usually small |
| Verifier-weighted | A separate scorer rates each reasoning path; votes weighted by score | You have or can train a reliable verifier |
| Universal self-consistency | The model is shown all candidates and asked which is most consistent with the others | Free-form outputs such as summaries or open answers |
| Semantic clustering | Embed answers, cluster near-duplicates, vote on clusters | Short free-text answers with paraphrase variation |
Universal self-consistency (Chen and colleagues, 2023) is the practical route for outputs that cannot be normalised: instead of counting, a final call selects the candidate that agrees most with the rest. It costs one more call and inherits that call's biases, such as preferring longer responses, so evaluate it rather than assuming it beats a single good sample. Verifier weighting is covered in the verifier article.
Agreement as a confidence signal
The most underused output of self-consistency is the vote share. When 9 of 10 samples agree, the answer is usually right; when the top answer holds 3 of 10, it often is not. That makes agreement a cheap confidence estimate for routing: accept high-agreement answers automatically, send low-agreement ones to a stronger model, retrieval, a tool call or a human.
Worked example. A support team extracts refund amounts from customer emails. On a labelled set of 500 emails, a single sample is right 88 percent of the time. Self-consistency at N of 7 lifts that to 93 percent. More useful, answers with agreement of at least 6 of 7 cover 81 percent of emails and are right 99 percent of the time; the other 19 percent go to a human queue. The automation now has a measured error rate instead of a hopeful one. These numbers are illustrative; the point is that the agreement threshold turns accuracy into a coverage-versus-precision dial you can set.
Calibrate the threshold on held-out data and recheck it whenever the model or prompt changes, because agreement levels shift with both.
Measuring it before paying for it
Never ship self-consistency on faith. Run the same labelled set at several values of N and record accuracy, calls per item, and coverage and precision at your acceptance threshold:
async def evaluate(dataset, generate, kind, ns=(1, 5, 10, 20)):
for n in ns:
correct = calls = accepted = acc_correct = 0
for item in dataset:
ans, agree, used = await self_consistent(generate, item["prompt"], kind, n=n)
calls += used
ok = ans == normalise(item["gold"], kind)
correct += ok
if agree >= 0.6:
accepted += 1
acc_correct += ok
print(f"N={n:>2} acc={correct/len(dataset):.3f} calls/item={calls/len(dataset):.1f} "
f"coverage@0.6={accepted/len(dataset):.2f} "
f"acc_when_accepted={acc_correct/max(accepted,1):.3f}")Compare against cheaper alternatives at equal cost: a stronger model with one sample, a better prompt, or retrieval. For models that already reason at length internally, the marginal gain from external voting is often smaller and the per-sample cost larger, so the comparison matters more. Integrate the harness with your existing prompt evaluation setup so regressions are caught in CI.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| No gain over one sample | Temperature too low, samples identical | Raise temperature; check answer diversity |
| Consensus on a wrong answer | Systematic error or trap shared by all paths | Better prompt, retrieval or tools; voting cannot fix bias |
| Votes split across equivalent answers | Weak normalisation | Type-aware normalisation; log raw answers |
| Confident answer from mostly broken output | Abstentions ignored in agreement | Count abstentions in the denominator |
| Cost blows up | Fixed N on easy questions | Early stopping; batch sampling; prompt caching |
| Latency spikes | Sequential sampling or one slow call | Parallel calls with a timeout; vote on what returned |
Where it fits
Self-consistency sits between single chain-of-thought and search methods such as tree of thoughts. It is embarrassingly parallel, needs no change to the model and composes with almost anything: retrieval, tool use and verifiers. Use it when answers are checkable, errors are varied rather than systematic and the cost of a wrong answer exceeds a few extra model calls.
What to do next
- Pick one task with a checkable answer and build a labelled set of at least 200 items.
- Add a fixed final-answer line to the prompt and a type-aware extractor that abstains on parse failure.
- Run the evaluation harness at N of 1, 5, 10 and 20 and plot accuracy against calls per item.
- If the curve justifies it, deploy with parallel batched sampling and margin-based early stopping.
- Choose an agreement threshold from held-out data and route low-agreement answers to escalation.
- Monitor abstention rate, average samples used and agreement distribution, and re-run the harness after every model or prompt change.