Debate prompting asks several model instances, or several personas inside one model, to answer the same question, read each other's answers, argue, and then converge or hand the disagreement to a judge. The idea is attractive: one confident model can be wrong in a way that a second model, asked to attack the answer, can catch. In practice a debate is only as good as its prompts. A weak rebuttal prompt turns debate into polite agreement, and a weak judge prompt rewards the most fluent answer rather than the correct one.
This article is about the prompt text. The system design, including topologies, cost curves and aggregation strategies, is covered in the multi-agent debate architecture guide, and the theory of why debate can amplify truth is in the maths of multi-agent debate. Here you get the four prompts every debate needs, a worked transcript, a runnable loop, and an evaluation that tells you whether debate is worth paying for on your task.
Where debate prompting comes from
The modern line of work starts with Irving, Christiano and Amodei (2018), who proposed debate as an alignment technique: two agents argue and a weaker judge decides, on the theory that lies are harder to defend than truths. Du and colleagues (2023) applied a simpler version to factuality and arithmetic: several instances answer, then repeatedly read each other's answers and revise, which improved accuracy on their benchmarks. Liang and colleagues (2023) added assigned opposing roles and a judge to fight what they called degeneration of thought, where a single model keeps defending its first idea. Khan and colleagues (2024) found that non-expert judges reached more accurate answers when reading a debate than when reading one consultant's argument.
The sceptical result matters just as much. Smit and colleagues (2024) compared debate strategies against simpler methods and found that debate did not reliably beat self-consistency, and that results were sensitive to settings such as agreement level and number of rounds. Treat that as the default expectation: debate is a tool to test on your data, not a guaranteed upgrade.
The four prompts
Every debate has the same parts. The format spec is shared and fixes how the final answer is written, for example a last line FINAL: <answer>. The opening prompt produces independent first answers. The rebuttal prompt runs in each later round and shows one debater the others' answers. The judge prompt decides when the debaters do not agree. Most failures trace back to one of these templates, so version them like code and keep them in your evaluation harness.
Opening prompts: independence first
The value of debate comes from errors that are not correlated. If all three debaters make the same mistake, no amount of arguing will find it. The opening prompt should therefore maximise independence: debaters must not see each other's answers in round one, and you should vary something real between them. Options, from weakest to strongest, are sampling temperature, a different reasoning instruction (work forwards, work backwards, estimate first), an assigned stance, or a different model family.
OPENING = '''You are Debater {name}. Answer the question below on your own.
{stance}
Show the reasoning that a careful checker would need to verify you.
State any assumption you make explicitly.
End with exactly one line in the form:
FINAL: <answer>
Question:
{question}'''
STANCES = {
"forward": "Work forwards from the given facts, step by step.",
"backward": "Start from what a correct answer must satisfy and work backwards.",
"skeptic": "Assume the obvious answer is wrong and look for the trap before answering.",
}Assigned personas help when the persona changes the method, not just the tone. "You are a sceptical auditor" works because it changes what the model checks; "you are a friendly professor" mostly changes the wording. The role prompting guide covers that distinction. For open-ended decisions rather than questions with one right answer, assign explicit sides ("argue for option A") so both options get their strongest case.
The rebuttal prompt: change your mind only for a reason
This is the prompt that decides whether debate works. Models tend to agree with whatever text is in front of them, so a naive "here are the other answers, update yours" produces fast convergence on whichever answer is most common or most confident, right or wrong. The rebuttal prompt has to make changing an answer cost something: the debater must name the specific step it now believes is wrong.
REBUTTAL = '''You are Debater {name}. This is round {round} of a debate.
Your previous answer:
<yours>
{own}
</yours>
Answers from other debaters (order is random; they may be wrong):
{peers}
Do the following in order:
1. For each other answer, find the first step you believe is wrong, or write
"no error found".
2. Re-check your own answer the same way.
3. Keep your answer unless you found a concrete error in it. Agreement from
others is not evidence. If you change, quote the step you are correcting.
End with exactly one line in the form:
FINAL: <answer>'''
def format_peers(peers):
return "\n".join(f"<answer id='{i + 1}'>\n{p}\n</answer>" for i, p in enumerate(peers))Three details matter. Peers are anonymised and shuffled, so debaters cannot defer to a name or to the first position. Peer answers are wrapped in delimiters, so instructions inside one cannot take over another debater. And the prompt explicitly says agreement is not evidence, which reduces conformity, although it never removes it.
The judge prompt: a rubric, not a vote
When debaters still disagree after the last round, you need a decision. A majority vote is cheap and often good. A judge is better when the question has checkable steps, because it can find the step where answers diverge and check only that step. The judge must not grade style, length or confidence.
JUDGE = '''You are the judge. Several debaters answered the question below.
Do not trust any of them. Do not reward length, confidence or fluency.
Question:
{question}
Debate transcripts:
{transcripts}
Procedure:
1. List the distinct final answers.
2. Find the first step where the reasoning behind them diverges.
3. Check that step yourself from the question's facts.
4. Choose the answer whose reasoning survives the check. If none does,
give your own answer and say why.
End with exactly one line:
FINAL: <answer>'''Use a different model, or at least temperature 0 and a fresh context, for the judge. A judge that shares the debaters' blind spot simply ratifies their error. If you cannot afford a separate judge call, use majority vote and log disagreement as a signal for human review.
Worked example: a time-arithmetic question
Question: A job is scheduled 1,000 minutes after 22:30 on Friday. When does it start? In round one, Debater A works forwards: 1,000 minutes is 16 hours 40 minutes, and 22:30 plus 16:40 is 15:10 on Saturday. Debater B, using the backward stance, converts too early and writes 14:10 on Saturday. Debater C, the sceptic, checks the day boundary: 90 minutes takes 22:30 to midnight, leaving 910 minutes, which is 15 hours 10 minutes, so 15:10 on Saturday.
In round two, B receives A and C anonymously. Following the rebuttal procedure, B re-checks its own conversion, finds that 16 hours 40 minutes added to 22:30 gives 39:10, which is 15:10 the next day, quotes that step as its correction, and changes to 15:10. A and C report "no error found" in each other and keep their answers. All three final lines match, so the loop stops early and the judge is never called. The log records one flip, from wrong to right, and the reason for it. That record is what you will aggregate in evaluation.
A runnable loop
The loop below works with any chat API. complete(prompt, temperature) is your wrapper that returns text. It stops early on unanimity, falls back to majority vote without a judge, and returns the history so you can measure flips.
import random, re
from collections import Counter
FINAL_RE = re.compile(r"^FINAL:\s*(.+?)\s*$", re.M)
def final(text):
hits = FINAL_RE.findall(text or "")
return hits[-1] if hits else None
def debate(question, complete, names=("A", "B", "C"), rounds=2, judge=None):
stance_keys = list(STANCES)
answers = [complete(OPENING.format(name=n, stance=STANCES[stance_keys[i % 3]],
question=question), temperature=0.8)
for i, n in enumerate(names)]
history = [answers]
for r in range(2, rounds + 1):
finals = [final(a) for a in answers]
if None not in finals and len(set(finals)) == 1:
break # unanimous: stop paying
new = []
for i, n in enumerate(names):
peers = [a for j, a in enumerate(answers) if j != i]
random.shuffle(peers)
new.append(complete(REBUTTAL.format(name=n, round=r, own=answers[i],
peers=format_peers(peers)), temperature=0.3))
answers = new
history.append(answers)
finals = [f for f in (final(a) for a in answers) if f]
if len(set(finals)) == 1:
return finals[0], history
if judge is not None:
verdict = judge(JUDGE.format(question=question, transcripts="\n\n".join(answers)))
return final(verdict), history
return (Counter(finals).most_common(1)[0][0] if finals else None), history
Single-call debate versus multi-call
You can also simulate a debate inside one prompt: "Three experts with these roles discuss the question for two rounds, then agree on an answer." Wang and colleagues (2023) studied a version of this, called Solo Performance Prompting, in which one model plays several personas. It costs one call, but the personas share one context and one set of blind spots, so their disagreement is weaker than real independent samples.
| Single-call simulated debate | Multi-call debate | |
|---|---|---|
| Calls per question | 1 | agents x rounds, plus an optional judge |
| Independence of errors | Low: one context, one sample | Higher: separate samples, optionally separate models |
| Latency | One long generation | Rounds run in sequence; debaters in a round run in parallel |
| Good for | Surfacing considerations for a decision | Questions with a checkable right answer |
| Main risk | Personas agree immediately | Cost, and conformity in later rounds |
Evaluating it honestly
Debate must beat the cheapest strong alternative at the same token budget. That alternative is usually self-consistency: sample N independent answers and take the majority. Three debaters for two rounds is six calls, and each rebuttal reads its peers' answers, so compare against at least six to nine self-consistency samples.
- Accuracy at equal tokens. Plot accuracy against total input plus output tokens for single answer, self-consistency and debate.
- Flip matrix. Count right-to-wrong and wrong-to-right changes between rounds. Debate is working only if wrong-to-right clearly outnumbers right-to-wrong.
- Unanimity rate and early stops. These set your real cost, because unanimous first rounds skip later rounds.
- Judge agreement with ground truth on split cases. This is the only place the judge matters, so measure it there.
Use a held-out set from your own domain with known answers. Public benchmarks may already be in the models' training data, and debate gains on them do not transfer automatically.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Everyone agrees in round one, sometimes wrongly | Openings are not independent: same model, same prompt, low temperature | Vary stance and method, raise opening temperature, mix model families |
| Right answers flip to wrong in round two | Conformity: the rebuttal prompt rewards agreement | Require a quoted error before any change; shuffle and anonymise peers |
| Judge picks the longest answer | Rubric grades fluency, or the judge shares the debaters' error | Divergence-step rubric, separate judge model, cap transcript length |
| Parser returns None | Model wrote the final answer in prose | Strict FINAL line, one retry with a format-only reminder |
| Costs several times a single answer for no gain | Task has no checkable steps, or errors are shared | Compare with self-consistency at equal tokens and switch if it wins |
| Prompt injection spreads between debaters | Peer text treated as instructions | Delimit peer answers and tell debaters they are data, not instructions |
What to do next
- Pick 100 to 200 questions from your own domain with verified answers and record single-answer accuracy and token cost.
- Implement the four templates above with a strict
FINAL:line, and store them with version numbers. - Run self-consistency and three-debater, two-round debate at matched token budgets and compare accuracy.
- Log the flip matrix per round; if right-to-wrong flips are frequent, tighten the rebuttal prompt before adding rounds.
- Add a judge only for split cases, and measure its accuracy on exactly those cases.
- Ship debate only for the question types where it wins, and route the rest to the cheaper method.