Multi-agent debate is a simple idea: instead of asking one model once, ask several model instances, show each of them the others' answers and reasoning, let them revise over a few rounds, and then aggregate. The hope is that agents catch each other's mistakes the way a good review meeting does, and that the final answer is more accurate and better justified than any single attempt.
Sometimes it works and sometimes it is an expensive way to reach the same answer as voting over independent samples, or a worse one when a confident wrong agent talks the others round. This article explains the architecture from first principles: the loop, the roles, how answers are aggregated and when to stop, what it costs, what published research supports, how to evaluate it honestly against cheaper baselines, and the failure modes to watch for in production. The code is provider-neutral; llm.complete stands for whatever client you use.
Where debate sits among multi-agent patterns
Most multi-agent systems split work: a planner hands sub-tasks to specialists, or agents bid for tasks, as described in collaboration and negotiation and the broader multi-agent systems overview. Debate does not split the work. Every agent attempts the whole problem, and the value comes from disagreement between attempts. That makes it a verification and error-correction pattern, closer to reflection and evaluator loops than to orchestration.
The distinction matters for design. In a division-of-labour system you add agents to cover more ground; in debate you add agents to get more independent opinions, and an extra agent is only useful if its errors are not the same as the others'. Everything below follows from that.
The core loop
A debate has three phases. In round zero every debater answers independently, with no view of the others. If they all agree, there is nothing to debate. Otherwise, in each later round every debater sees its own previous answer plus the other debaters' answers and reasoning, and produces a revised answer. After a fixed number of rounds, or once answers converge, an aggregator produces the final answer.
Three implementation details carry most of the weight. First, answers must be extracted into a normalised form (a number, a letter, a canonical entity) so agreement can be checked mechanically; free-text answers that say the same thing differently will look like disagreement forever. Second, round zero must truly be independent, which means no shared scratchpad and no peer context; the diversity you are paying for is created there. Third, the revision prompt should demand a specific reason to change, otherwise models tend to agree with whatever they have just read.
Roles: debaters, devil's advocate and judge
The symmetric version uses identical debaters. Variants assign roles. A devil's advocate is instructed to argue against the current majority, which counters premature convergence. In the formulation of Liang et al. (2023), two debaters argue opposing sides and a separate judge decides when the debate is settled and which answer wins; their motivation was that a single model reflecting on its own answer tends to lock in (they call it degeneration of thought), and an opposing voice breaks that.
Diversity can also come from the agents themselves rather than their instructions: different models, different temperatures, different retrieved context, or different tools. Agents built on the same model with the same prompt at low temperature mostly make the same mistakes, so the debate has nothing to correct. In practice, mixing two model families or giving one agent a code-execution tool does more than adding a fourth identical debater.
| Role | Instruction | Use it when |
|---|---|---|
| Symmetric debater | Answer, then revise against peers | Questions with a checkable answer |
| Devil's advocate | Find the strongest case against the majority | Agents converge too quickly |
| Judge | Read the transcript, choose and justify | Answers are not mechanically comparable |
| Tool-using verifier | Run code, query a source, report evidence | Claims can be tested externally |
An implementation
The code below implements the symmetric loop with early stopping and an optional judge. Agents within a round run in parallel; rounds are sequential, so wall-clock latency is roughly rounds multiplied by the slowest call.
from collections import Counter
from concurrent.futures import ThreadPoolExecutor
from dataclasses import dataclass
@dataclass
class Turn:
agent: str
round: int
answer: str # normalised final answer used for voting, e.g. "17" or "B"
reasoning: str # free text shown to peers in the next round
def ask(llm, agent, question, peers, round_no):
if not peers:
prompt = f"{agent.persona}\n\nQuestion: {question}\nReason step by step, then give FINAL: <answer>."
else:
shown = "\n\n".join(f"[{t.agent}] FINAL: {t.answer}\n{t.reasoning}" for t in peers)
prompt = (f"{agent.persona}\n\nQuestion: {question}\n\nYour previous answer: {agent.last.answer}\n"
f"{agent.last.reasoning}\n\nOther agents said:\n{shown}\n\n"
"Check each argument for errors. Change your answer only if you find a concrete "
"flaw in your own reasoning; say which step. End with FINAL: <answer>.")
text = llm.complete(model=agent.model, prompt=prompt, temperature=agent.temperature)
turn = Turn(agent.name, round_no, extract_final(text), text)
agent.last = turn
return turn
def debate(question, agents, llm, max_rounds=3, judge=None):
rounds = []
with ThreadPoolExecutor(len(agents)) as pool: # agents in a round run in parallel
turns = list(pool.map(lambda a: ask(llm, a, question, [], 0), agents))
rounds.append(turns)
for r in range(1, max_rounds):
if len({t.answer for t in turns}) == 1: # unanimous: stop early
break
turns = list(pool.map(
lambda a: ask(llm, a, question, [t for t in turns if t.agent != a.name], r), agents))
rounds.append(turns)
if judge is not None:
return judge.decide(question, rounds), rounds
answer, _ = Counter(t.answer for t in turns).most_common(1)[0]
return answer, roundsNote that ask shows each agent its own previous reasoning and its peers' reasoning in full. That gives the most information, and it also makes the prompt grow with both the number of agents and the number of rounds. Two common compressions are showing only peers' final answers and a short justification, or having a summariser condense each round. Both cut cost and both can strip out the one step that revealed an error, so measure accuracy when you add them.
Aggregation and stopping
Majority vote over the final round is the simplest aggregator and works whenever answers are comparable. Ties need a rule: fall back to the agent with the strongest track record, the judge, or the round-zero majority. A judge model reads the transcript and decides; it is needed for open-ended outputs such as a design recommendation, but judges have known biases (towards longer answers, towards the first or last position, towards their own model family), so randomise order and keep the judge on a different model where you can. The evaluator agent article covers judge design in more depth.
Stopping rules control cost. Stop at unanimity; stop when no agent changed its answer between two rounds, since further rounds rarely move a stable split; and enforce a hard cap of two or three rounds. A particularly cheap design is disagreement-triggered debate: run round zero, and only if the independent answers disagree spend anything on further rounds. For questions the agents find easy, this costs exactly the same as sampling N answers and voting.
What the research supports
The idea has two roots. Irving, Christiano and Amodei (2018) proposed debate as an AI-safety technique: two agents argue and a weaker judge decides, on the theory that it is easier to judge an argument than to produce the answer. Du et al. (2023) applied it to accuracy, having several instances of a language model debate over rounds, and reported gains over a single model on arithmetic, grade-school maths and factual biographies, with more agents and more rounds generally helping.
Later work tests both claims. Khan et al. (2024) gave debaters access to a text the judge could not see and found that debate helped weaker judges, including humans, reach correct answers more often than a single consultant arguing one side, and that making debaters more persuasive improved judge accuracy further. On the other hand, Smit et al. (2024) compared several debate protocols with other prompting strategies and found that debate did not reliably beat simpler methods such as self-consistency, with results sensitive to configuration such as how strongly agents are pushed to agree.
The practical reading: debate is not a free accuracy boost. It helps most when agents have different information or tools, when there is a checkable answer, and when a judge benefits from seeing arguments on both sides. On tasks where independent samples already vote correctly, it mainly adds tokens.
Cost model and a worked example
Let N be the number of agents, R the rounds actually run, q the question length in tokens and a the length of one answer with reasoning. Round zero costs N calls with about q input tokens each. Every later round costs N calls with about q + N times a input tokens each, since each agent sees its own previous answer and N minus 1 peers'. Output is about N times R times a in total.
Take N = 3, q = 300 and a = 400, with two rounds. Round zero: 900 input and 1,200 output tokens. Round one: three calls of 300 + 1,200 = 1,500 input tokens each, so 4,500 input and 1,200 output. Total: 5,400 input and 2,400 output tokens, versus 300 and 400 for one call. The fair baseline is not one call but six independent samples with a majority vote: 1,800 input and the same 2,400 output tokens, and one round of latency instead of two. Debate has to beat that baseline to justify the extra input and the doubled latency.
Now apply disagreement triggering. If the three round-zero answers agree on 80 percent of questions, round one runs on only 20 percent, and the expected cost drops to 900 + 0.2 times 4,500 = 1,800 input tokens per question. That is usually the configuration worth shipping.
Evaluating it honestly
Compare against compute-matched baselines on the same question set: a single call, self-consistency with the same total output tokens, and single-agent self-refinement. Report accuracy, tokens and latency for each. Then look inside the debates, because the aggregate number hides the mechanism. The most useful statistic is the flip count: how often an agent moved from wrong to right versus from right to wrong.
def flip_stats(rounds, gold):
# Count answer changes between the first and last round, per agent.
first = {t.agent: t.answer for t in rounds[0]}
last = {t.agent: t.answer for t in rounds[-1]}
fixed = sum(first[a] != gold and last[a] == gold for a in first) # wrong -> right
broke = sum(first[a] == gold and last[a] != gold for a in first) # right -> wrong
return {"fixed": fixed, "broke": broke, "rounds_used": len(rounds)}If right-to-wrong flips are common, agents are conforming rather than checking, and the revision prompt or the agent mix needs work. Also track the round-zero agreement rate (it tells you how often debate even runs) and accuracy split by whether round zero agreed. The agent evaluation at scale article covers building the harness.
Failure modes
- Conformity cascades. A fluent, confident wrong answer persuades the others; right-to-wrong flips rise. Counter with a revision prompt that requires naming a concrete flaw, heterogeneous agents, and a devil's advocate.
- Correlated errors. Identical agents share blind spots, so unanimity is not evidence of correctness. Measure accuracy on unanimous cases separately.
- Answer extraction drift. Formatting differences look like disagreement and trigger needless rounds, or two different answers normalise to the same string.
- Oscillation. Agents swap answers each round without converging; the no-change stopping rule and a round cap handle it.
- Judge bias. Position, length and self-preference biases decide close calls. Randomise order and audit a sample by hand.
- Context growth. Prompts grow with agents times rounds until they hit limits or degrade attention; compress, but verify that compression does not hide the evidence.
- Injected content. If agents read retrieved documents, a malicious document can argue through one agent into the others; treat tool output as untrusted.
What to do next
- Pick a task with a checkable answer and build a labelled set of a few hundred questions.
- Measure the baselines first: one call, and self-consistency at matched output tokens.
- Implement the loop with normalised answer extraction, independent round zero, disagreement triggering and a cap of two or three rounds.
- Diversify agents (two model families, a tool-using verifier or a devil's advocate) before adding more of the same agent.
- Log every turn and compute flip statistics, agreement rate and cost per question.
- Ship debate only for the slice of traffic where it beats the matched baseline, and route the rest to the cheaper method.