A multi-agent system runs several model loops, each with its own instructions, tools and context window, and connects them through messages. The most common form has a coordinator that splits the work and specialists that each do a piece of it. The pattern is popular because it looks like a team. That is the wrong reason to use it. The real reason is narrower and more mechanical: each agent gets a fresh context window, so work that would overflow or pollute one window can be spread across several and run in parallel.

This article is about deciding whether to split, choosing a topology, and defining the contract between agents. The orchestration runtime itself, with its task graph, scheduler and budgets, is covered in agent orchestrator architecture. The prompts are covered in multi-agent orchestration prompts.

Advertisement

What you gain and what you give up

A single agent accumulates everything in one context: every tool result, every dead end, every intermediate thought. Quality degrades as that context fills with material irrelevant to the current step, and each step re-reads everything that came before. Splitting the work buys three things.

  • Context isolation. A specialist sees only its brief and its own tool results. It can read fifty pages and return a two-hundred-word summary, so the coordinator's context stays small and relevant.
  • Parallelism. Independent branches run at the same time, so wall-clock time is set by the longest branch rather than the sum of all of them.
  • Specialisation. Each agent can have a narrower prompt, a smaller tool set and even a different model, which improves reliability and allows least-privilege access.

The cost is equally concrete. Agents do not share what they learned unless you send it explicitly, so decisions made in one branch are invisible to the others. Two specialists can make choices that are each reasonable but incompatible, such as different naming conventions, different assumptions about an API or overlapping searches. Cognition's widely read essay arguing against multi-agent designs makes exactly this point. Actions carry implicit decisions, and without the full trace an agent cannot respect decisions made elsewhere. The pattern works best when the branches are genuinely independent, as in research, broad search and review, and worst when they must make many tightly coupled decisions, as when several agents write parts of the same program.

What the evidence says

The most detailed public account is Anthropic's engineering write-up of its multi-agent research system, in which a lead agent plans, spawns parallel subagents and later passes the results to a separate citation step. Its reported numbers are useful for calibration. A multi-agent configuration with Claude Opus 4 as lead and Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2% on the team's internal research evaluation. In their BrowseComp analysis, token usage alone explained 80% of the variance in performance. Agents used about 4× the tokens of a chat interaction, and multi-agent systems about 15×. Letting subagents call tools in parallel cut research time by up to 90% for complex queries.

Read those numbers together. Much of the quality gain comes from spending more tokens, and the architecture is mainly a way to spend them productively, across many fresh contexts instead of one overloaded one. That only pays off when the task is valuable enough to justify roughly an order of magnitude more spend, and broad enough to split. The same write-up reports that vague delegation caused subagents to duplicate one another's work. That leads directly to the contract design below.

Advertisement

Topologies

TopologyShapeGood forWatch out for
Coordinator and specialistsHub plans, spokes execute, hub synthesisesResearch, analysis, review across many sourcesHub becomes a bottleneck; briefs too vague
Handoff (router)Agent passes the whole conversation to anotherCustomer support triage, domain switchingLost state at handoff; ping-pong loops
PipelineFixed stages, each consumes the previous outputExtract, transform, check workflowsEarly errors propagate; no backtracking
Parallel fan-out and reduceSame task, many inputs or many attempts, then mergeBulk classification, best-of-n, map-reduce over documentsMerge step is where quality is lost
Debate and critiqueAgents argue, a judge decidesHard judgement callsCost; agreement is not correctness

Start with the simplest topology that isolates the context you need to isolate. A pipeline is often just a workflow with model calls in it, and should be built as one. Debate has its own trade-offs, covered in multi-agent debate. The rest of this article uses coordinator and specialists, the shape most production systems converge on.

Coordinator and specialists: each agent owns its own context windowCoordinatorplans, writes briefs, synthesisesSpecialist Aown context + toolsSpecialist Bown context + toolsSpecialist Cown context + toolsVerifierchecks claims, no new workbrieftyped resultArtifact storefull outputs by reference: files, documents, tablesBriefs go down, short typed results come up, bulky material goes to the store.The coordinator never reads a specialist's whole transcript.
Coordinator-specialist topology with a verifier and an artifact store. Only briefs and compact typed results cross agent boundaries.

The contract between agents

Most multi-agent failures are interface failures. A brief written as a sentence ("look into pricing") gives a specialist no boundary, no output format and no stopping rule, so it duplicates other branches, wanders and returns prose the coordinator has to reinterpret. Treat the boundary like an API: typed in both directions and validated.

from dataclasses import dataclass, field
from typing import Literal

@dataclass
class Brief:
    task_id: str
    objective: str                 # one outcome, stated so success is checkable
    scope: str                     # what is in bounds, and what another agent owns
    tools: list[str]               # least privilege: only what this task needs
    output_schema: dict            # JSON schema the result must satisfy
    max_steps: int = 12
    max_tokens: int = 150_000

@dataclass
class Result:
    task_id: str
    status: Literal["done", "partial", "blocked"]
    summary: str                   # what the coordinator reads; keep it under ~300 words
    claims: list[dict] = field(default_factory=list)      # each with a source or artifact ref
    artifacts: list[str] = field(default_factory=list)    # references into the store
    open_questions: list[str] = field(default_factory=list)
    tokens_used: int = 0

Several fields carry more weight than they appear to. scope should say what the agent does not own, because overlapping branches are the most common source of wasted tokens. output_schema lets you validate a result mechanically and retry once with the validation error before the coordinator sees it. claims with sources make verification possible later. artifacts keep bulky material, such as fetched documents, generated code and tables, out of the coordinator's context. It receives a reference and fetches content only if it needs it. status lets a specialist say "blocked" honestly instead of inventing an answer to satisfy the schema.

If the agents live in different services or organisations, put the same contract on the wire with a protocol such as A2A rather than inventing one.

A coordinator, provider-agnostic

The sketch below shows the coordinator's control flow. llm is any function that sends a system prompt and input to a model and returns text. run_agent runs one specialist's tool loop and returns a Result. Both are deliberately abstract, so the structure carries over to any SDK.

import asyncio, json

async def run_team(question, llm, run_agent, max_parallel=4):
    # 1. Plan: the coordinator turns the question into disjoint briefs (JSON, validated).
    plan = json.loads(await llm(COORDINATOR_PLAN_PROMPT, question))
    briefs = [Brief(**b) for b in plan["briefs"]]
    assert len({b.task_id for b in briefs}) == len(briefs)

    # 2. Fan out with a concurrency cap; each specialist runs its own tool loop.
    sem = asyncio.Semaphore(max_parallel)
    async def one(brief):
        async with sem:
            try:
                return await asyncio.wait_for(run_agent(brief), timeout=600)
            except Exception as e:     # a failed branch is data, not a crash
                return Result(brief.task_id, "blocked", f"failed: {e!r}")
    results = await asyncio.gather(*(one(b) for b in briefs))

    # 3. Gap check: one bounded follow-up round for blocked or partial branches.
    retry = [i for i, r in enumerate(results) if r.status != "done"][:2]
    redo = await asyncio.gather(*(one(briefs[i]) for i in retry))
    for i, r in zip(retry, redo):
        results[i] = r                 # replace, so the digest has one result per brief

    # 4. Synthesise from summaries and claims only, then verify before returning.
    digest = [{"task": r.task_id, "status": r.status, "summary": r.summary,
               "claims": r.claims} for r in results]
    draft = await llm(COORDINATOR_SYNTH_PROMPT, json.dumps(digest))
    return await llm(VERIFIER_PROMPT, json.dumps({"draft": draft, "claims": digest}))

Notice what the coordinator never does: it never reads a specialist's transcript, and it never lets one failed branch fail the whole request. The follow-up round is bounded, and so is concurrency, because model rate limits and tool rate limits will bind before your thread pool does. The verifier at the end checks claims against their sources and does no new research, so its cost stays predictable.

Worked example: tokens and latency for one research task

Suppose a question needs about 30 tool calls. Each returns about 3,000 tokens, and the model writes about 300 tokens per step. The system prompt and task take 2,000 tokens.

Single agent. Before step i the context holds about 2,000 + 3,300 × i tokens, and every step re-reads it. Summed over 30 steps, input tokens come to 30 × 2,000 + 3,300 × (0 + 1 + ... + 29) = 60,000 + 3,300 × 435 ≈ 1.5 million. By the last step the context is about 100,000 tokens, most of it stale tool output. At about 10 seconds per step the run takes around 5 minutes.

Coordinator and four specialists. Each specialist does 8 steps, 32 in total, in its own context: 8 × 2,000 + 3,300 × (0 + ... + 7) = 16,000 + 92,400 ≈ 108,000 input tokens each, or about 434,000 for four. The coordinator plans with about 2,000 tokens and then synthesises and verifies over four summaries of about 1,500 tokens each, roughly 20,000 tokens in total. That is about 450,000 input tokens, under a third of the single agent's, because no context ever grows past about 30,000 tokens. Wall-clock time is the longest branch, 8 steps or about 80 seconds, plus perhaps 30 seconds of coordination.

So why do production systems report far higher token use? Because nobody stops at the same 30 tool calls. Once branches are cheap and parallel, teams give each specialist more breadth, and that extra exploration is where the quality comes from. Prompt caching also narrows the gap, because the single agent's re-read prefix is mostly cacheable. Budget for the system you will actually run, not for this idealised comparison, and enforce max_steps and max_tokens per brief.

Sharing state without sharing context

Specialists still need some shared facts. Put them in three places, each with a clear owner. Global constraints, such as the user's goal, deadlines and conventions, go in every brief and are written once by the coordinator. Artifacts go to a store that one agent writes and others read, never two writers on the same object. Decisions that later branches must respect go into a short decision log the coordinator maintains and appends to later briefs. This is the multi-agent version of context engineering.

If the work needs many shared decisions, for example several agents editing one codebase, that is a sign to use a single agent with good context management, or a strict pipeline, rather than parallel specialists.

Failure modes

  • Duplicate work. Vague briefs send two specialists to the same sources. Fix it with explicit scope and exclusions.
  • Incompatible outputs. Branches make conflicting assumptions. Fix it with a decision log, schemas and a reconcile step before synthesis.
  • Telephone-game loss. A key caveat disappears from a summary. Fix it with claims plus sources, and let the verifier pull the artifact.
  • Runaway fan-out. The coordinator spawns twenty agents for a simple question. Scale agent count to task complexity in the planning prompt and cap it in code.
  • Confident fabrication under schema pressure. Allow "blocked" and "partial" as first-class statuses.
  • Unbounded retries. A failing branch is re-run until the budget is gone. Allow one follow-up round, then report the gap.
  • Invisible failures. Without tracing you cannot tell which agent went wrong. Log every brief, result, token count and tool call with a shared request ID.

Evaluating and operating it

Evaluate the multi-agent system against the single-agent baseline on the same tasks with the same budget cap. Otherwise you are measuring spend, not architecture. Score final outputs with a rubric and also track per-branch metrics: schema-valid rate, blocked rate, duplicate-source rate and tokens per branch. A sudden rise in tokens per branch usually means a brief has become vague or a tool has started returning bloated output. In production, put cost and latency budgets per request in code, alert on fan-out counts and keep sampled full traces, because emergent behaviour in these systems shows up first as odd delegation patterns.

What to do next

  1. Write down why one agent is not enough. If the answer is not context overflow, parallelism or tool isolation, stay with one agent.
  2. Define Brief and Result types with scope, schema, status and budgets before writing any prompt.
  3. Build the coordinator with bounded fan-out, one follow-up round, an artifact store and a verifier.
  4. Run the same task set through the single-agent baseline and the multi-agent system under equal budgets and compare quality, tokens and latency.
  5. Add tracing with a request ID across all agents, and alert on fan-out and tokens per branch.
  6. Revisit the split whenever branches start needing each other's decisions.
Key takeaway: The multi-agent pattern is a way to spend more tokens productively: each specialist gets a fresh context window, works in parallel and returns a compact typed result, so no single context fills with stale material. It works when the work splits into genuinely independent branches and fails when the branches must share many decisions. Choose the simplest topology that isolates what you need, make briefs and results typed contracts with explicit scope and budgets, keep bulky material in an artifact store, bound fan-out and retries, and compare against a single-agent baseline at equal budget before you commit.