A multi-agent system is several model calls, each with its own instructions, context and tools, coordinated so that together they finish a task one call could not. The usual shape is an orchestrator that plans and delegates to workers, then combines what they return. The appeal is real: workers run in parallel, each starts with a clean context focused on one subproblem, and the orchestrator sees only compact results rather than every page a worker read.

The difficulty is that every boundary between agents is a lossy interface written in natural language. Most multi-agent failures are not model failures; they are bad briefs, missing contracts and synthesis steps that average away disagreement. This article is about those interfaces: the prompts and schemas that pass work between agents. For the runtime side, such as state stores, queues and retries, see agent orchestration architecture.

Advertisement

When several agents beat one

Splitting a task helps when it decomposes into parts that are independent, broad and context-heavy. Researching ten vendors, auditing twenty modules of a codebase, or checking a contract against five regulations all fit: each part needs a lot of reading, the reading for one part is irrelevant to the others, and the parts can run at the same time. A single agent would fill its context with material from part one while working on part seven.

It hurts when the parts are tightly coupled. Writing one function, editing one document or debugging one failure needs a shared, evolving picture; splitting it means each agent works from a stale snapshot of the others' decisions, and the merge step inherits the conflicts. It also costs more: every worker re-reads its brief and system prompt, and total token use for a multi-agent run is commonly several times that of a single agent. Use it when the extra breadth pays for that.

SignalSingle agentMulti-agent
Subtasks share state or edit the same artifactyesavoid
Subtasks are independent and each needs heavy readingcontext overflowsgood fit
Latency matters and parts can run in parallelserialgood fit
Budget is tightgood fitcostly

Topologies

Four shapes cover most systems. Orchestrator-worker: a lead agent plans, spawns workers with briefs, and synthesises; the default for research and audit tasks. Pipeline: fixed stages such as extract, transform and review, each one agent; simple and predictable, essentially prompt chaining with tools. Router: one classifier call picks a specialist agent; cheap, but only as good as the routing. Generator-evaluator loop: one agent drafts and another critiques against explicit criteria until the critique passes or a round limit is hit.

Orchestrator-worker: briefs go down, structured results come back, one synthesis step decidesOrchestratorplan, delegate, synthesiseWorker Aown context, own toolsWorker Bown context, own toolsWorker Cown context, own toolsbriefbriefbriefSchema validation + budget checkreject, retry or dropresult JSONSynthesis + verifierresolve conflicts, cite evidencegaps: new roundArtifacts (documents, files) live in a shared store and are passed by reference, not pasted into every prompt.
Orchestrator-worker with explicit contracts: briefs down, validated JSON up, synthesis and verification before anything reaches the user.

Start with the simplest topology that could work. A pipeline you can trace beats a dynamic swarm you cannot, and many systems that begin as orchestrator-worker turn out to have a fixed plan that a pipeline expresses better.

Advertisement

The orchestrator prompt

The orchestrator's job is to decide what work exists, how much of it, and who does which part. Its system prompt should say so explicitly, including rules for scaling effort. Without them, models tend to either spawn many workers for a simple question or do everything themselves.

You are the lead of a research team. You never do the research yourself.

1. Restate the user's goal in one sentence and list what a complete answer must contain.
2. Decide the effort:
   - one fact or definition: 1 worker, at most 5 tool calls
   - comparison of a few named items: 1 worker per item, at most 10 tool calls each
   - open survey: 4-6 workers with non-overlapping scopes
   Never exceed 6 workers in one round.
3. For each worker, write a brief using the BRIEF format exactly.
4. When results return, list which required parts are covered, which conflict,
   and which are missing. Start another round only for missing parts, at most 2 rounds.
5. Hand covered results to synthesis. Do not paraphrase worker findings from memory.

The explicit effort table and round limit are what keep cost bounded. The instruction not to paraphrase from memory matters too: an orchestrator that summarises worker output in its own words reintroduces the hallucinations the workers' sources were supposed to prevent.

The worker brief: a contract, not a request

A worker starts with nothing except its brief. It does not see the user's conversation, the other workers or the orchestrator's reasoning. Everything it needs must be in the brief, and the most common failure is a brief like "research competitor pricing" that leaves the worker to guess scope, depth and format. Two workers given vague briefs will overlap; one given a vague brief will stop early or wander.

Treat the brief as a typed contract with six fields: objective, scope boundaries, inputs, output schema, budget and stop conditions.

BRIEF
objective:   Find the list price per seat for Vendor B's team plan, as of this month.
why:         Part of a cost comparison of 4 vendors for a 200-seat team.
in_scope:    Vendor B's public pricing page and official docs. Annual and monthly billing.
out_of_scope: Other vendors (handled by other workers). Enterprise negotiated pricing.
inputs:      artifact://run42/vendor_list.json (entry "vendor_b")
output:      JSON matching schema worker_result_v1 (below). No prose outside the JSON.
budget:      at most 8 tool calls, 4,000 output tokens.
stop_when:   price found with a source URL, OR budget spent, OR source says pricing is unpublished.
if_unsure:   report status "partial" with what you found; never guess a number.

The why line is not decoration. A worker that knows the price feeds a 200-seat comparison will look for per-seat tiers and volume breaks; one that does not will return the headline number. The out-of-scope line prevents duplicate work. The stop conditions and the if unsure rule give the worker permission to return an honest partial result instead of padding.

Handoffs: schemas and references

What comes back must be machine-checkable. A free-text summary forces the orchestrator to re-interpret it, which is where facts drift. Require structured output with a fixed schema that separates claims from evidence and states the worker's own confidence and gaps.

{
  "status": "complete | partial | failed",
  "claims": [
    {"text": "Team plan is $X per seat per month, billed annually",
     "evidence": [{"source": "https://...", "quote": "exact sentence from the page"}],
     "confidence": "high | medium | low"}
  ],
  "gaps": ["Monthly billing price not published"],
  "artifacts": ["artifact://run42/vendor_b_pricing_page.md"],
  "tool_calls_used": 6
}

Two rules make this work. Every claim carries a verbatim quote, which makes verification cheap later. And large material, such as fetched pages, generated files or long tables, goes into a shared artifact store and is passed by reference. Pasting a worker's raw reading into the orchestrator's context defeats the point of having workers, and copying it through several agents is the telephone game in its purest form.

Worked example: the orchestration loop

Here is the control loop in Python. The model call is a placeholder for whichever provider you use; the structure is what matters. Workers run concurrently, every result is validated against the schema, invalid results get one repair attempt, and budgets are enforced by the code rather than trusted to the prompt.

import asyncio, json
from jsonschema import validate, ValidationError

MAX_WORKERS, MAX_ROUNDS = 6, 2

async def call_model(system: str, user: str, tools=None) -> str:
    ...  # provider SDK call; returns the final text of an agent run

async def run_worker(brief: dict) -> dict:
    raw = await call_model(WORKER_SYSTEM, json.dumps(brief), tools=brief["tools"])
    for attempt in range(2):
        try:
            result = json.loads(raw)
            validate(result, WORKER_RESULT_SCHEMA)
            return result
        except (json.JSONDecodeError, ValidationError) as e:
            if attempt == 1:
                break
            raw = await call_model(REPAIR_SYSTEM, f"Error: {e}\nFix this output:\n{raw}")
    return {"status": "failed", "claims": [], "gaps": [brief["objective"]], "artifacts": []}

async def orchestrate(goal: str) -> str:
    plan = json.loads(await call_model(ORCHESTRATOR_SYSTEM, goal))
    results = []
    for _ in range(MAX_ROUNDS):
        briefs = plan["briefs"][:MAX_WORKERS]           # enforce the cap in code
        results += await asyncio.gather(*(run_worker(b) for b in briefs))
        plan = json.loads(await call_model(
            ORCHESTRATOR_SYSTEM,
            json.dumps({"goal": goal, "results": results, "mode": "review"})))
        if not plan.get("briefs"):                       # no gaps left
            break
    return await call_model(SYNTHESIS_SYSTEM, json.dumps({"goal": goal, "results": results}))

In the vendor comparison, the orchestrator writes four briefs, one per vendor. Three workers return complete results; one returns partial with a gap for monthly pricing. The review step issues one more brief for that gap only, and the second round closes it. Synthesis then receives sixteen claims, each with a quote, rather than four essays.

The synthesis prompt

Synthesis is where multi-agent systems quietly lose quality. Given conflicting results, a model's default is to blend them into something plausible. Tell it not to.

You combine findings from several researchers into one answer.
- Use only the claims provided. Every sentence with a fact must cite a claim's source.
- If two claims conflict, show both with their sources and say which is more
  recent or more authoritative, and why. Do not average numbers.
- Claims with confidence "low" may be mentioned only as unconfirmed.
- List remaining gaps explicitly at the end.

Conflict surfacing is the single most valuable instruction here. A disagreement between workers is information: one source is stale, or the question was ambiguous. Averaging it away produces confident, wrong answers that no single worker would have given.

Verification

Add a verifier when errors are expensive. It is a separate call with a narrow job: for each claim in the draft, check that the cited quote exists in the artifact and supports the claim. Because quotes were captured at the source, this is a lookup and a judgement, not new research, and it can run on a smaller model. A generator-evaluator loop goes further, sending specific failures back to the writer for a bounded number of rounds. See prompt evaluation for building the test set that tells you whether the verifier actually catches errors.

Failure modes

  • Overlapping workers. Vague scopes make two workers research the same thing. Fix with explicit out-of-scope lines and an orchestrator check that briefs do not overlap.
  • Runaway spawning. The orchestrator keeps finding gaps and launching rounds. Cap workers and rounds in code, not only in the prompt.
  • Telephone-game drift. Facts mutate as each agent paraphrases the last. Require quotes, pass artifacts by reference, and forbid paraphrasing from memory.
  • Injection propagation. A worker reads a web page containing instructions and returns them inside a claim, and the orchestrator follows them. Treat worker output as data, never as instructions, and give workers the least privilege they need; see prompt injection defence.
  • Silent failure. A worker fails and the synthesis omits its part without saying so. Make status and gaps mandatory and surface them in the final answer.
  • Unbounded cost. Several workers, each with many tool calls, multiply spend. Log tokens per agent and per run, and alert on outliers.

Evaluating the system

Judge the end result and the trace. End-to-end, score final answers against a set of tasks with known answers, including tasks where the correct output is a reported gap. Per trace, check that briefs were non-overlapping, that worker results validated, that every final claim traces to a quote, and that cost stayed inside budget. Compare against a single-agent baseline on the same tasks; if the multi-agent version is not clearly better on your metric, the extra cost is not justified.

What to do next

  1. Pick one task and run it with a single agent first to get a quality and cost baseline.
  2. Write the worker brief format and the result schema before writing any orchestrator prompt.
  3. Add the effort-scaling table and hard worker and round caps, enforced in code.
  4. Require verbatim quotes for every claim and store large material as artifacts passed by reference.
  5. Write a synthesis prompt that surfaces conflicts and gaps instead of blending them.
  6. Build an evaluation set of 20 to 50 tasks, including some whose right answer is a gap, and compare against the baseline.
  7. Log every brief, result and token count per run so failures can be traced to a specific interface.
Key takeaway: Multi-agent orchestration pays off when a task splits into independent, reading-heavy parts that can run in parallel, and costs several times more tokens than one agent. Quality depends on the interfaces: an orchestrator prompt with explicit effort rules and caps, worker briefs written as contracts with scope, schema, budget and stop conditions, structured results that carry verbatim evidence, artifacts passed by reference, and a synthesis step that surfaces conflicts and gaps. Enforce limits in code and evaluate against a single-agent baseline.