The Agent2Agent (A2A) protocol defines how one agent sends work to another: discovery through an Agent Card, messages and tasks with a defined lifecycle, streaming, push notifications and cancellation. It deliberately does not define orchestration. Nothing in the specification says how to split a goal into steps, which agent should do each one, or what to do when the third of five agents fails. That is the job of an orchestrator, and it is where most of the engineering in a multi-agent system actually lives.
This page is about building that orchestrator on top of A2A. It covers the architecture, how plan steps map onto A2A tasks and contexts, the common patterns, how to choose an interaction mode for each call, a durable dispatch loop in code, and the failure modes that appear only when agents are remote, opaque and slow. Method and state names follow the A2A 1.0 specification. For orchestration inside a single process and framework choice, see multi-agent orchestration architecture.
Where orchestration sits in A2A
A2A has two roles per interaction: a client that sends a message and a remote agent that processes it. The remote agent is opaque: the client sees its Agent Card (name, skills, supported interfaces, capabilities such as streaming and push notifications, and security requirements), not its prompts, tools or memory. The protocol surface is small: SendMessage, SendStreamingMessage, GetTask, ListTasks, CancelTask, SubscribeToTask, and operations to create, get, list and delete task push-notification configurations.
An orchestrator is simply an agent that is a server to its caller and a client to several specialists. Upstream, it exposes its own Agent Card and accepts one task that represents the whole job. Downstream, it creates one or more A2A tasks per step. Because the specialists are opaque, the orchestrator cannot reach into them; everything it knows comes from task states, status messages and artifacts. That constraint shapes every design decision below. The protocol fundamentals are covered in the A2A protocol architecture.
Architecture of an A2A orchestrator
Four components are worth separating. The planner turns the caller's goal into a graph of steps with dependencies; it may be an LLM, a fixed workflow or a mix, but its output should be data, not free text. The card registry caches Agent Cards and answers "which agent can do this skill, over which interface, with which auth", and it refreshes cards on a schedule because skills and endpoints change. The task ledger is a durable store keyed by step, holding the remote agent, task id, context id, last known state and artifacts; it is what lets the orchestrator restart without losing or duplicating work. The dispatcher sends messages, tracks deadlines, retries, cancels, and converts remote outcomes into the orchestrator's own task state for its caller.
Keeping the ledger separate from the planner is the most important choice. LLM planners are non-deterministic; if a restart re-plans from scratch, it may create a different set of remote tasks and orphan the old ones. Persist the plan once, then execute it.
Mapping a plan onto tasks and contexts
A2A gives three identifiers to work with. A task is created by the remote agent in response to a message and has an id and a status. A context (contextId) groups related tasks and messages into one conversation on the remote side. A message may carry referenceTaskIds to point at earlier tasks it builds on. Every message also has a client-chosen messageId.
A workable mapping is: one context per step per agent, one task per attempt, and referenceTaskIds when a follow-up refines an earlier result on the same agent. Do not reuse one context across unrelated jobs: remote agents may keep conversation memory per context, and a shared context leaks one customer's material into another's answers. Make messageId deterministic, derived from job, step and attempt, so a resend after a lost response can be recognised; whether a given agent deduplicates on it is up to that agent, which is why the ledger also records task ids as soon as they arrive. The broader treatment is in A2A idempotency.
Task states drive the dispatcher. TASK_STATE_SUBMITTED and TASK_STATE_WORKING mean wait. TASK_STATE_COMPLETED, TASK_STATE_FAILED, TASK_STATE_CANCELED and TASK_STATE_REJECTED are terminal. TASK_STATE_INPUT_REQUIRED and TASK_STATE_AUTH_REQUIRED are interrupted states: the remote agent is paused waiting for the client. Rejected deserves its own handling, since it means the agent declined the work and retrying the same agent is pointless. The full lifecycle is in A2A task state.
The core patterns
| Pattern | Shape | Use when | Watch out for |
|---|---|---|---|
| Sequential pipeline | A then B then C, each output feeding the next | steps are genuinely dependent | latency adds up; one slow agent stalls all |
| Parallel fan-out, fan-in | independent steps sent together, joined by a synthesis step | research, checks and lookups that do not depend on each other | partial failure; the join needs a policy for missing inputs |
| Router | a classifier picks one specialist per request | many narrow agents behind one entry point | misroutes; keep a fallback generalist |
| Hierarchical delegation | a specialist is itself an orchestrator | large domains with their own sub-teams | compounding latency, hidden fan-out, deadline propagation |
| Human-in-the-loop | INPUT_REQUIRED surfaces to the caller or an operator | approvals, missing facts, ambiguous requests | tasks parked forever; set an expiry |
| Critic loop | a reviewer agent grades a draft; the writer revises | quality matters more than latency | unbounded loops; cap iterations and cost |
Most real orchestrations are a fan-out layer followed by a sequential tail. The planner's graph makes this explicit: each topological layer is a set of independent steps that can run in parallel, and each later layer starts only when its inputs exist.
Choosing an interaction mode per step
A2A offers four ways to wait for a result, and the right one depends on how long the step takes and what the orchestrator must do meanwhile.
| Mode | How | Good for | Cost |
|---|---|---|---|
| Blocking | SendMessage with the default configuration; the call returns at a terminal or interrupted state | steps of a few seconds | holds a connection; a dropped connection loses the result, not the task |
| Return immediately + poll | returnImmediately true, then GetTask with backoff | minutes-long steps, simple infrastructure | polling load and latency of the poll interval |
| Streaming | SendStreamingMessage, or SubscribeToTask to reattach | progress to show a user, partial artifacts | long-lived connections; must handle reconnect |
| Push notifications | create a push-notification config with a webhook URL and token | hours-long or human-gated steps | a public, authenticated webhook endpoint and verification of every callback |
A robust default is return-immediately plus push where the agent's card advertises push support, with polling as the fallback and as the source of truth: a webhook tells you to look, and GetTask tells you what is there. Treat push callbacks as untrusted input until verified, as described in A2A push notifications.
The dispatch loop in code
Below is a request that creates a finance step without waiting, followed by the orchestrator's core loop in Python-style pseudocode. The JSON uses the 1.0 method name and camelCase field names of the JSON mapping; check exact field names against the SDK you use, since early SDKs still target 0.x names such as message/send.
{
"jsonrpc": "2.0",
"id": "rpc-7",
"method": "SendMessage",
"params": {
"message": {
"messageId": "dd-4411:finance:attempt-1",
"role": "ROLE_USER",
"contextId": "ctx-dd-4411-finance",
"parts": [{"text": "Summarise Acme Ltd's last three filed annual accounts. Return JSON."}]
},
"configuration": {"returnImmediately": true}
}
}TERMINAL = {"TASK_STATE_COMPLETED", "TASK_STATE_FAILED", "TASK_STATE_CANCELED", "TASK_STATE_REJECTED"}
INTERRUPTED = {"TASK_STATE_INPUT_REQUIRED", "TASK_STATE_AUTH_REQUIRED"}
async def run_step(step, ledger, client, deadline):
rec = ledger.get(step.id) # durable: survives orchestrator restart
if rec is None:
msg_id = f"{step.job_id}:{step.id}:attempt-1" # deterministic -> safe to resend
task = await client.send_message(step.agent_url, msg_id, step.prompt,
context_id=step.context_id,
reference_task_ids=step.upstream_task_ids(ledger),
return_immediately=True)
rec = ledger.put(step.id, agent=step.agent_url, task_id=task["id"],
state=task["status"]["state"])
backoff = 1.0
while rec.state not in TERMINAL:
if now() > deadline:
await client.cancel_task(rec.agent, rec.task_id) # best effort
ledger.update(step.id, state="TIMED_OUT")
raise StepTimeout(step.id)
if rec.state in INTERRUPTED:
return await escalate(step, rec) # ask caller/human; never guess answers
await wait_for_push_or(backoff) # webhook wakes us early if configured
task = await client.get_task(rec.agent, rec.task_id)
rec = ledger.update(step.id, state=task["status"]["state"], task=task)
backoff = min(backoff * 2, 30.0)
if rec.state != "TASK_STATE_COMPLETED":
raise StepFailed(step.id, rec.state)
return validate_artifacts(step, rec.task["artifacts"]) # schema check before use
async def run_plan(plan, ledger, client):
for layer in plan.topological_layers(): # steps in a layer are independent
results = await asyncio.gather(
*(run_step(s, ledger, client, s.deadline()) for s in layer),
return_exceptions=True)
for s, r in zip(layer, results):
if isinstance(r, Exception):
plan.apply_policy(s, r) # retry, substitute agent, degrade or abortFour properties make this loop safe. The ledger write happens as soon as a task id is known, so a crash after that point resumes by polling rather than resending. Deadlines are per step and derived from the caller's deadline, so a slow specialist cannot silently consume the whole budget. Interrupted states escalate instead of being answered by the orchestrator's own model, which would invent facts. And artifacts are schema-validated before any downstream step consumes them, because a completed task is not the same as a correct result.
Worked example: a supplier due-diligence report
A procurement tool asks the orchestrator for a due-diligence report on a supplier, with a 10-minute deadline. The planner produces two layers. Layer one runs three independent steps: a research agent that gathers news and ownership information (streams progress, typically 90 seconds), a finance agent that analyses filed accounts (blocking-capable but slow, typically 3 minutes, so it is called with return-immediately and polled), and a legal agent that checks sanctions and litigation databases (supports push, typically 2 minutes but sometimes needs a human analyst). Layer two is a writer agent that takes the three artifacts, referenced by task id, and produces the report.
Budgeting: the writer needs about 60 seconds and the orchestrator reserves 60 seconds for validation and retries, so layer one gets 8 minutes. Each step's deadline is 8 minutes from start. Layer one's wall-clock time is the slowest step, about 3 minutes, rather than the 6.5-minute sum a sequential pipeline would take.
Now the legal agent returns TASK_STATE_INPUT_REQUIRED asking whether a similarly named subsidiary should be included. The orchestrator does not guess. It moves its own upstream task to input-required with the question, the procurement user answers, and the orchestrator sends the answer as a new message with the same taskId and contextId. Meanwhile the research and finance results sit in the ledger. If the user never answers, the step expires at its deadline, the orchestrator cancels the legal task and the policy decides: this job marks the legal section "not completed" and still delivers the report, because partial due diligence with an explicit gap is more useful than none. A different job, such as a payment approval, would abort instead. That policy is a business decision written into the plan, not something the orchestrator should improvise.
Failure modes
- Lost response, duplicate task. The send succeeds remotely but the response is lost; a blind retry creates a second task. Use deterministic messageIds, record task ids immediately, and list tasks in the context before resending.
- Orphaned tasks. The orchestrator times out or crashes and the remote task keeps running and billing. Always cancel on deadline, and sweep the ledger for non-terminal tasks older than their deadline.
- Parked interrupts. Tasks waiting in INPUT_REQUIRED or AUTH_REQUIRED forever. Give every interrupt an expiry and an owner.
- Deadline inversion. A child agent with a 30-minute internal timeout under a parent with 10 minutes. Propagate remaining time to every hop, in the message text or metadata the agent understands.
- Retry amplification. Retries at every level of a hierarchy multiply load on a struggling agent. Retry at one level, with a budget, and honour rejection.
- Context leakage and confused deputy. Reused contexts leak data between jobs; forwarding the caller's full credentials lets a specialist act beyond its role. Use per-job contexts and per-agent, least-privilege credentials.
- Trusting artifacts. A completed task can carry malformed, oversized or prompt-injecting content. Validate schemas and treat artifact text as data, never as instructions to the orchestrator.
Operations and trade-offs
Trace every job end to end: propagate a trace context on every A2A call and record the job id, step id, agent, task id and state transitions, so "why did this take nine minutes" has an answer. Track per-agent success rate, rejection rate, latency percentiles, interrupt rate and cost; a card registry that includes these observed numbers lets the router prefer healthy agents. Refresh cards, and pin to the interface version you tested against.
The central trade-off is orchestration versus choreography. A central orchestrator gives one place that knows the plan, the deadlines and the state of every step, which makes failures explainable and policies enforceable. The cost is a component that must be highly available and a potential bottleneck. Choreography, where agents hand off to each other directly, removes the centre but makes deadlines, cancellation and auditing much harder, and with opaque third-party agents it is rarely acceptable. Most teams should orchestrate, keep the orchestrator stateless apart from its durable ledger, and scale it horizontally.
What to do next
- Inventory the agents you call, read their Agent Cards, and record which support streaming and push and which security schemes they require.
- Separate planning from execution: persist the plan as a graph before dispatching anything.
- Add a durable task ledger that records the task id the moment it is returned, and make messageIds deterministic.
- Give every step a deadline derived from the caller's, and cancel remote tasks when it passes.
- Route INPUT_REQUIRED and AUTH_REQUIRED to a person or the caller, with an expiry, never to the orchestrator's own model.
- Validate every artifact against a schema before a downstream step reads it, and write down a partial-failure policy per plan type.