Agent2Agent (A2A) is an open protocol that lets one AI agent hand work to another agent that it does not control: built by another team, running on another stack, possibly owned by another company. It defines how a client agent discovers what a remote agent can do, sends it a request, tracks the resulting work as a task, receives results as artifacts, and gets told about progress without polling forever. This overview explains the mental model, names every core object and operation in A2A 1.0, walks one delegation with working client code, and ends with failure modes and a decision guide.
Facts here were checked against the A2A specification, released version 1.0.0, on 2026-10-01. If you learned A2A from 0.3-era material, expect renamed methods (SendMessage rather than slash-separated names), parts without a kind field, and enum values written like TASK_STATE_COMPLETED.
The problem A2A solves
Inside one process, agents cooperate by sharing memory. Across organisational boundaries that fails: the receipts team will not expose its prompts, model or tool credentials to the expense team. What both sides need is a contract at the boundary: how to ask, how to follow progress, how to get structured results back, how to authenticate, and what errors mean.
A2A's central design choice is that the remote agent is opaque. The client sees an advertised description, a task with a status, messages and artifacts. It never sees how the work is done. That makes A2A closer to a service protocol than to a framework: the remote side could be a single LLM call, a multi-agent graph, or a human queue behind an agent facade, and the client code does not change.
Opacity has consequences. Work can be long-running, so there are tasks with states, not just responses; interactive, so a task can pause for input; and multi-modal, so results are typed parts rather than strings.
A2A and MCP are different layers
The question everyone asks first is how A2A relates to the Model Context Protocol. They solve different problems and are designed to be used together. MCP connects an agent to its tools and context: a server exposes functions, resources and prompts, and the model decides when to call them. The calling agent is in charge and the tool is passive. A2A connects an agent to another agent: the remote side has its own reasoning, may take minutes, may ask questions back, and decides for itself how to do the work.
| Question | MCP | A2A |
|---|---|---|
| Who is on the other side? | A tool or data source | An autonomous agent |
| Unit of work | One tool call with arguments | A task with a lifecycle |
| Can the callee ask questions back? | Limited (elicitation) | Yes: input-required state |
| Typical duration | Milliseconds to seconds | Seconds to hours |
| Discovery | Client configuration | Agent Card at a well-known URL or registry |
A common shape is an orchestrator that uses MCP for its own tools and A2A to delegate to specialist agents, each of which uses MCP internally. The orchestration patterns article covers composing many remote agents.
The architecture at a glance
There are two roles. The client (often itself an agent) initiates; the server (the remote agent) exposes an A2A endpoint. The server publishes an Agent Card, a JSON document at https://{domain}/.well-known/agent-card.json or in a registry, describing its name, skills, supported interfaces, capabilities, accepted media types and security schemes. Some agents also offer an authenticated extended card with more detail, fetched with GetExtendedAgentCard.
Requests travel over one of three protocol bindings that carry the same abstract operations: JSON-RPC 2.0 over HTTP, gRPC, and HTTP+JSON (REST-style paths such as /message:send). The card's supportedInterfaces list says which bindings exist at which URLs; the client picks one it supports. Every request carries an A2A-Version header, and the specification says agents must interpret an empty value as 0.3, which produces confusing errors from a 1.0 server if you forget it.
Authentication is not reinvented. The card declares standard schemes (API keys, HTTP bearer, OAuth 2.0, OpenID Connect, mutual TLS) and the client obtains credentials out of band and sends them as ordinary HTTP headers. See A2A security for the threat model.
The five core objects
| Object | What it is | Key fields |
|---|---|---|
| Agent Card | The agent's public self-description | name, skills, supportedInterfaces, capabilities, securitySchemes |
| Message | One turn of communication from a user or agent | messageId, role (ROLE_USER/ROLE_AGENT), parts, optional taskId/contextId |
| Part | One piece of content inside a message or artifact | exactly one of text, raw, url, data; optional mediaType, filename |
| Task | The stateful unit of work the server creates | id, contextId, status (state, message, timestamp), artifacts, history |
| Artifact | An output the task produced | artifactId, name, parts |
Two identifiers do most of the organising. A task id names one unit of work. A context id groups related tasks and messages into one conversation, so a correction after a task finishes becomes a new task in the same context rather than a mutation of a finished one. Messages carry their own messageId, generated by the sender, which lets servers recognise a redelivered message.
Messages are conversation; artifacts are deliverables. Store artifacts as results and treat messages as transient. Field-level detail lives in the Agent Card specification article and the Task article.
The eleven operations
| Group | Operation | Purpose |
|---|---|---|
| Messaging | SendMessage | Send a message; the result holds exactly one of a Message or a Task |
| Messaging | SendStreamingMessage | Same, but the server streams task status and artifact updates (SSE on HTTP bindings) |
| Task management | GetTask | Read a task's current state, artifacts and optionally history |
| Task management | ListTasks | List tasks, for example within a context |
| Task management | CancelTask | Request cancellation; fails with TaskNotCancelableError on a finished task |
| Task management | SubscribeToTask | Reattach to the event stream of an existing task |
| Push | CreateTaskPushNotificationConfig | Register a webhook for a task |
| Push | GetTaskPushNotificationConfig, ListTaskPushNotificationConfigs, DeleteTaskPushNotificationConfig | Inspect and remove webhook registrations |
| Discovery | GetExtendedAgentCard | Fetch the authenticated, more detailed card |
Optional operations are gated by the card. If capabilities.streaming is not declared, calling the streaming method fails with UnsupportedOperationError; push operations on an agent without push support fail with PushNotificationNotSupportedError. Check the card before you call, and cache it with a sensible expiry so a capability change is picked up.
A2A also defines its own error types, such as TaskNotFoundError, ContentTypeNotSupportedError and VersionNotSupportedError. Map each to a client action (retry, new task, fix the request, give up) rather than logging generic failures.
The task state machine
Every task is in one of eight states (plus an unspecified zero value). They fall into three groups, and the grouping matters more than the names.
- Active:
TASK_STATE_SUBMITTED(accepted, not started) andTASK_STATE_WORKING. - Interrupted:
TASK_STATE_INPUT_REQUIRED(the agent needs more information; the status message says what) andTASK_STATE_AUTH_REQUIRED(the agent needs additional credentials). The task is alive and holding resources; the client continues it by sending a new message with the sametaskIdandcontextId. - Terminal:
TASK_STATE_COMPLETED,TASK_STATE_FAILED,TASK_STATE_CANCELEDandTASK_STATE_REJECTED. Nothing moves a terminal task again; a follow-up needs a new task, optionally pointing back withreferenceTaskIds.
By default SendMessage blocks until the task reaches a terminal or interrupted state. Setting returnImmediately in the request configuration makes it return the task as soon as it exists, which you then follow by polling, streaming or push. Choose by expected duration: blocking for seconds, streaming when a user is watching, push for minutes to hours or when the client cannot hold a connection open. The annotated examples show each pattern on the wire.
Worked example: a minimal client
You can speak A2A with nothing but an HTTP library, which is the best way to learn it before adopting an SDK. The client below fetches the card, selects the JSON-RPC interface, sends messages and waits for a task with capped exponential backoff, cancelling it if a deadline passes.
import uuid, time, httpx
A2A_HEADERS = {"A2A-Version": "1.0", "Content-Type": "application/json"}
class A2AError(Exception):
pass
class A2AClient:
def __init__(self, base, token):
self.http = httpx.Client(timeout=60, headers={**A2A_HEADERS, "Authorization": f"Bearer {token}"})
card = self.http.get(f"{base}/.well-known/agent-card.json").raise_for_status().json()
jsonrpc = [i for i in card["supportedInterfaces"] if i["protocolBinding"] == "JSONRPC"]
if not jsonrpc:
raise A2AError("agent offers no JSON-RPC interface")
self.url, self.card, self.next_id = jsonrpc[0]["url"], card, 0
def call(self, method, params):
self.next_id += 1
r = self.http.post(self.url, json={"jsonrpc": "2.0", "id": self.next_id,
"method": method, "params": params})
body = r.raise_for_status().json()
if "error" in body:
raise A2AError(f"{method}: {body['error']}")
return body["result"]
def send(self, text, task_id=None, context_id=None, background=False):
msg = {"messageId": str(uuid.uuid4()), "role": "ROLE_USER", "parts": [{"text": text}]}
if task_id:
msg["taskId"] = task_id
if context_id:
msg["contextId"] = context_id
params = {"message": msg}
if background:
params["configuration"] = {"returnImmediately": True}
return self.call("SendMessage", params) # holds exactly one of "message" or "task"
def wait(self, task_id, deadline_s=300):
delay, start = 1.0, time.monotonic()
while time.monotonic() - start < deadline_s:
task = self.call("GetTask", {"id": task_id})
state = task["status"]["state"]
if state in TERMINAL or state in INTERRUPTED:
return task
time.sleep(delay)
delay = min(delay * 2, 15) # back off; never hammer a busy agent
self.call("CancelTask", {"id": task_id})
raise A2AError(f"task {task_id} exceeded {deadline_s}s and was cancelled")
TERMINAL = {"TASK_STATE_COMPLETED", "TASK_STATE_FAILED", "TASK_STATE_CANCELED", "TASK_STATE_REJECTED"}
INTERRUPTED = {"TASK_STATE_INPUT_REQUIRED", "TASK_STATE_AUTH_REQUIRED"}Using it for the receipts scenario looks like this. Note that the code handles both result shapes, because the server, not the client, decides whether a request deserves a task.
client = A2AClient("https://receipts.example.com", token=get_token())
result = client.send("Extract this receipt total", background=True)
if "message" in result: # agent answered directly, nothing to track
print(result["message"]["parts"])
else:
task = client.wait(result["task"]["id"])
if task["status"]["state"] == "TASK_STATE_INPUT_REQUIRED":
question = task["status"]["message"]["parts"][0]["text"]
follow = client.send(ask_user(question), task_id=task["id"], context_id=task["contextId"])
elif task["status"]["state"] == "TASK_STATE_COMPLETED":
for art in task.get("artifacts", []):
store(art["artifactId"], art["parts"])Three habits carry over to production: a fresh messageId for every new turn; taskId and contextId persisted before waiting, so a restart resumes with GetTask instead of orphaning remote work; and a deadline with cancellation, because abandoned tasks still cost the remote agent money.
What the server must get right
On the server side the protocol layer should own the task store and the state machine, and the agent logic should only emit events. That separation keeps illegal transitions, such as completing a cancelled task, out of reach of model output.
# Server skeleton: the protocol layer owns the task store and state machine,
# the agent logic only produces events. Pseudocode; any web framework works.
ALLOWED = {
"SUBMITTED": {"WORKING", "REJECTED", "CANCELED", "FAILED"},
"WORKING": {"COMPLETED", "FAILED", "CANCELED", "INPUT_REQUIRED", "AUTH_REQUIRED"},
"INPUT_REQUIRED": {"WORKING", "CANCELED", "FAILED"},
"AUTH_REQUIRED": {"WORKING", "CANCELED", "FAILED"},
}
def transition(task, new_state, message=None):
old = task.state
if new_state not in ALLOWED.get(old, set()):
raise IllegalTransition(old, new_state) # terminal states have no exits
task.state, task.status_message = new_state, message
store.save(task) # persist before notifying anyone
events.publish(task.id, status_update(task)) # SSE subscribers and push configs
def on_send_message(req, principal):
prior = store.task_for_message(req.message.message_id, owner=principal)
if prior: # redelivery: return the original task
return prior
if req.message.task_id:
task = store.get(req.message.task_id, owner=principal) # TaskNotFoundError if absent
if task.state in TERMINAL:
raise UnsupportedOperationError("task is terminal; start a new task")
else:
task = store.create(context_id=req.message.context_id or new_id(), owner=principal)
store.record_message(req.message.message_id, task.id)
worker.enqueue(task.id, req.message)
return task if req.configuration.return_immediately else wait_terminal_or_interrupted(task)Persist before you notify, or a subscriber's next GetTask reads stale data. Scope every lookup by the authenticated principal so tenants cannot read each other's tasks. Bound how long interrupted tasks may wait so abandoned conversations do not pin workers.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
Missing A2A-Version header | 1.0 server rejects or misparses requests as 0.3 | Set the header in one place in the HTTP client |
| Client assumes a Task is always returned | Crash on a direct Message reply | Branch on which result field is present |
| Orphaned remote tasks | Remote costs grow after client restarts | Persist task ids first; cancel on deadline |
| Interrupted task treated as failure | Users never see the agent's question | Route input-required to a human or to client logic |
| Streaming connection drop | Client misses artifact chunks | Reattach with SubscribeToTask, reconcile with GetTask |
| Unverified webhook | Forged push updates accepted | Authenticate pushes; re-read the task with GetTask |
The cross-cutting lesson is that the authoritative state of a task lives on the server. Streams and webhooks are fast paths for learning about changes; GetTask is the reconciliation path, and every client should be able to fall back to it.
When A2A is the right tool, and when it is not
A2A earns its overhead when the callee is genuinely an agent owned by someone else: work is long or interactive, results are multi-part, and you want discovery and authentication that other teams already understand. It is a poor fit for a deterministic internal function (use a plain API or MCP), for latency-critical inner loops where a task object per call is wasted state, or for agents inside one process that can share memory directly.
Start with one high-value delegation using blocking calls, and add streaming and push only when durations demand them.
What to do next
- Fetch the Agent Card of one agent you depend on, or write one for your own agent, and check the path is
/.well-known/agent-card.json. - Implement the minimal client above against a test agent, including the Message-or-Task branch and the input-required follow-up.
- Decide per delegation whether to block, stream or use push, based on expected duration and whether a user is watching.
- Persist task and context ids before waiting, set a deadline, and cancel abandoned tasks.
- Map each A2A error type to a concrete client action and alert on the unexpected ones.
- Read the security article and confirm authentication, tenant scoping and webhook verification before exposing an agent outside your network.