Agent2Agent (A2A) is an open protocol that lets one AI agent hand work to another agent that it does not control: built by another team, running on another stack, possibly owned by another company. It defines how a client agent discovers what a remote agent can do, sends it a request, tracks the resulting work as a task, receives results as artifacts, and gets told about progress without polling forever. This overview explains the mental model, names every core object and operation in A2A 1.0, walks one delegation with working client code, and ends with failure modes and a decision guide.

Facts here were checked against the A2A specification, released version 1.0.0, on 2026-10-01. If you learned A2A from 0.3-era material, expect renamed methods (SendMessage rather than slash-separated names), parts without a kind field, and enum values written like TASK_STATE_COMPLETED.

Advertisement

The problem A2A solves

Inside one process, agents cooperate by sharing memory. Across organisational boundaries that fails: the receipts team will not expose its prompts, model or tool credentials to the expense team. What both sides need is a contract at the boundary: how to ask, how to follow progress, how to get structured results back, how to authenticate, and what errors mean.

A2A's central design choice is that the remote agent is opaque. The client sees an advertised description, a task with a status, messages and artifacts. It never sees how the work is done. That makes A2A closer to a service protocol than to a framework: the remote side could be a single LLM call, a multi-agent graph, or a human queue behind an agent facade, and the client code does not change.

Opacity has consequences. Work can be long-running, so there are tasks with states, not just responses; interactive, so a task can pause for input; and multi-modal, so results are typed parts rather than strings.

A2A and MCP are different layers

The question everyone asks first is how A2A relates to the Model Context Protocol. They solve different problems and are designed to be used together. MCP connects an agent to its tools and context: a server exposes functions, resources and prompts, and the model decides when to call them. The calling agent is in charge and the tool is passive. A2A connects an agent to another agent: the remote side has its own reasoning, may take minutes, may ask questions back, and decides for itself how to do the work.

QuestionMCPA2A
Who is on the other side?A tool or data sourceAn autonomous agent
Unit of workOne tool call with argumentsA task with a lifecycle
Can the callee ask questions back?Limited (elicitation)Yes: input-required state
Typical durationMilliseconds to secondsSeconds to hours
DiscoveryClient configurationAgent Card at a well-known URL or registry

A common shape is an orchestrator that uses MCP for its own tools and A2A to delegate to specialist agents, each of which uses MCP internally. The orchestration patterns article covers composing many remote agents.

Advertisement

The architecture at a glance

A2A 1.0: one client agent delegating to one opaque remote agentClient agentyour orchestrator / appRemote agent (server)opaque: no shared memoryAgent Card/.well-known/agent-card.jsonBindingJSON-RPC / gRPC / HTTP+JSONTask storeid, contextId, status, artifactsAgent logicLLM, tools, MCP serversWebhook receiverpush notificationsClient-side statetaskId, contextId per job1. fetch card, check capabilities2. SendMessage3. push: StreamResponse4. Task / SSE eventsEvery request carries A2A-Version: 1.0; auth is ordinary HTTP auth declared in the card.The client sees tasks, messages and artifacts, never the remote agent's prompts or tools.
The client fetches the Agent Card, chooses a binding, sends a message and follows the resulting task by response, stream or push. Everything to the right of the binding is the remote agent's private business.

There are two roles. The client (often itself an agent) initiates; the server (the remote agent) exposes an A2A endpoint. The server publishes an Agent Card, a JSON document at https://{domain}/.well-known/agent-card.json or in a registry, describing its name, skills, supported interfaces, capabilities, accepted media types and security schemes. Some agents also offer an authenticated extended card with more detail, fetched with GetExtendedAgentCard.

Requests travel over one of three protocol bindings that carry the same abstract operations: JSON-RPC 2.0 over HTTP, gRPC, and HTTP+JSON (REST-style paths such as /message:send). The card's supportedInterfaces list says which bindings exist at which URLs; the client picks one it supports. Every request carries an A2A-Version header, and the specification says agents must interpret an empty value as 0.3, which produces confusing errors from a 1.0 server if you forget it.

Authentication is not reinvented. The card declares standard schemes (API keys, HTTP bearer, OAuth 2.0, OpenID Connect, mutual TLS) and the client obtains credentials out of band and sends them as ordinary HTTP headers. See A2A security for the threat model.

The five core objects

ObjectWhat it isKey fields
Agent CardThe agent's public self-descriptionname, skills, supportedInterfaces, capabilities, securitySchemes
MessageOne turn of communication from a user or agentmessageId, role (ROLE_USER/ROLE_AGENT), parts, optional taskId/contextId
PartOne piece of content inside a message or artifactexactly one of text, raw, url, data; optional mediaType, filename
TaskThe stateful unit of work the server createsid, contextId, status (state, message, timestamp), artifacts, history
ArtifactAn output the task producedartifactId, name, parts

Two identifiers do most of the organising. A task id names one unit of work. A context id groups related tasks and messages into one conversation, so a correction after a task finishes becomes a new task in the same context rather than a mutation of a finished one. Messages carry their own messageId, generated by the sender, which lets servers recognise a redelivered message.

Messages are conversation; artifacts are deliverables. Store artifacts as results and treat messages as transient. Field-level detail lives in the Agent Card specification article and the Task article.

The eleven operations

GroupOperationPurpose
MessagingSendMessageSend a message; the result holds exactly one of a Message or a Task
MessagingSendStreamingMessageSame, but the server streams task status and artifact updates (SSE on HTTP bindings)
Task managementGetTaskRead a task's current state, artifacts and optionally history
Task managementListTasksList tasks, for example within a context
Task managementCancelTaskRequest cancellation; fails with TaskNotCancelableError on a finished task
Task managementSubscribeToTaskReattach to the event stream of an existing task
PushCreateTaskPushNotificationConfigRegister a webhook for a task
PushGetTaskPushNotificationConfig, ListTaskPushNotificationConfigs, DeleteTaskPushNotificationConfigInspect and remove webhook registrations
DiscoveryGetExtendedAgentCardFetch the authenticated, more detailed card

Optional operations are gated by the card. If capabilities.streaming is not declared, calling the streaming method fails with UnsupportedOperationError; push operations on an agent without push support fail with PushNotificationNotSupportedError. Check the card before you call, and cache it with a sensible expiry so a capability change is picked up.

A2A also defines its own error types, such as TaskNotFoundError, ContentTypeNotSupportedError and VersionNotSupportedError. Map each to a client action (retry, new task, fix the request, give up) rather than logging generic failures.

The task state machine

Every task is in one of eight states (plus an unspecified zero value). They fall into three groups, and the grouping matters more than the names.

  • Active: TASK_STATE_SUBMITTED (accepted, not started) and TASK_STATE_WORKING.
  • Interrupted: TASK_STATE_INPUT_REQUIRED (the agent needs more information; the status message says what) and TASK_STATE_AUTH_REQUIRED (the agent needs additional credentials). The task is alive and holding resources; the client continues it by sending a new message with the same taskId and contextId.
  • Terminal: TASK_STATE_COMPLETED, TASK_STATE_FAILED, TASK_STATE_CANCELED and TASK_STATE_REJECTED. Nothing moves a terminal task again; a follow-up needs a new task, optionally pointing back with referenceTaskIds.

By default SendMessage blocks until the task reaches a terminal or interrupted state. Setting returnImmediately in the request configuration makes it return the task as soon as it exists, which you then follow by polling, streaming or push. Choose by expected duration: blocking for seconds, streaming when a user is watching, push for minutes to hours or when the client cannot hold a connection open. The annotated examples show each pattern on the wire.

Worked example: a minimal client

You can speak A2A with nothing but an HTTP library, which is the best way to learn it before adopting an SDK. The client below fetches the card, selects the JSON-RPC interface, sends messages and waits for a task with capped exponential backoff, cancelling it if a deadline passes.

import uuid, time, httpx

A2A_HEADERS = {"A2A-Version": "1.0", "Content-Type": "application/json"}

class A2AError(Exception):
    pass

class A2AClient:
    def __init__(self, base, token):
        self.http = httpx.Client(timeout=60, headers={**A2A_HEADERS, "Authorization": f"Bearer {token}"})
        card = self.http.get(f"{base}/.well-known/agent-card.json").raise_for_status().json()
        jsonrpc = [i for i in card["supportedInterfaces"] if i["protocolBinding"] == "JSONRPC"]
        if not jsonrpc:
            raise A2AError("agent offers no JSON-RPC interface")
        self.url, self.card, self.next_id = jsonrpc[0]["url"], card, 0

    def call(self, method, params):
        self.next_id += 1
        r = self.http.post(self.url, json={"jsonrpc": "2.0", "id": self.next_id,
                                           "method": method, "params": params})
        body = r.raise_for_status().json()
        if "error" in body:
            raise A2AError(f"{method}: {body['error']}")
        return body["result"]

    def send(self, text, task_id=None, context_id=None, background=False):
        msg = {"messageId": str(uuid.uuid4()), "role": "ROLE_USER", "parts": [{"text": text}]}
        if task_id:
            msg["taskId"] = task_id
        if context_id:
            msg["contextId"] = context_id
        params = {"message": msg}
        if background:
            params["configuration"] = {"returnImmediately": True}
        return self.call("SendMessage", params)   # holds exactly one of "message" or "task"

    def wait(self, task_id, deadline_s=300):
        delay, start = 1.0, time.monotonic()
        while time.monotonic() - start < deadline_s:
            task = self.call("GetTask", {"id": task_id})
            state = task["status"]["state"]
            if state in TERMINAL or state in INTERRUPTED:
                return task
            time.sleep(delay)
            delay = min(delay * 2, 15)          # back off; never hammer a busy agent
        self.call("CancelTask", {"id": task_id})
        raise A2AError(f"task {task_id} exceeded {deadline_s}s and was cancelled")

TERMINAL = {"TASK_STATE_COMPLETED", "TASK_STATE_FAILED", "TASK_STATE_CANCELED", "TASK_STATE_REJECTED"}
INTERRUPTED = {"TASK_STATE_INPUT_REQUIRED", "TASK_STATE_AUTH_REQUIRED"}

Using it for the receipts scenario looks like this. Note that the code handles both result shapes, because the server, not the client, decides whether a request deserves a task.

client = A2AClient("https://receipts.example.com", token=get_token())
result = client.send("Extract this receipt total", background=True)
if "message" in result:                       # agent answered directly, nothing to track
    print(result["message"]["parts"])
else:
    task = client.wait(result["task"]["id"])
    if task["status"]["state"] == "TASK_STATE_INPUT_REQUIRED":
        question = task["status"]["message"]["parts"][0]["text"]
        follow = client.send(ask_user(question), task_id=task["id"], context_id=task["contextId"])
    elif task["status"]["state"] == "TASK_STATE_COMPLETED":
        for art in task.get("artifacts", []):
            store(art["artifactId"], art["parts"])

Three habits carry over to production: a fresh messageId for every new turn; taskId and contextId persisted before waiting, so a restart resumes with GetTask instead of orphaning remote work; and a deadline with cancellation, because abandoned tasks still cost the remote agent money.

What the server must get right

On the server side the protocol layer should own the task store and the state machine, and the agent logic should only emit events. That separation keeps illegal transitions, such as completing a cancelled task, out of reach of model output.

# Server skeleton: the protocol layer owns the task store and state machine,
# the agent logic only produces events. Pseudocode; any web framework works.
ALLOWED = {
    "SUBMITTED": {"WORKING", "REJECTED", "CANCELED", "FAILED"},
    "WORKING": {"COMPLETED", "FAILED", "CANCELED", "INPUT_REQUIRED", "AUTH_REQUIRED"},
    "INPUT_REQUIRED": {"WORKING", "CANCELED", "FAILED"},
    "AUTH_REQUIRED": {"WORKING", "CANCELED", "FAILED"},
}

def transition(task, new_state, message=None):
    old = task.state
    if new_state not in ALLOWED.get(old, set()):
        raise IllegalTransition(old, new_state)      # terminal states have no exits
    task.state, task.status_message = new_state, message
    store.save(task)                                 # persist before notifying anyone
    events.publish(task.id, status_update(task))     # SSE subscribers and push configs

def on_send_message(req, principal):
    prior = store.task_for_message(req.message.message_id, owner=principal)
    if prior:                                        # redelivery: return the original task
        return prior
    if req.message.task_id:
        task = store.get(req.message.task_id, owner=principal)   # TaskNotFoundError if absent
        if task.state in TERMINAL:
            raise UnsupportedOperationError("task is terminal; start a new task")
    else:
        task = store.create(context_id=req.message.context_id or new_id(), owner=principal)
    store.record_message(req.message.message_id, task.id)
    worker.enqueue(task.id, req.message)
    return task if req.configuration.return_immediately else wait_terminal_or_interrupted(task)

Persist before you notify, or a subscriber's next GetTask reads stale data. Scope every lookup by the authenticated principal so tenants cannot read each other's tasks. Bound how long interrupted tasks may wait so abandoned conversations do not pin workers.

Failure modes

FailureSymptomFix
Missing A2A-Version header1.0 server rejects or misparses requests as 0.3Set the header in one place in the HTTP client
Client assumes a Task is always returnedCrash on a direct Message replyBranch on which result field is present
Orphaned remote tasksRemote costs grow after client restartsPersist task ids first; cancel on deadline
Interrupted task treated as failureUsers never see the agent's questionRoute input-required to a human or to client logic
Streaming connection dropClient misses artifact chunksReattach with SubscribeToTask, reconcile with GetTask
Unverified webhookForged push updates acceptedAuthenticate pushes; re-read the task with GetTask

The cross-cutting lesson is that the authoritative state of a task lives on the server. Streams and webhooks are fast paths for learning about changes; GetTask is the reconciliation path, and every client should be able to fall back to it.

When A2A is the right tool, and when it is not

A2A earns its overhead when the callee is genuinely an agent owned by someone else: work is long or interactive, results are multi-part, and you want discovery and authentication that other teams already understand. It is a poor fit for a deterministic internal function (use a plain API or MCP), for latency-critical inner loops where a task object per call is wasted state, or for agents inside one process that can share memory directly.

Start with one high-value delegation using blocking calls, and add streaming and push only when durations demand them.

What to do next

  1. Fetch the Agent Card of one agent you depend on, or write one for your own agent, and check the path is /.well-known/agent-card.json.
  2. Implement the minimal client above against a test agent, including the Message-or-Task branch and the input-required follow-up.
  3. Decide per delegation whether to block, stream or use push, based on expected duration and whether a user is watching.
  4. Persist task and context ids before waiting, set a deadline, and cancel abandoned tasks.
  5. Map each A2A error type to a concrete client action and alert on the unexpected ones.
  6. Read the security article and confirm authentication, tenant scoping and webhook verification before exposing an agent outside your network.
Key takeaway: A2A is a contract for delegating work to an opaque agent you do not control. The client discovers the agent through its Agent Card, sends messages, and follows the resulting task through active, interrupted and terminal states, receiving results as typed artifacts by response, stream or push over JSON-RPC, gRPC or HTTP+JSON. It complements MCP, which connects an agent to its tools. Treat the server's task as the source of truth, persist ids, send the version header, and cancel what you no longer need.