The Agent2Agent (A2A) protocol lets one agent hand work to another over HTTP. On its JSON-RPC binding every interaction is a JSON-RPC 2.0 request to a single endpoint, and every answer, whether one reply or a stream of events, is a JSON-RPC response tied to that request by its id. That sounds simple until a task needs the user's input halfway through, a stream drops at minute four, a retry sends the same instruction twice, or a cancel arrives at the moment the task finishes.
This page follows the flow over time, from the first byte of a request to the last event of a task, and covers what each side must do at every step. It uses A2A 1.0: PascalCase method names such as SendMessage and GetTask, and task states such as TASK_STATE_WORKING. The layers of a single envelope and the full error-code table are in the A2A message envelope; the task store and executor are in A2A task architecture.
Two clocks: the RPC and the task
Every A2A interaction runs on two clocks. The RPC clock is short: a request goes out, a response comes back, and the JSON-RPC id joins them. The task clock can run for hours: a task moves from TASK_STATE_SUBMITTED through TASK_STATE_WORKING to a terminal state, possibly pausing at TASK_STATE_INPUT_REQUIRED or TASK_STATE_AUTH_REQUIRED along the way.
One task usually spans several RPCs. The client creates it with SendMessage, answers a question with another SendMessage carrying the task id, watches progress with SubscribeToTask or GetTask, and may end it with CancelTask. The JSON-RPC id is the key for one exchange; the task id is the key for the work. Most bugs in this area come from confusing the two.
JSON-RPC rules that shape the flow
- The id is the correlation key. A response carries the same id as its request. On a stream, every Server-Sent Events data line is a full JSON-RPC response with the original id. A client must reject a response whose id it did not send.
- No id means no reply. A request without an id is a JSON-RPC notification, and the server must not answer it, not even with an error. A client that forgets the id gets silence, which looks like a hung server.
- Result or error, never both. A successful
SendMessageresult holds exactly one oftaskormessage. Branch on which is present. - Batches are not part of A2A. JSON-RPC 2.0 permits an array of requests in one body, but the A2A specification does not define batch behaviour. Send one request per HTTP call, and if you run a server, reject arrays with
-32600rather than half-supporting them. - The method name is the operation. Version 0.3 used slash-style names such as
message/send. A 1.0 server answers those with-32601Method not found, so a version mismatch often shows up first as that error.
The server pipeline: parse, validate, dispatch, map
A server should treat every request as passing through fixed stages, each with its own error. Doing them in order means the error a client sees names the first thing that was actually wrong. Parse the body; check the JSON-RPC frame; read the A2A-Version header (the specification says an empty value means 0.3); authenticate; look up the method; validate params; run the handler; map any exception to a JSON-RPC error with the request's id.
import json
TERMINAL = {"TASK_STATE_COMPLETED", "TASK_STATE_FAILED",
"TASK_STATE_CANCELED", "TASK_STATE_REJECTED"}
class RpcError(Exception):
def __init__(self, code, message, reason=None):
super().__init__(message)
self.code, self.message, self.reason = code, message, reason
def error(id_, e):
err = {"code": e.code, "message": e.message}
if e.reason:
err["data"] = [{"@type": "type.googleapis.com/google.rpc.ErrorInfo",
"reason": e.reason, "domain": "a2a-protocol.org"}]
return {"jsonrpc": "2.0", "id": id_, "error": err}
async def handle(body: bytes, headers, methods, auth):
try:
req = json.loads(body)
except ValueError:
return error(None, RpcError(-32700, "Parse error"))
if not isinstance(req, dict) or req.get("jsonrpc") != "2.0" \
or not isinstance(req.get("method"), str):
return error(req.get("id") if isinstance(req, dict) else None,
RpcError(-32600, "Invalid Request"))
id_ = req.get("id")
is_notification = "id" not in req
try:
version = headers.get("A2A-Version") or "0.3" # empty means 0.3, per the spec
if version != "1.0":
raise RpcError(-32009, f"Version {version} not supported", "VERSION_NOT_SUPPORTED")
principal = await auth(headers) # raises on bad credentials
fn = methods.get(req["method"])
if fn is None:
raise RpcError(-32601, "Method not found")
params = req.get("params")
if not isinstance(params, dict):
raise RpcError(-32602, "Invalid params")
result = await fn(principal, params)
except RpcError as e:
return None if is_notification else error(id_, e)
except Exception:
log_exception(req["method"], id_) # never leak internals
return None if is_notification else error(id_, RpcError(-32603, "Internal error"))
return None if is_notification else {"jsonrpc": "2.0", "id": id_, "result": result}Two details matter. A parse error has no recoverable id, so the response carries null. And unexpected exceptions become -32603 with a generic message, while the stack trace goes to your logs under the request id.
A multi-turn wire trace
A travel agent delegates a booking to an airline agent. The first request creates the task and waits for it (returnImmediately is unset, so false):
--> {"jsonrpc": "2.0", "id": "r1", "method": "SendMessage",
"params": {"message": {"messageId": "m-1", "role": "ROLE_USER",
"parts": [{"text": "Book one economy seat to New York on 14 October."}]}}}
<-- {"jsonrpc": "2.0", "id": "r1", "result": {"task": {
"id": "task-5d2e", "contextId": "ctx-91af",
"status": {"state": "TASK_STATE_INPUT_REQUIRED",
"message": {"messageId": "m-a1", "role": "ROLE_AGENT",
"parts": [{"text": "Which city are you departing from?"}]}}}}}The blocking call returned before the task finished. The specification says a blocking SendMessage waits until the task reaches a terminal state or an interrupted one, and input-required is interrupted. The client answers in a new request that names the task and context, with a fresh messageId and this time asks not to wait:
--> {"jsonrpc": "2.0", "id": "r2", "method": "SendMessage",
"params": {"message": {"messageId": "m-2", "role": "ROLE_USER",
"taskId": "task-5d2e", "contextId": "ctx-91af",
"parts": [{"text": "From San Francisco (SFO)."}]},
"configuration": {"returnImmediately": true, "historyLength": 0}}}
<-- {"jsonrpc": "2.0", "id": "r2", "result": {"task": {
"id": "task-5d2e", "contextId": "ctx-91af",
"status": {"state": "TASK_STATE_WORKING"}}}}
--> {"jsonrpc": "2.0", "id": "r3", "method": "SubscribeToTask", "params": {"id": "task-5d2e"}}Setting historyLength to 0 tells the server to omit history, which keeps the reply small; the client already has the conversation. The subscription's first event is a Task snapshot, as the specification requires, followed by artifact and status updates, all carrying id r3. When the task reaches TASK_STATE_COMPLETED the server sends the final status update and closes the stream; in 1.0 the close itself is the end-of-stream signal.
Blocking, returning immediately, or streaming
The client picks how long one HTTP request should live. The choice is about the RPC clock, not the task clock:
| Mode | Request | The HTTP request lasts | Use when |
|---|---|---|---|
| Blocking | SendMessage, returnImmediately false | until terminal or interrupted | tasks reliably finish in seconds |
| Fire and poll | SendMessage with returnImmediately true, then GetTask | milliseconds each | long tasks, clients behind strict proxies |
| Stream | SendStreamingMessage or SubscribeToTask | the life of the task | progress or partial artifacts matter |
| Push | push notification config | milliseconds; the server calls you back | very long tasks, clients that sleep |
Blocking is the trap. Load balancers and API gateways often cut idle HTTP requests after a fixed time, commonly tens of seconds. A blocking call to a task that takes longer dies at the proxy, the client sees a transport error rather than a JSON-RPC response, and the task keeps running on the server with no one waiting. Set an explicit client timeout below the shortest proxy timeout in the path, and after a timeout recover with GetTask or SubscribeToTask instead of resending. Streams have their own proxy hazards, such as response buffering; see A2A streaming.
Recovering a dropped stream
When a stream breaks, the task does not. The client reconnects with SubscribeToTask, and the first event it receives is the current Task, so it can reconcile what it missed: compare the state, and for artifacts, rebuild from the snapshot rather than appending chunks it may already have. If the task finished while the client was away, the subscription fails with UnsupportedOperationError (-32004), because subscriptions are not allowed on terminal tasks; the client then calls GetTask to read the final result.
Retries and de-duplication
A transport failure leaves the client not knowing whether the server acted. Reads (GetTask, ListTasks) and SubscribeToTask are safe to repeat. CancelTask is safe to repeat in effect: the second call either cancels nothing new or reports that the task cannot be cancelled. SendMessage is the dangerous one, because a repeat can start a second booking.
The specification says Send Message operations MAY be idempotent and that agents may use the messageId to detect duplicates. Treat that as a contract you must arrange explicitly: the client generates the messageId once per logical message and reuses it on every retry, and the server remembers (tenant, messageId) for a retention window and replays the existing task instead of acting again. Do not de-duplicate on the JSON-RPC id. It identifies one exchange, and a client may legitimately generate a new one per HTTP attempt.
async def send_message(principal, params, store, executor):
msg = params["message"]
key = (principal.tenant, msg["messageId"])
prior = await store.seen_message(key) # messageId -> taskId, with a TTL
if prior is not None:
return {"task": await store.snapshot(prior, params)} # replay, do not re-run
task_id = msg.get("taskId")
if task_id:
task = await store.get(task_id, principal) # TaskNotFoundError (-32001) if absent
if task.state in TERMINAL:
raise RpcError(-32004, "Task is in a terminal state", "UNSUPPORTED_OPERATION")
async with store.lock(task_id): # one turn at a time per task
await store.append_history(task_id, msg)
await store.remember_message(key, task_id)
await executor.resume(task_id, msg)
else:
task = await store.create(principal, msg)
await store.remember_message(key, task.id)
await executor.start(task.id, msg)
task_id = task.id
if params.get("configuration", {}).get("returnImmediately"):
return {"task": await store.snapshot(task_id, params)}
await executor.wait_until_terminal_or_interrupted(task_id)
return {"task": await store.snapshot(task_id, params)}The per-task lock serialises turns, so two concurrent follow-ups to the same task cannot interleave in the history. The client side keeps the messageId fixed and only retries what is safe:
import asyncio, json, random, uuid
import httpx
HEADERS = {"A2A-Version": "1.0", "Content-Type": "application/json"}
async def call(client, url, method, params, attempts=4):
req_id = str(uuid.uuid4()) # one id for this call; the messageId
for n in range(attempts): # inside params must also stay fixed
try:
r = await client.post(url, json={"jsonrpc": "2.0", "id": req_id,
"method": method, "params": params},
headers=HEADERS, timeout=30)
body = r.json()
if body.get("id") != req_id:
raise RuntimeError("response id does not match request id")
if "error" in body:
code = body["error"]["code"]
if code == -32603 and n + 1 < attempts and method in RETRY_SAFE:
raise httpx.TransportError("internal error, retrying")
raise A2AError(body["error"])
return body["result"]
except (httpx.TransportError, httpx.TimeoutException):
if n + 1 == attempts or method not in RETRY_SAFE:
raise
await asyncio.sleep(random.uniform(0, 0.5 * 2 ** n)) # full jitter
# SendMessage is retry-safe only because the caller fixes messageId once, before the loop,
# and the server de-duplicates on it. Without that, remove it from this set.
RETRY_SAFE = {"GetTask", "ListTasks", "CancelTask", "SendMessage"} # streams use an SSE readerOnly advertise SendMessage as retry-safe to callers once the server side de-duplicates. If you call agents you do not control, assume they do not, and recover with GetTask or ListTasks filtered by context before resending. See A2A idempotency for storage and retention choices.
Races at the edges of the task clock
- Cancel against completion. The client sends
CancelTaskjust as the task completes. The server must decide under the task lock; whichever transition commits first wins, and the loser getsTaskNotCancelableError(-32002). The client must read the returned task or error, not assume the cancel worked. - A message to a finished task.
SendMessagewith the id of a terminal task fails withUnsupportedOperationError. Start a new task in the same context instead. - An answer that arrives before the question. A client that answers an input-required prompt it inferred, before the server has paused, can find the task still working. Accept the message only in states your server policy allows, and reject the rest with a clear error.
Failure modes
- Missing
A2A-Versionheader, stripped by a gateway: the server treats the client as 0.3 and returns VersionNotSupportedError, or Method not found for 1.0 names. - Missing request id: the request becomes a notification and the server stays silent.
- Client retries
SendMessagewith a new messageId: duplicate side effects, invisible until reconciliation. - Blocking calls longer than a proxy timeout: orphaned tasks and duplicate work when users retry.
- Errors keyed on HTTP status rather than the JSON-RPC code: on this binding, A2A errors arrive in a JSON-RPC error object, so the HTTP status alone does not tell you which error occurred. See A2A error handling.
- Logs keyed only by task id: the RPC that failed cannot be found. Log the JSON-RPC id, method, task id and context id on every line.
Operating it
Measure per method: request rate, latency, and error rate by JSON-RPC code. Track blocking-call duration against your proxy timeout, open stream count, stream reconnects per task, and de-duplication hits; a rising hit rate means clients are retrying more, which is often the first symptom of a network or timeout problem. Put the JSON-RPC id and the task id on the trace span for every call so one delegation can be followed across agents.
What to do next
- Check that every client sends a unique id and the
A2A-Version: 1.0header, and that no gateway in the path strips it. - Make clients generate messageId once per logical message and reuse it on retry; add server-side de-duplication on tenant plus messageId with a retention window.
- Replace long blocking calls with
returnImmediatelyplus streaming or polling, and set client timeouts below your proxy timeouts. - Implement stream recovery: reconnect with
SubscribeToTaskand backoff, reconcile from the first Task event, fall back toGetTaskon-32004. - Serialise state changes per task, and test the cancel-versus-complete race with two concurrent clients.
- Add per-method metrics by JSON-RPC error code, and log the RPC id and task id together.