Error handling decides what a failure means: is this retryable, whose fault is it, what should the caller see. Error recovery is what you do next, often after your own process has lost track of what it was doing. A stream dropped mid-task. Your orchestrator was redeployed with twelve delegations in flight. A remote agent failed after charging a card. The questions are no longer about codes; they are about state: which remote tasks exist, what did they do, and how do you get the whole workflow back to a known, correct place.

This article builds a recovery design for an Agent2Agent (A2A) 1.0 client on top of what the protocol actually guarantees. Facts were checked against the A2A specification, version 1.0.0, on 2026-10-02. Error taxonomy and retry classification are covered in A2A error handling and deadlines in A2A timeout handling; this page assumes them and starts after the failure.

Advertisement

What the protocol gives you to recover with

A2A does not define a recovery protocol, but it gives a client enough primitives to build one. Knowing exactly what each one guarantees, and what it only permits, is the whole game.

PrimitiveWhat the 1.0 spec saysRecovery use
SubscribeToTaskReturns the current Task as the first event; returns UnsupportedOperationError on a terminal taskReattach to a running task after a dropped stream
GetTaskReturns the task; historyLength limits messages (0 means none)Read authoritative state cheaply
ListTasksFilters by contextId, status, statusTimestampAfter; cursor paging; newest firstFind tasks you created but never recorded
messageIdRequired on every message; agents MAY use it to detect duplicatesSafe resend if the agent dedups; never assume it does
Terminal statesCOMPLETED, FAILED, CANCELED, REJECTED accept no further messagesA retry is always a new task
Interrupted statesINPUT_REQUIRED, AUTH_REQUIRED wait for the clientResume on the same task id

Two consequences follow. Because duplicate detection is optional, recovery must look before it resends: query the remote store first, resend only if nothing is there. And because terminal tasks cannot be restarted, "retry" at the task level always means a new task with a new message, which your journal must link to the old one.

The architecture: a journal and a reconciler

Client-side recovery: a durable delegation journal reconciled against the remote agentOrchestratorplans steps, delegatesDelegation journalstep, messageId, contextId,taskId, state, effectsRemote agentA2A 1.0 serverRemote task storestatus, history, artifactsReconcilerruns on start + timerCompensatorundo or verify effectsAlternative agentsregistry for re-delegation1. write intent2. SendMessage (same messageId on retry)3. stream / push / GetTask4. unfinished rows5. GetTask / ListTasksThe journal is written before every send. Recovery never trusts memory: it re-reads the remote taskstate, then decides to resubscribe, resume, re-delegate or compensate.
The orchestrator records intent before sending; the reconciler compares journal rows to remote task state on startup and on a timer.

In-memory orchestrators cannot recover; they can only restart. The minimum durable state per delegation is: the workflow step, the messageId and contextId you generated, the taskId once known, the last observed state, and any side effects the step is known to have caused. Generate the ids before sending and write them first; that ordering is what makes the crash window between "sent" and "recorded" recoverable.

The reconciler is a loop over unfinished journal rows. It runs at process start and periodically, because push notifications and streams can both be lost silently. It never trusts the journal's last state; it re-reads the remote task and lets that decide the next action.

Advertisement

Recovering a dropped stream

Streams over SSE drop for mundane reasons: load balancer idle timeouts, deploys, mobile networks. The recovery is SubscribeToTask with the same task id. The specification requires the first event of the new stream to be the current Task, precisely so that a reconnecting client does not lose information. Treat that snapshot as authoritative and replace local state with it rather than merging.

def follow(client, task_id, on_event, max_reconnects=8):
    attempt = 0
    while True:
        try:
            for event in client.stream("SubscribeToTask", {"id": task_id}):
                attempt = 0                         # first event is the current Task snapshot
                on_event(event)                     # handler must be idempotent
                if is_terminal(event):
                    return
        except UnsupportedOperationError:           # task went terminal while we were away
            on_event({"task": client.call("GetTask", {"id": task_id})})
            return
        except (ConnectionError, TimeoutError):
            attempt += 1
            if attempt > max_reconnects:
                raise
            sleep_with_jitter(min(2 ** attempt, 30))

Two rules make this safe. Event handlers must be idempotent, because the snapshot repeats information you already processed; key artifact writes on artifactId and ignore status updates older than the one you hold. And reconnection needs backoff with jitter and a cap: a hundred clients reconnecting in lockstep after a server restart is a self-inflicted outage. A2A streaming covers the event types themselves.

Recovering from your own crash

The hard case is the orchestrator dying. On restart, every journal row falls into one of three buckets.

  • Intent written, no task id. The process died around the send. Either the request never left, or it created a task whose id you never saved. Query ListTasks for your contextId with statusTimestampAfter set to the row's creation time, and look for your messageId in each candidate's history. Found: record the task id. Not found: resend with the same messageId, so an agent that does dedup recognises it.
  • Task id known, task not terminal. Read it with GetTask. If working, resubscribe or poll. If INPUT_REQUIRED or AUTH_REQUIRED, the journal must have kept the pending question so you can re-ask the user or re-run the auth flow and resume on the same task id.
  • Task terminal. Record the outcome and hand the step back to the orchestrator's decision logic below.
import uuid

TERMINAL = {"TASK_STATE_COMPLETED", "TASK_STATE_FAILED", "TASK_STATE_CANCELED", "TASK_STATE_REJECTED"}
INTERRUPTED = {"TASK_STATE_INPUT_REQUIRED", "TASK_STATE_AUTH_REQUIRED"}

def delegate(journal, client, step, text):
    e = journal.get(step)
    if e is None:                                   # 1. intent is durable before any I/O
        e = journal.create(step, message_id=str(uuid.uuid4()),
                           context_id=str(uuid.uuid4()), state="INTENT")
    if e.state == "INTENT":
        msg = {"messageId": e.message_id, "contextId": e.context_id,
               "role": "ROLE_USER", "parts": [{"text": text}]}
        result = client.call("SendMessage", {"message": msg,
                             "configuration": {"returnImmediately": True}})
        if "task" in result:
            journal.update(step, task_id=result["task"]["id"], state="SENT")
        else:                                       # direct message reply: nothing to track
            journal.update(step, state="TASK_STATE_COMPLETED", result=result["message"])
    return journal.get(step)

def find_task_for_message(client, e):
    """Crash between send and journal update: find the task our messageId created."""
    params = {"contextId": e.context_id, "statusTimestampAfter": e.created_at}
    while True:
        page = client.call("ListTasks", params)
        for t in page.get("tasks", []):
            full = client.call("GetTask", {"id": t["id"]})
            if any(m.get("messageId") == e.message_id for m in full.get("history", [])):
                return t["id"]
        if not page.get("nextPageToken"):
            return None
        params["pageToken"] = page["nextPageToken"]

def reconcile(journal, client, subscribe):
    for e in journal.unfinished():
        if e.task_id is None:
            e.task_id = find_task_for_message(client, e)
            if e.task_id is None:                   # never landed: resend, SAME messageId
                journal.update(e.step, state="INTENT")
                continue
            journal.update(e.step, task_id=e.task_id, state="SENT")
        task = client.call("GetTask", {"id": e.task_id, "historyLength": 0})
        state = task["status"]["state"]
        if state in TERMINAL:
            journal.update(e.step, state=state, result=task)     # decide next step later
        elif state in INTERRUPTED:
            journal.update(e.step, state=state, pending=task["status"].get("message"))
        else:
            subscribe(e.task_id)                    # still running: reattach

The spec lets an agent accept a client-supplied contextId or reject the request with an error, but not silently replace it. So if the send succeeded, the context id you journaled is the task's, and with a fresh one per delegation the context match alone identifies the task. If an agent rejects client context ids, omit it and search on the time filter plus the messageId match. That match can miss on long tasks, because servers may return less history than requested. ListTasks returns only tasks visible to the authenticated client, so the reconciler must use the same principal that sent the original message.

Terminal outcomes: what to do with each

StateMeaning for recoveryTypical action
COMPLETEDWork done; artifacts are the resultPersist artifacts, verify effects if they matter, continue
FAILEDAgent tried and could not finishRead the status message; if transient, new task to same agent with backoff; else re-delegate or surface
REJECTEDAgent declined the workDo not resend the same request to the same agent; re-delegate or change the request
CANCELEDSomeone cancelled itIf you did (deadline), apply your degraded path; if not, alert and investigate before retrying

Re-delegation means choosing another agent whose Agent Card advertises a matching skill, from a registry you maintain. Carry the reason forward: include the previous task id and a summary of the failure in the new message's text or metadata, so the second agent does not repeat the first one's mistake and your audit trail links both tasks. Limit total attempts per step across all agents, not per agent; otherwise a workflow can loop through a registry indefinitely. Detailed classification of which failures are transient is in A2A error handling.

Compensating side effects

A remote agent can fail after doing something irreversible: booking a seat, sending an email, writing a record. A2A carries no transaction semantics, so you cannot roll back across agents. Use the saga pattern: for each step with a side effect, define a compensating action (cancel the booking, send a correction) and run compensations in reverse order when a later step fails permanently.

Knowing whether an effect happened is the hard part. A FAILED status does not prove nothing happened, and a COMPLETED status is the agent's claim, not proof. For effects that matter, verify through an independent read path you trust: query the booking system, the payment ledger, the mail log. Asking the same agent "did you really do it?" adds latency and verifies nothing; the agent that misreported once can misreport twice. Record verified effects in the journal so compensation is driven by facts, and make compensations idempotent, since they will be retried by the same reconciler. Exactly-once effect patterns are covered in A2A idempotency.

Worked example: a travel orchestrator redeployed mid-trip

An orchestrator plans a trip with three delegations: a flight agent, a hotel agent and a car agent. It journals all three intents, sends the flight request (task F, now WORKING via stream), sends the hotel request, and is killed by a deploy before it records the hotel's task id. The car intent was never sent.

On startup the reconciler finds three unfinished rows. Flight: task id known; SubscribeToTask returns UnsupportedOperationError, so the task ended while the process was down; GetTask shows COMPLETED with a booking artifact, which is stored and verified against the airline's booking API. Hotel: no task id; ListTasks on its context id finds one task whose history contains the journaled messageId; its state is INPUT_REQUIRED ("two rooms or one?"). The question is saved and shown to the user again; the reply is sent on the same task id. Car: still INTENT; it is sent normally.

Later the car agent returns REJECTED (no availability). The orchestrator re-delegates to a second car agent from the registry with the first task's id in metadata. If that also fails and the user cancels the trip, the compensator runs in reverse: cancel the hotel (verified via its confirmation lookup) and then the flight. Each compensation is a journaled step itself, so a crash during compensation is recovered the same way.

Failure modes

  • Duplicate work: resending without looking first, against an agent that does not dedup by messageId. Always reconcile before resend.
  • Lost interruptions: INPUT_REQUIRED observed but the question not journaled; after a crash the task waits forever. Persist the status message.
  • Reconnect storms: synchronized resubscription after a server restart. Jitter and cap.
  • Unbounded re-delegation: per-agent attempt limits but no per-step limit.
  • Trusting COMPLETED for effects: no independent verification of money or messages.
  • Push-only tracking: a lost webhook leaves a step pending forever; the reconciler's timer is the backstop. See A2A push notifications.
  • Retention gaps: the remote agent may delete old tasks; TaskNotFoundError during reconciliation means you must decide from your own verified effects.

Trade-offs

A journal costs a durable write before every delegation and a reconciler you must operate. For short, read-only delegations (look up a fact, summarise a document) it is often cheaper to restart the workflow from the top and accept occasional duplicate calls. Reserve the full design for long-running tasks, tasks behind human input, and anything with side effects. Polling as a backstop costs requests; push alone costs correctness. Most systems use streams or push for latency and a slow reconciler timer, in minutes, for safety.

What to do next

  1. List every delegation your system makes and mark which have side effects or run longer than one request.
  2. For those, add a journal table and write messageId, contextId and intent before sending.
  3. Implement the reconciler and run it at startup and on a timer; test it by killing the process between send and journal update.
  4. Make stream handlers idempotent and add jittered, capped resubscription.
  5. Define a per-step attempt budget and an alternative-agent list for REJECTED and persistent FAILED outcomes.
  6. For each side effect, write the compensating action and an independent verification read, and test compensation after a crash.
Key takeaway: A2A gives a client the tools for recovery but not the procedure: a snapshot-first SubscribeToTask, GetTask, filtered ListTasks, optional messageId dedup, terminal states that force new tasks and interrupted states that resume in place. Build the procedure yourself: journal intent before sending, reconcile against remote state before resending, reattach to running tasks, re-delegate terminal failures within a budget, and compensate verified side effects in reverse.