An A2A task that takes seconds can live in memory. A task that takes hours, or that stops to ask a person a question and waits until the next morning, cannot: in that time the client's connection will drop, the agent will be redeployed and a worker will crash. Resumability means the task carries on correctly through all of that, without losing progress and without doing anything twice.
The client half, reattaching to a stream with SubscribeToTask and falling back to GetTask, is covered in the guide to send versus streaming send. This page is about the server half: what the A2A 1.0 specification promises, what it leaves to you, and a concrete runtime design (a durable task store, an ordered event log, step checkpoints, leases with fencing tokens and an effect ledger) that makes the work itself resumable.
Three things that can die
Resumability is easier to design once you separate what can fail. The connection can drop: the client's stream is cut by a proxy, a deploy or a laptop lid. The process can die: the agent replica running the task is redeployed, scaled in or crashes mid-step. The conversation can stall: the task needs input from a person or another agent and nobody answers for a day. Each needs a different mechanism. Connection loss is solved by keeping task state outside the connection; process loss by keeping progress outside the process; stalls by making a waiting task cost nothing while it waits.
What A2A 1.0 gives you, and what it leaves out
The 1.0.0 specification gives you the vocabulary for long-running work. Every task carries a TaskState. Work moves through TASK_STATE_SUBMITTED and TASK_STATE_WORKING, can pause in TASK_STATE_INPUT_REQUIRED or TASK_STATE_AUTH_REQUIRED, and ends in one of four terminal states: completed, failed, canceled or rejected. The methods cover every way a client waits:
| Method | Role in a long-running task |
|---|---|
SendMessage | Start a task, or continue one by sending a message with the same taskId and contextId. |
SendStreamingMessage | Start or continue with a Server-Sent Events stream of updates. |
SubscribeToTask | Reattach to a live task; the first event must be the current Task. Fails with UnsupportedOperationError on a terminal task. |
GetTask | Poll the current state; historyLength limits how many recent messages come back (zero excludes history). |
ListTasks | Find tasks, for example after a client restart that lost its ids. |
CancelTask | Ask the server to stop the work. |
CreateTaskPushNotificationConfig and its Get, List and Delete siblings | Register a webhook (url, token, authentication) so the client need not stay connected. |
What the specification does not say matters as much. It makes no statement about how long tasks are retained, whether they survive a server restart, or how work resumes after a crash. Those are runtime properties of your agent. Decide them, implement them, and document them for callers, because a client cannot tell from the protocol alone whether your tasks survive a deploy. The lifecycle states guide separates what the specification fixes from what it leaves to server policy.
The runtime architecture
The design rests on one rule: the durable store is the only source of truth about a task. Gateway replicas translate A2A calls into reads and writes on the store and hold no task state. Workers claim tasks under a time-limited lease, execute one step at a time and commit each step's result and its event in a single transaction. Stream fan-out and push dispatch run after the commit and read from the ordered event log. Any component can be killed at any time; whatever was committed survives, and whatever was not will be redone.
The data model
CREATE TABLE tasks (
task_id TEXT PRIMARY KEY,
context_id TEXT NOT NULL,
tenant TEXT NOT NULL,
state TEXT NOT NULL, -- TASK_STATE_* values from the spec
step INT NOT NULL DEFAULT 0, -- index of the next step to run
checkpoint JSONB, -- everything the next step needs
lease_owner TEXT,
lease_until TIMESTAMPTZ,
fence BIGINT NOT NULL DEFAULT 0, -- bumped on every lease grant
next_seq BIGINT NOT NULL DEFAULT 1,
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
expires_at TIMESTAMPTZ -- your retention policy, not the spec's
);
CREATE TABLE task_events ( -- ordered log for streams and webhooks
task_id TEXT REFERENCES tasks, seq BIGINT, kind TEXT, body JSONB,
PRIMARY KEY (task_id, seq)
);
CREATE TABLE side_effects ( -- run-once ledger
task_id TEXT, effect_key TEXT, result JSONB,
PRIMARY KEY (task_id, effect_key)
);The checkpoint is the heart of resumability. It must hold everything the next step needs and nothing that only lives in memory: identifiers of fetched documents rather than open handles, partial results, and the answers to any questions the task asked. If a step cannot be restarted from its checkpoint, the task is not resumable, however durable the store is.
Leases, fencing and step commits
Workers claim tasks with a lease. A task whose lease has expired is claimable again, which is exactly how crash recovery happens: nobody has to detect the crash, the lease simply runs out. The fence counter increases on every grant, and every commit checks it, so a worker that paused for a long garbage collection and lost its lease cannot overwrite progress made by its successor.
UPDATE tasks
SET lease_owner = :me, lease_until = now() + interval '30 seconds', fence = fence + 1
WHERE task_id = (SELECT task_id FROM tasks
WHERE state IN ('TASK_STATE_SUBMITTED', 'TASK_STATE_WORKING')
AND (lease_until IS NULL OR lease_until < now())
ORDER BY updated_at LIMIT 1
FOR UPDATE SKIP LOCKED)
RETURNING task_id, fence, step, checkpoint;STEPS = [plan, fetch_grants, analyse, ask_approval, apply_revocations, write_report]
def run(task_id, fence, step, checkpoint):
while step < len(STEPS):
renew_lease(task_id, fence) # raises LostLease if the fence moved on
result = STEPS[step](checkpoint, effect=lambda key, fn: once(task_id, key, fn))
if result.needs_input: # pause: release the lease, cost nothing
commit(task_id, fence, step, result.checkpoint,
state="TASK_STATE_INPUT_REQUIRED", event=result.question)
return
checkpoint, step = result.checkpoint, step + 1
commit(task_id, fence, step, checkpoint,
state="TASK_STATE_WORKING", event=progress(step))
commit(task_id, fence, step, checkpoint, state="TASK_STATE_COMPLETED",
event=artifacts(checkpoint))
def commit(task_id, fence, step, checkpoint, state, event):
with db.transaction() as tx:
row = tx.fetchone(
"UPDATE tasks SET step=%(st)s, checkpoint=%(cp)s, state=%(s)s, updated_at=now(),"
" next_seq=next_seq+1,"
" lease_owner=CASE WHEN %(s)s = 'TASK_STATE_WORKING' THEN lease_owner END,"
" lease_until=CASE WHEN %(s)s = 'TASK_STATE_WORKING' THEN lease_until END"
" WHERE task_id=%(t)s AND fence=%(f)s RETURNING next_seq-1",
{"st": step, "cp": json.dumps(checkpoint), "s": state, "t": task_id, "f": fence})
if row is None:
raise LostLease(task_id) # a newer worker owns the task
tx.execute("INSERT INTO task_events VALUES (%s, %s, 'status', %s)",
(task_id, row[0], json.dumps(event)))
notify(task_id) # wake streams and push after commitTwo details matter. The commit writes the checkpoint and the event in one transaction, so a reconnecting client never sees an event whose progress was lost, and never misses progress that was saved. And the commit clears the lease whenever the task leaves the working state, so a paused or finished task holds no worker. Long steps should be split into sub-steps with their own commits, so a crash costs at most one sub-step of repeated work.
Side effects that must run once
A resumed step runs again from its checkpoint, so anything it did before the crash may happen twice. Reads are harmless; sending an email, charging a card or revoking an access grant is not. Give every side effect a stable key derived from the task and the step, record its result after it succeeds, and look the key up before acting. Pass the same key downstream as an idempotency key, because the crash can land between the downstream call succeeding and the ledger write.
def once(task_id, key, fn):
row = db.fetchone("SELECT result FROM side_effects WHERE task_id=%s AND effect_key=%s",
(task_id, key))
if row:
return row[0] # already done before the crash
result = fn(idempotency_key=f"{task_id}:{key}") # downstream must honour the key
db.execute("INSERT INTO side_effects VALUES (%s, %s, %s) ON CONFLICT DO NOTHING",
(task_id, key, json.dumps(result)))
return resultThe same reasoning applies one level up, between client and agent; A2A idempotency covers how a caller avoids starting a second task when a send times out.
Pauses that last for days
When a step needs a decision, the worker commits the task in TASK_STATE_INPUT_REQUIRED with the question as a status message, and releases its lease. The task now costs one row. Streams close or go quiet, and the push dispatcher sends the update to every webhook registered for the task, so a client that disconnected hours ago still learns that it is needed. The input-required pattern covers how to phrase the question so a calling agent can answer it.
The answer arrives as an ordinary SendMessage carrying the same taskId and contextId, which is how the specification says an interrupted task continues. The gateway appends the message to the task's history, stores the answer in the checkpoint, sets the state back to working and leaves the lease empty. The next free worker claims the task and resumes at the stored step. Nothing about the original worker, connection or replica needs to exist any more.
Add a deadline to every pause. A question nobody answers should eventually fail or cancel the task with a clear status message rather than sit forever, and the deadline belongs in the checkpoint so it survives restarts too.
Worked example: a quarterly access review
A compliance agent receives a task to review access to 40 internal systems. Its steps are: plan, fetch grants per system, analyse, ask the system owners to approve the proposed revocations, apply the approved revocations and write a report. Fetching grants commits once per system, so the checkpoint records the list of completed systems.
Two hours in, a deploy terminates the worker while it is fetching system 23. Its lease expires within 30 seconds and a worker on the new version claims the task with a higher fence. The checkpoint says 22 systems are done, so it fetches from system 23 onward. The old worker never commits again; if it had survived a long pause, its commit would fail the fence check. The caller's stream dropped during the deploy, so the caller reattached with SubscribeToTask and received the current Task as its first event.
At hour three the agent asks for approval and the task enters input-required. The owners answer the next morning through the calling agent, which sends the decision with the same taskId and contextId. While applying revocations, a network error interrupts the step after 9 of 14 revocations; on retry, the effect ledger returns the stored result for the first 9 and only the remaining 5 run. The task completes with a report artifact, and a webhook delivers the final state to the caller.
Failure modes
- State only in memory. A task that lives in a Python object dies with the process. Everything a step needs must be in the checkpoint.
- Event before commit. Emitting an event and then crashing before saving progress shows clients a state the task will never reach. Commit first, then notify.
- No fencing. Two workers believe they own a task after a long pause, and the slower one overwrites newer progress.
- Side effects without keys. Resumed steps repeat emails, charges or revocations.
- Unbounded waits. Input-required tasks accumulate for months. Give every pause a deadline and every task an expiry.
- Silent retention. Tasks are deleted after a short window and callers polling later receive a not-found error. Document retention for callers.
- Code changes mid-task. A new release renames a step while old tasks still point at its index. Version the step list and keep old versions runnable until their tasks drain.
Trade-offs
Every commit is a database write, so finer checkpoints mean less repeated work after a crash and more write load. Commit at natural boundaries, such as one per document or per external call batch, rather than per token. Short leases recover faster from crashes but need more frequent renewal and risk false expiry under load; a lease a few times longer than a typical step is a reasonable starting point. A dedicated workflow engine gives you durable execution, retries and timers ready-made, at the cost of another system and its programming model; the schema above is enough for a handful of task types and a good way to learn what such an engine does for you.
What to do next
- Write down your retention policy, restart behaviour and maximum pause length, and publish them for callers.
- Move every long-running task's state into a durable store and split the work into named steps with explicit checkpoints.
- Add leases with a fence counter and check the fence on every commit.
- Give each side effect an effect key and pass it downstream as an idempotency key.
- Register push notification configs for tasks that may pause, and add a deadline to every input-required state.
- Test by killing a worker mid-step, redeploying mid-task and answering a question a day late, and confirm nothing is lost or repeated.