Agents retry. A delegating agent sends work to a remote agent, the connection drops before the response arrives, and the client cannot tell whether the work was accepted. If it sends again, will the remote agent do the work twice? In the Agent2Agent (A2A) protocol the answer depends on three identifiers, one optional server behaviour, and how carefully both sides are written.

This article explains idempotency under A2A 1.0.0, the latest released version of the specification as of 2026-10-02. It covers who creates each identifier, the four situations that look like reusing a task ID but mean different things, how to recover when the first response is lost, and how to stop a retried task from repeating its side effects. Earlier drafts of A2A had clients mint task IDs, and our A2A idempotency architecture article describes that older model; in 1.0 the server mints them, which changes the retry story. Background on the Task object itself is in Agent-to-Agent Task Architecture.

Advertisement

Three identifiers, three owners

Idempotency means that performing an operation more than once has the same effect as performing it once. To make a retry idempotent, both sides need a shared name for the operation, created before the first attempt. A2A 1.0 has three candidate names.

IdentifierCreated byScopeRole in retries
messageIdThe message creator (client for requests)One messageThe only identifier that exists before the first send; the natural dedup key
taskIdThe server, when a message creates a new taskOne unit of work and its lifecycleUnknown to the client until a response arrives
contextIdThe server MAY generate it if the message has none; a client may send one to group workA conversation or session spanning tasksNarrows recovery searches; not a dedup key

The consequence is the central fact of this article: on the very first request, the client cannot name the task it is trying to create. The task ID is assigned by the server, so the only handle the client controls is the messageId. The specification says Send Message operations MAY be idempotent and that agents may use messageId to detect duplicate messages. It is permitted, not required. A client that assumes deduplication against a server that does not implement it will create duplicate tasks on every retry.

We found no standard Agent Card field that advertises whether an agent deduplicates on messageId. Treat it as part of the integration contract: agree it with the remote agent's owner, document it, and test it, or design the client so that duplicates are harmless.

Four things that look like task ID reuse

Once a task exists, the client will send more messages that mention it. The spec defines the task lifecycle precisely enough to tell four cases apart, and a server must handle each differently.

messageIdminted by the sender, per messagetaskIdminted by the server, per taskcontextIdgroups tasks in a conversationSame messageIdnetwork retryNew messageId + taskIdtask interruptedNew messageId + taskIdtask terminalcontextId + referencesno taskIdDuplicatereturn the same taskContinuationtask resumesIllegal reuseUnsupportedOperationErrorRefinementnew task, same contextFour things that look like reuse; only one of them is a retry, and dedup on messageId is a MAY in the spec
Classifying an incoming message: the identifiers it carries and the task's state decide whether it is a retry, a continuation, an illegal reuse or a refinement.
  1. Retry. The same messageId arrives again because the sender did not see a response. If the server deduplicates, it returns the task created by the first copy, in its current state, and does nothing else.
  2. Continuation. A new messageId carries an existing taskId whose task is in an interrupted state, TASK_STATE_INPUT_REQUIRED or TASK_STATE_AUTH_REQUIRED. This is the designed way to supply missing input; the task resumes under the same ID.
  3. Illegal reuse. A new message carries the taskId of a task in a terminal state: completed, failed, canceled or rejected. The spec says terminal tasks cannot accept further messages and the server returns UnsupportedOperationError. Tasks never restart.
  4. Refinement. The user wants a change to a finished result. The client sends a message with the same contextId, no taskId, and lists the earlier task in referenceTaskIds. The server creates a new task in the same conversation.

One more rule closes a loophole: agents MUST reject a message whose contextId and taskId do not match. Without it, a buggy client could attach a message to a task from someone else's conversation.

Advertisement

Server design: a dedup table in front of the task store

A server that wants retries to be safe needs a durable record mapping each message to the task it produced. The record must be written atomically with the decision, so two copies of the same message racing to different replicas still produce one task. Use whatever conditional write your store offers: a unique constraint in a relational database, a conditional put in a key-value store, or a set-if-absent with expiry in a cache that is backed up.

Three details matter. First, scope the key by the authenticated caller, (principal, messageId), because message IDs are chosen by clients and two clients may collide. Second, store a digest of the message body and reject a known messageId with a different body, since that is a client bug, not a retry. Third, keep the record for longer than any client will retry. Seventy-two hours is a common choice; the right number is the longest retry horizon of your slowest client, plus margin.

import hashlib, json, time

class DuplicateMismatch(Exception): ...

def body_digest(message):
    # Hash the semantic content, not the transport envelope.
    keep = {k: message.get(k) for k in ("parts", "taskId", "contextId", "referenceTaskIds")}
    return hashlib.sha256(json.dumps(keep, sort_keys=True).encode()).hexdigest()

def handle_send_message(principal, message, store, executor):
    key = (principal, message["messageId"])          # scope by caller: IDs are not global secrets
    digest = body_digest(message)

    row = store.dedup_insert_if_absent(key, digest=digest, ttl_s=72 * 3600)
    if not row.inserted:                             # we have seen this messageId before
        if row.digest != digest:
            raise DuplicateMismatch("messageId reused with a different body")
        if row.task_id is None:                      # first copy still being processed
            raise RetryLater("in progress; retry the same messageId")
        return store.get_task(row.task_id)           # replay: same task, current state

    task_id = message.get("taskId")
    if task_id:                                      # continuation of an existing task
        task = store.get_task(task_id)               # TaskNotFoundError if absent
        if message.get("contextId") and message["contextId"] != task.context_id:
            store.dedup_delete(key)
            raise ValueError("contextId does not match taskId")   # spec: MUST reject
        if task.state in TERMINAL:
            store.dedup_delete(key)
            raise UnsupportedOperationError(f"task {task_id} is {task.state}")
        if task.state not in INTERRUPTED:            # working: queue it or reject, per your policy
            store.dedup_delete(key)
            raise UnsupportedOperationError(f"task {task_id} is not awaiting input")
    else:
        task_id = store.create_task(context_id=message.get("contextId") or new_context_id())

    store.dedup_bind(key, task_id)                   # now retries resolve to this task
    store.append_history(task_id, message)
    executor.schedule(task_id)
    return store.get_task(task_id)

TERMINAL = {"COMPLETED", "FAILED", "CANCELED", "REJECTED"}
INTERRUPTED = {"INPUT_REQUIRED", "AUTH_REQUIRED"}

For a continuation that arrives while the task is still working, the sketch rejects it, the conservative choice; queueing is reasonable if your agent can merge input mid-run. Document whichever you choose.

The lost first response

The hardest case is the first SendMessage of a new task. The server created the task, but the response never arrived, so the client does not know the taskId. It has three ways to recover, from best to worst.

  1. Resend the identical message. Same messageId, same body. Against a deduplicating server this returns the existing task. This is why the client must mint the messageId once and persist it before the first send, in an outbox row or a workflow state record, rather than generating it inside the retry loop.
  2. Look the task up. ListTasks accepts a contextId filter. If the client sent a known contextId, it can list tasks in that context and match on the messageId in each task's history. This works even when the server does not deduplicate, provided history is retained and returned.
  3. Escalate. If neither works, the outcome is unknown. Do not mint a new message and try again; that is how duplicate work happens. Mark the delegation as unknown and let a reconciliation job or a human decide.

Shrink the window as well. By default SendMessage blocks until the task reaches a terminal or interrupted state, which for a long agent run can be minutes, all of it a window in which the connection can drop. Setting returnImmediately to true in the request configuration makes the server return as soon as the task is created. The client learns the taskId within one round trip, then follows progress with GetTask, SubscribeToTask or push notifications, all of which are safe to repeat.

# Pseudocode over a thin A2A client wrapper; field names follow the 1.0 spec text,
# but check your SDK for the exact Part and enum spellings.
def delegate(client, outbox, context_id, parts):
    msg = outbox.get_or_create(intent_key="reconcile-2026-09", make=lambda: {
        "messageId": new_uuid(),        # minted ONCE, persisted before the first send
        "contextId": context_id,
        "parts": parts,
    })
    for attempt in range(6):
        try:
            task = client.send_message(msg, return_immediately=True)   # returnImmediately
            outbox.record_task(msg["messageId"], task.id)
            return task.id
        except (Timeout, ConnectionError, ServerUnavailable):
            sleep(backoff_with_jitter(attempt))
            found = find_task_for(client, context_id, msg["messageId"])
            if found:
                outbox.record_task(msg["messageId"], found)
                return found
    raise DelegationUnknown(msg["messageId"])   # escalate; do NOT mint a new messageId

def find_task_for(client, context_id, message_id):
    for task in client.list_tasks(context_id=context_id):           # ListTasks, contextId filter
        if any(m.get("messageId") == message_id for m in task.history or []):
            return task.id
    return None

Deduplicating the task is not enough

A deduplicated message guarantees one task. It does not guarantee one side effect. An agent's executor can crash after calling a payment API and before recording that it did, then restart and run the step again. The fix is the same idea one layer down: derive a deterministic idempotency key for each external call from the taskId and the step, and pass it to every API that accepts one.

def charge_step(task_id, step_no, invoice, payments):
    # Deterministic: a re-run of the same step of the same task reuses the key.
    key = f"a2a:{task_id}:step:{step_no}"
    return payments.create_charge(amount=invoice.total, currency=invoice.currency,
                                  idempotency_key=key)

A random key per attempt defeats the purpose. Steps chosen by a language model are harder, because a re-run may plan differently; record each tool call in the task store before executing it, and replay recorded steps instead of re-planning. Cancellation interacts with this too: a cancel that arrives mid-step should let the step finish or roll back cleanly, as described in A2A task cancellation.

Notifications run the other way. Push notifications and streams can deliver the same update more than once, so consumers should treat updates as idempotent: apply a status update only if it moves the task forward, and deduplicate artifacts by ID. See A2A push notifications for delivery details.

Client agentoutbox: messageIdDedup table(principal, messageId)Task storetaskId, state, historySendMessageinsert if absentExecutorruns the agenttask createdTools and APIskey = taskId:stepside effectsRecoveryListTasks by contextIdresponse lostGetTask / SubscribeTwo dedup layers: messageId stops duplicate tasks, taskId-derived keys stop duplicate side effects
End to end: the client persists the messageId, the server binds it to one task, and the executor derives side-effect keys from the taskId.

Worked example: a reconciliation delegation

An orchestrator delegates a month-end reconciliation to a finance agent. It persists messageId m-71 and contextId c-recon-09 in its outbox, then sends with returnImmediately set.

The finance agent inserts (orchestrator, m-71), creates task t-5a, binds it, and returns. The response is lost when a load balancer restarts. Two seconds later the orchestrator resends m-71; a different replica finds the row, sees the same digest, and returns t-5a in state working. One task exists.

At step 3 the agent needs an approver for an unusual adjustment and moves to TASK_STATE_INPUT_REQUIRED. The orchestrator sends m-72 with taskId t-5a; the server accepts it as a continuation. At step 5 the executor issues a ledger posting with key a2a:t-5a:step:5, crashes, restarts and posts again; the ledger API returns the first result. The task completes.

A week later the controller asks for a revised report. A careless client sends m-90 with taskId t-5a and receives UnsupportedOperationError, because the task is terminal. The correct message carries contextId c-recon-09, no task ID, and referenceTaskIds containing t-5a, producing task t-8c.

Failure modes

FailureCausePrevention
Duplicate tasks on retryServer does not deduplicate, or client mints a new messageId per attemptPersist messageId before sending; agree dedup in the contract; recover via ListTasks
Wrong task returnedDedup key not scoped by callerKey on principal and messageId
Silent body changeClient reuses a messageId for different contentStore a digest and reject mismatches
Duplicate after expiryDedup TTL shorter than client retry horizonSet TTL from the slowest client; alert on hits near expiry
Repeated side effectsExecutor re-runs a step after a crashKeys derived from taskId and step; record steps before executing
Message to finished taskClient treats taskId as a conversation handleUse contextId plus referenceTaskIds for follow-ups
Cross-conversation attachcontextId and taskId mismatch acceptedReject mismatches, as the spec requires

Export the dedup hit rate, digest mismatches, UnsupportedOperationError counts by client, and delegations left unknown. Streaming clients that reconnect should resubscribe rather than resend, as covered in A2A streaming.

Trade-offs

Server-assigned task IDs keep identity under the server's control: clients cannot collide with or guess each other's tasks, and the server can choose ID formats that suit its store. The price is that the client has no handle on the first request, which is exactly why messageId deduplication and returnImmediately matter. A dedup table costs a write per message, small next to an agent run. Recovery through ListTasks depends on history retention, which some agents trim for privacy or cost; if they do, deduplication on the server is the only reliable path.

What to do next

  1. Mint each messageId once, persist it before the first send, and reuse it on every retry.
  2. On servers, add a dedup table keyed on caller and messageId with a body digest and a retention window longer than any client's retries.
  3. Document in the integration contract whether your agent deduplicates, and test it with forced timeouts.
  4. Send long-running work with returnImmediately set and follow it with GetTask or SubscribeToTask.
  5. Implement lost-response recovery via ListTasks filtered by contextId, and an unknown state that escalates instead of resending a new message.
  6. Derive downstream idempotency keys from the taskId and step number for every side-effecting call.
  7. Use contextId with referenceTaskIds for follow-ups to finished tasks, never their taskId.
Key takeaway: In A2A 1.0 the server mints task IDs, so on a first request the client's only handle is the messageId it created. Deduplication on messageId is optional in the spec, so persist the messageId before sending, agree dedup behaviour with the remote agent, and keep a ListTasks-by-contextId recovery path. A new message with a known taskId is a continuation only while the task is interrupted. Terminal tasks reject messages with UnsupportedOperationError; follow-ups use contextId and referenceTaskIds. Deduplicating the task does not deduplicate its side effects, so derive idempotency keys from taskId and step.