Delegation is the moment one agent stops doing a piece of work itself and asks another agent to do it. In the Agent2Agent (A2A) protocol that hand-off is concrete: the delegator sends a message, the remote agent creates a task, and the task moves through a defined set of states until it produces artifacts or stops. The protocol tells you how to send and observe. It does not tell you how to remain accountable for work you no longer control, and that is where most delegation bugs live.

This article treats a single delegation from the delegator's side. It defines a delegation record that survives crashes, maps each task state to an action, handles the two interrupted states, carries budgets and depth limits through chains of sub-delegation, and checks results before using them. How a planner decomposes a goal into many such delegations is covered in A2A orchestration patterns; this page is about doing one of them correctly.

Advertisement

A note on protocol versions

Everything here follows the A2A 1.0 specification published at a2a-protocol.org. Version 1.0 renamed the JSON-RPC methods to PascalCase (SendMessage, SendStreamingMessage, GetTask, ListTasks, CancelTask, SubscribeToTask) and the task states to TASK_STATE_* values. Older material, including some pages on this site, uses the pre-1.0 names such as message/send and input-required; the concepts map one to one. Check which version a peer speaks, and keep any translation in one adapter.

What delegation is, and what it is not

A delegation transfers execution, not responsibility. If a research agent delegates clause extraction to a legal agent and the legal agent returns nonsense, the research agent's user still receives a wrong answer. So the delegator keeps four obligations: it must know at every moment which remote task is doing its work, it must decide what to do in every state that task can enter, it must bound how much time and money the task may consume, and it must check the result before building on it.

Three things are easy to confuse with delegation. A synchronous tool call has no remote lifecycle: it returns or fails. An event broadcast has no single owner of the outcome. And a conversation turn inside a shared contextId is not a new delegation unless it creates a new task. Everything below is keyed on one record per delegated task.

Advertisement

The architecture: a record that outlives the process

The core component is a durable delegation store owned by the delegator. Each row says what was asked, of whom, under which identifiers, with which budget, and what the delegator last observed. The remote task is a mirror of that row's subject, not the source of truth for your workflow; if your process restarts, the row tells you which remote tasks to reconcile.

One delegation, seen from the delegator: the record is the source of truth, the remote task is a mirrorDelegating agentPlanner / callerdecides what to hand offDelegation storerecord: messageId, taskId, budgetWatcherstream / pollWebhookpush receiverAcceptance checksRemote agentAgent Cardskills, auth schemesTaskid, contextId, statusArtifactsparts: text, data, filesSub-delegates (optional)1. fetch card2. SendMessage (returnImmediately)3. Task id + state4. statusUpdate / artifactUpdate4b. pushWrite the record BEFORE sending; reconcile it with GetTask after any crash; accept nothing unchecked.
The delegator writes a record, sends with returnImmediately, then observes through a stream, push notifications or polling. Results pass acceptance checks before anything downstream uses them.
FieldPurpose
delegation_idYour key; also used to derive the A2A messageId so retries are recognisable
agent_url, skillWhich remote agent and advertised skill was chosen, and the card version seen
task_id, context_idFilled from the first response; null means the send may or may not have landed
stateLast observed TaskState, plus your own SENDING pre-state
deadline, budgetWall-clock limit and cost ceiling for this delegation
depth, chainHow many hops deep this request is, and the agents already in the chain

Sending: idempotent and non-blocking

In A2A 1.0 the SendMessage request carries a message and an optional configuration. The message's messageId is required and created by the sender; the specification says send operations may be idempotent and that agents may use the messageId to detect duplicates. That is permission, not a guarantee, so derive the id deterministically from your delegation id and still treat a retried send as possibly duplicated.

By default SendMessage waits until the task reaches a terminal or interrupted state. For anything longer than a few seconds that ties up a connection and turns a network blip into an unknown outcome. Set returnImmediately to true, persist the returned task id, and observe separately. The configuration can also carry a taskPushNotificationConfig with a webhook url and a token so the remote agent can call you back.

import uuid, httpx

def message_id(delegation_id: str, attempt_kind: str = "create") -> str:
    # Same delegation + same purpose -> same messageId, so a retry looks like a duplicate.
    return str(uuid.uuid5(uuid.NAMESPACE_URL, f"deleg:{delegation_id}:{attempt_kind}"))

async def delegate(store, rpc_url, auth, d):
    store.update(d.id, state="SENDING")                      # write intent first
    body = {
        "jsonrpc": "2.0", "id": d.id, "method": "SendMessage",
        "params": {
            "message": {
                "messageId": message_id(d.id),
                "role": "ROLE_USER",
                "parts": [{"text": d.instructions},
                          {"data": d.inputs, "mediaType": "application/json"}],
                "referenceTaskIds": d.related_task_ids,
            },
            "configuration": {
                "returnImmediately": True,
                "acceptedOutputModes": ["application/json", "text/plain"],
                "historyLength": 0,
                "taskPushNotificationConfig": {"url": d.webhook_url, "token": d.webhook_token},
            },
            # Local convention, not a spec field: budget, deadline and chain travel in metadata.
            "metadata": {"x-deleg": {"deadline": d.deadline_iso, "budget_usd": d.budget,
                                     "depth": d.depth + 1, "chain": d.chain + [d.self_url]}},
        },
    }
    async with httpx.AsyncClient(timeout=15) as http:
        r = await http.post(rpc_url, json=body, headers=auth.headers())
    r.raise_for_status()
    reply = r.json()
    if "error" in reply:
        store.update(d.id, state="SEND_FAILED", error=reply["error"])
        return
    task = reply["result"].get("task")
    if task is None:                                           # agent answered with a Message
        store.update(d.id, state="ANSWERED_DIRECT", result=reply["result"]["message"])
        return
    store.update(d.id, task_id=task["id"], context_id=task["contextId"],
                 state=task["status"]["state"])

Two details in that code matter. The response is a one-of: an agent may reply with a plain message instead of creating a task, and a delegator that assumes a task crashes on simple agents. And the budget, deadline and chain in metadata are a convention you agree with your peers; A2A does not define them, and an agent you do not control will ignore them.

Mapping every state to a delegator action

A delegation is only robust if every state has a handler, including the ones you think will never happen. The table is the policy most teams converge on.

Remote stateKindDelegator action
TASK_STATE_SUBMITTED, TASK_STATE_WORKINGActiveKeep observing; enforce your deadline locally
TASK_STATE_INPUT_REQUIREDInterruptedAnswer from your own context, or escalate to your caller
TASK_STATE_AUTH_REQUIREDInterruptedObtain the credential through your auth flow, never by forwarding your own
TASK_STATE_COMPLETEDTerminalRun acceptance checks on artifacts, then release downstream
TASK_STATE_FAILEDTerminalClassify; retry with a new task or fall back to another agent
TASK_STATE_REJECTEDTerminalThe agent declined; do not retry the same agent with the same input
TASK_STATE_CANCELEDTerminalExpected if you cancelled; otherwise treat as failure

Terminal is final. The specification says a task in a terminal state cannot accept further messages and returns UnsupportedOperationError (JSON-RPC code -32004). A retry or refinement is therefore a new task, ideally in the same contextId with the old id in referenceTaskIds so the agent can reuse what it learned.

Interrupted states: input-required and auth-required

An interrupted task is paused waiting for you. To continue, send a new message carrying the same taskId and contextId. The hard question is who answers. A good delegator answers from its own context when the question is about data it holds ("which jurisdiction?" when the jurisdiction is in the original request) and escalates when only a human or the upstream caller knows. Escalation means your own task moves to its interrupted state, so the pause propagates up the chain rather than being guessed through.

Guessing is the failure to avoid. A delegator that auto-answers every question with a language model produces confident, unaccountable inputs that the remote agent will treat as instructions. Log every answer you send with its source, and cap the number of input rounds per delegation; a task that asks for input five times is usually mis-scoped.

TASK_STATE_AUTH_REQUIRED deserves extra care. The remote agent needs authority it does not have. The right response is to obtain a credential scoped to that agent and that task, for example through an OAuth flow for the end user, as described in securing A2A endpoints with OAuth 2. Forwarding the delegator's own broad token turns one agent's compromise into everyone's.

Budgets, deadlines and depth across sub-delegation

A delegate may delegate again. Without limits, a chain of agents can loop (A asks B, B asks A), fan out without bound, or outlive the user's patience. Three rules contain it, and because none of them are protocol fields, they work only among agents that agree to them.

  • Deadlines shrink. Each hop passes its remaining time minus a safety margin for its own processing. An agent that receives a deadline already in the past should reject immediately rather than start work.
  • Budgets split. A delegator with a $2 ceiling that sends three sub-tasks assigns each a share and keeps a reserve for retries. The sum of child budgets never exceeds the parent's remainder.
  • Depth and chain are checked. Reject when depth exceeds a small limit, typically three or four, and when your own URL already appears in the chain. That turns a loop into a fast, explainable rejection.

Enforce all three locally as well. If a remote agent ignores the deadline, your watcher still fires at your own deadline and calls CancelTask. Cancellation is idempotent in A2A, and a cancel that loses a race with completion returns TaskNotCancelableError (-32002), which you handle by reading the final state. Task cancellation covers propagation to sub-tasks.

Observing and recovering after a crash

A2A offers three ways to learn about progress: polling with GetTask, streaming via SendStreamingMessage or SubscribeToTask, and push notifications to a webhook. Use streaming for interactive work, push for long jobs, and polling as the reconciliation path that always works. Push notifications explains how to verify the incoming webhook calls.

Recovery is a loop over your store at startup. Records in SENDING with no task id are ambiguous: the send may have landed. If the delegation went into an existing context, look for it with ListTasks filtered by that context; otherwise re-send with the same derived messageId and accept that an agent which does not deduplicate may run it twice. Records with a task id and a non-terminal state get a GetTask, then a fresh subscription. A TaskNotFoundError (-32001) on a task you created means it was purged or never existed for your identity; mark it failed and decide on retry by policy.

TERMINAL = {"TASK_STATE_COMPLETED", "TASK_STATE_FAILED",
            "TASK_STATE_CANCELED", "TASK_STATE_REJECTED"}

async def reconcile(store, client):
    for d in store.open_delegations():
        if d.task_id is None:                     # ambiguous send: same messageId
            await client.find_in_context_or_resend(d)
            continue
        task = await client.get_task(d.task_id, history_length=0)
        if task is None:                          # -32001 TaskNotFoundError
            store.update(d.id, state="LOST"); continue
        store.update(d.id, state=task["status"]["state"])
        if task["status"]["state"] not in TERMINAL:
            client.subscribe_in_background(d.task_id)

Accepting results

A completed task means the remote agent believes it is done. Before releasing artifacts downstream, check them. Validate structured data parts against the schema you expected, reject unexpected media types, bound sizes, and treat any text from another agent as untrusted input: it must not be concatenated into your own prompts as instructions. Record which agent, skill and task produced each artifact so a later bad result can be traced and that agent's reputation adjusted.

Streaming agents emit artifactUpdate events before completion; show them as progress, but commit only what a completed task delivers.

Worked example: delegating clause extraction

A contract-review agent receives a 60-page supply agreement with a ten-minute deadline and a $1.50 budget. It discovers a legal-extraction agent whose card advertises a clause-extraction skill, writes a delegation record with depth 1 and a seven-minute deadline, and sends the PDF reference as a file part plus a data part listing the clause types it wants. The reply is a task in TASK_STATE_SUBMITTED, stored immediately.

Two minutes in, the task moves to TASK_STATE_INPUT_REQUIRED asking which governing law applies. The original request named English law, so the delegator answers from context, with the same task and context ids, and logs the source. The legal agent sub-delegates translation of a German annex, passing depth 2 and a four-minute deadline. Five minutes in, the task completes with a JSON artifact. Validation finds one clause with an empty text field; the delegator sends a refinement as a new task in the same context, referencing the first task id, and merges the two results after both pass the schema.

Failure modes

  • Lost task ids. Sending before persisting means a crash leaves remote work nobody watches.
  • Blocking sends. Default-blocking SendMessage on long tasks turns a proxy timeout into an unknown outcome.
  • Assuming a task. Agents may answer with a message; code that dereferences the task crashes.
  • Messaging a terminal task. Refinements must be new tasks, or you collect -32004 errors.
  • Guessed inputs. Auto-answering input requests injects unverified instructions into another agent.
  • Credential forwarding. Passing your own token to satisfy auth-required widens the blast radius of every agent in the chain.
  • Unbounded chains. Without depth, chain and deadline checks, loops and fan-out storms burn budget silently.

Operations and trade-offs

Track per remote agent and skill: delegation count, completion rate, rejection rate, input rounds per task, time to terminal state, and cost per completed task. Alert on open delegations past their deadline, which almost always means a watcher died. Polling is simple but wastes calls, push needs a reachable, authenticated webhook, and streaming ties a connection to each active task. Most production delegators use push or streaming for speed and a slow reconciliation poll for correctness.

What to do next

  1. Create a delegation table with the fields above and make every send write it first.
  2. Derive messageId from the delegation id and set returnImmediately to true for anything long.
  3. Write a handler for every TaskState, including rejected and canceled, and a unit test for each.
  4. Decide, per skill, which input-required questions you may answer and which you escalate.
  5. Agree on a metadata convention for deadline, budget, depth and chain with the agents you control, and enforce all four locally.
  6. Add a startup reconciliation loop using GetTask and SubscribeToTask, then kill the process mid-delegation in a test and confirm it recovers.
  7. Validate every artifact against a schema before use and record its provenance.
Key takeaway: Delegating over A2A hands off execution, not responsibility. Persist a delegation record before you send, use a derived messageId and returnImmediately, and give every A2A 1.0 task state a handler. Answer interrupted tasks only from context you actually hold, and never forward your own credentials. Carry deadlines, budgets and depth through sub-delegation by agreed convention while still enforcing them locally. Reconcile with GetTask after a crash, and check every artifact before you use it.