Every A2A client makes the same decision on every call: send a message and wait for the answer in one response, or open a stream and watch the task progress. Early drafts of the protocol named the two methods tasks/send and tasks/sendSubscribe, and many tutorials, SDK samples and blog posts still use those names. The names have changed twice since then, but the decision has not, and getting it wrong shows up as gateway timeouts, duplicate work, silent hangs and users staring at a spinner for a minute.

This article explains the choice as it exists in A2A 1.0, where there are really three ways to wait rather than two: what each returns, a client that survives a dropped stream, what the server must guarantee, and the failure modes that decide the question in practice. For the full list of operations and objects, start with the A2A protocol overview.

Advertisement

The names, then and now

The method names have moved, so the first job is to translate whatever you are reading into the version you speak. The behaviour behind each row is the same idea: one method returns a single result, one returns a stream of events, and one reattaches to the stream of a task that already exists.

PurposeEarly drafts0.3 specification1.0 specification
Send, single responsetasks/sendmessage/sendSendMessage
Send, event streamtasks/sendSubscribemessage/streamSendStreamingMessage
Reattach to a task's streamnot confirmedtasks/resubscribeSubscribeToTask

In 1.0 the JSON-RPC method strings are the PascalCase names, and every request carries an A2A-Version header so the server knows which vocabulary you mean. Two 1.0 changes matter for code written against older samples. The send result now holds exactly one of a message or a task, so the agent may answer a trivial question with a plain Message and no task at all. And the final flag that older status-update events carried was removed in 1.0.0 as redundant: the stream ends when the task does, so the client reads the state instead of a flag. Code that waits for final: true never sees it from a 1.0 server and tends to mistake the normal close for a dropped connection.

Three ways to wait, not two

The old two-way framing hides a third option that is often the right one. A2A 1.0 gives the client three shapes for the same work.

Blocking SendMessage. By default the server must hold the request open until the task reaches a terminal state (completed, failed, canceled or rejected) or an interrupted state (input-required or auth-required), and then return the task. One request, one response, no progress in between. It is the simplest code you can write.

SendMessage with returnImmediately. Setting returnImmediately: true in the request configuration makes the server return the task as soon as it exists, typically in the submitted state. The client then learns the outcome by polling GetTask or by registering a webhook, either in the same request's taskPushNotificationConfig or separately; see configuring push notifications. This is the shape for work measured in minutes or hours, and for callers that should not hold a connection at all.

SendStreamingMessage. The same parameters, but the response is a Server-Sent Events stream. It follows one of two patterns: either exactly one Message and then the stream closes, or a Task first, followed by status-update and artifact-update events. The specification requires the stream to close when the task reaches a terminal state. It does not say the stream must close on input-required or auth-required, so a client should treat those states as the end of its turn whether or not the connection stays open.

One task, three ways to wait for itClientA2A 1.0Agenttask storeA. SendMessage (blocking, default)one request; response at terminal or interrupted stateB. SendMessage + returnImmediatelyClientpolls or waitsAgentworks in backgroundsend, Task returned at once (SUBMITTED)GetTask ... GetTaskor push notification to a registered webhookC. SendStreamingMessage (SSE)Clientreads eventsAgentemits events1. Task (first event)statusUpdate WORKINGartifactUpdate (append, lastChunk)statusUpdate COMPLETED, stream closesdropped? SubscribeToTask: first event is the current TaskAll three read and write the same task; only the delivery of progress differs.
Blocking returns once; return-immediately hands back a task and leaves waiting to the client; streaming delivers the task first and then every change until a terminal state.
Advertisement

What travels on the wire

Both send methods take identical parameters, which is what makes switching cheap. In the JSON-RPC binding each streamed event is a full JSON-RPC response carrying the original request id, whose result is a StreamResponse with exactly one of task, message, statusUpdate or artifactUpdate. Annotated end-to-end transcripts are in A2A examples; the essentials look like this.

# Blocking: one JSON-RPC request, one response
POST /a2a/jsonrpc
A2A-Version: 1.0
{"jsonrpc":"2.0","id":7,"method":"SendMessage","params":{
  "message":{"messageId":"m-7f2a","role":"ROLE_USER",
             "parts":[{"text":"Summarise the Q3 incident reports"}]}}}

-> {"jsonrpc":"2.0","id":7,"result":{"task":{"id":"task-91","contextId":"ctx-4",
      "status":{"state":"TASK_STATE_COMPLETED"},"artifacts":[ ... ]}}}

# Streaming: same params, different method, response is text/event-stream
{"jsonrpc":"2.0","id":8,"method":"SendStreamingMessage","params":{ ...same... }}

data: {"jsonrpc":"2.0","id":8,"result":{"task":{"id":"task-92","contextId":"ctx-4","status":{"state":"TASK_STATE_SUBMITTED"}}}}
data: {"jsonrpc":"2.0","id":8,"result":{"statusUpdate":{"taskId":"task-92","contextId":"ctx-4","status":{"state":"TASK_STATE_WORKING"}}}}
data: {"jsonrpc":"2.0","id":8,"result":{"artifactUpdate":{"taskId":"task-92","contextId":"ctx-4","artifact":{"artifactId":"summary","parts":[{"text":"Three incidents ..."}]}}}}
data: {"jsonrpc":"2.0","id":8,"result":{"statusUpdate":{"taskId":"task-92","contextId":"ctx-4","status":{"state":"TASK_STATE_COMPLETED"}}}}

Artifact updates can arrive in pieces. An update with append: true adds parts to an artifact already sent under the same artifactId, and lastChunk: true says that artifact is complete. A client that renders text as it arrives needs this; a client that only wants the final answer can ignore chunks and read the finished task.

Streaming is optional for the agent. If its card does not declare capabilities.streaming, calls to SendStreamingMessage or SubscribeToTask must fail with UnsupportedOperationError. Read the card first and pick the method from it, rather than catching the error on every call.

Choosing per call

There is no global answer, because the right shape depends on how long the work takes, whether a human is watching and what sits on the network path. The table is a starting rule; the failure modes section explains the exceptions.

SituationUseWhy
Answer in a few seconds, machine callerBlocking SendMessageLeast code; one round trip; easy retries
Human watching, output useful before it is completeSendStreamingMessagePartial artifacts and status as they happen
Tens of seconds, gateway with a short timeout in the pathStreaming or returnImmediatelyA silent blocking request gets cut by idle timeouts
Minutes to hours, or caller may go awayreturnImmediately plus push or pollingNo connection held; survives client restarts
Agent card has no streaming capabilityreturnImmediately plus pollingStreaming methods are rejected
Orchestrator fanning out to many agentsreturnImmediately, then GetTask or pushThousands of open streams cost memory and sockets

A useful default: stream when a person is waiting, return immediately otherwise, and keep blocking for short calls that are cheap to retry.

A client that survives a dropped stream

The hard part of streaming is not reading events; it is what happens when the connection breaks. A stream that dies after the Task event has told you the task id, and the task keeps running on the server regardless of your connection. The client should reattach with SubscribeToTask, whose first event is the current Task, so you rebuild your view from it instead of replaying missed events. If the task finished while you were disconnected, subscribing fails with UnsupportedOperationError because the task is terminal, and the right fallback is GetTask. If the stream dies before the Task event, you do not know whether a task was created, so treat it as a failed send.

import json, time, uuid
import httpx

URL = "https://agent.example.com/a2a/jsonrpc"
HDRS = {"A2A-Version": "1.0", "Authorization": "Bearer <token>"}
TERMINAL = {"TASK_STATE_COMPLETED", "TASK_STATE_FAILED",
            "TASK_STATE_CANCELED", "TASK_STATE_REJECTED"}
INTERRUPTED = {"TASK_STATE_INPUT_REQUIRED", "TASK_STATE_AUTH_REQUIRED"}

def rpc(method, params, rid):
    return {"jsonrpc": "2.0", "id": rid, "method": method, "params": params}

def sse_events(resp):
    """Yield the JSON-RPC result of each SSE data line."""
    for line in resp.iter_lines():
        if line.startswith("data:"):
            msg = json.loads(line[5:])
            if "error" in msg:
                raise RuntimeError(msg["error"])
            yield msg["result"]

def apply(task, ev):
    """Fold one StreamResponse event into a local task view."""
    if "task" in ev:
        return ev["task"]
    if "statusUpdate" in ev:
        task["status"] = ev["statusUpdate"]["status"]
    elif "artifactUpdate" in ev:
        u = ev["artifactUpdate"]; art = u["artifact"]
        arts = task.setdefault("artifacts", [])
        old = next((a for a in arts if a["artifactId"] == art["artifactId"]), None)
        if old and u.get("append"):
            old["parts"].extend(art["parts"])
        elif old:
            old["parts"] = art["parts"]
        else:
            arts.append(art)
    return task

def stream_task(client, text, max_reconnects=5):
    msg = {"messageId": str(uuid.uuid4()), "role": "ROLE_USER",
           "parts": [{"text": text}]}
    req = rpc("SendStreamingMessage", {"message": msg}, 1)
    task, attempt = None, 0
    while True:
        try:
            with client.stream("POST", URL, json=req, headers=HDRS,
                               timeout=httpx.Timeout(10, read=90)) as resp:
                for ev in sse_events(resp):
                    if "message" in ev:          # message-only stream
                        return ev["message"]
                    task = apply(task or {}, ev)
                    state = task["status"]["state"]
                    if state in TERMINAL or state in INTERRUPTED:
                        return task
            # stream ended without a final state: fall through and reattach
        except httpx.TransportError:
            pass
        if task is None:      # never saw the Task: we cannot resume, only retry
            raise RuntimeError("stream failed before the task was created")
        attempt += 1
        if attempt > max_reconnects:
            break
        time.sleep(min(2 ** attempt, 30))
        req = rpc("SubscribeToTask", {"id": task["id"]}, 1 + attempt)
    # last resort: read the task; it may have finished while we were away
    r = client.post(URL, json=rpc("GetTask", {"id": task["id"]}, 99), headers=HDRS)
    return r.json()["result"]

The read timeout applies per read, so a long stream stays open while events keep coming, and reconnects back off. The local view is replaced whenever a full Task arrives, which is what makes reconnecting safe: the server's task store is the truth and the stream only a view of it. Background on recovering from crashes on both sides is in A2A error recovery.

What the server must get right

  • The task store is the truth. Persist each state change before emitting its event. A reconnecting client calls SubscribeToTask and must receive a Task that reflects everything already streamed.
  • Task first. For a task-lifecycle stream the first event must be the Task. Clients depend on it to learn the id they need for reconnecting.
  • Broadcast in order. The specification allows several concurrent streams for one task and requires every stream to receive the same events in the same order. Fan out from one ordered event log per task, not from separate code paths.
  • Never let a slow reader block the work. Give each stream a bounded queue. If a client cannot keep up, close its stream and let it resubscribe; the work itself must not wait for a socket.
  • Close on terminal states. Leaving a stream open after completion leaks connections and leaves clients waiting for something that will never come.

Worked example: a ninety-second report

An orchestrator asks a research agent for an incident summary. The agent spends 5 seconds planning, then 80 seconds reading documents, and writes the summary in the last 5 seconds. Between them sits a load balancer with a 60-second idle timeout, a common default for managed HTTP load balancers.

With blocking SendMessage, the request is silent for 90 seconds. At second 60 the load balancer closes the connection and the client sees a gateway error. The task, however, is still running, and it completes at second 90 with nobody listening. A naive retry starts a second, identical task.

With SendStreamingMessage, the client gets the Task at about 100 milliseconds, a working status at second 5, and progress updates while documents are read, so the connection is never idle for long enough to be cut if the agent emits an update at least every few tens of seconds. The summary arrives in chunks between seconds 85 and 90 and the stream closes on the completed state. If a deploy restarts the proxy at second 40, the client reattaches with the task id and loses nothing.

With returnImmediately, the client has the task id within 100 milliseconds and polls GetTask every 10 seconds, making about nine extra small requests, or receives one webhook at second 90. Nothing is held open, so this scales to thousands of concurrent reports, at the cost of up to one polling interval of extra latency.

Failure modes

  • Buffering proxies. A reverse proxy that buffers responses holds every event until the stream ends, so streaming silently degrades to slow blocking. Disable response buffering for the A2A route (in nginx, proxy_buffering off or an X-Accel-Buffering: no response header) and test through the real path, not localhost.
  • Idle timeouts. Long silent periods inside a stream are cut just like silent blocking requests. Emit status updates during long steps; some servers also send SSE comment lines as keep-alives, which conforming SSE parsers ignore.
  • Retrying a blocking send that timed out. The task usually survived. Before retrying, look for it with ListTasks in the context, or reuse the same messageId so a server that deduplicates can recognise the retry.
  • Waiting for a removed flag. Clients ported from older samples loop until final is true. In 1.0 decide from the task state.
  • Hanging on interrupted states. A client that only stops on terminal states will wait forever when the agent asks for input. Treat input-required and auth-required as the end of your turn.
  • Resubscribing to a finished task. SubscribeToTask fails on terminal tasks; fall back to GetTask instead of retrying the subscription.

Trade-offs

DimensionBlockingreturnImmediatelyStreaming
Client complexityLowestPolling loop or webhook endpointSSE parsing, chunk assembly, reconnects
Time to first feedbackEnd of taskTask id at once, then polling intervalTask at once, then live
Connection heldWhole task, silentNoneWhole task, active
Survives client restartNoYesYes, with the task id
Agent requirementNoneNone (push needs capability)capabilities.streaming
Scales to many tasksPoorlyBestModerately

What to do next

  1. Search your code and SDK samples for tasks/send, tasks/sendSubscribe, message/stream and loops on final, and port them to the 1.0 names and state checks.
  2. Read each peer's Agent Card at startup and record whether it declares streaming and push notifications; choose methods from that, not from try-and-catch.
  3. Classify your calls by expected duration and audience using the table above, and set blocking only for short machine-to-machine calls.
  4. Put a stream reconnect path in your client: remember the task id, resubscribe with backoff, fall back to GetTask.
  5. Run a streaming call through your real load balancers and proxies and confirm events arrive as they are sent, not all at the end.
  6. On the server, persist before emitting, emit the Task first, close streams on terminal states and cap per-stream buffers. For lifecycle rules see A2A task lifecycle states.
Key takeaway: The old tasks/send and tasks/sendSubscribe names are SendMessage and SendStreamingMessage in A2A 1.0, with SubscribeToTask to reattach. There are three ways to wait: blocking until a terminal or interrupted state, returning the task immediately and polling or receiving a push, or streaming the Task followed by status and artifact updates until the task ends. Choose per call by duration, audience and the network path. Streaming clients must remember the task id, reconnect with SubscribeToTask and fall back to GetTask, and servers must persist before emitting, broadcast in order and never let a slow reader stall the work.