Running an agent in two regions is a deployment problem. Running an A2A agent in two regions is also a routing problem, because A2A is not a stateless request/response protocol. When an agent accepts a message that starts long-running work, it creates a task. The task has an ID the server generates, a lifecycle (submitted, working, input-required, completed, failed, canceled and so on), a history, artifacts, and optionally push-notification configurations that the server must keep until the task finishes. Every later call about that task (polling it, re-attaching a stream, cancelling it, registering a webhook) has to reach the one place where that state lives.

This article is about the topology that makes that work: how new tasks are placed, how follow-up calls find the owning region, what to publish in the Agent Card, how push notifications cross region boundaries, and what happens when the owning region dies mid-task. The code is plain Java 17 with java.net.http so it does not depend on any particular SDK version. For deploying the ADK agent itself across regions (home-region pinning, model quota, failover of the runtime), read Multi-Region Deployment for ADK Java first; this page assumes it.

Why A2A tasks have a home region

A2A across two regions: stateless entry, stateful task ownershipCalling agent / clientSendMessage, GetTask, SubscribeToTaskGlobal A2A routernew task: nearest healthy regionRegion us-eastRegion eu-westA2A server + ADK agentmints task id use1-...Task storestate, historyPush configspersist to terminalWebhook senderretries, idempotent idsA2A server + ADK agentmints task id euw1-...Task storestate, historyPush configspersist to terminalWebhook senderretries, idempotent idsnew taskfollow-ups by ownerTask locator (replicated)
New tasks go to the nearest healthy region; every follow-up about a task goes to the region that minted its ID, found from the prefix or the replicated locator.

Start from what the protocol guarantees. The A2A 1.0 specification says task IDs are server-generated when a new task is created in response to a message; clients cannot choose them. It says server-generated contextId values should be treated as opaque by clients, that an agent must reject a message whose contextId and taskId do not match, and that a push-notification configuration must persist until the task completes or the configuration is deleted. Put those together and you have three kinds of state with region affinity:

  • Task state: status, history, artifacts. Read by GetTask and ListTasks, changed by CancelTask.
  • Live streams: a SendStreamingMessage or SubscribeToTask connection is an open Server-Sent Events response held by one process in one region.
  • Push configurations: webhook URL, token and authentication details registered against a task, which the owning region's sender reads every time the task changes.

If a client creates a task through us-east and its next GetTask is geo-balanced to eu-west, eu-west answers "task not found" or, worse, answers from a stale replica. The rest of the design is about making that impossible.

Method names: 1.0 versus 0.x

A note on names before the code. The A2A 1.0 JSON-RPC binding uses PascalCase method names that match gRPC: SendMessage, SendStreamingMessage, GetTask, ListTasks, CancelTask, SubscribeToTask and CreateTaskPushNotificationConfig. The REST binding maps these to paths such as GET /tasks/{id} and POST /tasks/{id}:subscribe. The 0.x releases used slash-style JSON-RPC names (message/send, message/stream, tasks/get, tasks/resubscribe), which other pages on this site use. Check which version your SDK and peers speak (1.0 defines an A2A-Version header); the routing logic below is identical for both.

Three topologies

There are three workable shapes, and they differ in where the routing decision lives.

TopologyHow a follow-up finds its taskGood forCost
Regional endpoints, regional cardsThe client keeps talking to the URL that created the taskData residency; peers that already pin a regionClients must handle failover themselves
Global router, owner-encoded IDsRouter reads the region prefix from the task ID it mintedOne public URL; your own agentsOnly works for IDs you mint
Global router plus task locatorRouter looks the task ID up in a replicated tableFronting third-party or legacy agents with opaque IDsOne more store on the hot path

The diagram shows the second and third combined, which is what most teams end up with. New tasks go to the nearest healthy region; follow-ups go to the owner. The router is stateless and can run at the edge in every region, because the only state it needs is either inside the ID or in the replicated locator.

An owner-aware router in Java

The routing core decides a destination and forwards the bytes without rewriting the A2A payload, so it survives protocol additions.

import java.net.URI;
import java.net.http.*;
import java.time.Duration;
import java.util.*;

record Region(String code, URI a2aEndpoint) {}

interface TaskLocator {                       // replicated KV: taskId -> region code
    Optional<String> regionOf(String taskId);
    void record(String taskId, String regionCode);
}

final class TaskRouter {
    private final Map<String, Region> regions;
    private final TaskLocator locator;
    private final HttpClient http = HttpClient.newBuilder()
            .connectTimeout(Duration.ofSeconds(2)).build();

    TaskRouter(Map<String, Region> regions, TaskLocator locator) {
        this.regions = regions;
        this.locator = locator;
    }

    /** Owner of an existing task: prefix we minted first, locator second, never a guess. */
    Region ownerOf(String taskId) {
        int dash = taskId.indexOf('-');
        if (dash > 0) {
            Region r = regions.get(taskId.substring(0, dash));
            if (r != null) return r;
        }
        return locator.regionOf(taskId).map(regions::get)
                .orElseThrow(() -> new NoSuchElementException("unknown task " + taskId));
    }

    /** Forward a follow-up JSON-RPC call (GetTask, CancelTask, ...) to the owning region. */
    HttpResponse<String> forward(String taskId, String jsonRpcBody, String auth)
            throws Exception {
        Region owner = ownerOf(taskId);
        HttpRequest req = HttpRequest.newBuilder(owner.a2aEndpoint())
                .timeout(Duration.ofSeconds(10))
                .header("Content-Type", "application/json")
                .header("Authorization", auth)        // pass the caller's identity through
                .POST(HttpRequest.BodyPublishers.ofString(jsonRpcBody))
                .build();
        return http.send(req, HttpResponse.BodyHandlers.ofString());
    }
}

Three details carry the correctness. First, the prefix is trusted only when it names a region you operate; other agents' IDs are treated as opaque and go through the locator. Second, an unknown task is an error, never a fallback to the nearest region, which would tell the client a running task has vanished. Third, the router passes the caller's credentials through rather than calling with its own, because the specification requires task operations to check the caller's access rights, and that check has to happen in the region that owns the task.

To mint owner-encoded IDs, make the A2A server's task-ID generator prepend the region code, for example euw1- followed by a UUID. For the locator path, write the mapping when the SendMessage response comes back with a new task, before returning it to the caller, so a follow-up can never outrun the record.

What to put in the Agent Card

The Agent Card, served at /.well-known/agent-card.json, is the discovery document. In 1.0 its supportedInterfaces field is an ordered list of endpoints and protocol bindings, and the first entry is preferred. It is tempting to list one URL per region there and let clients pick. Don't: clients reasonably treat the entries as interchangeable ways to reach the same agent, and a client that starts a task on the first URL and polls the second has recreated the cross-region miss.

Two patterns work. With a global router, publish one card whose interface URL is the router; the client never needs to know about regions. With regional endpoints, publish a separate card per region (for example on us.agents.example.com and eu.agents.example.com) and make region choice an explicit client decision, which is what you want when an EU tenant's data must not be processed elsewhere. Either way, keep skills and capabilities identical across regions, or a failover becomes a silent capability change.

1.0 also defines a tenant field that a server serving many agents or tenants interprets for routing. Resolve the tenant to its home region first, then the task to its owner.

Push notifications across regions

Push notifications reverse the direction of traffic. The owning region's webhook sender POSTs task updates to a URL the client registered. That URL may be in another region or another company, so treat it like any outbound integration:

  • Send from the owner only. The sender reads the task's push configurations from the owner's store. A replica in another region must never send, or the client receives each update twice from two senders that disagree on ordering.
  • Expect duplicates and handle them. The spec tells clients to process notifications idempotently because duplicate deliveries may occur. On the receiving side, key on task ID plus the update's status and timestamp, and drop anything not newer than what you hold.
  • Validate before trusting. The spec requires the receiver to check the task ID is one it expects, and recommends verifying the source. Use the token you registered with the configuration and the authentication scheme you agreed, and treat the payload as a hint: on any doubt, call GetTask through the router and act on that answer.

When the owning region dies

Region loss is where A2A differs most from a stateless service. Suppose eu-west goes down while task euw1-7f3c is working. The router can still parse the prefix, so it knows exactly who owns the task, and it should say so: return a retryable error for follow-ups rather than "task not found". Then you choose, ahead of time, one of three policies:

  1. Wait. Tasks are lost only if the store is lost. If eu-west's task store survives (or is restored), tasks resume when the region returns. Simplest, and correct for most interactive work that is cheap to redo.
  2. Fail and restart. Have the client start a new task in another region with the same inputs. Because the new task gets a new server-generated ID, the client must not assume the work is deduplicated. The spec says agents may use the messageId to detect duplicate messages; it does not require it, so any operation with external side effects needs its own idempotency key carried in the message and checked by the tool that performs the side effect.
  3. Adopt. Replicate the task store and let the survivor take ownership via the locator, resuming from the last checkpoint. It needs a per-task fencing token so a returning old owner cannot keep writing, and only pays for long, expensive tasks.

Open streams break regardless. A client holding a SubscribeToTask stream to eu-west sees the connection drop and should re-subscribe through the router with backoff, then call GetTask to recover any updates missed while it was disconnected. The A2A client streaming recipe covers the reconnect loop.

Worked example: a delegated review across the Atlantic

Worked example. A procurement agent in us-east delegates contract review to a legal agent that runs in both regions. The client is in Frankfurt and the contract belongs to an EU entity.

  1. The procurement agent sends SendMessage to the router with the contract and its tenant. The router resolves the tenant's home to eu-west and forwards there.
  2. eu-west creates task euw1-2b91 in state submitted, registers the push configuration the caller attached, and returns the task.
  3. The router writes nothing for this task: the prefix is enough. A third-party agent behind the same router would have needed a locator record here.
  4. Ten minutes later the task enters input-required: the agent wants the governing-law clause clarified. The eu-west sender POSTs the update to the procurement agent's webhook in us-east, which checks the task ID and token, then calls GetTask via the router to confirm.
  5. The procurement agent replies with a message carrying task ID euw1-2b91 and the matching context ID. The router reads the prefix and forwards to eu-west; the agent resumes and later completes with an artifact.

Four calls crossed the ocean: the send, the webhook, the confirmation and the reply. That is negligible next to minutes of model time, but not for agents that chat in a tight loop: keep chatty pairs in one region and let only coarse delegations cross.

Operational guidance

  • Tag every span and log line with the owning region and the task ID. Cross-region bugs are invisible without them. See A2A error propagation for keeping failures legible across the boundary.
  • Alert on wrong-region hits. Count requests where the router had to forward to a non-local region for a task created locally; a non-zero rate means some client is bypassing the router.
  • Expire locator entries after the task is terminal plus your history retention window, or the table grows without bound.
  • Keep timeouts asymmetric. Short connect timeouts to the owner, long read timeouts for streams, and never blindly retry a SendMessage that may already have created a task.

Failure modes

SymptomCauseFix
Intermittent task not foundFollow-ups load-balanced away from the ownerRoute by owner; no geo balancing for task calls
Every update arrives twiceSenders in two regions read replicated push configsOnly the owner sends; replicas are read-only
Duplicate side effects after failoverClient restarted the task elsewhereIdempotency key checked by the side-effecting tool
Stream silently stallsOwner region gone, client never re-subscribesHeartbeat timeout, re-subscribe, then GetTask
EU data processed in us-eastRegion chosen by latency, not tenantResolve tenant home before nearest region

Trade-offs

Owner-encoded IDs cost nothing at runtime and survive locator outages, but only work for IDs you mint and leak the region to clients (harmless for most, a concern for some). A locator works for anyone's IDs but adds a replicated store to the critical path, and its replication lag must be shorter than the gap between a task's creation and its first follow-up. Regional cards push complexity to clients but make residency auditable. Task-store replication enables adoption but brings fencing and double execution risks; most teams should prefer wait or restart with idempotent side effects. For the protocol basics behind all of this, see A2A integration in ADK Java.

What to do next

  1. List every A2A operation your agents call and mark which ones carry a task ID; those are the ones that need owner routing.
  2. Decide the topology: regional cards for residency, a global router for one URL, a locator if you front agents you did not write.
  3. Make your server mint region-prefixed task IDs, and verify the router rejects unknown prefixes instead of guessing.
  4. Confirm which A2A version and method names your SDK and peers use, and test both spellings through the router if you have mixed peers.
  5. Make webhook receivers idempotent and have them confirm with GetTask.
  6. Pick wait, restart or adopt for region loss, write it down, and rehearse it with a task in flight.
Key takeaway: An A2A task lives in exactly one region, because its ID, state, streams and push configurations are created there. Send new tasks to the nearest healthy region allowed for the tenant, and route every follow-up to the owner using a region prefix on IDs you mint or a replicated locator for IDs you don't. Publish cards that never make clients guess between regions, send webhooks only from the owner, make receivers idempotent, and decide before an outage whether tasks wait, restart with idempotent side effects, or get adopted.