Running an agent in two regions is a deployment problem. Running an A2A agent in two regions is also a routing problem, because A2A is not a stateless request/response protocol. When an agent accepts a message that starts long-running work, it creates a task. The task has an ID the server generates, a lifecycle (submitted, working, input-required, completed, failed, canceled and so on), a history, artifacts, and optionally push-notification configurations that the server must keep until the task finishes. Every later call about that task (polling it, re-attaching a stream, cancelling it, registering a webhook) has to reach the one place where that state lives.
This article is about the topology that makes that work: how new tasks are placed, how follow-up calls find the owning region, what to publish in the Agent Card, how push notifications cross region boundaries, and what happens when the owning region dies mid-task. The code is plain Java 17 with java.net.http so it does not depend on any particular SDK version. For deploying the ADK agent itself across regions (home-region pinning, model quota, failover of the runtime), read Multi-Region Deployment for ADK Java first; this page assumes it.
Why A2A tasks have a home region
Start from what the protocol guarantees. The A2A 1.0 specification says task IDs are server-generated when a new task is created in response to a message; clients cannot choose them. It says server-generated contextId values should be treated as opaque by clients, that an agent must reject a message whose contextId and taskId do not match, and that a push-notification configuration must persist until the task completes or the configuration is deleted. Put those together and you have three kinds of state with region affinity:
- Task state: status, history, artifacts. Read by
GetTaskandListTasks, changed byCancelTask. - Live streams: a
SendStreamingMessageorSubscribeToTaskconnection is an open Server-Sent Events response held by one process in one region. - Push configurations: webhook URL, token and authentication details registered against a task, which the owning region's sender reads every time the task changes.
If a client creates a task through us-east and its next GetTask is geo-balanced to eu-west, eu-west answers "task not found" or, worse, answers from a stale replica. The rest of the design is about making that impossible.
Method names: 1.0 versus 0.x
A note on names before the code. The A2A 1.0 JSON-RPC binding uses PascalCase method names that match gRPC: SendMessage, SendStreamingMessage, GetTask, ListTasks, CancelTask, SubscribeToTask and CreateTaskPushNotificationConfig. The REST binding maps these to paths such as GET /tasks/{id} and POST /tasks/{id}:subscribe. The 0.x releases used slash-style JSON-RPC names (message/send, message/stream, tasks/get, tasks/resubscribe), which other pages on this site use. Check which version your SDK and peers speak (1.0 defines an A2A-Version header); the routing logic below is identical for both.
Three topologies
There are three workable shapes, and they differ in where the routing decision lives.
| Topology | How a follow-up finds its task | Good for | Cost |
|---|---|---|---|
| Regional endpoints, regional cards | The client keeps talking to the URL that created the task | Data residency; peers that already pin a region | Clients must handle failover themselves |
| Global router, owner-encoded IDs | Router reads the region prefix from the task ID it minted | One public URL; your own agents | Only works for IDs you mint |
| Global router plus task locator | Router looks the task ID up in a replicated table | Fronting third-party or legacy agents with opaque IDs | One more store on the hot path |
The diagram shows the second and third combined, which is what most teams end up with. New tasks go to the nearest healthy region; follow-ups go to the owner. The router is stateless and can run at the edge in every region, because the only state it needs is either inside the ID or in the replicated locator.
An owner-aware router in Java
The routing core decides a destination and forwards the bytes without rewriting the A2A payload, so it survives protocol additions.
import java.net.URI;
import java.net.http.*;
import java.time.Duration;
import java.util.*;
record Region(String code, URI a2aEndpoint) {}
interface TaskLocator { // replicated KV: taskId -> region code
Optional<String> regionOf(String taskId);
void record(String taskId, String regionCode);
}
final class TaskRouter {
private final Map<String, Region> regions;
private final TaskLocator locator;
private final HttpClient http = HttpClient.newBuilder()
.connectTimeout(Duration.ofSeconds(2)).build();
TaskRouter(Map<String, Region> regions, TaskLocator locator) {
this.regions = regions;
this.locator = locator;
}
/** Owner of an existing task: prefix we minted first, locator second, never a guess. */
Region ownerOf(String taskId) {
int dash = taskId.indexOf('-');
if (dash > 0) {
Region r = regions.get(taskId.substring(0, dash));
if (r != null) return r;
}
return locator.regionOf(taskId).map(regions::get)
.orElseThrow(() -> new NoSuchElementException("unknown task " + taskId));
}
/** Forward a follow-up JSON-RPC call (GetTask, CancelTask, ...) to the owning region. */
HttpResponse<String> forward(String taskId, String jsonRpcBody, String auth)
throws Exception {
Region owner = ownerOf(taskId);
HttpRequest req = HttpRequest.newBuilder(owner.a2aEndpoint())
.timeout(Duration.ofSeconds(10))
.header("Content-Type", "application/json")
.header("Authorization", auth) // pass the caller's identity through
.POST(HttpRequest.BodyPublishers.ofString(jsonRpcBody))
.build();
return http.send(req, HttpResponse.BodyHandlers.ofString());
}
}Three details carry the correctness. First, the prefix is trusted only when it names a region you operate; other agents' IDs are treated as opaque and go through the locator. Second, an unknown task is an error, never a fallback to the nearest region, which would tell the client a running task has vanished. Third, the router passes the caller's credentials through rather than calling with its own, because the specification requires task operations to check the caller's access rights, and that check has to happen in the region that owns the task.
To mint owner-encoded IDs, make the A2A server's task-ID generator prepend the region code, for example euw1- followed by a UUID. For the locator path, write the mapping when the SendMessage response comes back with a new task, before returning it to the caller, so a follow-up can never outrun the record.
What to put in the Agent Card
The Agent Card, served at /.well-known/agent-card.json, is the discovery document. In 1.0 its supportedInterfaces field is an ordered list of endpoints and protocol bindings, and the first entry is preferred. It is tempting to list one URL per region there and let clients pick. Don't: clients reasonably treat the entries as interchangeable ways to reach the same agent, and a client that starts a task on the first URL and polls the second has recreated the cross-region miss.
Two patterns work. With a global router, publish one card whose interface URL is the router; the client never needs to know about regions. With regional endpoints, publish a separate card per region (for example on us.agents.example.com and eu.agents.example.com) and make region choice an explicit client decision, which is what you want when an EU tenant's data must not be processed elsewhere. Either way, keep skills and capabilities identical across regions, or a failover becomes a silent capability change.
1.0 also defines a tenant field that a server serving many agents or tenants interprets for routing. Resolve the tenant to its home region first, then the task to its owner.
Push notifications across regions
Push notifications reverse the direction of traffic. The owning region's webhook sender POSTs task updates to a URL the client registered. That URL may be in another region or another company, so treat it like any outbound integration:
- Send from the owner only. The sender reads the task's push configurations from the owner's store. A replica in another region must never send, or the client receives each update twice from two senders that disagree on ordering.
- Expect duplicates and handle them. The spec tells clients to process notifications idempotently because duplicate deliveries may occur. On the receiving side, key on task ID plus the update's status and timestamp, and drop anything not newer than what you hold.
- Validate before trusting. The spec requires the receiver to check the task ID is one it expects, and recommends verifying the source. Use the token you registered with the configuration and the authentication scheme you agreed, and treat the payload as a hint: on any doubt, call
GetTaskthrough the router and act on that answer.
When the owning region dies
Region loss is where A2A differs most from a stateless service. Suppose eu-west goes down while task euw1-7f3c is working. The router can still parse the prefix, so it knows exactly who owns the task, and it should say so: return a retryable error for follow-ups rather than "task not found". Then you choose, ahead of time, one of three policies:
- Wait. Tasks are lost only if the store is lost. If eu-west's task store survives (or is restored), tasks resume when the region returns. Simplest, and correct for most interactive work that is cheap to redo.
- Fail and restart. Have the client start a new task in another region with the same inputs. Because the new task gets a new server-generated ID, the client must not assume the work is deduplicated. The spec says agents may use the
messageIdto detect duplicate messages; it does not require it, so any operation with external side effects needs its own idempotency key carried in the message and checked by the tool that performs the side effect. - Adopt. Replicate the task store and let the survivor take ownership via the locator, resuming from the last checkpoint. It needs a per-task fencing token so a returning old owner cannot keep writing, and only pays for long, expensive tasks.
Open streams break regardless. A client holding a SubscribeToTask stream to eu-west sees the connection drop and should re-subscribe through the router with backoff, then call GetTask to recover any updates missed while it was disconnected. The A2A client streaming recipe covers the reconnect loop.
Worked example: a delegated review across the Atlantic
Worked example. A procurement agent in us-east delegates contract review to a legal agent that runs in both regions. The client is in Frankfurt and the contract belongs to an EU entity.
- The procurement agent sends
SendMessageto the router with the contract and its tenant. The router resolves the tenant's home to eu-west and forwards there. - eu-west creates task
euw1-2b91in state submitted, registers the push configuration the caller attached, and returns the task. - The router writes nothing for this task: the prefix is enough. A third-party agent behind the same router would have needed a locator record here.
- Ten minutes later the task enters input-required: the agent wants the governing-law clause clarified. The eu-west sender POSTs the update to the procurement agent's webhook in us-east, which checks the task ID and token, then calls
GetTaskvia the router to confirm. - The procurement agent replies with a message carrying task ID
euw1-2b91and the matching context ID. The router reads the prefix and forwards to eu-west; the agent resumes and later completes with an artifact.
Four calls crossed the ocean: the send, the webhook, the confirmation and the reply. That is negligible next to minutes of model time, but not for agents that chat in a tight loop: keep chatty pairs in one region and let only coarse delegations cross.
Operational guidance
- Tag every span and log line with the owning region and the task ID. Cross-region bugs are invisible without them. See A2A error propagation for keeping failures legible across the boundary.
- Alert on wrong-region hits. Count requests where the router had to forward to a non-local region for a task created locally; a non-zero rate means some client is bypassing the router.
- Expire locator entries after the task is terminal plus your history retention window, or the table grows without bound.
- Keep timeouts asymmetric. Short connect timeouts to the owner, long read timeouts for streams, and never blindly retry a
SendMessagethat may already have created a task.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Intermittent task not found | Follow-ups load-balanced away from the owner | Route by owner; no geo balancing for task calls |
| Every update arrives twice | Senders in two regions read replicated push configs | Only the owner sends; replicas are read-only |
| Duplicate side effects after failover | Client restarted the task elsewhere | Idempotency key checked by the side-effecting tool |
| Stream silently stalls | Owner region gone, client never re-subscribes | Heartbeat timeout, re-subscribe, then GetTask |
| EU data processed in us-east | Region chosen by latency, not tenant | Resolve tenant home before nearest region |
Trade-offs
Owner-encoded IDs cost nothing at runtime and survive locator outages, but only work for IDs you mint and leak the region to clients (harmless for most, a concern for some). A locator works for anyone's IDs but adds a replicated store to the critical path, and its replication lag must be shorter than the gap between a task's creation and its first follow-up. Regional cards push complexity to clients but make residency auditable. Task-store replication enables adoption but brings fencing and double execution risks; most teams should prefer wait or restart with idempotent side effects. For the protocol basics behind all of this, see A2A integration in ADK Java.
What to do next
- List every A2A operation your agents call and mark which ones carry a task ID; those are the ones that need owner routing.
- Decide the topology: regional cards for residency, a global router for one URL, a locator if you front agents you did not write.
- Make your server mint region-prefixed task IDs, and verify the router rejects unknown prefixes instead of guessing.
- Confirm which A2A version and method names your SDK and peers use, and test both spellings through the router if you have mixed peers.
- Make webhook receivers idempotent and have them confirm with
GetTask. - Pick wait, restart or adopt for region loss, write it down, and rehearse it with a task in flight.