Running an ADK Java agent in two regions looks like a load-balancer exercise: deploy the same container twice and put an anycast IP in front. The compute part really is that easy. The service is stateless between requests, and Cloud Run or Kubernetes will happily run it anywhere. What makes agents different is everything around the container that has a region. The session history the Runner reads at the start of every turn lives in a regional store. The model endpoint is regional, or global with consequences. And the tools call backends that live somewhere specific.
This article designs a multi-region deployment around those facts. It covers what has a region and why, three topologies, an active-active design with home-region pinning and the Java code that routes a turn to its session's home, model endpoint and quota planning, failover semantics including a turn that dies halfway, the deployment commands, and a sizing example. It builds on the single-region service in deploying ADK Java to Cloud Run and reuses its server code.
What has a region
List everything a turn touches and write its location next to it. For a typical ADK Java agent on Google Cloud there are four entries.
| Component | Where it lives | Multi-region implication |
|---|---|---|
| Agent service (Runner, tools code) | Cloud Run service or GKE deployment, per region | Stateless; replicate freely |
| Session store | VertexAiSessionService: an Agent Engine resource in one location | Each region's sessions are separate; a session has one home |
| Model endpoint | Vertex AI regional endpoint, or the global endpoint | Regional pins processing location; global does not |
| Tool backends | Your databases and APIs, wherever they run | A tool calling a single-region database ties the agent to it anyway |
The session store is where designs go wrong. VertexAiSessionService stores sessions in a managed service attached to an Agent Engine resource, and that resource exists in one Google Cloud location. The API reference documents a constructor that takes the project and location explicitly. Prefer it in a multi-region service rather than relying on environment defaults, and verify the behaviour of your ADK version. Either way, a European instance and a US instance talk to two different session stores, and a session ID created in one is unknown to the other. If the load balancer sends turn two of a conversation to the other region, the Runner finds no session.
The model endpoint has its own trade-off. Vertex AI exposes Gemini models on regional endpoints and on a global endpoint. Google documents the global endpoint as more available than any single region, but you cannot control or know which region processes a request sent to it. Google's data residency guidance says not to use it if you have ML processing location requirements. Some preview models are offered only on the global endpoint and carry no residency guarantee until they reach general availability. Check the current model list for your region before you design around a particular model.
Three topologies
Three topologies cover nearly every real deployment:
| Topology | How sessions work | Region loss | Use when |
|---|---|---|---|
| Active-passive | One primary session store; standby region idle or minimal | Fail over compute; live sessions lost unless the store is replicated | Latency is not regional, recovery can take minutes |
| Active-active, home pinning | Each region owns the sessions it created; turns route to the home | New conversations unaffected; live ones in the lost region pause or restart | Users in several geographies, residency per region |
| Active-active, replicated store | Custom BaseSessionService over a multi-region database | Any region continues any conversation | Conversations are long and valuable, and you can operate the database |
Home pinning is the sensible default, and the rest of this article builds it. It needs no custom session service, keeps each user's history in the region that created it (which is usually what residency rules want), and its failure behaviour is easy to explain. The replicated store means implementing your own session service and owning cross-region write conflicts, replication lag and the database. Do it only when losing a live conversation is genuinely expensive.
Home-region pinning in Java
The trick is to make the session's home visible in the token the client carries. On turn one, the region that creates the session returns eu1.<session-id> instead of the bare ID. On later turns, whichever region receives the request reads the prefix. It runs the turn locally if the home is local, and otherwise forwards the request to the home region's own service URL. Forwarding in code, rather than load-balancer header rules, keeps the routing testable.
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
import java.util.Map;
final class HomeRouter {
static final String LOCAL = System.getenv("REGION_TAG"); // "eu1" or "us1"
static final Map<String, String> HOMES = Map.of( // regional run.app URLs
"eu1", System.getenv("EU1_URL"),
"us1", System.getenv("US1_URL"));
static final HttpClient HTTP = HttpClient.newBuilder()
.connectTimeout(Duration.ofSeconds(2)).build();
record Token(String home, String sessionId) {
static Token parse(String raw) {
int dot = raw.indexOf('.');
if (dot < 1) throw new IllegalArgumentException("bad session token");
return new Token(raw.substring(0, dot), raw.substring(dot + 1));
}
String encode() { return home + "." + sessionId; }
}
/** Returns null when the turn should run here, else the home region's response. */
static HttpResponse<String> forwardIfRemote(Token t, String user, String body,
String idToken) throws Exception {
if (t.home().equals(LOCAL)) return null;
String base = HOMES.get(t.home());
if (base == null) throw new IllegalArgumentException("unknown home " + t.home());
HttpRequest req = HttpRequest.newBuilder(URI.create(base + "/chat"))
.timeout(Duration.ofSeconds(120))
.header("Authorization", "Bearer " + idToken) // service-to-service
.header("X-User-Id", user)
.header("X-Session-Id", t.encode())
.header("X-Forwarded-Once", "1") // loop guard
.POST(HttpRequest.BodyPublishers.ofString(body)).build();
return HTTP.send(req, HttpResponse.BodyHandlers.ofString());
}
}Inside the /chat handler from the Cloud Run article, parse the token first. If forwardIfRemote returns a response, relay it. Otherwise call sessions.getSession(APP, user, t.sessionId(), Optional.empty()) and runner.runAsync(user, t.sessionId(), content) exactly as before, and when creating a session return new Token(LOCAL, s.id()).encode(). Reject any request that arrives with X-Forwarded-Once but whose home is not local. That means a misconfigured URL map, and without the guard two regions would bounce one request between them until the timeout. Authenticate the hop: give each runtime service account the Cloud Run Invoker role on the other region's service. The prefix is input, not a secret, so still authorise the user against the session.
Model endpoints and regional quota
Configure each regional deployment with its own GOOGLE_CLOUD_LOCATION, so model calls resolve to that region, and pass the same location to the session service explicitly. Do not copy the backend-selection environment variable from an old example without checking it. The Google Gen AI SDK documentation has renamed its Vertex switch over time, so use the one documented for the SDK version on your classpath.
Model capacity is the most common reason failover fails. Pay-as-you-go Gemini traffic on Vertex AI is served from dynamic shared quota, a pool shared with other customers, so a 429 can mean the pool is busy even when you are within your limits. During a failover the survivor suddenly carries all your traffic. If availability matters, reserve Provisioned Throughput in each region for the full load it must absorb, set its --max-instances to match, and back off and retry on 429s. Put a circuit breaker on the model client so a struggling endpoint sheds load quickly rather than holding threads; see circuit breakers for ADK Java and quota management.
Falling back from a regional endpoint to the global one is a policy decision, not a technical one: fine without residency obligations, wrong with them.
Failover semantics
Failover of new conversations is not automatic by default. Serverless network endpoint groups do not use classic health checks, so the global load balancer shifts traffic away from a failing region only if you enable outlier detection on the backend service or Cloud Run's service health with readiness probes. Otherwise remove the backend by hand. Once traffic shifts, new conversations lose nothing.
Live conversations whose home is the failed region are the real decision. Their history is unreachable, and with home pinning there is no copy. Fail closed means the surviving region returns 503 with a clear message and the client retries later, which suits agents where acting without full context is dangerous, such as anything that moves money. Fail open means it starts a new local session and seeds it with a summary. The summary can be one the client keeps (the last answer and the user's goal) or one the home region writes asynchronously to a replicated store every few turns. Fail open suits help desks. Either way, tell the user the assistant has lost the details.
Turns in flight are the subtle case. A turn may have executed a side-effecting tool before the region died, and the client's retry, now landing in the survivor, would run it again. Give every side-effecting tool an idempotency key derived from the session token and a client-generated turn ID, and have the backend deduplicate on it. This is the same discipline as retries in one region, just with a bigger blast radius.
Deploying two regions
Deploy the same image to each region with region-specific environment, then put a global external Application Load Balancer in front, with one serverless network endpoint group per region. The commands below are the shape of it. Add your service account, concurrency and limits from the single-region article.
for R in europe-west1:eu1 us-central1:us1; do
REGION=${R%%:*}; TAG=${R##*:}
gcloud run deploy agent --image "$IMAGE" --region "$REGION" \
--service-account "agent-runtime@$PROJECT.iam.gserviceaccount.com" \
--ingress internal-and-cloud-load-balancing --no-allow-unauthenticated \
--max-instances 20 --concurrency 20 \
--set-env-vars "GOOGLE_CLOUD_PROJECT=$PROJECT,GOOGLE_CLOUD_LOCATION=$REGION,REGION_TAG=$TAG"
gcloud compute network-endpoint-groups create "agent-neg-$TAG" \
--region "$REGION" --network-endpoint-type serverless --cloud-run-service agent
done
gcloud compute backend-services create agent-backend \
--global --load-balancing-scheme EXTERNAL_MANAGED
for R in europe-west1:eu1 us-central1:us1; do
gcloud compute backend-services add-backend agent-backend --global \
--network-endpoint-group "agent-neg-${R##*:}" \
--network-endpoint-group-region "${R%%:*}"
done
# then: URL map, target HTTPS proxy with a managed certificate, global forwarding ruleEach region also needs its own Agent Engine resource for sessions, and the app name passed to the Runner must match that region's resource. Pass it per region as an environment variable, as the sibling article does. Restricted ingress stops clients bypassing the load balancer, but it can also block the forwarding hop, which then looks exactly like a region outage, so test that path explicitly. Rolling a new version means deploying region by region, with each region draining cleanly; graceful shutdown covers the drain.
Worked example: sizing for survival
A support agent serves Europe and North America. At peak, 240 turns are in flight in Europe and 160 in the US, a total of 400. Instances run at concurrency 20.
| Quantity | Normal, EU | Normal, US | EU lost: US carries |
|---|---|---|---|
| Turns in flight | 240 | 160 | 400 |
| Instances at concurrency 20 | 12 | 8 | 20 |
| max-instances needed | 12 + headroom | 8 + headroom | 20 + headroom |
| Model capacity to reserve | EU share | US share | EU + US share in US |
| Live conversations affected | - | - | every open EU-homed session |
The naive sizing sets max-instances to 12 and 8 and asks for quota in the same ratio. It passes every normal-day load test, then fails the first real outage. Survivable sizing gives each region a max-instances of at least 20 plus headroom and model capacity for all 400 turns, about twice its normal need; on Cloud Run idle capacity scales down, so the cost is mostly reserved model capacity. Residency adds a constraint: if European users' data must stay in Europe, EU-homed traffic cannot fail over to the US at all, and the right answer is two European regions, not one European and one American.
Failure modes
- Session not found after a reroute. Bare session IDs with a global load balancer break whenever routing changes. Encode the home in the token.
- Forwarding loops. Two regions each believe the other is home. Carry a hop marker and reject a second hop.
- Survivor capacity exhaustion. Failover works on paper, then 429s everywhere. Size model capacity and instance ceilings for the full load in every region you fail over to.
- Duplicated side effects. A retried turn reruns a tool that already ran. Use idempotency keys tied to the session token and turn ID.
- Single-region tools. Both regions are healthy, but every tool calls one database in one of them. Map tool dependencies before claiming multi-region availability.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Home pinning vs replicated sessions | No custom store, clear residency | Live conversations in a lost region pause or restart |
| Forward in service vs LB header routing | Testable logic, survives path changes | Extra hop latency for misrouted turns |
| Regional vs global model endpoint | Known processing location | Lower availability than global, regional capacity |
| Survivable sizing vs normal sizing | Failover actually works | Reserved capacity and ceilings about twice normal need |
| Fail closed vs fail open | No action without context | Users wait instead of continuing |
What to do next
- Write down the location of every component a turn touches: service, session store, model endpoint and each tool backend.
- Decide residency rules per agent first; they decide which regions may fail over to which.
- Encode the home region in the session token and add in-service forwarding with a hop guard and service-to-service authentication.
- Size model capacity and max-instances so each region can absorb the full load of the regions that fail over to it.
- Choose fail closed or fail open per agent, and add idempotency keys to every side-effecting tool.
- Run a game day: disable one region's backend and watch new, live and in-flight conversations; compare with the single-region baseline in running ADK Java on Kubernetes if you deploy on GKE instead.