High availability for an ordinary web service mostly means running several identical replicas behind a load balancer. An agent built with ADK for Java needs that too, but it is the easy part. One agent turn can last minutes, reads and writes a session, calls a hosted model with its own outages and quotas, and may call tools that move money. When something fails halfway, the question is whether another replica knows what already happened and whether redoing the work is safe.
This article works through high availability (HA) from the outside in: the arithmetic that sets your ceiling, the failure domains, where state must live, what to do about a turn interrupted by a crash, provider failover, quota and regional recovery. Kubernetes probes, drain and autoscaling are covered in running ADK Java agents on Kubernetes and the shutdown sequence in ADK Java graceful shutdown, so they are referenced here rather than repeated.
Availability arithmetic: the chain you actually depend on
A turn succeeds only if every component on its path works. When dependencies are in series and fail independently, their availabilities multiply, so the product is always lower than the weakest link. When two components can each do the job, the turn fails only if both fail, so unavailabilities multiply instead. The trap is that the parallel formula assumes independent failures, and replicas sharing one database, zone or quota are not independent.
# Illustrative figures, not vendor SLAs. Replace them with your own measured monthly availability.
serial = {"load balancer": 0.9999, "replicas": 0.9995, "session store": 0.9995,
"model API": 0.999, "tool API": 0.999}
def chain(parts):
a = 1.0
for v in parts:
a *= v
return a
def redundant(a, copies): # only valid if the copies fail independently
return 1 - (1 - a) ** copies
single = chain(serial.values())
two_models = chain([*(v for k, v in serial.items() if k != "model API"), redundant(0.999, 2)])
for name, a in [("one model provider", single), ("two model providers", two_models)]:
minutes = (1 - a) * 30 * 24 * 60
print(f"{name}: {a:.5f} -> about {minutes:.0f} minutes of failed turns per 30 days")With these illustrative numbers one provider gives about 0.9969, roughly 135 minutes of failed turns a month, which misses a 99.9% objective before any bug or deployment is counted. Adding an independent second provider removes about 43 minutes of that. The model API and the tool APIs dominate, so that is where redundancy pays. Adding a fourth replica changes almost nothing.
Failure domains and the reference architecture
A failure domain is a set of things that fail together. Listing them tells you what each layer of redundancy protects against and what it does not.
| Failure domain | Typical cause | Protection | Recovery time target |
|---|---|---|---|
| JVM process or pod | OOM, crash, bad node | Several replicas, readiness probes | Seconds |
| Zone | Power, network partition | Replicas spread across zones, zonal or regional store | Seconds to a minute |
| Region | Regional control plane or network outage | Second region, replicated session store | Minutes |
| Model provider | Provider incident, model retired | Second provider or second region of the same provider | Per request |
| Quota | Rate limit exhausted, often during a failover | Pre-provisioned quota, load shedding | Per request |
| Bad release | Regression in prompt, tool or code | Canary rollout, fast rollback | Minutes |
Replication cannot help with the last row, because every replica runs the same broken build, so rollout discipline is part of HA.
Rule one: no state lives in a replica
ADK keeps a conversation as a session: an ordered list of events plus a state map, stored by an implementation of BaseSessionService. The InMemorySessionService used in quick starts holds that data in the JVM heap, so a pod restart erases every conversation it was serving, and a load balancer that sends turn two to a different pod finds no session at all. Sticky sessions hide the second problem and make the first worse.
For HA, every replica must read and write sessions in a durable store that outlives it. ADK Java offers VertexAiSessionService for the managed Vertex AI Agent Engine sessions, and a Firestore integration in the separate artifact google-adk-firestore-session-service, which provides FirestoreSessionService and a FirestoreDatabaseRunner. You can also implement BaseSessionService over your own database. The same rule applies to artifacts and long-term memory: if a replica writes them locally, they are lost with it.
// Wiring a Runner to a durable session store (Firestore integration artifact).
Firestore firestore = FirestoreOptions.getDefaultInstance().getService();
FirestoreDatabaseRunner runner = new FirestoreDatabaseRunner(supportAgent, "support", firestore);
// Any replica can now serve any turn of any session.
Flowable<Event> events = runner.runAsync(userId, sessionId,
Content.fromParts(Part.fromText(userText)), RunConfig.builder().build());Caches inside a replica must be rebuildable from the store. A streaming response ties a client to one replica, so if that replica dies the client must reconnect and resume from the session, as the next section describes.
When a replica dies in the middle of a turn
A turn is a sequence of steps, and a crash can land between any two of them. What a retry does depends on which step was the last to be persisted.
| Crash point | What the store contains | Effect of a naive retry |
|---|---|---|
| Before the user message is appended | Nothing new | Safe: the turn simply runs |
| After the user message, before the model answers | A dangling user event | Duplicate user message in history |
| After a tool with side effects ran, before its result was appended | A function call with no response | The tool runs twice: two refunds, two emails |
| After the final answer was appended | The complete turn | A second answer to the same question |
The fix has two parts. The client sends a turn id with every request and reuses it on retry, and the server keeps a small turn ledger keyed by session and turn id. A retry of a completed turn returns the stored answer. A retry of an in-progress turn whose owner has stopped heartbeating is taken over. Tools with side effects take an idempotency key derived from the turn id and the call id, so the backend refuses a second execution; ADK Java idempotency shows that half in detail.
// Turn ledger: claim a turn, or learn that another replica owns or finished it.
// Table: turns(session_id, turn_id, owner, status, heartbeat_at, answer), PK (session_id, turn_id)
enum Claim { RUN, ALREADY_DONE, OWNED_ELSEWHERE }
Claim claim(String sessionId, String turnId, String me) throws SQLException {
try (var ps = db.prepareStatement(
"INSERT INTO turns (session_id, turn_id, owner, status, heartbeat_at) "
+ "VALUES (?, ?, ?, 'RUNNING', now()) "
+ "ON CONFLICT (session_id, turn_id) DO UPDATE "
+ "SET owner = EXCLUDED.owner, heartbeat_at = now() "
+ "WHERE turns.status = 'RUNNING' AND turns.heartbeat_at < now() - interval '30 seconds' "
+ "RETURNING owner")) {
ps.setString(1, sessionId); ps.setString(2, turnId); ps.setString(3, me);
try (var rs = ps.executeQuery()) {
if (rs.next()) return Claim.RUN; // new turn, or a stale owner was replaced
}
}
return isDone(sessionId, turnId) ? Claim.ALREADY_DONE : Claim.OWNED_ELSEWHERE;
}The owner refreshes its heartbeat while the turn runs and writes DONE with the answer at the end. A takeover after a crash should inspect the session for a function call without a response and resolve it by asking the tool backend for the outcome under the idempotency key, not by calling the tool again blindly.
Model provider failover with a BaseLlm decorator
Every model call in ADK Java goes through BaseLlm, whose generateContent(LlmRequest, boolean stream) returns a Flowable<LlmResponse>. Because an LlmAgent accepts a BaseLlm instance as its model, a decorator that wraps two models is the cleanest place for failover; implementing a custom LLM in ADK Java covers the contract in depth.
public final class FallbackLlm extends BaseLlm {
private final BaseLlm primary, secondary;
public FallbackLlm(BaseLlm primary, BaseLlm secondary) {
super(primary.model());
this.primary = primary; this.secondary = secondary;
}
@Override
public Flowable<LlmResponse> generateContent(LlmRequest req, boolean stream) {
return Flowable.defer(() -> {
AtomicBoolean emitted = new AtomicBoolean(false);
return primary.generateContent(req, stream)
.timeout(45, TimeUnit.SECONDS)
.doOnNext(r -> emitted.set(true))
.onErrorResumeNext(e -> !emitted.get() && retryable(e)
? secondary.generateContent(req, stream) // nothing reached the caller yet
: Flowable.error(e)); // mid-stream: surface the error
});
}
@Override
public BaseLlmConnection connect(LlmRequest req) { return primary.connect(req); }
static boolean retryable(Throwable e) {
return e instanceof TimeoutException || e instanceof IOException
|| ProviderErrors.status(e).map(s -> s == 429 || s >= 500).orElse(false);
}
}
LlmAgent agent = LlmAgent.builder()
.name("support")
.model(new FallbackLlm(primaryModel, secondaryModel))
.instruction("...")
.build();The emitted flag carries the key limitation. Once partial output has reached the caller, switching providers would splice two different answers together, so a mid-stream failure must surface as an error and be handled by the turn retry. ProviderErrors.status is your own helper that reads an HTTP status from your adapter's exceptions. Do not fall back on a 400. The secondary receives the same LlmRequest, so its adapter must send its own model name rather than one carried in the request.
The secondary model is a different model. Tool-calling accuracy, refusal behaviour and output length all change, so run your evaluation set against the fallback path before you trust it, and record which model produced each answer. The same model in a second provider region narrows that gap but shares more failure modes.ADK Java rate limiting covers the per-provider token buckets that keep fallback traffic within the secondary's quota.
Quota is an availability limit
Model quota is usually set per project, region and model. Normally it is a cost control; during a failover it decides whether you stay up. If region B suddenly receives all of region A's traffic without the quota for it, a regional outage becomes a global wave of 429 errors.
- Provision quota in each region for the full load you would move there.
- Prioritise traffic so interactive turns are served and batch work waits under a squeeze.
- Shed load early with a fast error instead of queueing until timeout.
- Alert on utilisation against the post-failover requirement, not today's limit.
Regional failover: RPO, RTO and the session store
Two numbers define regional recovery. The recovery point objective (RPO) is how much recently written data you can lose. The recovery time objective (RTO) is how long service may be down. For an agent, RPO is measured in conversation turns. With asynchronous replication, the last few seconds of events never reach region B, so users who are failed over find their latest turn missing. Synchronous multi-region storage gives an RPO of zero at the price of cross-region latency on every write, and a turn writes several events.
Active-passive designs keep region B warm and promote it on failure, a step that must be drilled. Active-active designs serve from both regions continuously, which proves region B works, but need a store that accepts writes in both regions or a home region per session. Either way, failover should be automatic for the stateless tier and deliberate, with a fencing step, for the store, so two primaries never accept writes for the same session. multi-region LLM serving covers the inference side of the same decision.
Worked example: a support agent at 99.9%
A support agent must succeed on 99.9% of turns per month, an error budget of about 43 minutes of full outage. Its path is a load balancer, ADK replicas, a session store, one model and an order-lookup tool. Run the arithmetic first. With one provider and a single-zone database, the illustrative chain lands near 99.7%, so the design cannot meet the objective even before releases are considered.
The changes, in order of value: a zone-redundant session store; a FallbackLlm with a second provider, accepted after tool-call accuracy on the 300-case suite stayed within two points; a timeout and degraded answer on the order tool; three replicas across three zones; and a turn ledger with idempotency keys on the refund tool. Region B stays warm, with an RPO of 10 seconds of asynchronous replication and an RTO of 15 minutes for a drilled promotion. The team accepts that a regional event spends a third of the monthly budget and writes that decision down.
Failure modes to design and test for
| Failure mode | Symptom | Mitigation |
|---|---|---|
| Sessions in memory | Conversations vanish after deploys | Durable BaseSessionService everywhere except tests |
| Duplicate side effects | Two refunds after a pod restart | Turn ledger plus tool idempotency keys |
| Fallback never exercised | Second provider fails on first real use | Route a small share of traffic through it continuously |
| Failover exhausts quota | 429 errors in the surviving region | Quota sized for post-failover load, priority shedding |
| Split brain in the store | Two answers to one turn, diverged history | Fencing before promotion, single writer per session |
| Correlated release bug | All replicas fail at once | Canary, automated rollback on turn success rate |
Test each row on purpose. A game day kills a pod mid-tool-call, blocks egress to the primary model, forces a quota error and promotes region B while someone watches turn success and duplicate side effects.
What to do next
- Write down the serial chain for one turn and multiply your measured availabilities; compare the result with your objective.
- Replace every
InMemorySessionServiceoutside tests with a durable session service, and check artifacts and memory too. - Add a client turn id and a turn ledger, and put idempotency keys on every tool with side effects.
- Wrap your model in a fallback decorator, run your evaluation suite against the secondary, and send it a little live traffic.
- Size model quota in each region for the load it would carry after a failover.
- Choose and document RPO and RTO for the session store, and drill the promotion.
- Schedule a quarterly game day that exercises each failure mode in the table.