High availability for an ordinary web service mostly means running several identical replicas behind a load balancer. An agent built with ADK for Java needs that too, but it is the easy part. One agent turn can last minutes, reads and writes a session, calls a hosted model with its own outages and quotas, and may call tools that move money. When something fails halfway, the question is whether another replica knows what already happened and whether redoing the work is safe.

This article works through high availability (HA) from the outside in: the arithmetic that sets your ceiling, the failure domains, where state must live, what to do about a turn interrupted by a crash, provider failover, quota and regional recovery. Kubernetes probes, drain and autoscaling are covered in running ADK Java agents on Kubernetes and the shutdown sequence in ADK Java graceful shutdown, so they are referenced here rather than repeated.

Advertisement

Availability arithmetic: the chain you actually depend on

A turn succeeds only if every component on its path works. When dependencies are in series and fail independently, their availabilities multiply, so the product is always lower than the weakest link. When two components can each do the job, the turn fails only if both fail, so unavailabilities multiply instead. The trap is that the parallel formula assumes independent failures, and replicas sharing one database, zone or quota are not independent.

# Illustrative figures, not vendor SLAs. Replace them with your own measured monthly availability.
serial = {"load balancer": 0.9999, "replicas": 0.9995, "session store": 0.9995,
          "model API": 0.999, "tool API": 0.999}

def chain(parts):
    a = 1.0
    for v in parts:
        a *= v
    return a

def redundant(a, copies):          # only valid if the copies fail independently
    return 1 - (1 - a) ** copies

single = chain(serial.values())
two_models = chain([*(v for k, v in serial.items() if k != "model API"), redundant(0.999, 2)])
for name, a in [("one model provider", single), ("two model providers", two_models)]:
    minutes = (1 - a) * 30 * 24 * 60
    print(f"{name}: {a:.5f} -> about {minutes:.0f} minutes of failed turns per 30 days")

With these illustrative numbers one provider gives about 0.9969, roughly 135 minutes of failed turns a month, which misses a 99.9% objective before any bug or deployment is counted. Adding an independent second provider removes about 43 minutes of that. The model API and the tool APIs dominate, so that is where redundancy pays. Adding a fourth replica changes almost nothing.

Failure domains and the reference architecture

An ADK Java agent service across two regions: what fails, and what takes overGlobal load balancerhealth-checked backendsRegion A (serving)Region B (serving or warm standby)Replica zone 1stateless RunnerReplica zone 2stateless RunnerReplica zone 1stateless RunnerReplica zone 2stateless RunnerSession store (primary)events, state, turn ledgerSession store (replica)sync or async copyreplicationFallbackLlmBaseLlm decoratorModel provider 1primary region, quotaModel provider 2different failure domainTool backendsidempotency keyson errorReplicas are cheap to replace. The session store, the model quota and the tool side effects decide availability.
Figure 1. A two-region ADK Java deployment. Replicas hold no state; the session store, model access and tool calls are the components that decide whether a turn survives a failure.

A failure domain is a set of things that fail together. Listing them tells you what each layer of redundancy protects against and what it does not.

Failure domainTypical causeProtectionRecovery time target
JVM process or podOOM, crash, bad nodeSeveral replicas, readiness probesSeconds
ZonePower, network partitionReplicas spread across zones, zonal or regional storeSeconds to a minute
RegionRegional control plane or network outageSecond region, replicated session storeMinutes
Model providerProvider incident, model retiredSecond provider or second region of the same providerPer request
QuotaRate limit exhausted, often during a failoverPre-provisioned quota, load sheddingPer request
Bad releaseRegression in prompt, tool or codeCanary rollout, fast rollbackMinutes

Replication cannot help with the last row, because every replica runs the same broken build, so rollout discipline is part of HA.

Advertisement

Rule one: no state lives in a replica

ADK keeps a conversation as a session: an ordered list of events plus a state map, stored by an implementation of BaseSessionService. The InMemorySessionService used in quick starts holds that data in the JVM heap, so a pod restart erases every conversation it was serving, and a load balancer that sends turn two to a different pod finds no session at all. Sticky sessions hide the second problem and make the first worse.

For HA, every replica must read and write sessions in a durable store that outlives it. ADK Java offers VertexAiSessionService for the managed Vertex AI Agent Engine sessions, and a Firestore integration in the separate artifact google-adk-firestore-session-service, which provides FirestoreSessionService and a FirestoreDatabaseRunner. You can also implement BaseSessionService over your own database. The same rule applies to artifacts and long-term memory: if a replica writes them locally, they are lost with it.

// Wiring a Runner to a durable session store (Firestore integration artifact).
Firestore firestore = FirestoreOptions.getDefaultInstance().getService();
FirestoreDatabaseRunner runner = new FirestoreDatabaseRunner(supportAgent, "support", firestore);

// Any replica can now serve any turn of any session.
Flowable<Event> events = runner.runAsync(userId, sessionId,
        Content.fromParts(Part.fromText(userText)), RunConfig.builder().build());

Caches inside a replica must be rebuildable from the store. A streaming response ties a client to one replica, so if that replica dies the client must reconnect and resume from the session, as the next section describes.

When a replica dies in the middle of a turn

A turn is a sequence of steps, and a crash can land between any two of them. What a retry does depends on which step was the last to be persisted.

Crash pointWhat the store containsEffect of a naive retry
Before the user message is appendedNothing newSafe: the turn simply runs
After the user message, before the model answersA dangling user eventDuplicate user message in history
After a tool with side effects ran, before its result was appendedA function call with no responseThe tool runs twice: two refunds, two emails
After the final answer was appendedThe complete turnA second answer to the same question

The fix has two parts. The client sends a turn id with every request and reuses it on retry, and the server keeps a small turn ledger keyed by session and turn id. A retry of a completed turn returns the stored answer. A retry of an in-progress turn whose owner has stopped heartbeating is taken over. Tools with side effects take an idempotency key derived from the turn id and the call id, so the backend refuses a second execution; ADK Java idempotency shows that half in detail.

// Turn ledger: claim a turn, or learn that another replica owns or finished it.
// Table: turns(session_id, turn_id, owner, status, heartbeat_at, answer), PK (session_id, turn_id)
enum Claim { RUN, ALREADY_DONE, OWNED_ELSEWHERE }

Claim claim(String sessionId, String turnId, String me) throws SQLException {
  try (var ps = db.prepareStatement(
      "INSERT INTO turns (session_id, turn_id, owner, status, heartbeat_at) "
      + "VALUES (?, ?, ?, 'RUNNING', now()) "
      + "ON CONFLICT (session_id, turn_id) DO UPDATE "
      + "SET owner = EXCLUDED.owner, heartbeat_at = now() "
      + "WHERE turns.status = 'RUNNING' AND turns.heartbeat_at < now() - interval '30 seconds' "
      + "RETURNING owner")) {
    ps.setString(1, sessionId); ps.setString(2, turnId); ps.setString(3, me);
    try (var rs = ps.executeQuery()) {
      if (rs.next()) return Claim.RUN;          // new turn, or a stale owner was replaced
    }
  }
  return isDone(sessionId, turnId) ? Claim.ALREADY_DONE : Claim.OWNED_ELSEWHERE;
}

The owner refreshes its heartbeat while the turn runs and writes DONE with the answer at the end. A takeover after a crash should inspect the session for a function call without a response and resolve it by asking the tool backend for the outcome under the idempotency key, not by calling the tool again blindly.

Model provider failover with a BaseLlm decorator

Every model call in ADK Java goes through BaseLlm, whose generateContent(LlmRequest, boolean stream) returns a Flowable<LlmResponse>. Because an LlmAgent accepts a BaseLlm instance as its model, a decorator that wraps two models is the cleanest place for failover; implementing a custom LLM in ADK Java covers the contract in depth.

public final class FallbackLlm extends BaseLlm {
  private final BaseLlm primary, secondary;

  public FallbackLlm(BaseLlm primary, BaseLlm secondary) {
    super(primary.model());
    this.primary = primary; this.secondary = secondary;
  }

  @Override
  public Flowable<LlmResponse> generateContent(LlmRequest req, boolean stream) {
    return Flowable.defer(() -> {
      AtomicBoolean emitted = new AtomicBoolean(false);
      return primary.generateContent(req, stream)
          .timeout(45, TimeUnit.SECONDS)
          .doOnNext(r -> emitted.set(true))
          .onErrorResumeNext(e -> !emitted.get() && retryable(e)
              ? secondary.generateContent(req, stream)   // nothing reached the caller yet
              : Flowable.error(e));                      // mid-stream: surface the error
    });
  }

  @Override
  public BaseLlmConnection connect(LlmRequest req) { return primary.connect(req); }

  static boolean retryable(Throwable e) {
    return e instanceof TimeoutException || e instanceof IOException
        || ProviderErrors.status(e).map(s -> s == 429 || s >= 500).orElse(false);
  }
}

LlmAgent agent = LlmAgent.builder()
    .name("support")
    .model(new FallbackLlm(primaryModel, secondaryModel))
    .instruction("...")
    .build();

The emitted flag carries the key limitation. Once partial output has reached the caller, switching providers would splice two different answers together, so a mid-stream failure must surface as an error and be handled by the turn retry. ProviderErrors.status is your own helper that reads an HTTP status from your adapter's exceptions. Do not fall back on a 400. The secondary receives the same LlmRequest, so its adapter must send its own model name rather than one carried in the request.

The secondary model is a different model. Tool-calling accuracy, refusal behaviour and output length all change, so run your evaluation set against the fallback path before you trust it, and record which model produced each answer. The same model in a second provider region narrows that gap but shares more failure modes.ADK Java rate limiting covers the per-provider token buckets that keep fallback traffic within the secondary's quota.

Quota is an availability limit

Model quota is usually set per project, region and model. Normally it is a cost control; during a failover it decides whether you stay up. If region B suddenly receives all of region A's traffic without the quota for it, a regional outage becomes a global wave of 429 errors.

  • Provision quota in each region for the full load you would move there.
  • Prioritise traffic so interactive turns are served and batch work waits under a squeeze.
  • Shed load early with a fast error instead of queueing until timeout.
  • Alert on utilisation against the post-failover requirement, not today's limit.

Regional failover: RPO, RTO and the session store

Two numbers define regional recovery. The recovery point objective (RPO) is how much recently written data you can lose. The recovery time objective (RTO) is how long service may be down. For an agent, RPO is measured in conversation turns. With asynchronous replication, the last few seconds of events never reach region B, so users who are failed over find their latest turn missing. Synchronous multi-region storage gives an RPO of zero at the price of cross-region latency on every write, and a turn writes several events.

Active-passive designs keep region B warm and promote it on failure, a step that must be drilled. Active-active designs serve from both regions continuously, which proves region B works, but need a store that accepts writes in both regions or a home region per session. Either way, failover should be automatic for the stateless tier and deliberate, with a fencing step, for the store, so two primaries never accept writes for the same session. multi-region LLM serving covers the inference side of the same decision.

Worked example: a support agent at 99.9%

A support agent must succeed on 99.9% of turns per month, an error budget of about 43 minutes of full outage. Its path is a load balancer, ADK replicas, a session store, one model and an order-lookup tool. Run the arithmetic first. With one provider and a single-zone database, the illustrative chain lands near 99.7%, so the design cannot meet the objective even before releases are considered.

The changes, in order of value: a zone-redundant session store; a FallbackLlm with a second provider, accepted after tool-call accuracy on the 300-case suite stayed within two points; a timeout and degraded answer on the order tool; three replicas across three zones; and a turn ledger with idempotency keys on the refund tool. Region B stays warm, with an RPO of 10 seconds of asynchronous replication and an RTO of 15 minutes for a drilled promotion. The team accepts that a regional event spends a third of the monthly budget and writes that decision down.

Failure modes to design and test for

Failure modeSymptomMitigation
Sessions in memoryConversations vanish after deploysDurable BaseSessionService everywhere except tests
Duplicate side effectsTwo refunds after a pod restartTurn ledger plus tool idempotency keys
Fallback never exercisedSecond provider fails on first real useRoute a small share of traffic through it continuously
Failover exhausts quota429 errors in the surviving regionQuota sized for post-failover load, priority shedding
Split brain in the storeTwo answers to one turn, diverged historyFencing before promotion, single writer per session
Correlated release bugAll replicas fail at onceCanary, automated rollback on turn success rate

Test each row on purpose. A game day kills a pod mid-tool-call, blocks egress to the primary model, forces a quota error and promotes region B while someone watches turn success and duplicate side effects.

What to do next

  1. Write down the serial chain for one turn and multiply your measured availabilities; compare the result with your objective.
  2. Replace every InMemorySessionService outside tests with a durable session service, and check artifacts and memory too.
  3. Add a client turn id and a turn ledger, and put idempotency keys on every tool with side effects.
  4. Wrap your model in a fallback decorator, run your evaluation suite against the secondary, and send it a little live traffic.
  5. Size model quota in each region for the load it would carry after a failover.
  6. Choose and document RPO and RTO for the session store, and drill the promotion.
  7. Schedule a quarterly game day that exercises each failure mode in the table.
Key takeaway: An ADK Java agent is highly available when a failure anywhere on a turn's path is either absorbed or made safe to retry. Multiply serial dependencies to find your real ceiling and put redundancy where the arithmetic says: usually the model and the tools, not the replicas. Keep every session in a durable BaseSessionService, make interrupted turns resumable with a turn ledger and idempotency keys, fail over between models only before output has streamed, size quota for the post-failover load, and choose RPO and RTO for the session store deliberately and drill them.