For a stateless web service, a zero-downtime deployment means no request is refused or cut off while the new version replaces the old. An ADK Java agent raises the bar in two ways. A single turn can run for tens of seconds while the model reasons, calls tools and streams tokens, so draining is slower and the cost of cutting a turn is higher: the tool may already have acted. And a conversation spans many turns stored in a session, so during any rollout one conversation can be served by the old release on turn three and the new release on turn four, with the new code reading events and state that the old code wrote.
This article covers both halves. The first is mechanical and mostly solved by the platform: surge capacity, readiness, draining and grace periods, which the Kubernetes and graceful shutdown articles on this site cover in detail and which are summarized here. The second is specific to agents and is where real outages come from: keeping session state, tool names and instructions compatible across adjacent releases, pinning conversations when they cannot be, and choosing between rolling, blue-green and canary strategies with that in mind.
Four kinds of downtime for an agent
It helps to name the four ways an agent deploy causes downtime, because each needs a different fix. Refused requests happen when traffic reaches a pod that is not ready or has already stopped listening. Cut turns happen when a pod is killed while a model call or a streaming response is in flight. Broken sessions happen when the new release cannot interpret what the old one stored: a renamed state key, a tool the model is told to call that the other release does not register, or an event shape that fails to deserialize. Behaviour discontinuity happens when a user's conversation changes tone, policy or capability mid-stream because the instruction changed between turns. The first two are infrastructure problems; the last two are application design problems, and no deployment strategy alone fixes them.
A rolling update that never loses capacity
The baseline is a rolling update that never reduces serving capacity: maxUnavailable: 0 so no old pod stops before a new one is ready, and maxSurge to bring new pods up alongside. The readiness probe must mean 'can serve a turn now', which for an agent includes having loaded configuration and reached the session store, not merely that the JVM is up. The pod's terminationGracePeriodSeconds must exceed the longest turn you allow, plus the time for the load balancer to stop routing to the pod, and a short preStop sleep covers the window in which endpoints are still being removed (distroless images have no sleep binary; use the native sleep action on clusters that support it, or an HTTP hook). The arithmetic for choosing these numbers, and probe design in general, is worked through in the Kubernetes article linked at the end; the manifest below shows the shape.
apiVersion: apps/v1
kind: Deployment
metadata:
name: support-agent
spec:
replicas: 6
selector:
matchLabels: {app: support-agent} # never include the release label
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0 # never drop below 6 ready pods
maxSurge: 2 # bring up 2 new pods at a time
template:
metadata:
labels: {app: support-agent, release: "2026-10-06.1"}
spec:
terminationGracePeriodSeconds: 150 # longest turn (120s) + routing lag + margin
containers:
- name: agent
image: registry.example.com/support-agent:2026-10-06.1
env:
- {name: AGENT_RELEASE, value: "2026-10-06.1"}
readinessProbe:
httpGet: {path: /ready, port: 8080}
periodSeconds: 5
failureThreshold: 1
lifecycle:
preStop:
exec: {command: ["sleep", "10"]} # let endpoints converge first
Draining long turns inside the JVM
Inside the JVM, draining needs three things: stop reporting ready as soon as termination begins, refuse new turns that arrive anyway, and wait for in-flight turns up to a deadline. The gate below wraps the turn handler of a server built around ADK's Runner, in the same style as the server in the Cloud Run article. It counts in-flight turns, flips readiness on SIGTERM, and returns 503 with a Retry-After header for new work, which a well-behaved client or load balancer retries on another pod.
final class DrainGate {
private final AtomicBoolean draining = new AtomicBoolean(false);
private final AtomicInteger inFlight = new AtomicInteger();
boolean ready() { return !draining.get(); }
/** Returns false if the caller must reject the turn with 503. */
boolean enter() {
inFlight.incrementAndGet();
if (draining.get()) { inFlight.decrementAndGet(); return false; }
return true;
}
void exit() { inFlight.decrementAndGet(); }
/** Called from the shutdown hook; blocks until idle or the deadline passes. */
void drain(Duration deadline) throws InterruptedException {
draining.set(true);
long end = System.nanoTime() + deadline.toNanos();
while (inFlight.get() > 0 && System.nanoTime() < end) Thread.sleep(200);
}
}
// Wiring, inside main():
DrainGate gate = new DrainGate();
server.createContext("/ready", ex -> reply(ex, gate.ready() ? 200 : 503, ""));
server.createContext("/chat", ex -> {
if (!gate.enter()) {
ex.getResponseHeaders().set("Retry-After", "1");
reply(ex, 503, "draining");
return;
}
try {
handleTurn(ex); // getSession + runner.runAsync(...) as usual
} finally {
gate.exit();
}
});
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
try { gate.drain(Duration.ofSeconds(130)); } catch (InterruptedException ignored) {}
server.stop(5);
}));Streaming responses deserve one more rule. If a turn cannot finish within the grace period, the client will see the stream end early, and whether the turn's effects persisted depends on how far the run got. Make that recoverable rather than hoping it never happens: give each turn a client-generated idempotency key, record it in the session state when the turn starts, and on reconnect let the client fetch the session with getSession and read the latest events before deciding whether to resend. Tools with side effects should use the same key, so a resent turn does not issue a second refund or a second ticket.
Adjacent-release compatibility
The rule that makes rolling deployments safe for agents is the same one databases use: any two adjacent releases must be able to serve the same session. Because a rollout always has exactly two releases live, N and N+1, it is enough for each release to be compatible with its neighbours; changes that break that rule are split across releases with an expand, migrate, contract sequence.
Session state keys. Session state persists between turns, and both releases read it. To rename or reshape a key, the expand release reads the new key with a fallback to the old one and writes both; the migrate release writes only the new key; the contract release stops reading the old key once sessions written by the expand release have expired or been migrated. The same discipline applies to the database tables behind a persistent session service, and the migrations article linked below covers it at the schema level.
Tool names and signatures. The session's event history contains the function calls the model made, by name. If release N+1 drops a tool that release N's model was told about, a turn served by N+1 may see a history full of calls to a tool it does not declare, and a turn served by N may try to call a tool that N+1 removed from the backend. Keep the old tool registered as an alias for at least one release, change the instruction to prefer the new name, and remove the alias only in the contract release. Add parameters as optional; never change the meaning of an existing one.
Instructions and model settings. Prompt changes are code changes with behavioural effects. A conversation that switches instruction mid-stream can contradict itself. For small edits that is acceptable; for policy changes, pin the conversation to the release that started it, as described below.
// Expand release: register both names; the old one delegates to the new one.
public final class OrderTools {
@Annotations.Schema(description = "Fetch an order by id.")
public static Map<String, Object> getOrder(
@Annotations.Schema(name = "orderId", description = "Order id") String orderId) {
return OrderClient.fetch(orderId);
}
@Annotations.Schema(description = "Deprecated alias of getOrder.")
public static Map<String, Object> lookupOrder(
@Annotations.Schema(name = "orderId", description = "Order id") String orderId) {
return getOrder(orderId);
}
}
LlmAgent agent = LlmAgent.builder()
.name("support_agent")
.model("gemini-2.5-flash")
.instruction("Use lookupOrder to fetch orders.") // release B switches to getOrder
.tools(FunctionTool.create(OrderTools.class, "getOrder"),
FunctionTool.create(OrderTools.class, "lookupOrder")) // removed in release C
.build();
// Reading a renamed state key with a fallback during the expand phase.
static Object cart(Session s) {
Object v = s.state().get("cart_v2");
return v != null ? v : s.state().get("cart");
}Choosing a strategy
With compatibility handled, the strategy is a question of blast radius and cost. A rolling update needs no extra infrastructure beyond the surge pods and is the right default when releases are adjacent-compatible. Blue-green brings up a full second fleet and switches traffic at once; it gives a fast rollback, but the switch moves every live conversation to the new release at the same moment, so it needs the same compatibility work, and doubles serving capacity during the transition, which is costly when model quota or GPU-backed tools are attached per pod. A canary sends a small share of turns to the new release and compares error rates, latency, tool failure rates and an evaluation score before widening; for agents, compare quality on a fixed evaluation set too, since a prompt regression returns HTTP 200.
When a change cannot be made adjacent-compatible, for example a policy change that must not apply halfway through a conversation, use session pinning. Store the release that created the session in its initial state, for instance by passing new ConcurrentHashMap<>(Map.of("agent_release", RELEASE)) as the state argument of createSession, echo it to the client in a response header, and have the gateway route each request to the matching release. The old fleet then serves only existing conversations and can be scaled down as they go idle. Pinning costs a longer coexistence period and a router that understands releases, so reserve it for changes that need it.
| Strategy | Extra capacity | Rollback speed | Mixed versions per session | Use when |
|---|---|---|---|---|
| Rolling | surge pods only | minutes (roll back) | yes, briefly | adjacent-compatible changes |
| Blue-green | full second fleet | seconds (switch back) | yes, at the switch | fast rollback needed |
| Canary | a few pods | seconds (stop canary) | yes, for canary sessions | behavioural risk |
| Session-pinned | old fleet until sessions idle | seconds for new sessions | no | policy or breaking changes |
Worked example: renaming a tool and a state key
A team wants to rename lookupOrder to getOrder and move the cart from a flat cart key to a structured cart_v2. Sessions expire after seven days of inactivity. Release A, deployed Monday by rolling update, registers both tools, reads cart_v2 with a fallback and writes both keys; release N, still serving during the rollout, never sees anything it cannot read. Release B, deployed Wednesday, instructs the model to use only getOrder and writes only cart_v2; A and B coexist safely because A already reads the new key. Release C ships the following Thursday, after every session last touched before Wednesday has expired, and removes the alias and the fallback. A dashboard counting calls to the alias tool should be at zero before C is approved; if it is not, some client is replaying old histories and C waits.
Failure modes
- Grace period shorter than a turn. Long tool calls are killed mid-flight, leaving half-applied side effects. Measure the 99th percentile turn duration and size from it.
- Readiness equals liveness. New pods take traffic before they can reach the session store, and the first turns of the rollout fail.
- One-step rename. A tool or key renamed in a single release breaks every conversation that crosses the rollout boundary.
- In-memory sessions in production. With the in-memory session service, every replacement pod loses its conversations; no strategy is zero-downtime with it.
- Canary judged on HTTP errors only. A prompt regression passes every infrastructure check. Gate on an evaluation score as well.
- Retries without idempotency. A client retries a turn cut by draining and the tool acts twice.
What to do next
- Measure turn duration percentiles and set the grace period and drain deadline from them, following graceful shutdown for ADK Java.
- Make readiness check the session store and configuration, as in running ADK Java on Kubernetes.
- Add the drain gate and an idempotency key per turn, and pass the key to side-effecting tools.
- List your compatibility surface: state keys, tool names and parameters, instructions, event shapes.
- Adopt the expand, migrate, contract rule for every change to that surface, and track the persistent schema with migrations and schema versioning.
- Record the release in session state and in logs, and tag releases as described in ADK Java versioning.
- Add an evaluation gate to canary promotion, not just error rate and latency.
- On Cloud Run, apply the same rules to revisions; see deploying ADK Java to Cloud Run.