Blue-green deployment keeps two complete copies of a service. One copy, blue, serves all traffic. The other, green, runs the new release with no traffic. You verify green, move traffic to it with one switch, and keep blue running so that switching back is just as fast. The idea is simple for a stateless web service. An ADK Java agent service is not stateless, so the same idea needs more care.
An agent turn reads a session that may be hours old, calls a model through a shared quota, and runs tools that change other systems. This article covers the switch on Cloud Run and Kubernetes, checking the idle colour before cutover, session compatibility in both directions, warm-up against model quota, and how long to keep the way back open. Gradual percentage rollouts are covered in ADK Java canary deployment, and pinning sessions to a release bundle is covered in ADK Java versioning. This page links to those rather than repeating them.
What blue-green buys an agent service, and what it complicates
Blue-green suits releases you cannot split by percentage, framework or JDK upgrades you want verified as a whole runtime before anyone uses them, and any change where rollback must take seconds rather than a rolling restart.
Agents add four complications. Sessions outlive deployments: after a rollback, green's events must make sense to blue. Model behaviour is part of the release: a new model or instruction can pass every health check and still answer worse. Capacity is shared: two warm colours draw on one model quota. Tools act on the world: switching back does not undo green's tool calls.
Architecture: two colours, shared state
The architecture has a switch in front of two colours, and both colours share the stateful dependencies behind them. Nothing below the switch is duplicated except the agent processes themselves.
- The switch is a Cloud Run traffic split between tagged revisions, or a Kubernetes Service whose selector names a colour label. It should be one command or one manifest change, and it should be reversible the same way.
- The session store is shared. If each colour had its own store, every switch would lose every conversation. In ADK Java that means a
BaseSessionServiceimplementation backed by something durable: the Vertex AI session service, or your own service over PostgreSQL.InMemorySessionServicecannot be used with blue-green at all. - The verifier drives green through its own address before any user does.
- Shared external capacity (model quota, tool APIs, databases) takes load from both colours during verification and drain.
The compatibility contract runs in both directions
Rollback is only instant if blue can carry on every conversation that green touched. So the release contract has two directions. Green must read sessions written by blue, and blue must read sessions written by green. Three things cross the boundary: session state keys, the event history the model sees on the next turn, and the tool names and argument shapes recorded in that history.
So: add state keys freely but never rename or retype one in the same release, have older code ignore unknown keys, keep removed tools registered as stubs for one release, and record the writing release in state. The guard below runs before every turn and refuses a session it cannot read safely, instead of corrupting it.
public final class SessionCompat {
// No "app:"/"user:" prefix: those scopes are shared across sessions; these keys must be per session.
static final String WRITER_KEY = "writer_release"; // set on every turn
static final String SCHEMA_KEY = "state_schema"; // integer, bumped on breaking change
private final int readableMin; // oldest schema this build understands
private final int readableMax; // newest schema this build understands
private final String release;
public SessionCompat(int readableMin, int readableMax, String release) {
this.readableMin = readableMin; this.readableMax = readableMax; this.release = release;
}
/** Throws before the turn runs if this build would misread the session. */
public void check(Session session) {
Object raw = session.state().getOrDefault(SCHEMA_KEY, readableMin);
int schema = ((Number) raw).intValue();
if (schema < readableMin || schema > readableMax) {
throw new IncompatibleSessionException(session.id(), schema, release);
}
}
/** Stamped into the state delta of every turn this build runs. */
public Map<String, Object> stamp(int writtenSchema) {
return Map.of(WRITER_KEY, release, SCHEMA_KEY, writtenSchema);
}
}Blue at schema 3 is built with readableMax = 4 before green ships schema 4. That is the expand step, and it is the same discipline as database expand/contract, applied to session state. Green writes schema 4 only for keys blue already knows how to ignore. The contract step, dropping schema 3 support, waits until blue is gone. The database side of the same pattern is covered in adk-postgres migrations and schema versioning.
Verifying green before anyone uses it
Green has to prove more than "the process started". Run three gates in order against green's private address, and stop at the first failure.
Readiness. Your existing readiness probe, called through the tag URL: process up, model client authenticated, tool backends reachable.
Golden replays. Twenty to two hundred recorded conversations with outcome checks: the right tool with the right key arguments, required facts present, nothing forbidden. Use sandboxed tool doubles, because a replay that books a real meeting is a side effect.
Continuation. Copy a sample of recent real sessions into a scratch namespace and run one more turn on green. It is the only gate that tests the compatibility contract against real data, and the one most often skipped.
// Runs against green's tag URL, e.g. https://green---agent-svc-abcdef.a.run.app
public Verdict verify(URI green, List<Golden> goldens, List<String> sampledSessions) {
if (!http.get(green.resolve("/readyz")).ok()) return Verdict.fail("readiness");
int passed = 0;
for (Golden g : goldens) {
String sid = api.createSession(green, "verifier", g.initialState());
TurnTrace trace = null;
for (String userMsg : g.userTurns()) trace = api.turn(green, sid, userMsg);
if (g.check(trace)) passed++; // tool names, args, required facts
}
double passRate = (double) passed / goldens.size();
if (passRate < baselinePassRate - 0.02) return Verdict.fail("goldens " + passRate);
for (String sid : sampledSessions) {
String copy = api.cloneSession(sid, "verify-scratch"); // never mutate the real one
try { api.turn(green, copy, "continue"); }
catch (IncompatibleSessionException e) { return Verdict.fail("continuation " + sid); }
}
return Verdict.pass(passRate);
}Compare green's golden pass rate with blue's on the same set, measured the same day. Model output varies between runs, and the question is whether green is worse than what users already have.
Warm-up and the shared model quota
A green with cold connection pools, cold JIT and empty caches answers its first minute slowly, which can look like a failed release. Warm it with the golden replays plus a short synthetic load at about a tenth of production rate, and on Cloud Run set minimum instances on the green revision before cutover.
Model quota is the constraint people forget. Token and request limits usually apply per project and model, not per deployment, so verification, warm-up and the drain overlap all count against the same limit. If production uses 70% of quota at peak, cut over off-peak or get headroom first, or the switch will produce 429s that look like a green bug.
The switch on Cloud Run and Kubernetes
On Cloud Run, colours are revisions, and tags give each one a stable address. Deploy green with no traffic, verify it at its tag URL, then move traffic with one command. The commands below follow the Cloud Run traffic-management documentation.
# 1. Deploy green: new revision, zero traffic, reachable only at its tag URL
gcloud run deploy agent-svc --image "$IMAGE" --no-traffic --tag green
# 2. Verify against https://green---agent-svc-<hash>.a.run.app (verifier above)
# 3. Cut over: all traffic to the revision tagged green
gcloud run services update-traffic agent-svc --to-tags green=100
# Rollback: back to the previous revision by name
gcloud run services update-traffic agent-svc --to-revisions agent-svc-00041-abc=100On Kubernetes, run two Deployments that differ only in a color label, and point the Service selector at one of them.
# Two Deployments, agent-blue and agent-green, labelled color=blue / color=green.
kubectl patch service agent-svc -p '{"spec":{"selector":{"app":"agent","color":"green"}}}'
# Rollback is the same patch with "blue".On Cloud Run, new requests move at once. A Kubernetes selector flip moves only new connections, so keep-alive and HTTP/2 clients stay on blue until blue closes them. Streaming turns can run for tens of seconds, so blue must stay up and drain. Shutdown and drain behaviour is covered in ADK Java graceful shutdown. Never scale blue to zero as part of the cutover command.
Sessions in the middle of a task
The switch moves requests, not conversations, so you need a policy for sessions that are in the middle of a task. There are two reasonable choices.
Hard cut. Every session's next turn goes to green. This is the simplest option and relies entirely on the compatibility contract. Use it for releases that only change code paths and add state keys.
Drain-pinned. Sessions active in the last N minutes stay on blue until they go idle or reach a deadline, and new sessions go to green. This needs a session-aware router in front of both colours, which is the same component the canary article builds, run with only two states. Use it when green changes model or instruction enough that switching mid-task would confuse the user. The router sends a session to blue only while it is pinned there, its last turn is within the idle timeout, and the drain deadline has not passed.
Keep the drain deadline short, typically 30 to 60 minutes. A drain that never ends turns blue-green into running two versions indefinitely, with none of the comparison data a canary would give you.
The rollback window and what makes rollback unsafe
The rollback window is how long blue stays warm. Longer windows catch slower regressions and cost more compute, but the real limit is what makes rollback unsafe even while blue runs.
| Green did this | Rollback is | What to do |
|---|---|---|
| Added state keys blue ignores | Safe | Nothing; this is the normal case |
| Wrote a schema blue cannot read | Unsafe for those sessions | Ship blue with the wider readable range first |
| Called a new tool with side effects | Safe for routing, not for effects | Make the tool idempotent; log effects for manual reversal |
| Started long-running tool jobs | Callbacks may land on blue | Version the callback payload; blue must accept or park it |
| Ran a contracting DB migration | Unsafe | Never contract in the same release as the switch |
A common default: blue at full size for an hour after cutover, reduced until the next day's quality review, then deleted. Record each release as "rollback-safe until" a specific time, so on-call knows whether flipping back is allowed.
Worked example: a tool rename and a model change
Here is a worked example. A support agent's release 42 replaces the lookup_order tool with get_order, which takes an order id plus an optional customer id, and moves to a newer model. Blue is release 41.
First, release 41.1 ships to blue as an ordinary rolling deploy. It registers get_order as an alias, so blue can read histories that mention it, and widens readableMax to schema 4. Release 42 then deploys as green with no traffic, keeping a lookup_order stub that forwards to get_order. Green scores 94.2% on 120 goldens against blue's 93.8% the same morning, and 50 cloned production sessions pass a continuation turn. Warm-up runs at 10% of production rate; model quota peaks at 61%, so no extra headroom is needed.
Cutover is one update-traffic command. Twenty minutes later, the order-lookup error rate on green is 0.4% against blue's 0.1%. Calls without a customer id hit an authorisation check that blue never performed. The on-call engineer flips back. Blue reads every green-touched session because of 41.1, and no conversation breaks. The fix ships as 42.1 through the same pipeline the next day. Release 43 removes the lookup_order stub and drops schema 3, after blue has been deleted.
Failure modes
- Per-colour session stores. Each switch silently starts every conversation from scratch. Use one shared store.
- Health-check-only verification. Green is up and answers worse. Gate on goldens compared with blue the same day.
- One-way compatibility. Green reads blue's sessions, but rollback strands green-touched sessions. Ship the expand step to blue first.
- Quota collision. Warm-up plus peak traffic exceeds the shared quota, and the 429s get blamed on the release.
- Verifier side effects. Golden replays send real emails or tickets. Use sandboxed tool doubles.
- Blue scaled to zero at cutover. In-flight streams are cut, and rollback lands on cold instances.
Trade-offs against canary and rolling updates
Blue-green doubles compute during the window and gives all-or-nothing exposure: if the verifier misses a defect, everyone sees it until you flip back. A canary exposes a small share of sessions and gives you statistical evidence, but it needs session-aware routing and a long enough run to collect that evidence. A rolling update is the cheapest, but it runs mixed versions with no instant way back.
Many teams use blue-green for the runtime (JDK, framework, ADK library upgrades) and canary for behaviour (model, instructions, tools). The session contract is the same either way. Platform specifics are in deploying ADK Java to Cloud Run and running ADK Java on Kubernetes.
What to do next
- Confirm that every environment uses one durable shared session store, and remove any in-memory session service from production.
- Add a schema marker and writer release to session state, plus a guard that refuses to run unreadable sessions.
- Adopt the two-step rule: widen what blue can read before green writes anything new.
- Record 50 or more golden conversations with tool-level checks, and run them against blue daily to get a baseline.
- Build a continuation gate that clones real sessions and runs one turn on green.
- Measure peak model quota use and decide the cutover window from it.
- Script the switch and the rollback as single commands, and rehearse both in staging.
- Write "rollback-safe until" into each release record, and delete blue only after that time.