A canary release sends a small share of real traffic to a new version, compares it with the current version, and either widens the share or rolls back. For a stateless web service that is a solved problem: split requests by weight and watch error rates and latency. An agent built with ADK Java breaks three of those assumptions. A conversation spans many requests and must stay on one version. The important failures are behavioural, such as a wrong tool call or an unhelpful answer, rather than HTTP errors. And the signal is noisy, because the same input can produce different outputs.
ADK Java does not ship a traffic-splitting or canary feature; as of version 1.11.0 it gives you agents, runners, sessions and an event stream, and you build the canary around them. This article shows how, from session assignment to the controller that promotes or rolls back. The release pipeline that produces the candidate, and the full table of release gates, are covered in CI/CD for ADK Java; this page is about the canary itself.
Two kinds of canary for an agent
There are two ways to run a canary for an agent, and most teams need both.
| Infrastructure canary | In-process variant | |
|---|---|---|
| What differs | The build: code, libraries, JVM, tools | Configuration: instruction, model id, tool set |
| How it runs | Two deployments, a router in front | One deployment holding two agent definitions |
| Session pinning | Router sets a header or cookie at creation | Lookup in the assignment store per turn |
| Rollback speed | Router change, then drain canary pods | Flip a config flag, effective on the next turn |
| Typical use | Framework upgrades, new tool code | Prompt edits, model upgrades |
Prompt and model changes are the most frequent changes to an agent and the most likely to shift behaviour, and they need no new build, so the in-process variant gives the fastest, cheapest loop. Code changes need the infrastructure form. The assignment logic and the analysis are the same in both; only where the routing decision is enforced differs.
One warning for the infrastructure form. Weighted routing in load balancers and service meshes generally splits per request. If you use it directly, a ten-turn conversation hits both versions. The router must make the decision once, when the session is created, record it, and route every later turn by that record, for example by a header your gateway matches on.
Architecture of the canary loop
Both variants share one session service, so a session's history lives in one place, which is what makes rollback by reassignment possible. The controller's percentage affects only sessions created afterwards; existing sessions keep their variant until they end.
Assigning sessions, once
Assignment is a deterministic bucket for new sessions plus a persisted record. Do not recompute the variant from a hash on every turn: when the canary share rises from 5 to 25 percent, a recomputed hash moves sessions that started on the baseline onto the canary in the middle of a conversation.
public final class VariantRouter {
public enum Variant { BASELINE, CANARY }
private final AssignmentStore store; // e.g. a table keyed by session id
private volatile int canaryBasisPoints; // 0..10000, written by the controller
private final String salt; // change per experiment
public VariantRouter(AssignmentStore store, String salt, int initialBasisPoints) {
this.store = store; this.salt = salt; this.canaryBasisPoints = initialBasisPoints;
}
/** Called once, when the session is created. */
public Variant assign(String sessionId) {
int bucket = Math.floorMod((salt + ":" + sessionId).hashCode(), 10_000);
Variant v = bucket < canaryBasisPoints ? Variant.CANARY : Variant.BASELINE;
store.putIfAbsent(sessionId, v); // first write wins on races
return store.get(sessionId);
}
/** Called on every turn. Unknown sessions (pre-canary) stay on baseline. */
public Variant variantFor(String sessionId) {
Variant v = store.get(sessionId);
return v != null ? v : Variant.BASELINE;
}
public void setCanaryBasisPoints(int bp) { this.canaryBasisPoints = bp; }
}Bucket by user rather than session when users return, so one person does not alternate between versions on consecutive days. Use a fresh salt per experiment, and in production a well-mixed hash such as truncated SHA-256 rather than String.hashCode. Exclude internal test traffic and accounts you cannot experiment on before bucketing.
Running two agent configurations in one service
For a configuration canary, build both agents at startup from configuration and keep one runner per variant. Each turn looks up the session's variant and calls that runner. The builder methods below are the documented ones; the model id comes from configuration rather than being hard-coded, because model names change faster than code.
LlmAgent buildAgent(AgentConfig cfg, List<BaseTool> tools) {
return LlmAgent.builder()
.name(cfg.name()) // same name in both variants
.description(cfg.description())
.model(cfg.modelId()) // e.g. read from release manifest
.instruction(cfg.instruction())
.tools(tools)
.build();
}
Map<Variant, Runner> runners = Map.of(
Variant.BASELINE, runnerFactory.create(buildAgent(baselineCfg, tools)),
Variant.CANARY, runnerFactory.create(buildAgent(canaryCfg, tools)));
TurnResult handleTurn(String userId, String sessionId, Content message) {
Variant v = router.variantFor(sessionId);
long start = System.nanoTime();
int toolCalls = 0;
StringBuilder answer = new StringBuilder();
for (Event e : runners.get(v).runAsync(userId, sessionId, message).blockingIterable()) {
toolCalls += e.functionCalls().size();
if (e.finalResponse()) answer.append(textOf(e));
}
outcomes.recordTurn(sessionId, v, toolCalls, System.nanoTime() - start);
return new TurnResult(answer.toString(), v);
}runnerFactory stands for however your service builds its production runner today, with the same app name and the same session, artifact and memory services for both variants; sessions are looked up by app name, user and session id, so a different app name strands every reassigned session.textOf is your existing helper that pulls text parts out of an event's content; check the accessors it uses against the ADK Java version you run. If the canary changes tool definitions, both runners need tools that can read each other's session state, which is the expand-and-contract discipline described in the CI/CD article.
Measuring per session, not per call
The unit of randomisation is the session, so the unit of measurement must be the session too. Tool calls and turns within one conversation are correlated: an agent that misreads a tool schema once usually misreads it on every call in that session. Counting 3,000 tool calls as 3,000 independent trials overstates your confidence badly. Reduce each finished session to a few outcomes and compare those:
| Session outcome | Defined as | Gate direction |
|---|---|---|
| Tool failure | At least one tool call failed or had invalid arguments | Canary not worse |
| Escalated or abandoned | Handed to a human, or user left mid-task | Canary not worse |
| Turns to resolution | Turns in sessions that completed | Within a set margin |
| Cost | Model tokens and tool spend per session | Within a budget ratio |
| Turn latency | 95th percentile per turn | Within a set margin |
Close a session for measurement after a fixed idle period, say 30 minutes, so both arms are measured the same way. Always compare against the baseline sessions created in the same window, never against last week: traffic mix, tool backends and the model provider's behaviour all drift, and a concurrent comparison cancels most of that out. Instrumentation for tagging every span and metric with the variant is covered in ADK Java observability.
How many sessions each stage needs
A canary that runs until it feels long enough decides on noise. Work out how many sessions each arm needs to detect the smallest regression you care about. For a proportion such as the tool-failure rate, with baseline rate p1, a regression to p2 worth catching, a one-sided significance level alpha and power 1 minus beta:
n per arm = (z_alpha + z_beta)^2 * (p1(1 - p1) + p2(1 - p2)) / (p2 - p1)^2Worked example. The baseline has a tool failure in 4 percent of sessions, and you want to catch a rise to 6 percent. With alpha 0.05 one-sided (z = 1.645) and 80 percent power (z = 0.842), the sum is 2.487 and its square 6.185. The variances add to 0.04 x 0.96 + 0.06 x 0.94 = 0.0948. The difference squared is 0.0004. So n = 6.185 x 0.0948 / 0.0004, about 1,466 sessions per arm.
Now the schedule. With 20,000 new sessions a day, a 5 percent canary gets 1,000 a day and needs about a day and a half to reach 1,466. A 25 percent stage gets there in about seven hours. Detecting a smaller shift, from 4 to 5 percent, needs roughly four times as many sessions, because the required n grows with the inverse square of the difference. This is why agent canaries take days rather than minutes, and why the earliest stages should rely on guardrails that catch gross breakage rather than on statistics.
The second trap is repeated looks. If you run the same test at 5, 25 and 50 percent and stop at the first significant result, your real false-alarm rate is well above 5 percent. The simplest valid fix is to plan the looks in advance and split alpha across them: with three looks, use 0.05 / 3, about 0.0167 per look, which raises the one-sided threshold to z = 2.13. It is conservative, and that is acceptable for a release gate. Sequential tests designed for continuous monitoring are more efficient, but use a library that implements one rather than improvising.
The controller
The controller is a small state machine. Guardrails run every minute at every stage and can roll back immediately; they use absolute thresholds, such as a crash rate or a tool error spike that is obviously broken, and need no statistics. The statistical gate runs only when a stage has collected its planned sessions.
record Arm(long sessions, long failures) {
double rate() { return sessions == 0 ? 0 : (double) failures / sessions; }
}
enum Verdict { PROMOTE, HOLD, ROLLBACK }
static double zWorse(Arm base, Arm canary) {
double pooled = (double) (base.failures() + canary.failures())
/ (base.sessions() + canary.sessions());
double se = Math.sqrt(pooled * (1 - pooled)
* (1.0 / base.sessions() + 1.0 / canary.sessions()));
return se == 0 ? 0 : (canary.rate() - base.rate()) / se;
}
Verdict evaluate(Stage stage, Arm base, Arm canary, Guardrails g) {
if (g.tripped()) return Verdict.ROLLBACK; // any time, any stage
if (canary.sessions() < stage.requiredSessions()) return Verdict.HOLD;
double z = zWorse(base, canary);
if (z > stage.zThreshold()) return Verdict.ROLLBACK; // e.g. 2.13 for 3 looks
return Verdict.PROMOTE; // next stage or 100%
}Run one comparison per gated outcome and count every gate in the alpha split. The base arm is baseline sessions created in the same window. Record each verdict with its inputs so a rollback can be explained afterwards.
Rolling back with sessions in flight
Rolling back stops new assignments immediately: set the canary share to zero. The harder question is what happens to sessions already on the canary. There are two choices. Drain: leave them on the canary until they go idle, which is safe for state but leaves some users on the bad version for a while. Reassign: move them to the baseline on their next turn, which ends exposure but requires the baseline to read whatever the canary wrote into session state. Choose per incident: for a harmful behaviour, reassign; for a cost or latency regression, drain.
In the infrastructure form, draining means the canary deployment keeps running until its pinned sessions end, so a rollback is not complete when the router changes. Termination grace periods and drain behaviour for long streaming turns are covered in ADK Java on Kubernetes. Before the canary starts, decide which option each failure class gets, and rehearse a reassignment in staging.
Failure modes
- Per-request splitting. Mid-conversation switches mix two instructions in one transcript and make both arms' metrics meaningless. Pin at creation.
- Recomputed hashes. Raising the share moves live sessions. Persist assignments.
- Counting tool calls as trials. Correlated calls inflate confidence. Measure per session.
- Peeking. Testing at every stage at full alpha produces false rollbacks and false promotions. Split alpha across planned looks.
- Historical baselines. Comparing with last week confuses drift with regression. Use concurrent baseline sessions.
- Shared quota. Both variants draw on the same model quota; a canary that loops on tool calls can throttle the baseline. Put a per-variant budget in the guardrails.
- Silent state drift. The canary writes new session keys the baseline ignores, and rollback by reassignment loses them. Test cross-reading before starting.
Trade-offs
The in-process variant is cheap and quick but only covers configuration changes, and a crash in shared code takes down both arms. The infrastructure canary isolates builds at the cost of a second deployment and a router that understands sessions. Larger stages reach a decision sooner but expose more users to a regression; guardrails make early stages safe, and statistics make later ones meaningful. Offline evaluation before the canary, described in the ADK Java evaluation framework, catches most regressions more cheaply; the canary exists for what replayed cases miss, such as real users' phrasing, real tool backends and real cost.
What to do next
- Write down your gated session outcomes and the smallest regression in each that you would refuse to ship.
- Compute sessions per arm for each, and from your daily session volume, the time each stage needs.
- Add an assignment store and a router that assigns once at session creation and routes every turn by the stored variant.
- Build both agent configurations at startup from the release manifest and tag every turn's metrics with the variant.
- Define absolute guardrails that roll back at any stage without statistics, including a per-variant model budget.
- Plan your looks, split alpha across them, and encode thresholds in the controller rather than in a runbook.
- Decide drain or reassign for each failure class, and rehearse a reassignment in staging.
- Run the first canary on a harmless instruction change to prove the plumbing before you depend on it.