A canary sends a small share of real traffic to a new version and promotes it only if it behaves. For an ADK Java agent that serves people through a web front end, the session-pinned canary covers assignment, per-session metrics and how many sessions each stage needs. This page covers the harder case: an agent that other agents call over A2A, deployed as its own service on Kubernetes. Its users are programs that read its agent card, hold contextId values across turns, poll tasks by ID and retry. A canary that ignores those habits breaks callers in ways that look like random flakiness.
The plan here is four parts: check the new agent card against the old one before any traffic moves; route by conversation, not by request, so tasks stay on the version that created them; judge the canary on task outcomes rather than HTTP status; and roll back without orphaning tasks that are still running. The examples use an Istio mesh and a2a-java 0.3.x method names, as pinned by google-adk-a2a; the ideas carry to any L7 router.
What changes when your callers are agents
Three things differ from a canary of a stateless HTTP API.
The contract is the card. Calling models choose to delegate by reading the card's description and skills, and calling code chooses a transport, streaming and input modes from it. A canary that drops a skill or flips capabilities.streaming changes caller behaviour before a single turn runs, and a canary serving a different card than stable is serving two contracts at one URL.
Requests address state. With in-memory task stores, a tasks/get that reaches the other version returns -32001 TaskNotFoundError for a task that exists. A follow-up turn that crosses versions runs against a session the new version may not even read the same way. Per-request weighted routing guarantees both for a fraction of conversations.
Success is not a status code. The reference JSON-RPC server answers protocol errors with HTTP 200, and a task that ends failed is a successful HTTP exchange. A canary judged on 5xx rate will promote a version that fails half its tasks.
Gate on the agent card
Gate the rollout on a card diff. Fetch both cards (the canary's through a port-forward or a header-routed request) and classify every change:
import json, sys, urllib.request
def card(url, headers=None):
req = urllib.request.Request(url + "/.well-known/agent-card.json", headers=headers or {})
with urllib.request.urlopen(req, timeout=5) as r:
return json.load(r)
def diff(old, new):
breaking, review = [], []
old_sk = {s["id"]: s for s in old.get("skills", [])}
new_sk = {s["id"]: s for s in new.get("skills", [])}
for sid in old_sk.keys() - new_sk.keys():
breaking.append(f"skill removed: {sid}")
for sid in old_sk.keys() & new_sk.keys():
if old_sk[sid].get("description") != new_sk[sid].get("description"):
review.append(f"skill description changed: {sid}") # changes delegation
for k in ("defaultInputModes", "defaultOutputModes"):
if set(old.get(k, [])) - set(new.get(k, [])):
breaking.append(f"{k} narrowed")
if old.get("capabilities", {}).get("streaming") and not new.get("capabilities", {}).get("streaming"):
breaking.append("streaming withdrawn")
for k in ("url", "preferredTransport", "securitySchemes"):
if old.get(k) != new.get(k):
breaking.append(f"{k} changed")
if old.get("description") != new.get("description"):
review.append("agent description changed")
return breaking, review
base = "http://research-agent.agents.svc.cluster.local"
b, r = diff(card(base, {"x-agent-cohort": "stable"}), card(base, {"x-agent-cohort": "canary"}))
print("\n".join(["BREAKING " + x for x in b] + ["REVIEW " + x for x in r]))
sys.exit(1 if b else 0)Breaking changes do not belong in a canary at all: they need a new card version served side by side, with callers migrating explicitly. Description changes are allowed but flagged, because they shift which requests calling models send you, which also shifts the canary's traffic mix and makes the comparison unfair.
Route conversations, not requests
Route whole conversations. The caller decides a cohort once per contextId, remembers it, and sends it as a header on every call in that conversation, including task polls. The mesh routes on the header. Deciding once matters: if the cohort were recomputed from a percentage on each call, raising the canary from 5% to 20% would move live conversations from stable to canary mid-task.
public final class CohortInterceptor extends ClientCallInterceptor {
private final Map<String, String> byContext = new ConcurrentHashMap<>(); // bounded LRU in production
private final Map<String, String> byTask = new ConcurrentHashMap<>(); // filled from response events
private volatile int canaryPercent; // from dynamic config
@Override
public PayloadAndHeaders intercept(String method, Object payload, Map<String, String> headers,
AgentCard card, ClientCallContext ctx) {
String cohort = "stable";
if (payload instanceof JSONRPCRequest<?> req) {
Object params = req.getParams();
if (params instanceof MessageSendParams msp && msp.message().getContextId() != null) {
String cid = msp.message().getContextId();
cohort = byContext.computeIfAbsent(cid,
k -> Math.floorMod(k.hashCode(), 100) < canaryPercent ? "canary" : "stable");
} else if (params instanceof TaskQueryParams q) {
cohort = byTask.getOrDefault(q.id(), "stable");
} else if (params instanceof TaskIdParams t) {
cohort = byTask.getOrDefault(t.id(), "stable");
}
}
Map<String, String> h = new HashMap<>(headers);
h.put("x-agent-cohort", cohort);
return new PayloadAndHeaders(payload, h);
}
/** Call from the client's event consumer when a task appears. */
public void recordTask(String taskId, String contextId) {
byTask.put(taskId, byContext.getOrDefault(contextId, "stable"));
}
public void setCanaryPercent(int pct) { canaryPercent = pct; }
}This needs callers to mint contextId values before the first send, so turn one already has a key. When you cannot change callers, put the same logic in an ingress gateway that owns the maps instead. The mesh side is plain Istio:
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata: { name: research-agent }
spec:
hosts: [research-agent.agents.svc.cluster.local]
http:
- match: [{ headers: { x-agent-cohort: { exact: canary } } }]
route: [{ destination: { host: research-agent.agents.svc.cluster.local, subset: v15 } }]
- route: [{ destination: { host: research-agent.agents.svc.cluster.local, subset: v14 } }]
---
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata: { name: research-agent }
spec:
host: research-agent.agents.svc.cluster.local
subsets:
- { name: v14, labels: { version: v14 } }
- { name: v15, labels: { version: v15 } }Requests without the header fall to stable, which is the safe default for card fetches and for callers that have not adopted the interceptor. Avoid weighted routes for this service: per-request weights are exactly the cross-version scatter the header prevents.
Signals that decide promotion
Judge on what callers experience, split by version. Emit these from the server, using the ADK executor's after-execute and after-event callbacks, with a version label; the metric names below are your own, not built-ins:
| Signal | Why | Example threshold |
|---|---|---|
| Terminal state mix: completed / failed / canceled / input-required | the real success rate; input-required spikes mean the agent stopped understanding requests | failed share no more than 1 point above stable |
| Time to terminal state, p50 and p95 | callers budget for it; streaming hides it | p95 within 20% of stable |
| JSON-RPC error codes by method | protocol regressions, routing leaks (-32001) | no new codes; -32001 near zero |
| Model tokens and tool calls per task | cost and loops | within 15% of stable |
| Golden-task eval score | quality the metrics above cannot see | no drop beyond the eval's noise band |
The golden-task score comes from replaying a fixed set of recorded A2A requests against both versions with the header set, and scoring with the same evaluator you use in CI. Live traffic gives volume and real distribution; replays give a paired comparison on identical inputs. Use both. For how many conversations each stage needs before a difference is real, use the sample-size method in the session-pinned canary guide rather than eyeballing dashboards.
Rollback and promotion with tasks in flight
Rolling back means setting the canary percentage to zero, and that only stops new conversations. Existing canary conversations still carry the canary header and their tasks still live in canary pods. If the canary failure is severe, also clear the caller's cohort maps so new turns of those conversations go to stable, accepting that they start without their session unless sessions are in a shared store that both versions read. Either way, keep canary pods running until their in-flight tasks reach a terminal state: scaling the deployment to zero immediately turns every running canary task into -32001 on the next poll.
The same rule governs promotion. Promote by moving stable to the new image, then clear cohorts, then drain the old canary subset, in that order. A version that writes session state in a new shape must be readable by the old version for the length of the rollout, the expand-and-contract discipline from zero-downtime deploys.
Worked example: a description change that moved traffic
Version 15 of a research agent switches its summarisation model and rewrites one skill description. The card diff reports one REVIEW item and no BREAKING items. Callers handle about 4,000 conversations a day.
Stage 1, 5% for 4 hours: about 33 canary conversations. Nothing crashes, and HTTP 5xx is zero on both versions. The terminal-state metric tells a different story: input-required is 9% on v15 against 2% on v14. Logs show the new description attracts requests for a kind of report v15 asks clarifying questions about, a traffic-mix change the card diff warned of.
Decision: 33 conversations is too few to call a 7-point difference with confidence, so the controller holds at 5% instead of ramping. The golden-task replay of 200 recorded requests shows quality equal within noise and input-required at 8% on the subset of report-style requests. The team reverts the description, ships v15.1, and the next 5% stage shows input-required at 2.4%, which ramps to 25% and then 100% over two days.
Without cohort routing the same stage would have scattered polls across versions: with 5% weights and in-memory stores, about 5% of stable tasks' polls would hit canary pods and fail with -32001, a failure the team would have blamed on the canary when it was the router.
Failure modes
- Judging on HTTP status. JSON-RPC errors and failed tasks are HTTP 200. Fix: terminal-state and error-code metrics by version.
- Per-request weights. Polls and follow-ups cross versions. Fix: per-conversation cohorts carried in a header.
- Card drift in the canary. Two contracts at one URL. Fix: card diff as a pipeline gate; breaking changes get a new versioned endpoint.
- Orphaned tasks on rollback. Canary pods removed with tasks running. Fix: drain to terminal states before scaling down.
- Unequal traffic mix. A changed description changes who delegates to you. Fix: compare on replayed golden tasks as well as live traffic.
- Unbounded cohort maps. The caller's maps grow forever. Fix: LRU bounded by your longest conversation lifetime.
Trade-offs
Caller-side cohorts give exact conversation affinity, but every calling team must adopt the interceptor; a gateway that owns the cohort maps centralises it at the cost of one more hop and a stateful component. Shared session and task stores make rollback and promotion clean, because any version can serve any conversation, but they force both versions to agree on stored formats. Shadow traffic, replaying requests to the new version without returning its answers, avoids caller impact entirely but cannot measure how callers react to the new agent's questions or artifacts, and duplicates any side effects the agent's tools perform. Most teams replay golden tasks first, then run a cohort canary on live traffic.
What to do next
- Add the card diff to your deploy pipeline and fail on BREAKING changes.
- Have callers mint
contextIdvalues and add a cohort interceptor, or put the cohort logic in a gateway. - Create header-matched Istio routes with stable as the default; remove weighted routes for A2A services.
- Emit terminal-state, time-to-terminal, error-code and cost metrics with a version label.
- Record 100 to 300 representative A2A requests as a golden replay set and score both versions on it.
- Write the rollback runbook: percentage to zero, optional cohort reset, drain, then scale down.
- Go deeper with the session-pinned canary, the agent deploy pipeline, progressive rollouts and kill switches, zero-downtime deploys and agent versioning.