A shadow deployment runs a candidate version of an agent on real production requests without letting it affect anyone. Users get the production reply; the candidate's reply, tool calls and costs are recorded and compared. For an ordinary service, shadowing is mostly a traffic-mirroring problem. For an agent it is harder in three specific ways: the candidate's answer depends on the whole conversation so far, its tools have side effects that must not happen twice, and its behaviour diverges as soon as it calls a tool production did not call.
This article builds shadowing for ADK Java: capturing each production turn, forking the session into a separate store, quarantining tools with recorded reads and stubbed writes, scoring divergence, controlling cost and privacy, and reading the results. It assumes you already run offline regression checks such as those in the regression detection article; shadowing adds coverage of the traffic you did not think to put in a dataset.
What shadowing can and cannot tell you
Shadowing answers one question well: on today's real requests, where does the candidate behave differently from production, and are those differences better or worse? It is the right tool for a new model, a rewritten instruction, a changed tool set or an ADK upgrade, where you cannot enumerate the inputs that matter.
It cannot measure what users do with an answer, because they never see it, so satisfaction and task completion still need a canary. It also cannot observe true multi-turn behaviour: each shadow turn is forked from production's history, so the candidate never sees the consequences of its own previous replies. Shadow evaluates single decisions in real context; canaries evaluate whole conversations.
Architecture
The production path is unchanged apart from emitting a turn record. Everything the candidate touches is either a copy (the forked session), a recording (reads production already made) or a stub (writes), and the scorer compares its output to production's record.
Capturing the production turn
Capture in the serving path, after the turn completes, and never block the reply. Collect the events the runner emits, build a turn record and offer it to a bounded queue; if the queue is full, drop the record and count the drop. The record holds what the shadow needs to fork and what the scorer needs to compare.
record TurnRecord(String userId, String sessionId, String invocationId, Content userMessage,
Map<String, Object> stateAtStart, // session state read before runAsync
List<ToolCall> toolCalls, // name, canonical args, result, read or write
String finalText, long tokens, long latencyMs) {}
Flowable<Event> serve(String userId, String sessionId, Content msg) {
List<Event> seen = new CopyOnWriteArrayList<>();
long start = System.nanoTime();
return runner.runAsync(userId, sessionId, msg)
.doOnNext(seen::add)
.doOnComplete(() -> {
if (!sampler.take(userId, sessionId)) return; // per-session sampling, see below
TurnRecord r = TurnRecords.from(userId, sessionId, msg, seen, toolClasses,
(System.nanoTime() - start) / 1_000_000);
if (!shadowQueue.offer(r)) metrics.increment("shadow.dropped");
});
}Tool results are not all present on events in a convenient form, so pair each function call with its function response by call id when building the record. Canonicalise arguments, with sorted keys and normalised numbers, so the replay lookup later matches calls that are equal in meaning.
Also snapshot the session's state before calling the runner and store it in the record. A session read after the turn already includes this turn's writes, and state set when the session was created never appears as an event, so neither can be reconstructed later. The snapshot includes the merged user: and app: values production saw, which is exactly what the fork needs.
Forking the session
The shadow worker builds a fresh session holding production's history up to, but not including, this turn, then runs the candidate on the same user message. Use a separate session service instance, never production's. Keys prefixed user: and app: are shared across sessions; a candidate writing them through a shared service would change production behaviour, which defeats the point of a shadow.
List<Event> runShadowTurn(TurnRecord r) {
Session prod = prodSessions.getSession("helpdesk", r.userId(), r.sessionId(), Optional.empty())
.blockingGet(); // read-only use of production
Map<String, Object> seed = r.stateAtStart(); // snapshot taken before the turn;
Session fork = shadowSessions.createSession("helpdesk-shadow", r.userId(), seed, null)
.blockingGet(); // deltas re-applied on top are idempotent
for (Event e : eventsBefore(prod, r.invocationId())) {
shadowSessions.appendEvent(fork, e).blockingGet(); // same history the prod agent saw
}
QuarantinePlugin q = new QuarantinePlugin(r.toolCalls(), toolClasses, readBudget);
Runner shadow = Runner.builder()
.agent(candidateFactory.build()).appName("helpdesk-shadow")
.artifactService(new InMemoryArtifactService())
.sessionService(shadowSessions)
.plugins(q)
.build();
List<Event> out = shadow.runAsync(r.userId(), fork.id(), r.userMessage())
.toList().timeout(60, TimeUnit.SECONDS).blockingGet();
shadowSessions.deleteSession("helpdesk-shadow", r.userId(), fork.id()).blockingAwait();
return out;
}Keep the candidate's agent names the same as production's where you can. History events carry authors, and a candidate whose router expects different sub-agent names may misread who said what. If names must change, rewrite authors while copying. The session context article explains which parts of history reach the model request.
Quarantining tools
Every tool call the candidate makes goes through a quarantine plugin. It has three rules, applied in order.
- Recorded read: if production made a read call with the same name and canonical arguments in this turn, return production's recorded result. This keeps the candidate's view of the world identical to production's and costs nothing.
- Unrecorded read: if the tool is on a read allowlist and the per-turn read budget allows, call it live; otherwise return
{"status":"unavailable_in_shadow"}and mark the turn as diverged. - Write: never execute. Record the arguments and return a synthetic success shaped like the real response.
@Override
public Maybe<Map<String, Object>> beforeToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx) {
String key = tool.name() + "|" + Canonical.json(args);
if (classes.isWrite(tool.name())) {
shadowWrites.add(new ToolCall(tool.name(), args));
return Maybe.just(syntheticSuccess.forTool(tool.name(), args)); // never reaches the API
}
Map<String, Object> recorded = recordedReads.get(key);
if (recorded != null) return Maybe.just(recorded);
diverged = true;
return readBudget.tryAcquire() && allowlist.contains(tool.name())
? Maybe.empty() // live, read-only call
: Maybe.just(Map.of("status", "unavailable_in_shadow"));
}The write rule is a deliberate choice. Returning unavailable for writes would make the candidate apologise or retry, so its trajectory would differ from production's for reasons that have nothing to do with its quality. Synthetic success keeps the trajectory comparable, at the cost of making everything after the write slightly fictional. So score write tools on their arguments only, and treat text after a synthetic write as lower-confidence. Default-deny: a tool missing from the class registry is treated as a write.
Scoring divergence
| Signal | How it is computed | What a difference means |
|---|---|---|
| Tool sequence | Edit distance between production and candidate tool name lists | Different plan; inspect |
| Write arguments | Exact match after canonicalisation, per write tool | Different action in the world; most important |
| Divergence flag | Any unrecorded read | Candidate explored; later steps less comparable |
| Final answer | Pairwise judge: better, same or worse | Quality change; needs judge calibration |
| Cost and latency | Tokens and wall time per turn | Budget impact at full traffic |
| Errors | Error events, exhausted retries | Robustness regression |
Use a pairwise judge that sees both answers in random order, as described in the LLM judge article, and calibrate it on a few hundred human-labelled pairs before trusting its verdicts. Report the share of turns where the write arguments differ separately from everything else, and read those turns by hand first.
Aggregate per intent or tool, not only overall. A candidate that is better on average can still be consistently worse on one kind of request, such as group changes, and that pattern disappears in a single headline number. Keep every scored turn joined to its turn record so a reviewer can open production's reply, the candidate's reply and both tool sequences side by side; most of the value of a shadow run comes from those side-by-side reads.
Worked example: a model migration
Illustrative example: a helpdesk team moves to a new model and shadows 5% of sessions for a week, about 9,000 turns. In 8,100 turns the candidate calls the same tools in the same order. 610 turns diverge on reads only, mostly an extra directory lookup before answering. 290 turns are ones in which either version wrote; in 262 the arguments match production's exactly, and of the 28 that differ, review finds 19 better (a narrower group), 6 equivalent and 3 worse, all three resetting a password when the user asked for an unlock. The judge prefers the candidate on 41% of turns, production on 33% and calls the rest even, and candidate tokens per turn are 18% higher.
The decision follows from the write diffs, not the judge score: the unlock confusion is fixed in the tool descriptions, the shadow is rerun on the same week's records, and the three cases disappear. The candidate then goes to a canary, where user outcomes can finally be measured.
Operating it: sampling, cost, privacy
- Sampling. Sample by session, not turn, so the turns you shadow have comparable histories, and start at 1 to 5%.
- Budget. The candidate's model calls cost real money; cap tokens per day and stop the worker when the cap is reached.
- Privacy. Turn records and shadow sessions are copies of user data. Apply the same retention and redaction as production, using the approach in the PII redaction article, and delete forks after scoring.
- Isolation. Run the worker in a separate deployment with credentials that cannot call write APIs at all, so the quarantine plugin is not the only barrier.
- Lag. Records that wait too long see a newer production state; drop records older than a few minutes.
Failure modes
- Shared session service. Shadow writes to
user:orapp:state change production. Use a separate store. - Write tools not classified. A new tool defaults to execute. Default-deny in the plugin and fail the build on unclassified tools.
- Comparing whole conversations. Forked turns never see the candidate's own earlier replies; do not report multi-turn success from shadow data.
- Trusting the judge alone. An uncalibrated judge rewards length. Read write diffs by hand.
- Blocking the reply. Capturing synchronously adds latency. Offer to a bounded queue and drop on overflow.
Trade-offs
Shadowing roughly doubles model spend for the sampled share of traffic, adds a worker deployment and a second session store, and needs people to read write diffs. In exchange it finds behaviour changes on inputs nobody wrote down, before any user is exposed, which offline datasets cannot do and canaries do only by exposing users.
The quarantine design trades realism for safety. Recorded reads keep the candidate honest only while it asks the same questions production asked; once it explores, live reads under a budget give it real data but load downstream systems, and the unavailable marker keeps load at zero but bends the trajectory. Synthetic writes keep trajectories comparable but make everything after them partly fictional. There is no setting that removes all three distortions, so report divergence and post-write turns separately and weigh them less.
Shadowing is also a poor fit for some agents. If most turns are writes with immediate, visible consequences, such as a booking agent whose next step depends on the confirmation number, the fiction compounds quickly and a small, carefully watched canary on internal users teaches more. If most turns are reads and answers, shadowing is cheap and very informative. Decide per agent, and per release: a prompt wording change may need only offline evaluation, while a model migration usually justifies a week of shadow traffic.
What to do next
- Classify every tool as read or write in one registry, defaulting to write.
- Build turn records in the serving path and offer them to a bounded queue with a drop counter.
- Stand up a separate session service and a worker without write credentials.
- Implement the quarantine plugin with recorded reads, a read budget and synthetic writes.
- Score tool sequence, write arguments, cost and errors first; add a calibrated pairwise judge later.
- Shadow 1 to 5% of sessions for a week before the next model or instruction change, review every write diff, then canary.