You want to change one component of a production agent: a new model version, a cheaper model, a rewritten retrieval tool. Offline evaluation tells you how it does on the cases you thought of. A canary tells you how it does on real users, by letting some of them receive its answers. Shadow testing sits between the two: the candidate runs on real production inputs, its outputs are recorded and compared, and no user ever sees them.
This article builds component-level shadow testing inside one ADK Java service: a BaseLlm wrapper that sends the identical LlmRequest to a candidate model, a tool shadow for read-only tools, a comparison that classifies disagreements, and the statistics for deciding whether the candidate is ready. Whole-agent shadowing, which forks sessions and runs a complete second agent, is covered in a separate article; the component version is smaller, cheaper and answers a narrower question very precisely.
What component shadowing can tell you
Component shadowing asks: given exactly what production gave the current component, does the candidate make the same decision, and where it differs, which one was right? Because both receive the same request object, every difference is caused by the component. There is no history drift, no difference in retrieved data, no second session.
The narrowness is the price. A shadow model call returns a decision, such as calling lookup_order with certain arguments, or a text answer. The shadow never executes its tool calls, so you cannot observe what would have happened three steps later. If the change alters multi-step behaviour (planning, recovery from tool errors), use whole-agent shadowing or a canary. If it alters single decisions (model swap for routing, a prompt tweak, a faster model for extraction), component shadowing is the sharpest instrument you have.
| Method | Inputs | User impact | Observes |
|---|---|---|---|
| Offline eval | curated cases | none | full trajectory on known cases |
| Component shadow | live requests, identical | none | single decisions, paired |
| Whole-agent shadow | live turns, forked session | none | multi-step behaviour, drifting |
| Canary | live users, split | real | outcomes and user reaction |
Shadowing a model with a BaseLlm wrapper
ADK Java agents accept a model object: LlmAgent.Builder.model(BaseLlm). BaseLlm has one abstract method, generateContent(LlmRequest, boolean stream), returning a Flowable<LlmResponse>. Wrapping it gives you the identical request at the one place every model call passes through. LlmRequest is immutable, so both calls can share its contents and config safely.
One field must differ. ADK's flow stamps the agent's model name onto the request, and the bytecode of ADK Java 1.11.0's Gemini class shows it calling request.model().orElse(model()): the name in the request wins over the model object's own. Pass the production request unchanged to a Gemini candidate and it calls the production model, and the shadow reports perfect agreement. So the wrapper derives a copy with toBuilder().model(candidate.model()) that differs only in the model name.
public final class ShadowLlm extends BaseLlm {
private final BaseLlm primary, candidate;
private final Sampler sampler;
private final PairSink sink;
private final Semaphore permits = new Semaphore(16); // max concurrent shadow calls
public ShadowLlm(BaseLlm primary, BaseLlm candidate, Sampler sampler, PairSink sink) {
super(primary.model());
this.primary = primary; this.candidate = candidate;
this.sampler = sampler; this.sink = sink;
}
@Override
public Flowable<LlmResponse> generateContent(LlmRequest req, boolean stream) {
if (!sampler.take(req) || !permits.tryAcquire()) {
return primary.generateContent(req, stream); // the common path: untouched
}
LlmRequest shadowReq = req.toBuilder().model(candidate.model()).build(); // see below
Single<LlmResponse> shadow = candidate.generateContent(shadowReq, false)
.lastOrError()
.timeout(30, TimeUnit.SECONDS)
.subscribeOn(Schedulers.io())
.cache();
shadow.subscribe(r -> {}, e -> {}); // start now, in parallel
List<LlmResponse> seen = new CopyOnWriteArrayList<>();
return primary.generateContent(req, stream)
.doOnNext(seen::add)
.doFinally(() -> shadow.subscribe(
cand -> { permits.release(); sink.offer(req, Responses.merge(seen), cand); },
err -> { permits.release(); sink.shadowFailed(req, err); }));
}
}
LlmAgent agent = LlmAgent.builder()
.name("support_router")
.model(new ShadowLlm(LlmRegistry.getLlm(PROD_MODEL), LlmRegistry.getLlm(CANDIDATE_MODEL),
Sampler.sessions(0.05), pairStore))
.instruction(ROUTER_INSTRUCTIONS)
.tools(lookupOrder, issueRefund, escalate)
.build();Five details carry the safety. The method only ever returns the primary's Flowable; the candidate's response goes to the sink and nowhere else. The candidate runs on the IO scheduler with a timeout, so it neither blocks nor outlives the request for long. The semaphore caps concurrent shadow calls, and when it is exhausted the call simply is not shadowed: under load, the shadow sheds itself first. The shadow is non-streaming because only its final response is compared; Responses.merge concatenates the primary's streamed partial responses into one for comparison. And a shadow failure is recorded, not thrown: a candidate that errors is a finding, not an outage.
Sample by session, not by call, so that every model call within a sampled session is shadowed and you can read disagreements in context. Hash the session ID into [0, 1) and compare it with the rate; that keeps the decision stable without storing it.
Shadowing read-only tools
The same idea works for tools, with one hard rule: shadow only tools that read. A plugin's afterToolCallback(BaseTool, Map args, ToolContext, Map result) sees the production result and the arguments that produced it, which is everything a candidate implementation needs.
@Override
public Maybe<Map<String, Object>> afterToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx, Map<String, Object> result) {
ShadowTool cand = readOnlyCandidates.get(tool.name()); // allowlist, reads only
if (cand != null && sampler.take(ctx) && permits.tryAcquire()) {
Single.fromCallable(() -> cand.call(args))
.timeout(5, TimeUnit.SECONDS)
.subscribeOn(Schedulers.io())
.doFinally(permits::release)
.subscribe(c -> toolPairs.offer(tool.name(), args, result, c),
e -> toolPairs.shadowFailed(tool.name(), e));
}
return Maybe.empty(); // never alter the real result
}Keep the allowlist explicit and reviewed. A tool that "just reads" but writes an access log, warms a cache that changes production behaviour, or consumes a rate-limited quota is not side-effect free. Point the candidate at a replica where you can, and give it its own credentials so its traffic is identifiable.
Comparing decisions, not bytes
Agreement must be measured on the decision, not the bytes. Classify each pair by what the model chose to do:
| Primary | Candidate | Comparison | Class |
|---|---|---|---|
| function call | same function | canonical args equal? | agree / arg-diff |
| function call | different function | - | tool-diff (most important) |
| function call | text | - | skipped-tool |
| text | function call | - | extra-tool |
| text | text | rubric or judge on meaning | agree / text-diff |
Canonicalise arguments before comparing: sort keys, normalise numbers and dates, trim strings. Without that, {"qty": 2} and {"qty": 2.0} count as a disagreement and the noise drowns the signal. For text, exact match is useless; use a short rubric (same facts, same commitments, same refusal decision) scored by a judge model, and check the judge against a few dozen human-labelled pairs before trusting it.
Record latency and token usage for both sides on every pair. A candidate that agrees 99% of the time at twice the latency is a different decision from one that agrees 97% at half the cost.
From agreement to a decision
Agreement is not correctness: when the two disagree, either could be right. Two statistical tools turn shadow data into a decision.
Bounding a rare failure: the rule of three. If you observe zero occurrences of a failure class in n shadowed calls, the 95% upper confidence bound on its rate is about 3/n. To claim the candidate's skipped-tool rate is below 0.5%, you need about 600 shadowed calls with none. One occurrence changes the arithmetic, so plan for more.
Comparing correctness on paired data: McNemar's test. Label a sample of pairs with the right answer, from later outcomes (the user accepted the routing, the order lookup was correct) or from reviewers. Because both models answered the same request, only the discordant pairs carry information: b, where the primary was right and the candidate wrong, and c, the reverse. The statistic is:
chi2 = (|b - c| - 1)^2 / (b + c) # continuity-corrected, 1 degree of freedom
reject "no difference" at 5% if chi2 > 3.84Pairing is the reason shadow testing is so efficient. An unpaired comparison of two accuracy rates needs far more samples, because the variance of the easy cases both models get right is counted twice. Here those cases cancel out.
Oversampling disagreements for labelling is legitimate for this test, because concordant pairs do not enter the statistic. It is not legitimate for estimating either model's absolute accuracy: if you labelled every disagreement but only one agreement in ten, weight each labelled agreement by ten before computing accuracy, or the candidate looks far worse than it is. Keep the sampling weights with the labels so nobody recomputes a headline number from the raw table later.
Finally, fix the decision rule before you look. Write down the blocking classes, the bound each must meet, and the significance level, then collect. Choosing thresholds after seeing the data is how a candidate that someone wants to ship gets shipped.
Worked example: a cheaper router model
A team wants to replace the model behind a support router with a cheaper one. They shadow 5% of sessions for a week: 2,000 model calls.
Decision agreement is 96.1%: 78 calls disagree. Of those, 41 are tool-diff (mostly escalate versus lookup_order), 22 arg-diff (date formats the canonicaliser missed, fixed and re-scored to 9 real ones) and 15 skipped-tool. Latency at the median falls from 1.9 s to 0.8 s.
Reviewers label 600 pairs, oversampling disagreements. On those, b = 31 (primary right, candidate wrong) and c = 14. McNemar gives (|31 - 14| - 1)^2 / 45 = 256 / 45 = 5.69, above 3.84: the candidate is worse, and the difference is not noise. Reading the 31 shows a pattern: the candidate escalates whenever the user mentions a refund, even when the order lookup would answer it.
The team adds two sentences to the router instruction for the candidate only, shadows again for four days, and gets b = 12, c = 15, chi-squared 0.15: no detectable difference. Skipped-tool count is zero in 800 calls, bounding it below about 0.4%. Only now does a canary start, with the cost saving already measured.
Operating the shadow
- Cost. Every shadowed call is a second model call. At a 5% session rate the spend increase is about 5% of model cost times the candidate's relative price; set the rate from the sample size you need, not from habit, and stop when you have it.
- Quota. If both models share a project quota, shadow traffic can cause production rate-limit errors. Use a separate quota or project for the candidate.
- Data handling. Sending production prompts to a different provider or region is a data transfer. Clear it the same way you would for production.
- Observability. Export shadow calls attempted, skipped for lack of permits, failed and timed out. A shadow that silently skips 80% of calls under load produces a biased sample from quiet hours.
Failure modes
- Returning the wrong stream. One misplaced variable sends candidate answers to users. Unit-test that the wrapper's output equals the primary's for a candidate that returns different text.
- Shadowing production with itself. Forwarding the request without replacing its model name runs the primary model twice. Agreement near 100% on the first day is a symptom; assert the model name recorded on each shadow response.
- Shadowing a write. A tool on the allowlist that sends an email or charges a card does it twice. Keep writes out by construction, not by review.
- Latency leaking. Subscribing to the shadow on the request thread, or awaiting it before returning, adds its latency to production.
- Noisy comparison. Uncanonicalised arguments or a strict text match report disagreement that is not there, and teams learn to ignore the dashboard.
- Agreement as the goal. A candidate that agrees 99% but is wrong on the 1% that matter is worse. Label disagreements.
Trade-offs
Component shadowing is precise, cheap and invisible to users, and blind to multi-step consequences. Whole-agent shadowing sees trajectories but drifts after the first divergence and costs a full second agent. Canaries measure real outcomes and expose real users. Use them in that order for risky changes: shadow to find and fix disagreements, then canary to confirm outcomes.
Related reading: whole-agent shadow deployment with forked sessions, session-pinned canary releases, running evaluations in CI and unit testing ADK Java agents.
What to do next
- Pick one decision-shaped change, such as a model swap for routing or extraction, and write down the disagreement classes that would block it.
- Implement
ShadowLlmwith session sampling, a permit cap, a timeout and a sink; add a unit test proving the user always receives the primary's response. - Build the comparator with canonical arguments and a judged text rubric; validate the judge on human-labelled pairs.
- Compute the sample size from the failure rates you need to bound, using the rule of three, and set the sampling rate from that.
- Label a few hundred pairs, oversampling disagreements, and run McNemar's test before any canary.
- Add read-only tool shadowing only after the allowlist has been reviewed for hidden writes.