A shadow deployment runs a candidate version of an agent on real production requests without letting it affect anyone. Users get the production reply; the candidate's reply, tool calls and costs are recorded and compared. For an ordinary service, shadowing is mostly a traffic-mirroring problem. For an agent it is harder in three specific ways: the candidate's answer depends on the whole conversation so far, its tools have side effects that must not happen twice, and its behaviour diverges as soon as it calls a tool production did not call.

This article builds shadowing for ADK Java: capturing each production turn, forking the session into a separate store, quarantining tools with recorded reads and stubbed writes, scoring divergence, controlling cost and privacy, and reading the results. It assumes you already run offline regression checks such as those in the regression detection article; shadowing adds coverage of the traffic you did not think to put in a dataset.

What shadowing can and cannot tell you

Shadowing answers one question well: on today's real requests, where does the candidate behave differently from production, and are those differences better or worse? It is the right tool for a new model, a rewritten instruction, a changed tool set or an ADK upgrade, where you cannot enumerate the inputs that matter.

It cannot measure what users do with an answer, because they never see it, so satisfaction and task completion still need a canary. It also cannot observe true multi-turn behaviour: each shadow turn is forked from production's history, so the candidate never sees the consequences of its own previous replies. Shadow evaluates single decisions in real context; canaries evaluate whole conversations.

Architecture

Shadow deployment: the candidate replays real turns but never touches the worlduserturn nproduction runnerserves the replyprod session storereal toolsreads and writesturn recordshadow queuesampled, budgetedshadow workerfork, run candidateshadow session storeseparate instancequarantine pluginreplay reads, stub writesscorerdiff, judge, costNothing on the dashed path can reach the user, the production store or a write API.
Production serves the user; a sampled turn record drives a forked candidate run whose tools are quarantined.

The production path is unchanged apart from emitting a turn record. Everything the candidate touches is either a copy (the forked session), a recording (reads production already made) or a stub (writes), and the scorer compares its output to production's record.

Capturing the production turn

Capture in the serving path, after the turn completes, and never block the reply. Collect the events the runner emits, build a turn record and offer it to a bounded queue; if the queue is full, drop the record and count the drop. The record holds what the shadow needs to fork and what the scorer needs to compare.

record TurnRecord(String userId, String sessionId, String invocationId, Content userMessage,
                  Map<String, Object> stateAtStart,  // session state read before runAsync
                  List<ToolCall> toolCalls,         // name, canonical args, result, read or write
                  String finalText, long tokens, long latencyMs) {}

Flowable<Event> serve(String userId, String sessionId, Content msg) {
  List<Event> seen = new CopyOnWriteArrayList<>();
  long start = System.nanoTime();
  return runner.runAsync(userId, sessionId, msg)
      .doOnNext(seen::add)
      .doOnComplete(() -> {
        if (!sampler.take(userId, sessionId)) return;          // per-session sampling, see below
        TurnRecord r = TurnRecords.from(userId, sessionId, msg, seen, toolClasses,
                                        (System.nanoTime() - start) / 1_000_000);
        if (!shadowQueue.offer(r)) metrics.increment("shadow.dropped");
      });
}

Tool results are not all present on events in a convenient form, so pair each function call with its function response by call id when building the record. Canonicalise arguments, with sorted keys and normalised numbers, so the replay lookup later matches calls that are equal in meaning.

Also snapshot the session's state before calling the runner and store it in the record. A session read after the turn already includes this turn's writes, and state set when the session was created never appears as an event, so neither can be reconstructed later. The snapshot includes the merged user: and app: values production saw, which is exactly what the fork needs.

Forking the session

The shadow worker builds a fresh session holding production's history up to, but not including, this turn, then runs the candidate on the same user message. Use a separate session service instance, never production's. Keys prefixed user: and app: are shared across sessions; a candidate writing them through a shared service would change production behaviour, which defeats the point of a shadow.

List<Event> runShadowTurn(TurnRecord r) {
  Session prod = prodSessions.getSession("helpdesk", r.userId(), r.sessionId(), Optional.empty())
      .blockingGet();                                        // read-only use of production
  Map<String, Object> seed = r.stateAtStart();              // snapshot taken before the turn;
  Session fork = shadowSessions.createSession("helpdesk-shadow", r.userId(), seed, null)
      .blockingGet();                                        // deltas re-applied on top are idempotent
  for (Event e : eventsBefore(prod, r.invocationId())) {
    shadowSessions.appendEvent(fork, e).blockingGet();       // same history the prod agent saw
  }
  QuarantinePlugin q = new QuarantinePlugin(r.toolCalls(), toolClasses, readBudget);
  Runner shadow = Runner.builder()
      .agent(candidateFactory.build()).appName("helpdesk-shadow")
      .artifactService(new InMemoryArtifactService())
      .sessionService(shadowSessions)
      .plugins(q)
      .build();
  List<Event> out = shadow.runAsync(r.userId(), fork.id(), r.userMessage())
      .toList().timeout(60, TimeUnit.SECONDS).blockingGet();
  shadowSessions.deleteSession("helpdesk-shadow", r.userId(), fork.id()).blockingAwait();
  return out;
}

Keep the candidate's agent names the same as production's where you can. History events carry authors, and a candidate whose router expects different sub-agent names may misread who said what. If names must change, rewrite authors while copying. The session context article explains which parts of history reach the model request.

Quarantining tools

Every tool call the candidate makes goes through a quarantine plugin. It has three rules, applied in order.

  1. Recorded read: if production made a read call with the same name and canonical arguments in this turn, return production's recorded result. This keeps the candidate's view of the world identical to production's and costs nothing.
  2. Unrecorded read: if the tool is on a read allowlist and the per-turn read budget allows, call it live; otherwise return {"status":"unavailable_in_shadow"} and mark the turn as diverged.
  3. Write: never execute. Record the arguments and return a synthetic success shaped like the real response.
@Override
public Maybe<Map<String, Object>> beforeToolCallback(
    BaseTool tool, Map<String, Object> args, ToolContext ctx) {
  String key = tool.name() + "|" + Canonical.json(args);
  if (classes.isWrite(tool.name())) {
    shadowWrites.add(new ToolCall(tool.name(), args));
    return Maybe.just(syntheticSuccess.forTool(tool.name(), args));  // never reaches the API
  }
  Map<String, Object> recorded = recordedReads.get(key);
  if (recorded != null) return Maybe.just(recorded);
  diverged = true;
  return readBudget.tryAcquire() && allowlist.contains(tool.name())
      ? Maybe.empty()                                                    // live, read-only call
      : Maybe.just(Map.of("status", "unavailable_in_shadow"));
}

The write rule is a deliberate choice. Returning unavailable for writes would make the candidate apologise or retry, so its trajectory would differ from production's for reasons that have nothing to do with its quality. Synthetic success keeps the trajectory comparable, at the cost of making everything after the write slightly fictional. So score write tools on their arguments only, and treat text after a synthetic write as lower-confidence. Default-deny: a tool missing from the class registry is treated as a write.

Scoring divergence

SignalHow it is computedWhat a difference means
Tool sequenceEdit distance between production and candidate tool name listsDifferent plan; inspect
Write argumentsExact match after canonicalisation, per write toolDifferent action in the world; most important
Divergence flagAny unrecorded readCandidate explored; later steps less comparable
Final answerPairwise judge: better, same or worseQuality change; needs judge calibration
Cost and latencyTokens and wall time per turnBudget impact at full traffic
ErrorsError events, exhausted retriesRobustness regression

Use a pairwise judge that sees both answers in random order, as described in the LLM judge article, and calibrate it on a few hundred human-labelled pairs before trusting its verdicts. Report the share of turns where the write arguments differ separately from everything else, and read those turns by hand first.

Aggregate per intent or tool, not only overall. A candidate that is better on average can still be consistently worse on one kind of request, such as group changes, and that pattern disappears in a single headline number. Keep every scored turn joined to its turn record so a reviewer can open production's reply, the candidate's reply and both tool sequences side by side; most of the value of a shadow run comes from those side-by-side reads.

Worked example: a model migration

Illustrative example: a helpdesk team moves to a new model and shadows 5% of sessions for a week, about 9,000 turns. In 8,100 turns the candidate calls the same tools in the same order. 610 turns diverge on reads only, mostly an extra directory lookup before answering. 290 turns are ones in which either version wrote; in 262 the arguments match production's exactly, and of the 28 that differ, review finds 19 better (a narrower group), 6 equivalent and 3 worse, all three resetting a password when the user asked for an unlock. The judge prefers the candidate on 41% of turns, production on 33% and calls the rest even, and candidate tokens per turn are 18% higher.

The decision follows from the write diffs, not the judge score: the unlock confusion is fixed in the tool descriptions, the shadow is rerun on the same week's records, and the three cases disappear. The candidate then goes to a canary, where user outcomes can finally be measured.

Operating it: sampling, cost, privacy

  • Sampling. Sample by session, not turn, so the turns you shadow have comparable histories, and start at 1 to 5%.
  • Budget. The candidate's model calls cost real money; cap tokens per day and stop the worker when the cap is reached.
  • Privacy. Turn records and shadow sessions are copies of user data. Apply the same retention and redaction as production, using the approach in the PII redaction article, and delete forks after scoring.
  • Isolation. Run the worker in a separate deployment with credentials that cannot call write APIs at all, so the quarantine plugin is not the only barrier.
  • Lag. Records that wait too long see a newer production state; drop records older than a few minutes.

Failure modes

  • Shared session service. Shadow writes to user: or app: state change production. Use a separate store.
  • Write tools not classified. A new tool defaults to execute. Default-deny in the plugin and fail the build on unclassified tools.
  • Comparing whole conversations. Forked turns never see the candidate's own earlier replies; do not report multi-turn success from shadow data.
  • Trusting the judge alone. An uncalibrated judge rewards length. Read write diffs by hand.
  • Blocking the reply. Capturing synchronously adds latency. Offer to a bounded queue and drop on overflow.

Trade-offs

Shadowing roughly doubles model spend for the sampled share of traffic, adds a worker deployment and a second session store, and needs people to read write diffs. In exchange it finds behaviour changes on inputs nobody wrote down, before any user is exposed, which offline datasets cannot do and canaries do only by exposing users.

The quarantine design trades realism for safety. Recorded reads keep the candidate honest only while it asks the same questions production asked; once it explores, live reads under a budget give it real data but load downstream systems, and the unavailable marker keeps load at zero but bends the trajectory. Synthetic writes keep trajectories comparable but make everything after them partly fictional. There is no setting that removes all three distortions, so report divergence and post-write turns separately and weigh them less.

Shadowing is also a poor fit for some agents. If most turns are writes with immediate, visible consequences, such as a booking agent whose next step depends on the confirmation number, the fiction compounds quickly and a small, carefully watched canary on internal users teaches more. If most turns are reads and answers, shadowing is cheap and very informative. Decide per agent, and per release: a prompt wording change may need only offline evaluation, while a model migration usually justifies a week of shadow traffic.

What to do next

  1. Classify every tool as read or write in one registry, defaulting to write.
  2. Build turn records in the serving path and offer them to a bounded queue with a drop counter.
  3. Stand up a separate session service and a worker without write credentials.
  4. Implement the quarantine plugin with recorded reads, a read budget and synthetic writes.
  5. Score tool sequence, write arguments, cost and errors first; add a calibrated pairwise judge later.
  6. Shadow 1 to 5% of sessions for a week before the next model or instruction change, review every write diff, then canary.
Key takeaway: Shadowing an agent means replaying real turns through a candidate that sees production's history and production's tool results but can change nothing. Fork each turn into a separate session store, replay recorded reads, stub writes with synthetic success and score writes on their arguments, then read the write diffs before trusting any judge. Shadow tells you where the candidate decides differently; a canary still has to tell you whether users are better off.