A shadow deployment sends a copy of real production traffic to a candidate model, records what it would have said, and throws the answer away before any user sees it. For classic ML models this is almost boring: score the same feature vector twice, compare two numbers. For large language models it is not. The candidate burns real GPU time on every mirrored prompt, its outputs are free text that will never match byte for byte, and a modern LLM endpoint calls tools that can send email or move money.

This page is about doing shadow deployment for LLM serving specifically: where to mirror, how to keep the shadow from causing side effects, how to compare outputs that are supposed to differ, why latency numbers from a shadow pool lie unless you control for cache state, and how much GPU capacity the whole exercise costs. The general pattern for any model is covered in Shadow Deployment for ML models, and the next step after a shadow passes, exposing real users to a slice of traffic, is covered in LLM Canary Releases.

What a shadow can and cannot tell you

Start from what a shadow can and cannot tell you. It sees the real prompt distribution: long system prompts, odd languages, pasted logs, users who type one word, agents that loop. Offline eval sets rarely capture that tail, and the tail is where a new model, a new quantisation or a new serving engine usually breaks. A shadow also exercises the real request shapes against the candidate's serving stack, so you find out that the new tokenizer chokes on a control character, or that the new engine rejects a parameter your clients send, before anyone depends on it.

What a shadow cannot tell you is how users react, because nobody reads the shadow answer. In a multi-turn chat, from turn two onward the conversation history contains the primary model's replies, not the shadow's. You are measuring how model B continues a conversation that model A started, which is a fair test of single-turn quality and a poor test of how B would steer a whole session. Treat a shadow as a gate for correctness, safety, format and cost, and leave preference and engagement questions to a canary or an A/B test. Also decide what kind of change you are shadowing: a new model changes behaviour on purpose, while a new quantisation or engine should change almost nothing, so their thresholds differ by an order of magnitude.

Where the mirror lives

The mirror belongs in the gateway or router that already terminates client requests, not in the client and not inside the inference engine. The gateway can sample, strip or hash identifiers, attach a pair ID, and fire the copy asynchronously so the primary path never waits on it. The shadow pool must be separate capacity: separate replicas, ideally separate GPUs. If you put the candidate on the same replicas as production, its prefill work competes for the same compute and KV cache memory, and the experiment itself degrades the latency you are trying to protect.

Shadow path: the user only ever sees the primary; the shadow's output goes to a comparison storeClientchat / APIGatewaysample + mirrorrequestPrimary pool (model A)real tools, real side effectsShadow pool (model B)separate GPUs, low priority1. forward2. async copyresponseTool stubs / recorded replayno writes, no emails, no chargestool callsPair storeA and B outputs, redactedB outputA outputComparatorsrules, embeddings, judgePromotion reportdiff rates, latency, cost
The gateway forwards to the primary synchronously and copies a sample to the shadow pool asynchronously. Shadow tool calls hit stubs, and both outputs land in a pair store for comparison.

The shadow path is fire-and-forget: never awaited, never retried, with its own timeout, and the first thing shed under load. Drop rather than queue, because queued requests measure your queue, not the model.

Envoy's route-level request mirroring can do the copying. Mirrored requests go to a separate cluster, responses are ignored, and Envoy appends -shadow to the Host header:

routes:
- match: { prefix: "/v1/chat/completions" }
  route:
    cluster: llm-primary
    timeout: 120s
    request_mirror_policies:
    - cluster: llm-shadow
      runtime_fraction:
        default_value: { numerator: 5, denominator: HUNDRED }
        runtime_key: shadow.chat.fraction

Envoy's mirror does not store the primary's response next to the shadow's, so most teams add a small mirror service behind it that applies the rules below and writes the pair.

Keeping the shadow inert

The shadow must be inert. If the request includes tool definitions and the shadow model calls send_refund, something has to stop that call. Removing the tools changes the task, so the shadow needs the tools and needs them to be harmless.

There are three workable patterns. Stub everything: every tool on the shadow side returns a fixed, plausible response and logs the arguments. Cheap, and enough to compare the first tool decision. Record and replay: the primary's real tool results are recorded with the pair ID; when the shadow calls the same tool with equivalent arguments, it gets the recorded result, and any call the primary did not make gets a stub. This lets multi-step agent runs proceed realistically as long as the two models agree. Read-only passthrough: tools classified as pure reads, such as search or a catalogue lookup, call the real backend with a shadow credential; anything else is stubbed. That classification has to be an explicit allowlist, never a guess from the tool name.

READ_ONLY = {"search_docs", "get_order_status", "lookup_product"}   # explicit allowlist

class ShadowToolRouter:
    def __init__(self, recorded, real_reads, stub):
        self.recorded = recorded      # {(tool, canonical_args): result} from the primary run
        self.real_reads = real_reads  # callables using a read-only, shadow-tagged credential
        self.stub = stub

    def call(self, pair_id, tool, args):
        key = (tool, canonical(args))
        log_tool_call(pair_id, side="shadow", tool=tool, args=args)
        if key in self.recorded:
            return self.recorded[key]           # replay: same call, same world
        if tool in READ_ONLY:
            return self.real_reads[tool](**args)
        return self.stub(tool, args)            # never executes a write

The same applies downstream: no writes to the conversation store, billing, rate limits, memory or webhooks. Use a header such as x-shadow: 1 that every service treats as read-only, plus a test asserting that a shadow request moves no write-side metric.

A minimal mirror service

A minimal mirror service has four jobs: decide whether to sample, prepare the shadow request, run it with a budget, and store the pair. Sampling should be deterministic on a hash of the conversation or user ID, so a sampled conversation stays sampled across turns and you can study whole sessions. Prepare means applying the same template the candidate would get in production, swapping the model name, and pinning decoding parameters you want to compare.

import asyncio, hashlib, time

SAMPLE_PERCENT = 5
SHADOW_TIMEOUT_S = 60
inflight = asyncio.Semaphore(32)          # hard cap on concurrent shadow requests

def sampled(conversation_id: str) -> bool:
    h = int(hashlib.sha256(conversation_id.encode()).hexdigest()[:8], 16)
    return h % 100 < SAMPLE_PERCENT

async def mirror(pair_id, request, primary_future):
    if not sampled(request["conversation_id"]) or inflight.locked():
        metrics.inc("shadow.skipped")
        return
    async with inflight:
        shadow_req = dict(request, model="candidate-b", stream=False)
        t0 = time.monotonic()
        try:
            shadow = await asyncio.wait_for(shadow_client.chat(shadow_req), SHADOW_TIMEOUT_S)
            status = "ok"
        except Exception as e:                   # never propagates to the user path
            shadow, status = None, type(e).__name__
        latency = time.monotonic() - t0
        primary = await primary_future           # already resolved for the user
        await pair_store.write(pair_id, redact(request), redact(primary), redact(shadow),
                               status=status, shadow_latency=latency)

The shadow request is sent non-streaming, which hides time to first token. If TTFT is a promotion criterion, stream inside the mirror service, record the first chunk's timestamp and discard the rest.

Comparing outputs that are supposed to differ

Exact-match comparison fails for reasons that have nothing to do with quality. With sampling enabled, two calls to the same model differ. With greedy decoding and a fixed seed, outputs can still differ, because the floating point reductions inside attention and matrix kernels depend on batch size and batch composition, and a continuously batched engine never runs your prompt in the same batch twice. A tiny difference in one logit flips a near-tie, and every token after that diverges. So even an engine-only change, same weights and same settings, needs a comparison that tolerates divergence.

Use a ladder of comparators, cheapest first, and only escalate the pairs that need it:

  1. Structural checks. Valid JSON, same tool called, refusal, hit the token limit. Cheap, exact, and they catch most quantisation and template regressions.
  2. Field-level equality. For extraction and classification, compare parsed fields.
  3. Semantic similarity. Embed both answers and flag pairs below a threshold calibrated by running the primary against itself, which shows what normal divergence looks like.
  4. Judge review. Send flagged pairs to an LLM judge in randomised order, and a sample to humans. Judges have position and length biases, which LLM-as-judge calibration covers in detail.
def compare(pair, self_agreement_p05):
    a, b = pair.primary, pair.shadow
    out = {
        "a_json_ok": is_valid_json(a.text) if pair.expects_json else None,
        "b_json_ok": is_valid_json(b.text) if pair.expects_json else None,
        "tool_match": first_tool(a) == first_tool(b),
        "a_refused": is_refusal(a.text), "b_refused": is_refusal(b.text),
        "b_truncated": b.finish_reason == "length",
        "len_ratio": len(b.text) / max(1, len(a.text)),
    }
    out["sim"] = cosine(embed(a.text), embed(b.text))
    out["needs_judge"] = (out["sim"] < self_agreement_p05) or not out["tool_match"] \
        or out["a_refused"] != out["b_refused"]
    return out

Report rates per slice, by endpoint, tenant, language and prompt length, not one average score.

Latency and capacity on GPUs

Latency from a shadow pool is easy to measure and easy to misread. Three confounders matter on GPUs. Cache warmth: the primary pool serves every turn of every conversation, so its prefix cache usually holds the shared system prompt and the earlier turns. The shadow sees only a sample, so its hit rate is lower and its prefill slower, which makes the candidate look worse than it is. Prefix caching explains why a hit can cut prefill time by most of the prompt. Conversation-level sampling helps, because a sampled conversation brings all its turns.

Batch occupancy: at 5 percent of traffic, the shadow engine runs small batches. Decode steps are memory-bandwidth bound, so a small batch gives each request a fast inter-token time that will not survive full load. The primary at full load is batching dozens of sequences per step. Comparing inter-token latency across these two regimes compares batch sizes, not models. Continuous batching in production covers how batch size trades against per-token latency.

Queueing: an undersized shadow pool adds wait time to the measurement, so record queue time separately. For a real capacity number, run a full-rate replay load test of recorded traffic on the target hardware.

Worked example: shadowing an FP8 build

Suppose a support assistant runs an 8B model in BF16 and the team wants to ship an FP8 quantised build of the same weights on a newer engine version. This is an engine-and-quantisation change, so the bar is near-equivalence. Production peaks at 40 requests per second with a mean prompt of 1,800 tokens and a mean output of 250 tokens; about a third of requests carry tool definitions.

At 5 percent conversation sampling, the shadow sees roughly 2 requests per second, or about 3,600 prompt tokens and 500 output tokens per second. A load test shows one candidate replica sustains about 12 requests per second at this mix within the latency target, so one replica is plenty, and the team adds a second only so a crash does not stop data collection. The semaphore caps concurrent shadow calls at 32. Over a week they collect about a million pairs.

Results: JSON validity is 99.6 percent on both sides. Tool choice agrees on 98.9 percent of tool requests; the self-agreement run of BF16 against itself agrees on 99.1 percent, so the gap is small but real. Slicing by prompt length shows it: above 6,000 tokens, tool agreement drops to 95 percent. The judge confirms that on long prompts the FP8 build more often misses a constraint stated early in the system prompt. The team keeps the new engine, tries a different quantisation recipe for the attention projections, reruns the shadow for three days, and gets long-prompt agreement back within the self-agreement band. Only then does the build go to a 1 percent canary. Without the length slice, the headline 98.9 percent would have passed.

Failure modes

The failures below recur across teams; each has a cheap guard.

  • The shadow causes a side effect. A tool that was assumed read-only writes an audit row or sends a notification. Guard with the explicit allowlist, the shadow header and an automated test that fails if any write-side metric moves.
  • The mirror slows the primary. Shared connection pools or GPUs. Compare primary p99 with mirroring on and off.
  • Pairs leak personal data. The pair store holds prompts and two answers. Redact before writing, keep retention short and restrict access.
  • Uncalibrated thresholds. A fixed similarity cut-off without an A versus A baseline floods reviewers with false disagreements.
  • Shadow drops are biased. Load shedding drops the longest requests first, because those time out, so the comparison silently excludes the hardest prompts. Track the length distribution of completed pairs against the sampled requests.

Shadow versus replay, canary and A/B

ApproachSees real trafficUser exposureGPU costBest for
Offline eval setNoNoneLow, one-offCapability and known regressions
Recorded replayPast trafficNoneBursty, controllableThroughput and capacity tests
ShadowLive, sampledNoneContinuous, sized by sample rateFormat, safety, tool and slice regressions
CanaryLive, slicedSmall sliceShares serving fleetUser-facing metrics and real errors
A/B testLive, splitLarge sliceFullPreference and business outcomes

A shadow is the last point where a broken candidate harms no one while still seeing today's prompts. For an engine or quantisation change, shadow plus replay is often enough before a canary; for a new model family, plan on all stages.

What to do next

  • Decide whether your change is a model change, an engine or quantisation change, or a prompt change, and write the promotion thresholds for that type before collecting data.
  • Add a mirror at the gateway with conversation-level sampling, a concurrency cap and a separate timeout; verify primary p99 is unchanged with mirroring on.
  • Inventory every tool and downstream write; build stubs, a read-only allowlist and a shadow header, and test that a shadow request moves no write-side metric.
  • Run the current model against itself to get self-agreement baselines for each comparator.
  • Store redacted pairs with a short retention and report disagreement rates per endpoint, language and prompt-length bucket.
  • Measure capacity separately with a full-rate replay test on the candidate's target hardware.
  • Set an end date and GPU budget for the shadow, then promote to a canary or shut it down.
Key takeaway: An LLM shadow is worth its GPU bill when it is inert, sampled per conversation, compared against a self-agreement baseline and sliced by prompt type. Mirror at the gateway without awaiting, stub or replay every tool, distrust shadow latency until you control for cache warmth and batch size, and give every shadow an end date that leads to a canary or a shutdown.