Blue-green deployment keeps two complete production environments. Blue serves all traffic; green is built, loaded and verified alongside it with no users; then the router moves all traffic to green in one step, and blue stays up for a while so that rollback is the same step in reverse. For web services, that is a well-worn pattern. For a fleet of GPU replicas serving a large language model, three things change: the second environment is expensive enough that its lifetime has to be budgeted, the new environment starts with cold caches that real traffic will expose, and the number of things that must match between blue and green is larger than most teams expect.

This article covers the blue-green specific parts: what parity means for a model server, how to load, warm and verify green, how to make the switch one atomic change, what happens to latency in the minutes after cutover, and how long to keep blue. Gradual traffic shifting is a different technique, covered in LLM canary deployment, and mirroring traffic without serving it is covered in LLM shadow deployment.

The architecture

Blue-green for an LLM fleet: two full environments, one routing decisionclientschat, API, agentsgateway / routerweights: blue 100, green 0BLUE (live)replica x6model v1prefix cachewarmdrains in-flight streams after the switchGREEN (candidate)replica x6model v2prefix cachecold until warmedloaded, warmed, verified before any user trafficlive trafficprobes, replayartifact storeweights, tokenizer, configparity checkhashes must match manifestswitch controllergates, flip, watch, rollback
Figure: two complete environments behind one router. Green gets synthetic probes and replayed prompts, never user traffic, until a single weight change moves everything.

The pieces are deliberately few. An artifact store holds an immutable release: weights, tokenizer files, the chat template, the generation config and the serving engine image. A switch controller, which can be a pipeline job or a small service, reads a release manifest, brings up green from exactly those artifacts, runs the gates, changes the router weights and then watches. The router is the only component that knows which colour is live. Clients keep one stable endpoint and never learn that the backend changed.

Blue-green fits when the change is all-or-nothing: a new model version where you do not want two answers to the same question coexisting, an engine upgrade with a different KV-cache layout, or a change of tensor-parallel layout that cannot share replicas. It fits less well when you need statistical evidence from real traffic before committing, which is what a canary is for.

Parity: everything except the change must match

The promise of blue-green is that green differs from blue only in what you meant to change. With a model server, unintended differences hide in files nobody reviews. A model that is identical in its weights can produce different outputs because the chat template changed a newline, because the default temperature moved from a config file into an environment variable, or because a stop sequence went missing. Treat the following as a manifest with a hash for every entry, and refuse to switch if anything other than the declared change differs.

ItemWhy it driftsHow to check
WeightsWrong revision pulled, partial downloadHash every shard against the manifest
Tokenizer filesRe-exported tokenizer with different special tokensHash tokenizer files; round-trip a fixed text
Chat templateTemplate stored in tokenizer config edited separatelyRender a fixed conversation and diff the string
Generation defaultsTemperature, top-p, max tokens, stop sequences set in different placesDump effective defaults from the server and diff
Engine and kernelsImage rebuilt with a new engine or CUDA versionPin image digest; record engine version string
Parallel layoutTensor-parallel degree or quantization flag changedRecord launch arguments in the manifest
Context limitMaximum model length lowered to fit memoryProbe with a prompt near the advertised limit

Engine and kernel upgrades deserve a note. Changing the engine changes floating-point reduction order, so even with greedy decoding the token sequence can diverge after some number of tokens. That is not a defect, but it means exact-match comparisons between blue and green are the wrong gate for engine changes. Compare task-level scores instead.

Bringing green up: load, warm, verify

Green comes up in three stages, and the controller should treat each as a gate it can fail. Load: pull weights and start the engine. For a 70-billion-parameter model in 16-bit precision that is about 140 GB per replica, so six replicas pulling simultaneously ask the storage tier for roughly 840 GB at once; stage the pull or use node-local caches so blue's own restarts are not starved. Warm: engines that capture CUDA graphs or compile kernels do that at startup or on the first request of each shape, so send traffic across the batch sizes and sequence lengths you serve before measuring anything. Verify: run a fixed suite against green directly, bypassing the router.

import hashlib, json, time, requests

def sha256(path):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        for chunk in iter(lambda: f.read(1 << 20), b""):
            h.update(chunk)
    return h.hexdigest()

def parity_failures(manifest, release_dir, declared_change):
    """Every artifact must match the manifest; only the declared ones may differ from blue."""
    bad = []
    for name, entry in manifest["artifacts"].items():
        if sha256(f"{release_dir}/{name}") != entry["sha256"]:
            bad.append(name)
        if entry["sha256"] != entry["blue_sha256"] and name not in declared_change:
            bad.append(name + " (undeclared change)")
    return bad

def verify_green(base_url, suite, slo):
    """Hit green directly. Each case: messages, a checker, and a latency budget."""
    failures, ttfts = [], []
    for case in suite:
        t0 = time.monotonic()
        r = requests.post(f"{base_url}/v1/chat/completions", timeout=120, json={
            "model": case["model"], "messages": case["messages"],
            "max_tokens": case["max_tokens"], "temperature": 0, "stream": True},
            stream=True)                     # without this, requests waits for the whole body
        first = None
        chunks = []
        for line in r.iter_lines():
            if line and first is None:
                first = time.monotonic() - t0
            chunks.append(line)
        ttfts.append(first if first is not None else float("inf"))
        if r.status_code != 200 or not case["check"](chunks):
            failures.append(case["name"])
    ttfts.sort()
    p95 = ttfts[int(0.95 * (len(ttfts) - 1))]
    return failures, p95, p95 <= slo["ttft_p95_s"]

The suite should contain three kinds of case: functional checks such as valid JSON for tool calls, refusal behaviour on a few policy prompts and correct stop-sequence handling; quality checks scored against a reference set, compared with blue's score on the same set; and latency checks at a realistic concurrency, which you can drive with the approach in LLM load testing. Run latency checks only after warm-up, or the first-request compilation cost will fail the gate for the wrong reason.

The switch: one change in the routing tier

The switch should be one change to one object, applied by the routing tier, and reversible by the same change. With the Kubernetes Gateway API that is the weight on two backend references of an HTTPRoute. A service-selector flip works similarly. DNS is the weakest option, because resolvers and clients cache records beyond their TTL and some long-lived clients never re-resolve; DNS failover for LLM serving covers that timeline in detail.

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: chat-llm
spec:
  parentRefs:
  - name: inference-gateway
  rules:
  - backendRefs:
    - name: llm-blue
      port: 8000
      weight: 0        # was 100
    - name: llm-green
      port: 8000
      weight: 100      # was 0

Two details make the switch actually atomic for users. First, conversation affinity: if your router pins sessions to replicas for cache reuse, the pin must be per colour, and the switch should move new requests even for existing sessions, otherwise some users stay on blue indefinitely. Second, in-flight requests: the router change affects new requests only. A streaming response already running on blue keeps running there until it finishes, which is what you want.

After cutover: the cold prefix cache

Blue has been serving for days and its prefix cache holds the system prompts, tool schemas and few-shot preambles that most requests share. Green starts with none of that. In the first minutes after cutover every shared prefix is prefilled from scratch, so time to first token rises and GPU compute per request rises with it, sometimes enough to trip latency alerts or autoscaling. The effect is largest for agent and RAG workloads with long shared preambles and smallest for short, unique chat turns.

The fix is to warm the cache before the switch with the prefixes that matter. Sample recent requests from blue's logs, extract their shared prefixes (system prompt plus tool definitions is usually enough), and send them to every green replica with a maximum of one output token. Because prefix caches are per replica, warming must reach each replica, not just the service. Expect the residual bump to be short; if it lasts longer than the cache-fill time, look for a real regression rather than a cold cache.

Draining blue and paying for the rollback window

After the flip, blue has two jobs: finish its in-flight streams, and remain ready to take traffic back. Finishing streams takes as long as your longest allowed generation. With a 4,096-token cap and a decode rate of 30 tokens per second for a single stream, that is about 137 seconds, so pod termination grace periods and any pre-stop hook must exceed it, or the last users on blue get cut off mid-answer.

Then decide the rollback window. Everything about blue-green's instant rollback depends on blue still being loaded; once you scale it down, rollback means a cold start, which for large models is minutes, not seconds. The cost of the window is simple arithmetic. Six replicas of eight GPUs is 48 GPUs; keeping them for four hours is 192 GPU-hours, which at an assumed $2.50 per GPU-hour is $480 per release. That is cheap insurance for a weekly model release and expensive for a team that ships several times a day, which is why frequent shippers tend to use canaries instead. A common compromise is to shrink blue to one or two replicas after the first hour: rollback then restores service quickly at reduced capacity while blue scales back up.

Worked example: a release that rolled back

A team moves a chat assistant from model v1 to v2 on six replicas of a 70B model. The manifest declares weights and generation config as the only changes. Parity fails on the first run: the chat template hash differs because v2's tokenizer config carries a template that adds a default system message. The team either declares it, after reviewing the rendered output, or overrides it to keep v1's behaviour. They choose to override, re-run parity and pass.

Green loads in about nine minutes with staged pulls. Warm-up replays 2,000 sampled prompts across batch shapes and then the 40 most common prefixes on each replica. The verification suite passes functional checks, scores within one point of blue on the reference set and meets the p95 TTFT budget. The flip happens at a low-traffic hour; TTFT p95 rises for about three minutes as uncommon prefixes fill, then settles. Twenty minutes later, a support dashboard shows more truncated answers: v2's generation config sets a lower max-token default. The controller flips the weights back, blue takes traffic immediately because it is still loaded, and the fix ships in the next release. Without the rollback window that same incident would have needed a cold start of v1.

Failure modes

  • Undeclared parity drift. Template, defaults or stop sequences differ. Gate on hashes and on rendered prompts, not on version strings.
  • Green verified cold. Latency gates fail on compilation and cache fill, or pass on an idle green and fail under load. Warm first, then measure at realistic concurrency.
  • Split brain. Some clients keep resolving blue through DNS or a pinned connection pool. Switch in the routing tier, and watch blue's request rate fall to zero after the flip.
  • Cut-off streams. Grace periods shorter than the longest generation. Size them from the token cap and decode rate.
  • Rollback to nothing. Blue scaled down too early or its nodes reclaimed by the scheduler. Hold its capacity with a reservation for the whole window.
  • Shared state migrated one way. Conversation stores, feature flags or vector indexes updated for v2 that v1 cannot read. Keep state changes backward compatible until the window closes.

Trade-offs against other release patterns

ApproachExtra GPU costRollback speedEvidence before full exposure
Blue-greenUp to a full second fleet for the windowSeconds while blue is loadedSynthetic and replayed traffic only
CanaryA small slice at a timeSeconds for the sliceReal traffic on a fraction of users
ShadowFull or sampled duplicate computeNot applicable, no users servedReal inputs, outputs never shown
In-place rollingLittle or noneMinutes, needs another rollLittle; mixed versions during the roll

Many teams combine them: shadow to compare outputs, blue-green to cut over all at once, and the rollback window as the safety net. Active-passive serving uses a similar standby but for failure, not release, and its standby tiers are a useful model for sizing blue after the first hour.

What to do next

  1. Write a release manifest that hashes weights, tokenizer files, chat template, generation config and engine image, with a field naming the declared change.
  2. Build the parity check and make the controller refuse to switch on any undeclared difference.
  3. Assemble a verification suite with functional, quality and latency cases, and run it against green directly after warm-up.
  4. Extract your most common shared prefixes from logs and warm every green replica with them.
  5. Make the switch a single routing-tier change, and alert if blue still receives requests after it.
  6. Set termination grace periods from your token cap and decode rate.
  7. Price the rollback window per release and decide how long blue stays fully loaded.
Key takeaway: Blue-green for LLMs means a second complete GPU fleet, built from a hashed release manifest, warmed and verified before it sees users, and enabled with one routing-tier change. Check parity on templates and defaults, not only weights. Warm prefix caches per replica, size drain periods from generation length, and budget how long blue stays loaded, because instant rollback lasts only as long as blue does.