CI/CD for an LLM service looks like CI/CD for any service until the first time an engine upgrade passes every unit test, deploys, and drops throughput by a fifth or shifts answers on a slice of prompts. The failure modes that matter live on the GPU: kernels, numerics, memory and batching. A pipeline that never runs the model on the hardware it serves on is testing the wrong thing.

This article covers the GPU side of the pipeline: what the release unit is, how to tier tests so most changes never wait for a GPU, how to check numerical parity and performance without being fooled by noise, and how to run a GPU runner fleet that is not a money pit. The general shape of ML pipelines is in CI/CD patterns for ML and LLM applications; here we go deeper on the hardware-facing parts.

The release unit: what you are actually shipping

An LLM service's behaviour is fixed by more than its code. Five things determine what a request returns and how fast: the serving code and engine version, the container image with its CUDA user-space libraries, the host GPU driver, the weights, and the serving configuration such as tensor-parallel degree, maximum batch size, KV-cache fraction and quantisation. Change any one and you have a new release.

The driver is the trap. The image carries CUDA user-space libraries, but the kernel-mode driver lives on the host, and an image built against a newer CUDA than the host driver supports fails at startup or falls back in surprising ways. So the release manifest records the minimum driver, and the deploy step refuses nodes that do not meet it.

# release.yaml - the unit that is tested and promoted
service: chat-serving
image: registry.example.com/chat-serving@sha256:4be1...   # never a tag
engine: {name: <engine>, version: <pinned>}
cuda_runtime: "12.x"            # user-space, inside the image
min_host_driver: "<from engine release notes>"
weights:
  uri: gs://models/chat-8b/rev-0193/
  digest: sha256:9c0e...         # hash of the shard manifest
serving:
  tensor_parallel: 2
  max_num_seqs: 256
  kv_cache_fraction: 0.90
  dtype: bfloat16
gpu: {type: <type>, count: 2}

Everything downstream keys off this file. Tests record which manifest they ran, the registry stores it next to the image, and promotion copies the manifest without rebuilding anything. Digests rather than tags mean the bytes that passed the gates are the bytes that ship.

Pipeline architecture

Changecode, engine, weightsBuildimage digest + manifestTier 0: CPUunit, tiny modelTier 1: 1 GPUsmoke + parityTier 2: perffixed load profileTier 3: qualityeval suite, nightlyRegistrypromote by digestCanarythen full rolloutGPU runner poolephemeral, labelled by GPU type and driverWeight cacheby digest, local NVMeruns tiers 1-3The digest that passed the gates is the digest that ships; nothing is rebuilt after testing.
Figure 1. The pipeline. Cheap CPU tiers gate access to GPU tiers; GPU tiers run on an ephemeral runner pool with a local weight cache; promotion moves the tested digest, not a rebuild.

The pipeline is a funnel. Tier 0 runs on CPU for every commit and should finish in minutes. Tier 1 starts the real engine with real weights on one GPU and checks that it serves and agrees with a reference. Tier 2 measures performance under a fixed load profile on the production GPU type. Tier 3 runs the quality evaluation suite, usually nightly or on release candidates because it is the most expensive. Each tier only runs if the cheaper one passed, and the rules for which changes trigger which tiers are explicit: a prompt-template change needs tier 3 but not tier 2; an engine or config change needs all of them.

Tier 0: everything that does not need a GPU

Most bugs in serving code have nothing to do with GPUs: request parsing, stop sequences, streaming framing, tokenizer edge cases, tool-call JSON, timeouts. Test them on CPU with a tiny randomly initialised model of the same architecture, or with a fake engine that returns scripted tokens. The tiny model has the same tokenizer and chat template, so template rendering and detokenisation bugs show up, while a forward pass costs milliseconds.

Also check the manifest here: the image reference is a digest, the weights digest matches the shard list, and the serving config fits the declared GPU memory by a simple estimate of weights plus KV cache. Catching a config that cannot fit before booking a GPU saves the most expensive kind of red build.

Tier 1: numerical parity on one GPU

Tier 1 boots the candidate on a GPU and compares it with a reference: the currently deployed release, run on the same GPU type over a fixed prompt set. Exact token equality is the wrong test. GPU kernels are not bit-reproducible across batch sizes, kernel choices or engine versions, because floating-point reductions in a different order give slightly different results; two runs of the same release with different batching can diverge after a near-tie token, and once one token differs the rest of the greedy continuation differs too.

So compare distributions at the first few positions, where divergence has not compounded, and treat the reference's own run-to-run variation as the noise floor.

import numpy as np

def parity(ref, cand, k=5):
    """ref, cand: per prompt, list over the first N positions of
    {token_id: logprob} for the top-k tokens, from teacher-forced scoring
    of the SAME reference continuation, so positions line up."""
    top1, overlap, max_diff = [], [], []
    for r_pos, c_pos in zip(ref, cand):
        for r, c in zip(r_pos, c_pos):
            r_top = sorted(r, key=r.get, reverse=True)[:k]
            c_top = sorted(c, key=c.get, reverse=True)[:k]
            top1.append(r_top[0] == c_top[0])
            overlap.append(len(set(r_top) & set(c_top)) / k)
            shared = set(r) & set(c)
            max_diff.append(max(abs(r[t] - c[t]) for t in shared) if shared else np.inf)
    return {"top1_agree": float(np.mean(top1)),
            "topk_overlap": float(np.mean(overlap)),
            "p99_logprob_diff": float(np.percentile(max_diff, 99))}

# thresholds come from noise, not from a blog post:
noise = parity(ref_run_a, ref_run_b)          # reference vs itself, different batching
cand  = parity(ref_run_a, candidate_run)
fail = (cand["top1_agree"] < noise["top1_agree"] - 0.01 or
        cand["p99_logprob_diff"] > 3 * noise["p99_logprob_diff"])

Teacher forcing is the key trick: feed both engines the same continuation and read the log probabilities they assign, rather than letting each generate freely. That makes positions comparable and isolates numerics from compounding. A quantisation or kernel change is expected to move the numbers somewhat; the gate's job is to distinguish that from a broken RoPE scaling, a wrong chat template or a mis-sharded weight, all of which show up as a collapse in top-1 agreement rather than a small shift. When a change is supposed to alter numerics, the parity tier reports and tier 3 decides.

Tier 2: the performance gate

Performance regressions are the other class unit tests never see. The gate replays a fixed load profile: a recorded distribution of prompt and output lengths, a fixed concurrency or arrival rate, fixed sampling parameters and a warm-up period that is discarded. It reports time to first token, inter-token latency at p50 and p99, and output tokens per second per GPU. Running the profile on a different GPU type, driver or power cap than production makes the numbers meaningless, so the runner label pins all three.

One run is not a measurement. Run baseline and candidate several times each, interleaved on the same node class, and compare medians against the spread of the baseline. The statistics of doing this well, bootstrap intervals and how many requests you need to see a given change, are covered in LLM performance regression analysis. A minimal gate:

import statistics as st

def perf_gate(base_runs, cand_runs, metric, worse_if_higher, budget=0.03):
    b = [r[metric] for r in base_runs]
    c = [r[metric] for r in cand_runs]
    noise = (max(b) - min(b)) / st.median(b)      # relative spread of the baseline
    delta = (st.median(c) - st.median(b)) / st.median(b)
    if not worse_if_higher:
        delta = -delta
    allowed = max(budget, 2 * noise)              # never gate inside the noise
    return {"metric": metric, "delta": round(delta, 4),
            "allowed": round(allowed, 4), "pass": delta <= allowed}

checks = [perf_gate(base, cand, "ttft_p99_ms", True),
          perf_gate(base, cand, "itl_p50_ms", True),
          perf_gate(base, cand, "tok_s_per_gpu", False)]

Record a peak-memory figure as well. An engine upgrade that reserves more memory for activations shrinks the KV cache, which looks fine at test concurrency and falls over at production concurrency when sequences start being preempted.

The GPU runner fleet

GPU runners are where pipeline cost goes. Four practices keep it sane. Make runners ephemeral: one job per machine or pod, then the node is reset, so a crashed test cannot leave a zombie process holding GPU memory for the next job. Label runners by GPU type, GPU count and driver version, and have jobs request labels from the manifest rather than hard-coding them. Cache weights by digest on local NVMe, because downloading tens of gigabytes per job dominates runtime; key the cache on the digest so a stale copy can never be served under a new name. And scale the pool to zero outside working hours, with nightly tier 3 jobs packed into one window.

Before every GPU job, run a short health check: the expected number of GPUs is visible, the driver version matches the label, no other process holds memory, and for multi-GPU jobs a quick collective test passes. A test that fails because of a sick node wastes more engineering time than the GPU-hour it ran on, because someone investigates a code regression that does not exist.

Promotion and rollback

Promotion moves the manifest, image digest and weights digest that passed the gates into the production registry; nothing is rebuilt. The deploy step checks node driver versions against the manifest, starts the new version on a slice of capacity and hands control to the canary process described in LLM canary deployment, with automatic abort on latency, error rate and quality signals. Rollback is a promotion of the previous manifest, and it must be rehearsed, because weights for the old version may already have been evicted from node caches.

Worked example: an engine upgrade

A team bumps the serving engine by one minor version. Tier 0 passes. Tier 1 boots on one GPU and parity against the deployed release shows top-1 agreement over the first 16 positions of 400 prompts at 0.991, against a reference-versus-itself noise of 0.993; the p99 log-probability difference is within twice the noise. Parity passes.

Tier 2 runs five interleaved baseline and candidate runs on the production GPU type. Throughput is up 4 percent, which the team is pleased about, but p99 time to first token is up 11 percent against a baseline spread of 3 percent, so the gate fails. The changelog shows the new version enables a larger default prefill chunk. The fix is a one-line config change that restores the old value, which creates a new manifest; the pipeline reruns tiers 1 and 2, both pass, tier 3 passes overnight, and the canary proceeds the next morning. Without tier 2 the regression would have reached users as a latency complaint, with five changes in flight to blame.

These figures are illustrative; what matters is that each threshold was relative to measured noise.

Failure modes

FailureSymptomFix
Image tag reusedProduction runs bytes nobody testedDigest-only references, enforced by the manifest check
CUDA newer than host driverStartup failure on some nodes onlymin_host_driver in the manifest; scheduler refuses mismatched nodes
Exact-match output testsFlaky failures on every kernel changeTeacher-forced top-k parity with a measured noise floor
Perf test on a different GPU or power capGreen in CI, slow in productionRunner labels pin GPU type, driver and power settings
Dirty runnerOut-of-memory in an unrelated jobEphemeral runners and a pre-job health check
Stale weight cacheTests pass against old weightsCache keyed by weights digest, verified on load
Tier 3 on every commitGPU bill and queue time explodeTrigger rules by change type; nightly batches

Trade-offs

More tiers catch more regressions but lengthen the path to production; most teams accept a same-day path for code and a next-day path for engine and weight changes. Teacher-forced parity is cheap and robust but blind to changes that only appear in free generation, such as sampling bugs, which is why tier 3 still generates. Interleaved repeated perf runs cost several GPU-hours but are the only way to see single-digit changes reliably. Owning a runner pool gives control over drivers and caches; renting per-job capacity avoids idle cost but makes weight caching harder. Decide which you value per tier, not once for the whole pipeline. For quality gating itself, an LLM evaluation harness for regression testing covers suite design and statistics.

What to do next

  1. Write a release manifest for your current production service, including the driver it actually runs on, and switch every image reference to a digest.
  2. Build a tier 0 suite on CPU with a tiny model of the same architecture and tokenizer.
  3. Record a teacher-forced reference run of a few hundred prompts on the deployed release, run it twice with different batching, and save both as your noise floor.
  4. Add a parity job that runs on one GPU for engine, config and weight changes.
  5. Capture a production load profile and build a tier 2 job that runs interleaved repeats on the production GPU type.
  6. Make GPU runners ephemeral, labelled by type and driver, with a digest-keyed weight cache and a pre-job health check.
  7. Write trigger rules mapping change types to tiers, and rehearse a rollback by promoting the previous manifest.
Key takeaway: An LLM release is code plus engine, image digest, host driver, weights digest and serving config, so pin all of them in a manifest and promote that manifest without rebuilding. Run most tests on CPU with a tiny model, gate GPU changes with teacher-forced top-k parity against a measured noise floor, and gate performance with interleaved repeated runs of a fixed load profile on the production GPU type. Keep GPU runners ephemeral, labelled and backed by a digest-keyed weight cache, and let canary analysis make the final call.