Continuous integration was designed around one assumption: the same input produces the same output, so a test either passes or fails. Machine learning and LLM applications break that assumption in three ways. Their behavior depends on artifacts that are not code, such as data snapshots, trained weights and prompts. Their outputs are often nondeterministic, or deterministic but sensitive to changes nobody expects to matter. And the tests that actually measure quality, evaluations over hundreds or thousands of examples, are slow and cost real money per run.

The result in many teams is a CI system that runs unit tests on code and nothing else, while prompts are edited in a console and models are swapped by hand. This article describes a pipeline shape that treats every change stream as a first-class input, puts cheap deterministic checks before expensive statistical ones, makes evaluation gates robust to noise, and promotes immutable artifacts through staging, shadow and canary stages. Reversing a bad release is covered separately; the focus here is on catching problems before they reach users.

Advertisement

Four change streams

An ML or LLM application changes when any of four things change, and each needs a trigger. Code changes are the familiar case. Prompt and configuration changes, including templates, tool schemas, sampling parameters and guardrail thresholds, are the most frequent and the most under-tested. Data changes cover new training snapshots, updated eval sets and refreshed retrieval corpora. Model changes include a new fine-tune, a new base model version, or a hosted provider silently updating the model behind an alias.

Put prompts and configuration in the repository, or in a store whose changes open a pull request, so they pass through the same pipeline as code. Represent data and models by digest in a lock file that the pipeline reads, so bumping a model or dataset is a one-line diff that triggers the right suites. For hosted models, pin a dated version identifier rather than a floating alias, and schedule a job that detects when a provider announces a new version or a deprecation, so model changes arrive as pull requests instead of surprises.

# ai.lock - every non-code input the application depends on
model:      provider/model-2026-08-14        # pinned, never "latest"
judge:      provider/judge-2026-06-02        # eval judge pinned separately
embedder:   registry://e5-v3@sha256:9e77...
prompts:    prompts/support/                 # in-repo, hashed by CI
eval_sets:
  smoke:    s3://evals/support-smoke@v12     # ~100 cases, runs on every PR
  full:     s3://evals/support-full@v12      # ~2,000 cases, runs on merge
  redteam:  s3://evals/support-redteam@v5

Pipeline shape

The pipeline is a funnel ordered by cost. The first tier runs on every commit and takes minutes: linting, type checks, unit tests with the model mocked, schema validation of prompt templates and tool definitions, and contract tests. The second tier runs on every pull request that touches a prompt, a model reference or retrieval code: a smoke evaluation of around a hundred representative cases against the real model, with results cached. The third tier runs on merge to the main branch: the full evaluation suite, compared against the current production release, producing a report that becomes part of the build.

After that, the output is a release manifest pinning every artifact by digest together with its evaluation report. That manifest, not a rebuilt artifact, is what moves through staging, shadow traffic and canary into production. Building once and promoting by digest ensures that what was evaluated is exactly what ships, which is harder than it sounds when prompts are fetched at runtime and models are referenced by name.

Change streams enter one pipeline; every artifact is built once and promoted by digestcode changeprompt / configdata snapshotmodel / adapterTier 1: fast checkslint, unit, mocked LLMTier 2: smoke eval~100 cases, cachedTier 3: full evalpaired vs baseline, CImergebuild manifestdigests + eval reportstagingintegration + safetyshadowreal traffic, no userscanary 1-5%guardrail metricspromote prodalias flipNightly: full regression + red-team suites on maincatches slow drift and judge noiseResponse cache keyed by (model digest, prompt digest, input, params)reruns cost nothing unless an input to the call changed
Four change streams feed a cost-ordered funnel: fast deterministic checks on every commit, a cached smoke eval on relevant pull requests, a full paired eval on merge, then a digest-pinned manifest promoted through staging, shadow and canary.
Advertisement

Tests that do not need a model

Most bugs in LLM applications are not model quality problems. They are template variables that fail to render, tool schemas that do not match the function they describe, parsers that crash on a missing field, retries that double-charge, and context assembly that exceeds the window. All of these can be tested deterministically by replacing the model with a fake that returns scripted outputs, including malformed ones.

import json, pytest
from app.agent import answer_ticket
from app.llm import FakeLLM

def test_parser_survives_truncated_json():
    llm = FakeLLM(responses=['{"category": "billing", "reply": "We have iss'])
    result = answer_ticket("I was charged twice", llm=llm)
    assert result.status == "fallback"          # no exception, no partial send
    assert llm.calls == 2                       # exactly one repair retry

def test_prompt_renders_within_budget():
    ticket = load_fixture("longest_ticket.txt")
    prompt = render("support/answer", ticket=ticket, docs=load_fixture("top8_docs.json"))
    assert count_tokens(prompt) <= 12_000       # leaves room for the answer

@pytest.mark.parametrize("tool", load_tool_schemas())
def test_tool_schema_matches_handler(tool):
    handler = TOOL_HANDLERS[tool["name"]]
    assert set(tool["parameters"]["required"]) <= set(signature_params(handler))

Recorded fixtures extend this further. Capture real model responses for a set of inputs once, store them keyed by model and prompt digest, and replay them in tests of downstream logic. When the prompt or model changes, the keys no longer match and the fixtures are re-recorded, which is itself a useful signal of which tests depend on the change.

Making evaluation gates trustworthy

An evaluation gate that fails randomly will be overridden until nobody respects it. Three habits make gates stable. First, compare paired results, not absolute scores: run the candidate and the baseline on the same items and look at the per-item difference, which removes most of the variance due to item difficulty. Second, report uncertainty: with a few hundred items, a two-point change in pass rate is often inside the noise, and a gate should say so rather than flip. Third, gate on regressions with a tolerance rather than on hitting an absolute bar, since the absolute bar drifts as eval sets change.

import random

def paired_gate(base: list[float], cand: list[float], tolerance=0.01, n_boot=5000, seed=7):
    """Per-item scores in [0,1] for the same items. Fail only on a confident regression."""
    assert len(base) == len(cand)
    diffs = [c - b for b, c in zip(base, cand)]
    rng = random.Random(seed)
    means = sorted(
        sum(rng.choice(diffs) for _ in diffs) / len(diffs) for _ in range(n_boot)
    )
    lo, hi = means[int(0.025 * n_boot)], means[int(0.975 * n_boot)]
    mean = sum(diffs) / len(diffs)
    if hi < -tolerance:
        return "fail", mean, (lo, hi)            # confidently worse
    if lo < -tolerance:
        return "inconclusive", mean, (lo, hi)    # needs more items or a human
    return "pass", mean, (lo, hi)

The inconclusive outcome is important: it routes borderline changes to a person with the report in hand instead of forcing a binary answer from insufficient data. When model-graded evaluation is used, pin the judge model and its prompt as strictly as the model under test, validate the judge against human labels periodically, and sample at temperature zero where the provider allows it. A judge that changes silently turns every comparison against an old baseline into noise. Slice results by category as well as reporting the aggregate, because a change that improves the average while breaking one small but important category, such as refunds or safety escalations, is exactly what an aggregate hides.

Cost control through caching

Evaluation cost scales with items times runs times model price, and a busy repository can burn through a surprising budget re-running identical calls. Cache every model call in CI keyed by the model digest, the rendered prompt, the input and the sampling parameters. A pull request that changes only application code then re-runs the smoke eval almost for free, while one that changes a prompt pays only for the calls whose rendered prompt actually differs. Path filters help as well: changes under a directory that no eval depends on should not trigger evals at all.

Caching interacts with nondeterminism. If sampling is stochastic, a cache freezes one sample and hides variance. For gates, that is usually acceptable and even desirable because it makes reruns reproducible. For measuring variance itself, use a separate nightly job that deliberately bypasses the cache and runs several samples per item.

Training pipelines in CI

For classical ML and fine-tuning, CI should not train production models on every pull request. It should validate the pieces: data schema and distribution checks on the training snapshot, such as null rates, label balance, value ranges and leakage between train and eval splits; a smoke training run on a tiny sample that confirms the pipeline executes end to end and the loss decreases; and deterministic checks on feature code. Full training runs are triggered by merges or by new data snapshots, record the data digest, code commit, configuration and random seeds, and emit a candidate model that enters the same evaluation and promotion path as any other model change.

Testing retrieval and tool use

Retrieval-augmented and agentic applications need two more kinds of test. Retrieval changes, whether a new chunker, embedding model or ranking rule, should be evaluated on retrieval metrics directly, such as recall of known relevant chunks for a labeled query set, before any end-to-end answer quality is measured. A drop in answer quality with stable retrieval points at the prompt or model; a drop in retrieval explains the answer drop without spending judge calls. Run the retrieval suite against an index built in CI from a small frozen corpus, so the test is reproducible and does not depend on production data that changes daily.

Tool-using flows should run against sandboxed tool implementations that record calls instead of executing them. Assert on the trajectory as well as the final answer: the right tools were called with valid arguments, no forbidden tool was called, and the number of steps stayed within budget. Trajectory assertions catch a model that reaches the right answer through an unsafe or expensive route, which answer-only grading misses entirely.

Promotion, shadow and canary

Staging runs integration tests against real dependencies and the safety suite. Shadow deployment then replays a copy of production traffic through the candidate without returning its outputs to users, which surfaces latency, cost, error rate and output-format problems on the real input distribution. Shadowing costs a full second copy of inference for its duration, so sample traffic rather than mirroring all of it, and never shadow requests whose tools have side effects. Canary exposes a small share of users, compares guardrail metrics against the concurrent baseline, and either promotes by moving the production alias or stops.

# .github/workflows/ai.yml (abridged)
on: [pull_request, push]
jobs:
  fast:
    runs-on: ubuntu-latest
    steps: [{uses: actions/checkout@v4}, {run: make lint unit contracts}]
  smoke-eval:
    needs: fast
    if: github.event_name == 'pull_request'
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: python -m evals.run --suite smoke --cache s3://ci-cache/llm --baseline prod
      - run: python -m evals.gate --report out/smoke.json --tolerance 0.02
  full-eval-and-build:
    needs: fast
    if: github.ref == 'refs/heads/main'
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: python -m evals.run --suite full --suite redteam --baseline prod
      - run: python -m evals.gate --report out/full.json --tolerance 0.01
      - run: python -m release.build_manifest --lock ai.lock --eval out/full.json

Trade-offs

Every tier trades feedback speed against confidence. Running the full eval on every pull request catches more, but slows iteration and costs money on changes that are abandoned. Smaller smoke sets are faster but must be curated so they cover the categories that actually break. Strict tolerances block more regressions and more good changes; loose ones let small regressions accumulate across many merges, which is why the nightly run should compare main against the last production release and not only against yesterday. Paying for shadow traffic is justified when outputs are hard to evaluate offline, and wasteful when offline evals already predict online behavior well.

Failure modes

  • Unpinned model aliases. A hosted model changes under a floating name and every baseline comparison is invalid.
  • Console-edited prompts. Prompt changes bypass CI entirely because they are not in the repository.
  • Absolute-threshold gates. Gates flip on noise and teams learn to override them.
  • Drifting judge. An unpinned grading model makes scores incomparable across runs.
  • Rebuild on promote. Staging and production build separately and ship something that was never evaluated.
  • Aggregate-only reports. A regression in a small critical category hides inside an improved average.
  • Side-effecting shadows. Shadow traffic executes tools that send email or write records.
Treat code, prompts, data and models as four change streams that all enter one cost-ordered pipeline: deterministic tests with a fake model first, a cached smoke eval on pull requests, a paired full eval with confidence intervals and an inconclusive outcome on merge, then a digest-pinned manifest promoted unchanged through staging, sampled shadow traffic and canary.