A single fine-tuning run is a project. A fine-tuned model that has to stay current is a pipeline. Once you retrain on new data every week, support more than one task, or have to answer "which data produced the model that misrouted this ticket?", you need machinery around the training loop. That means versioned data snapshots, a manifest that pins everything, an orchestrated DAG whose steps can be cached and re-run, and evaluation gates against the model currently in production. It also means a registry that turns promotion and rollback into a pointer move, and triggers that decide when to train again.

This article is about that machinery. Choosing a fine-tuning method and its memory math is covered in SLM fine-tuning, in depth. Running one job on a GPU, with checkpointing and monitoring, is covered in the fine-tuning ops runbook. Here the training step is one box in a DAG, and the questions are lineage, gating, promotion and when to run again.

Advertisement

The pipeline at a glance

Ten steps repeat on every cycle: snapshot, curate, split with a leakage check, train, export, evaluate, register, roll out, monitor and trigger. Each step takes content-addressed inputs and writes content-addressed outputs, so any artifact can be traced back to the exact data and code that produced it.

A fine-tuning pipeline: every arrow carries a content hash, every box is a cacheable stepSnapshotimmutable, hashedCuratededup, PII, formatSplit + leakagegroup by customerTrainmanifest-pinnedExportmerge, quantizeEvaluate vs productiontask, regression, safetyRegistrycandidate or rejectedgateShadow, then canarylive traffic, guardrailsProduction pointerrollback = move pointerMonitor + triggersschedule, volume, drift, base updateretrainLineage from every artifact back to its snapshot answers the audit question.
Data flows left to right into training and export. The artifact is evaluated against production before it can enter the registry, and promotion moves a pointer. Monitoring feeds triggers back into the next snapshot.

The manifest: one file that pins a model

Every candidate model is defined by a manifest that is written before training and stored with the artifact. Two runs from the same manifest should produce statistically equivalent models. If a model misbehaves, its manifest tells you exactly what went into it.

# Illustrative manifest; the URI schemes are placeholders for your own stores.
run_id: tickets-router-2026w40
base_model:
  repo: org/base-1.5b-instruct
  revision: 3f9c2a1                # a commit hash, never a branch name
tokenizer_sha256: 9b1e...          # tokenizer and chat template files
data:
  train: snap://tickets/2026-09-28/train@sha256:41ac...
  val:   snap://tickets/2026-09-28/val@sha256:c07d...
eval_sets:
  task: evalset://tickets-gold/v12
  regression: evalset://general-instruct/v4
  safety: evalset://safety-format/v7
code:
  git_commit: a91d0e4
  image: registry.local/ft-train@sha256:77e2...
train:
  seed: 1234
  lr: 2.0e-4
  epochs: 2
  lora: {r: 16, lora_alpha: 32, lora_dropout: 0.05,
         target_modules: [q_proj, k_proj, v_proj, o_proj]}
export: {merge_adapter: true, quantize: int8}
gates:
  task_min_gain_vs_prod: 0.005
  regression_max_drop: 0.01
  safety_max_drop: 0.0

Three fields cause most irreproducibility: a base model referenced by branch instead of commit, a tokenizer or chat template that changed between runs, and a dataset path pointing at a mutable location. Content hashes fix all three. Bit-identical training is usually impossible on GPUs because some kernels are non-deterministic. Aim instead for reruns that land within the eval's noise band, and measure that band once by training the same manifest with three seeds.

Advertisement

Data: snapshots, lineage and leakage

Training data for an operational model comes from live systems: tickets, labelled corrections, chat logs. Freeze it. Each run reads an immutable snapshot whose files are content-addressed, and the snapshot records its source query, its time window and the version of the curation code. Curation (deduplication, PII scrubbing, filtering and formatting) is a pipeline step with its own cached output; training data curation covers it in depth.

Leakage is the failure that one-off projects rarely hit but retraining pipelines hit often. When you retrain weekly on fresh data, this week's eval examples, or near-duplicates of them, can reach next week's training set through the same customer, a templated reply or a reopened ticket. Split by a stable group key such as customer or thread, never by row, and check overlap mechanically before training. The exact-index version below is fine for tens of thousands of rows; for millions, use MinHash with locality-sensitive hashing.

import re
from hashlib import blake2b

def shingles(text, n=8):
    toks = re.findall(r"\w+", text.lower())
    return {blake2b(" ".join(toks[i:i + n]).encode(), digest_size=8).digest()
            for i in range(max(0, len(toks) - n + 1))}

def leakage_report(train_rows, eval_rows, threshold=0.5):
    index = {}
    for i, row in enumerate(train_rows):
        for s in shingles(row["text"]):
            index.setdefault(s, set()).add(i)
    leaked = []
    for j, row in enumerate(eval_rows):
        sh, hits = shingles(row["text"]), {}
        for s in sh:
            for i in index.get(s, ()):
                hits[i] = hits.get(i, 0) + 1
        best = max(hits.values(), default=0) / max(len(sh), 1)
        if best >= threshold:
            leaked.append((j, round(best, 2)))
    return leaked        # the DAG step fails if this is non-empty

Orchestration: a DAG of cacheable steps

Run the pipeline in whatever orchestrator you already have, whether that is Airflow, Argo, Kubeflow, Dagster or a CI system. The properties matter more than the tool. Each step should be idempotent, take content-addressed inputs, produce content-addressed outputs, and have a cache key derived from its inputs and code version. Then a re-run after fixing a bug in the eval step does not retrain the model.

# Orchestrator-neutral sketch: the properties matter, not the tool.
def step_key(name, code_version, input_hashes, params):
    blob = json.dumps([name, code_version, input_hashes, params], sort_keys=True)
    return hashlib.sha256(blob.encode()).hexdigest()

def run_step(name, fn, code_version, inputs, params, store):
    key = step_key(name, code_version, [i.sha for i in inputs], params)
    if store.exists(key):
        return store.get(key)                 # same inputs and code: reuse
    out = fn(*inputs, **params)               # must read nothing outside inputs
    store.put(key, out, lineage={"step": name, "code": code_version,
                                 "inputs": [i.sha for i in inputs], "params": params})
    return out

snap   = run_step("snapshot", snapshot_tickets, "v3", [], {"week": "2026-W40"}, store)
clean  = run_step("curate", curate, "v9", [snap], {}, store)
splits = run_step("split", group_split, "v2", [clean], {"key": "customer_id"}, store)
adapter = run_step("train", train_lora, "v5", [splits, base], m["train"], store)
art    = run_step("export", merge_and_quantize, "v4", [adapter, base], m["export"], store)
report = run_step("evaluate", evaluate_all, "v7", [art, prod_art, *eval_sets], {}, store)

The lineage written on every put answers the audit question: starting from a deployed artifact's hash, you can walk back to the snapshot and its source query. Put the GPU step on preemptible capacity only if training resumes from checkpoints. The other steps are cheap enough to simply re-run.

The training step

Inside the DAG, training is one container invocation that reads the manifest. For a LoRA run with Hugging Face peft, the manifest's fields map straight onto the library:

from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM, AutoTokenizer, set_seed

set_seed(m["train"]["seed"])
repo, rev = m["base_model"]["repo"], m["base_model"]["revision"]
tok = AutoTokenizer.from_pretrained(repo, revision=rev)
assert tokenizer_sha256(tok) == m["tokenizer_sha256"], "tokenizer or template drifted"
base = AutoModelForCausalLM.from_pretrained(repo, revision=rev, torch_dtype="auto")
model = get_peft_model(base, LoraConfig(task_type="CAUSAL_LM", **m["train"]["lora"]))
model.print_trainable_parameters()
# ... training loop or Trainer; save adapter, manifest and metrics side by side

Add fail-fast checks that the operations team can read: loss is not finite in the first 50 steps, gradient-norm spikes, or training loss keeps falling while validation loss rises after the first epoch. A bad run should fail the DAG loudly, not produce a model that the gate then has to catch.

Gates: against production, on the shipped artifact

Two rules separate a pipeline gate from a research eval. First, compare the candidate with the model currently in production, not with the base model; the question is whether shipping this one makes things better. Second, evaluate the exported artifact (merged, quantized and in the serving format), because merging and quantizing can erase small fine-tuning gains.

Gate on three sets, following the layout in SLM evaluation architecture: the task set, a general regression set, and a safety and format set. Thresholds come from the manifest. The seed noise band sets their floor, because a minimum gain smaller than seed-to-seed variation will pass and fail at random.

def gate(cand, prod, g):
    reasons = []
    if cand["task"] - prod["task"] < g["task_min_gain_vs_prod"]:
        reasons.append("task gain below threshold")
    if prod["regression"] - cand["regression"] > g["regression_max_drop"]:
        reasons.append("general-ability regression")
    if prod["safety"] - cand["safety"] > g["safety_max_drop"]:
        reasons.append("safety or format regression")
    return not reasons, reasons

A failed gate is a normal outcome, not an incident. The registry records the candidate as rejected along with its reasons, and production is untouched.

Registry, promotion and rollback

The registry stores each artifact with its manifest, its lineage, its eval report and a stage: candidate, shadow, canary, production or retired. Serving resolves a name and stage, such as tickets-router@production, to a hash. Promotion and rollback move that pointer, so rolling back takes seconds and needs no rebuild. Keep the previous production artifact loadable until the new one has run through a full traffic cycle.

Between the gate and production, run a shadow stage: the candidate gets a copy of live traffic, and its outputs are logged but not served, then compared with production on agreement rate and latency. Follow it with a canary on a small share of traffic, using business metrics as guardrails. If you serve LoRA adapters on a shared base, as in multi-LoRA serving, promotion can be just an adapter swap. In that case the registry must also pin the hash of the base model the adapter was trained on, because an adapter on the wrong base loads without error and produces garbage.

When to retrain

TriggerSignalWatch out for
ScheduleWeekly or monthlyRetrains when nothing changed; cheap if steps cache
Data volumeN new verified labels since the last snapshotThe label quality of the new data
DriftInput distribution shift; falling accuracy on sampled labelsUnlabelled drift cannot show the model got worse
Base model updateNew base revisionFull re-evaluation; adapters do not transfer
IncidentA systematic class of errorAdd the cases to the eval set first

Watch for two feedback hazards. If the model's own outputs become training labels, for example auto-accepted predictions, its errors reinforce themselves; train only on human-verified labels and corrections. And if the eval set never changes while the training data does, the pipeline slowly optimises toward a stale picture. Refresh the task set on a slower cadence than training, version it, and re-score production on each new version so the comparison stays like for like.

Worked example: a weekly ticket router

The numbers in this example are illustrative. A 1.5B-parameter model routes support tickets to 40 queues. Every Sunday the pipeline snapshots the week's tickets that agents rerouted or confirmed, about 6,000 rows, and appends them to a rolling 90-day window. Curation strips signatures and PII, the split groups rows by customer, and the leakage check typically flags a few dozen near-duplicate template tickets, which are dropped from training. The LoRA run takes under an hour on one GPU, and the export merges the adapter and quantizes to int8.

The gate compares the candidate with production on gold set v12. This week the candidate gains 0.9 points of macro-F1, which clears the 0.5-point threshold, set above a measured seed noise of 0.3; the regression and safety sets are flat. A day in shadow shows 96% agreement with production, and the disagreements cluster on a queue created mid-week, which is the intended change. After two days of canary at 10%, the model is promoted.

Three weeks later, a holiday spike brings a new kind of ticket, and accuracy on sampled labels drops. The drift trigger fires, but the gate rejects the retrain because only 40 labelled examples of the new kind exist. The team routes the new tickets with a rule and labels more data first.

Failure modes

  • Unpinned base: a branch moved, so the "same" run trains a different model.
  • Template drift: a tokenizer update changed the chat template, so the model was trained on one template and is served with another. Quality drops with no errors.
  • Leakage: eval scores rise every week while production metrics do not.
  • Gating the wrong artifact: the unquantized model passes, and the quantized one that ships was never measured.
  • Adapter on the wrong base: it loads cleanly and is silently wrong.
  • Label feedback loop: predictions become labels and errors compound.
  • Incomplete cache keys: a step leaves out a dependency and reuses stale output after a code change.

Trade-offs

ChoiceYou gainYou pay
Trigger-based retrainingTrains only when it mattersNeeds reliable drift and label signals
Adapters rather than merged weightsCheap storage; many tasks share one baseBase pinning and adapter-aware serving
Strict gatesFewer regressions reach usersFewer ships; good candidates rejected on noise
Rolling data windowTracks the current distributionForgets rare, older cases
Shadow stageCatches live-traffic surprisesExtra inference cost while it runs

What to do next

  1. Write a manifest for your current production model, and fill in every hash you can still recover; any you cannot recover show where your lineage has gaps.
  2. Move training data to immutable, content-addressed snapshots, split by a group key.
  3. Add the leakage check as a failing DAG step.
  4. Train one manifest with three seeds to measure your noise band, then set gate thresholds above it.
  5. Gate against the production model on the exported, quantized artifact.
  6. Put artifacts in a registry with stages, and rehearse a rollback by moving the pointer.
  7. Pick one retrain trigger beyond the schedule, and confirm that its labels are human-verified.
Key takeaway: Operating fine-tuned small models means treating each model as a build. Pin the base revision, tokenizer, data snapshot, code and hyperparameters in a manifest; freeze data in content-addressed snapshots, split by group and check for leakage; and run cacheable steps that write lineage. Gate the exported, quantized artifact against the production model with thresholds set above measured seed noise. Promote through shadow and canary by moving a registry pointer, which makes rollback instant. Retrain on deliberate triggers, using only human-verified labels.