Most security reviews of an AI system look at one moment: the model as deployed, behind an API, with a prompt filter in front. Attackers and accidents do not respect that moment. Poisoned rows enter months earlier during data collection, a compromised training node can tamper with a checkpoint, an evaluation set leaks into training and makes a weak model look safe, a stale replica of retired weights sits in a bucket long after anyone remembers it. Model lifecycle security treats the model as an artifact with a custody chain, and asks at each stage what can be tampered with, stolen or lost, and what evidence the next stage needs before it accepts the artifact.
This article walks the eight stages from data intake to retirement, gives runnable Python for a lineage manifest, a promotion gate, load-time verification and a retirement check, and ends with a worked example and a checklist. Third-party model intake is covered in depth in AI Supply Chain Security Program; organisational gates and decision rights in AI Governance Program Structure. Here the focus is the technical controls that make each hand-off verifiable.
The lifecycle as a chain of custody
Think of every artifact, whether a dataset snapshot, a checkpoint, an evaluation report or a container image, as identified by a content digest such as SHA-256 over its bytes. Names and version tags are mutable labels; digests are not. The core rule is simple: a stage may consume an artifact only if a lineage store holds a passing, attributable record for that exact digest from the previous stage. If anything is modified in between, the digest changes and the chain breaks loudly instead of silently.
Two properties matter. Integrity: the bytes you deploy are the bytes you evaluated. Attribution: each record says who or what produced it, from which inputs, in which environment. Frameworks such as SLSA and in-toto formalise these ideas for software builds, and the same thinking applies to training runs; you do not need to adopt a specific tool to get most of the value, but you do need records that cannot be edited by the people whose work they describe.
Stage 1: data intake
Threats at intake are poisoning, licence and privacy violations, and silent drift of a source. Poisoning is cheap when your crawler or labelling vendor accepts whatever arrives; Data Poisoning, in depth covers the attack families. The lifecycle controls are mostly about freezing and recording what you used.
- Snapshot, then hash. Train only from immutable snapshots. Record a manifest of file digests and row counts, not a URL to a live table.
- Source allowlist. Each source has an owner, a licence note and a sensitivity class. Unknown sources are rejected, not quarantined forever.
- Scans before admission. PII detection, deduplication against evaluation sets, and anomaly checks on label distributions per source.
- Write access is narrow. The people who can add a source are not the people who can approve a release.
Stage 2: training
A training cluster holds the most valuable artifact you have, often for days, on many machines. Treat it like a build farm for release binaries. Jobs run in ephemeral environments built from pinned images identified by digest. Outbound network access is denied except to the artifact store and package mirror. Base weights are verified against their recorded digest before the first step. Credentials are short-lived and scoped to the job.
Checkpoints deserve the same protection as final weights, because a mid-run checkpoint is nearly as useful to a thief and is far less likely to be watched. Encrypt them at rest with keys scoped to the project, and expire them on a schedule. Load checkpoints in a format that cannot execute code, such as safetensors, rather than pickle; the reasons are covered in Supply Chain Attacks on ML. The job ends by writing a manifest that ties the output digest to the inputs it consumed.
A lineage manifest you can run
The manifest below is plain Python with no external dependencies. It hashes every file in an output directory, records input digests and environment, and produces a single digest for the whole artifact, so later stages can refer to one identifier.
import hashlib, json, os, time
def file_digest(path, chunk=1 << 20):
h = hashlib.sha256()
with open(path, "rb") as f:
while block := f.read(chunk):
h.update(block)
return h.hexdigest()
def build_manifest(out_dir, inputs, image_digest, run_id):
files = {}
for root, _, names in os.walk(out_dir):
for n in sorted(names):
full = os.path.join(root, n)
files[os.path.relpath(full, out_dir)] = file_digest(full)
body = {
"run_id": run_id,
"created": int(time.time()),
"image": image_digest, # training container, by digest
"inputs": sorted(inputs), # dataset + base-model digests
"files": dict(sorted(files.items())),
}
canon = json.dumps(body, sort_keys=True, separators=(",", ":"))
body["artifact_digest"] = hashlib.sha256(canon.encode()).hexdigest()
return bodySign the canonical manifest with a key held by the pipeline, not by an engineer, and store the signature beside it. Sigstore-style keyless signing is one option; a KMS-held key is another. What matters is that the identity that signs is the automated stage, so a signature asserts the process ran, not that a person promised it did.
Stage 3: evaluation
Evaluation is where weak controls create false confidence. If evaluation items leaked into training, scores are inflated and nobody notices. Keep evaluation sets in a separate store whose read access the training job does not have, and run the deduplication check at intake in both directions. Bind every report to the artifact digest it measured: a report that names a model version string rather than a digest cannot prove which weights it scored.
Security evaluation belongs here too: jailbreak and prompt-injection suites, tool-misuse tests for agents, and probes for backdoor triggers when data provenance is weak. LLM red team architecture describes how to keep an attack corpus growing. The output of this stage is a signed evaluation record with pass or fail per required suite.
Stage 4: the promotion gate
Promotion to the registry is the point where all earlier evidence is checked mechanically. The gate is a function, not a meeting, though a human approval can be one of its inputs for high-risk tiers.
REQUIRED = {"low": ["functional"],
"high": ["functional", "safety", "injection", "privacy"]}
def promotion_gate(manifest, records, tier, approvals):
d = manifest["artifact_digest"]
errors = []
for inp in manifest["inputs"]:
if not records.get(("intake", inp), {}).get("passed"):
errors.append(f"input {inp[:12]} has no passing intake record")
for suite in REQUIRED[tier]:
r = records.get(("eval:" + suite, d))
if not r or not r["passed"]:
errors.append(f"suite {suite} missing or failed for {d[:12]}")
if tier == "high" and len({a["by"] for a in approvals}) < 2:
errors.append("high tier needs two distinct approvers")
if any(a["by"] == manifest.get("submitted_by") for a in approvals):
errors.append("submitter cannot approve own model")
return errors # empty list means promoteReturn every error at once; a gate that stops at the first failure turns a release into a slow loop of fix-and-retry. Store the gate result as another lineage record, signed by the registry.
Stage 5: deployment and load-time verification
Deployment closes the loop by verifying, at load time, that the weights on disk match the promoted digest. Registries can be wrong, caches can be stale, and an operator can copy the wrong directory.
def verify_before_load(model_dir, promoted):
for rel, expected in promoted["files"].items():
actual = file_digest(os.path.join(model_dir, rel))
if actual != expected:
raise RuntimeError(f"{rel}: digest mismatch, refusing to serve")
extra = set(os.listdir(model_dir)) - {r.split(os.sep)[0] for r in promoted["files"]}
if extra:
raise RuntimeError(f"unexpected files present: {sorted(extra)}")Rejecting unexpected files matters: a stray loader script or pickle beside clean safetensors is a classic route to code execution. Serving processes should run without write access to the model directory and without credentials that could read the registry broadly. For proprietary weights, keep them encrypted at rest, decrypt only into the serving process, and make sure no debug endpoint or support tool can export them.
Stages 6 and 7: monitoring and the retrain loop
Once live, the threats shift to abuse and extraction. Monitor for request patterns that look like systematic querying to distil the model, for spikes in refusals or policy hits that suggest a new jailbreak, and for quality drift that could indicate the wrong artifact is serving. Log the serving digest with every request batch so an incident can be scoped to exactly the traffic a bad artifact touched.
Rollback is a lifecycle control, not an afterthought. Keep the previous promoted digest warm, make the switch a single configuration change, and rehearse it. Feedback data collected in production, such as thumbs-down ratings or corrected answers, re-enters at stage one as an untrusted source; users can and do try to steer fine-tuning by flooding feedback.
Stage 8: retirement
Retirement is the stage teams skip. A retired model still has weights in object storage, replicas on inference nodes, cached copies on developer laptops, adapters derived from it, evaluation outputs that may contain training data, and API keys that still route to it. Each is either a leak or a liability.
def retirement_findings(digest, inventory):
"""inventory: list of dicts {kind, location, digest, ...} from scans."""
live = [i for i in inventory if i["digest"] == digest or i.get("parent") == digest]
out = []
for i in live:
if i["kind"] in ("weights", "checkpoint", "adapter", "cache"):
out.append(f"delete or crypto-shred {i['kind']} at {i['location']}")
elif i["kind"] in ("endpoint", "route", "api_key"):
out.append(f"revoke {i['kind']} {i['location']}")
return outCrypto-shredding, destroying the encryption key for stored weights, is the practical way to retire copies in backups you cannot edit. Keep the manifests, gate records and evaluation reports: they are small, contain no weights, and are what an auditor or incident responder will ask for later.
Worked example: a support-model fine-tune
Consider a team fine-tuning an open-weight model on support transcripts. Intake: transcripts are exported nightly, so the team snapshots one export, runs PII redaction, and records the digest D1. The base weights are verified against the digest recorded when the supply chain process admitted them, B1. Training runs in an ephemeral job with no internet, producing an adapter and a manifest whose artifact digest is A1, inputs [D1, B1].
Evaluation finds that 0.4 percent of evaluation prompts appear verbatim in the training snapshot, a contamination result: two-way deduplication at intake was not yet in place, which is exactly the gap this control closes. The team removes the overlapping items, producing a new snapshot D2, and retrains, producing A2. The high-tier gate then passes: intake records exist for D2 and B1, all four suites pass for A2, and two approvers who did not submit the run have signed. In production, the serving node refuses to start on one host because an old adapter file sits in the directory; the extra-file check caught a stale deploy. Eight months later the model is retired: the inventory scan lists two endpoints, one API key, three replicas and a cached copy in a notebook bucket, and each is closed with a record.
Failure modes
| Failure | What it looks like | Control |
|---|---|---|
| Version tags instead of digests | Report says v3 passed; v3 was rebuilt since | Bind every record to a digest |
| Self-attested lineage | Engineer edits the manifest after a failed eval | Pipeline identity signs; humans cannot write records |
| Eval contamination | Scores jump after a data refresh | Separate eval store, two-way dedup at intake |
| Checkpoint sprawl | Mid-run checkpoints in a shared bucket for a year | Encrypt, scope keys, expire on schedule |
| Pickle on the load path | Loader runs code from a model file | Safetensors only, reject extra files |
| Forgotten replicas | Retired model still answers on an old route | Inventory scans keyed by digest, retirement check |
| Feedback poisoning | Coordinated ratings steer the next fine-tune | Treat feedback as a new untrusted source |
Trade-offs
Rigor costs iteration speed. Hashing large checkpoints takes minutes, egress-free training breaks convenient package installs, and two-approver gates add a day. Scale controls to risk: a low-tier internal classifier needs digests and a functional gate, while a customer-facing model with tool access needs the full chain. Centralising lineage in one store simplifies audits but creates a high-value target; keep it append-only and replicate it. Signing gives strong attribution but only if signing identities are protected better than the artifacts they vouch for.
What to do next
- List your production models and, for each, whether you can name the exact data snapshot and base-weight digests it came from.
- Add a manifest step to the end of every training job using the code above or an equivalent.
- Move evaluation sets into a store the training job cannot read, and add two-way deduplication at intake.
- Write the promotion gate as a function with tier-specific required suites, and make it the only path into the registry.
- Add load-time digest verification and extra-file rejection to every serving image.
- Log the serving digest with request batches and rehearse a rollback.
- Pick one model to retire on paper: run an inventory scan by digest and count what you would have missed.