An AI system pulls in more outside material than most software. It loads model weights trained by someone else, fine-tunes them on datasets assembled by someone else, runs them on a Python stack of hundreds of packages, and calls tools, plugins and hosted APIs. Every one of those is a supplier. Several have already been the way in for real attacks: a malicious package on PyTorch's nightly channel, a compromised release of a popular vision library, and model files built to run code when opened.

A scanner alone does not fix this. What fixes it is a program: a single, enforced path by which external AI artifacts enter the organisation, an inventory that records where each one runs, and a practised way to respond when one turns out to be bad. This article designs that program. It covers the threat model with real incidents, the intake pipeline, the inventory, signing and safe loading, risk tiers, metrics and incident response. It ends with a worked onboarding of an open-weight model.

What the program covers

Start by listing what counts as an AI supply-chain artifact, because programs fail at the boundary of their scope. A useful list has six classes.

ClassExamplesMain risk
Model weightsOpen-weight checkpoints, adapters, quantised files, tokenizersCode execution on load; hidden behaviour
DatasetsPre-training dumps, fine-tuning sets, evaluation setsPoisoning; licence and privacy violations
ML librariesTraining frameworks, serving engines, convertersMalicious releases; dependency confusion
Containers and runtimesBase images, CUDA layers, serving imagesOutdated or tampered layers
Tools and connectorsAgent tools, plugins, MCP servers, retrieval connectorsExcess permissions; untrusted input paths
Hosted modelsThird-party model APIsSilent version changes; data handling

Hosted models need contracts and due diligence more than file scanning; that side is covered in third-party LLM risk. The rest of this article focuses on artifacts you download and run yourself.

What has actually gone wrong

Real incidents show what the program has to stop. In December 2022, PyTorch's nightly builds pulled a malicious package named torchtriton from the public index instead of the project's own index. It was a dependency-confusion attack, and the package collected data from the machines that installed it. In December 2024, an attacker abused the GitHub Actions workflow of the Ultralytics project to publish versions 8.3.41 and 8.3.42 to PyPI with a cryptocurrency miner inside. Researchers reported that later releases were also affected. In February 2025, ReversingLabs described models on Hugging Face that used a technique they called nullifAI. The pickle data was packed in a 7z archive instead of the usual ZIP, and the stream was deliberately broken after a malicious call. The platform's scanner reported an error but did not flag the dangerous code, which still ran when the file was loaded.

These cases share a pattern. The attacker did not break any cryptography. They put a malicious artifact where a trusted name pointed, or relied on a loader that runs code. That gives four threat classes to design against.

  • Code execution on load. Python pickle can call any function while unpickling, and many checkpoint formats are pickle inside an archive.
  • Name and source confusion. Typosquats, dependency confusion between internal and public indexes, renamed or deleted accounts whose names are taken over, and package names that a coding assistant invented and an attacker later registered.
  • Compromised publishers. A real maintainer's build pipeline or token is used to ship a real name with bad contents.
  • Behavioural tampering. Weights or data altered to produce a backdoor that passes normal evaluation; see model backdoors and data poisoning.

The architecture

The architecture is a gated path with an inventory alongside it. Every external artifact enters through quarantine, is fetched by content hash, and is inspected without being executed. It is then evaluated, approved, signed with an internal key and published to an internal registry. Deployment systems pull only from that registry, and an admission check verifies the signature before anything is loaded. The inventory records every step, so when an advisory arrives you can ask which services run the affected artifact.

One path for every model, dataset and ML package into productionRequestowner, use, tierQuarantinefetch by hash, no execScan and inspectformat, opcodes, licenseEvaluatequality, safety, probesApprove and signinternal signatureInternal registryonly source for deploysAdmission checkverify before loadRuntimepinned, egress limitedInventory (AIBOM)what runs where, from which artifact and sourcerecordreportAdvisoriesnew CVE or bad modelqueryNothing reaches the runtime except through the registry, and everything in the registry is in the inventory.
The gated path. The control that matters most is the last one: runtimes must refuse artifacts that did not come from the internal registry with a valid internal signature.

The single most important rule is that the runtime refuses anything else. Without enforcement, the pipeline becomes a recommendation, and teams under deadline will pull directly from public hubs. Enforce it with network egress rules that block public model hubs and package indexes from production, plus an admission check at load time.

Intake as code

Intake decisions should be code, so they are consistent and can be reviewed. A minimal policy looks like this.

import hashlib, json, zipfile
from pathlib import Path

SAFE_FORMATS = {".safetensors", ".gguf", ".json", ".txt", ".model"}   # no code paths on load
PICKLE_FORMATS = {".bin", ".pt", ".pth", ".ckpt", ".pkl"}
ALLOWED_LICENSES = {"apache-2.0", "mit", "bsd-3-clause"}             # legal sets this list
IGNORED_NAMES = {"README.md", ".gitattributes", "LICENSE"}           # repo metadata, never loaded

def sha256(path, chunk=1 << 20):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        while block := f.read(chunk):
            h.update(block)
    return h.hexdigest()

def intake_decision(model_dir: Path, meta: dict, tier: str) -> dict:
    findings, files = [], {}
    for f in sorted(model_dir.rglob("*")):
        if not f.is_file() or f.name in IGNORED_NAMES:
            continue
        files[str(f.relative_to(model_dir))] = sha256(f)
        ext = f.suffix.lower()
        if ext in PICKLE_FORMATS:
            findings.append(("block", f.name, "pickle-based format; convert or reject"))
        elif ext == ".py":
            findings.append(("block", f.name, "custom code shipped with weights"))
        elif ext not in SAFE_FORMATS:
            findings.append(("review", f.name, "unknown format"))
    if meta.get("license", "").lower() not in ALLOWED_LICENSES:
        findings.append(("review", "license", meta.get("license")))
    if meta.get("source_revision") is None:
        findings.append(("block", "source", "no pinned upstream revision"))
    blocked = any(level == "block" for level, *_ in findings)
    needs_review = tier == "high" or any(level == "review" for level, *_ in findings)
    return {"decision": "reject" if blocked else ("review" if needs_review else "accept"),
            "files": files, "findings": findings}

Note what the policy does not trust. It does not trust a scanner's clean result on a pickle file, because nullifAI showed that scanners can be evaded. It blocks pickle formats unless a person converts the weights to safetensors in an isolated sandbox with no network, and the converted file then goes through intake again. It blocks Python files shipped with a model. In the Hugging Face transformers library those files run only when a caller passes trust_remote_code=True, so the program should forbid that flag in production code and check for it in review. It also requires a pinned upstream revision, a commit hash and not a branch name, so the artifact you reviewed is the one you get.

The inventory

An inventory answers two questions quickly: what is this artifact, and where does it run? For models and datasets, use a standard format rather than a spreadsheet. CycloneDX added machine-learning support, called ML-BOM, in version 1.5 in 2023, and SPDX 3.0 has AI and dataset profiles. Pick one and generate it automatically at intake, not by hand. A minimal record for one approved model needs these fields.

{
  "artifact_id": "acme/llm-base-8b@2026-09-30",
  "type": "model",
  "upstream": {"source": "huggingface", "repo": "example-org/base-8b",
               "revision": "3f2c9e1d0b7a...", "license": "apache-2.0"},
  "files": {"model-00001-of-00004.safetensors": "sha256:9b1e...",
            "tokenizer.json": "sha256:41ac..."},
  "derived_from": [],
  "training_data": ["dataset:acme/support-tickets@v7"],
  "evaluations": ["eval:quality-suite@2026-09-30", "eval:safety-probes@2026-09-30"],
  "approved_by": "ml-security", "tier": "high",
  "signature": "registry://models/acme/llm-base-8b@2026-09-30.sig",
  "deployments": ["svc:support-assistant/prod-eu", "svc:support-assistant/prod-us"]
}

The derived_from and deployments fields are what make incident response fast. A fine-tuned adapter points at its base model. A quantised file points at the full-precision checkpoint. When a base model is found to be bad, a graph query over these links lists every derived artifact and every service that runs one. Fill deployments from the admission check, not from team declarations, so the inventory reflects what actually loaded.

Signing and safe loading

Signing ties the artifact to a decision. OpenSSF's model signing project, published as the model-signing Python package, signs a whole model directory and stores the result as a Sigstore bundle. It supports keyless Sigstore identities, key pairs and certificates. With a key pair, the commands from the project's documentation look like this.

pip install model-signing

# sign at the end of intake, with the program's signing key (writes model.sig)
model_signing sign key ./llm-base-8b --private-key key.priv

# verify in the admission check before the model is loaded
model_signing verify key ./llm-base-8b --signature model.sig --public-key key.pub

Check the flags against the version you install; the project is active and its interface has changed between releases. Verification answers one question: is this the exact set of files the program approved? It does not say the model is safe. That came from intake and evaluation. Signing the upstream author's files with your key also does not prove who the upstream author was. Record the upstream revision separately, and verify the upstream signature too if the publisher provides one.

Then load safely, even after verification, as defence in depth. Prefer safetensors, which stores raw tensors and has no code path. If you must load a PyTorch checkpoint, note that torch.load has used weights_only=True by default since PyTorch 2.6, which restricts unpickling to tensors and a short allowlist of types. Never turn it off for external files. Container images get the same treatment, as described in container supply chain for LLM.

Risk tiers and exceptions

Not every artifact deserves the same scrutiny, and a program that treats them all as high risk will be bypassed. Tier by what the artifact can touch.

TierTypical useRequired before approval
LowOffline research, no production dataHash pinning, safe format, licence check
MediumInternal tools, non-sensitive dataLow tier plus signing, inventory and a quality evaluation
HighCustomer-facing, sensitive data, or tools with write accessMedium tier plus safety probes, backdoor-trigger testing, security review and two approvers

Write down how exceptions work: who can grant one, for how long, and what compensating control applies, such as running a pickle-format model only in an isolated job with no credentials. Exceptions that expire automatically stop a temporary workaround from becoming permanent.

Worked example: onboarding an open-weight model

Here is a worked onboarding. A team wants a new open-weight 8B model, fine-tuned for customer support. They file a request naming the use, the data it will see and the owner. That makes it tier high. The intake job fetches the repository at a pinned commit into quarantine. It finds safetensors weights, a tokenizer, a config and one modeling_custom.py file. The Python file blocks intake. The team confirms the architecture is supported by the standard library without custom code, so the file is excluded and the model loads with trust_remote_code=False.

Evaluation runs the quality suite and the safety probes, plus a trigger scan that compares outputs with and without a list of rare token sequences. Results are attached to the inventory record. Two approvers sign off, and the pipeline signs the files and publishes them to the registry. Fine-tuning reads only from the registry. The resulting adapter is a new artifact with derived_from pointing at the base, and it goes through the same steps. At deploy time, the admission check verifies both signatures, and the inventory gains two deployment entries.

Six weeks later, an advisory reports a tokenizer flaw in the upstream family. One inventory query returns the base model, the adapter and two production services. The owners know who to call within minutes, not days.

Metrics

Measure the program by what it prevents and how fast it responds.

  • Coverage: share of production model loads that pass a signature check. The target is 100 percent; anything less means a path around the program exists.
  • Bypass attempts: blocked egress to public hubs and indexes from production, by team.
  • Intake lead time: median days from request to approval. If it grows, teams will bypass the program.
  • Time to answer: minutes from an advisory to a complete list of affected deployments. Practise it.
  • Exception age: number of open exceptions and the oldest one.

Map the program to the frameworks your auditors use. NIST published SP 800-218A in July 2024, an SSDF profile for generative AI and foundation-model development, and the NIST AI RMF covers governance of third-party AI components.

Failure modes

  • Pipeline without enforcement: the gate exists, but production can still reach public hubs.
  • Scanner as verdict: trusting a clean pickle scan instead of refusing or converting pickle files.
  • Mutable references: approving a branch name or tag, so the artifact changes after review.
  • Manual inventory: a spreadsheet that drifts from reality within a quarter.
  • Derived artifacts skipped: adapters, merges and quantised files deployed without lineage.
  • Slow intake: a three-week approval queue that teaches teams to work around the program.

Trade-offs

Strict format rules reject some useful models, and converting pickle weights costs effort. Running a private registry costs storage and operations. Heavy evaluation for every artifact slows research. The usual balance is to keep the low tier light and automated, reserve human review for the high tier, and make the fast path the safe path, with one command that requests, scans and registers an artifact.

What to do next

  1. List every place your systems load models, datasets and ML packages from today, including notebooks and CI.
  2. Block production egress to public model hubs and package indexes, and point all builds at internal mirrors.
  3. Stand up a registry and an intake job that pins revisions, hashes files and rejects pickle and bundled code.
  4. Generate an ML-BOM record at intake, and fill deployments from the admission check.
  5. Sign approved artifacts and verify the signature at load time; alert on any unsigned load.
  6. Search your code for trust_remote_code=True and weights_only=False and remove them or file exceptions.
  7. Run a tabletop: a base model is declared malicious. Time how long the affected-service list takes.
Key takeaway: An AI supply chain program is one enforced path: quarantine, inspection, evaluation, signing and an internal registry, with runtimes that refuse anything else. Reject pickle files and bundled code rather than trusting scanners, pin exact revisions, record lineage and deployments in an ML-BOM, and verify signatures at load. Then measure coverage and practise answering which services run a bad artifact.