A model file is not inert data. Depending on the format and the loader, opening it can import Python modules, call functions, render templates or execute code shipped alongside the weights. A dataset is not inert either: if it is a list of URLs, the bytes behind those URLs can change between the day someone labelled them and the day you train on them. Supply chain attacks on machine learning exploit exactly these two facts, and they are attractive because ML teams download third-party artifacts all day, often onto machines holding cloud credentials and GPU quota.

This article explains the mechanics of each attack at the artifact level: how a pickle runs code, why scanners miss some payloads, which loaders execute code by design, how chat templates became an injection point, how hub identities are abused, and how web-scale datasets are poisoned for the price of a domain name. Every example here is benign and aimed at detection. The organisational programme around it, intake, inventory and approval, is covered in the AI supply chain security program.

The attack surface

Where malicious models and datasets execute: every loader that interprets bytes is an attack surfaceModel hub reponame, revision, filesDataset indexURLs + labelsDownloadtag or commit hashFetch at train timecontent may have changedpickle / torch.loadopcodes run codeKeras Lambdamarshalled bytecodetrust_remote_coderepo Python filesChat templateJinja in metadataCode executionat load timePoisoned weightsat training timeswapped contentDefences in the pathpin commits, scan opcodes, safetensors, hash data, sandbox loadsLoad-time attacks run code on your machine; training-time attacks change what the model learns.
Two families: artifacts that execute when loaded, and data whose content changes after it was vetted. Defences sit on the path between download and use.

Two families of attack share the diagram. Load-time attacks hide executable behaviour in an artifact so that the act of loading it compromises the host. Training-time attacks change the bytes a model learns from, or the weights themselves, so that the compromise is in the model's behaviour rather than on the machine. The first is a classic remote code execution problem with an unusual file format; the second is closer to data integrity and is much harder to detect after the fact.

How a pickle runs code

Python's pickle format is a small stack-based program. The unpickler reads opcodes and executes them: push a string, look up a global by module and name, build a tuple, call the thing on top of the stack with those arguments. Any object can control how it is pickled by defining __reduce__, which returns a callable and its arguments; the unpickler simply calls it. That is a feature, used by PyTorch to rebuild tensors, and it is also why loading an untrusted pickle is equivalent to running untrusted code. PyTorch .pt and .bin checkpoints are zip archives whose data.pkl member is a pickle.

The benign example below shows the mechanism without a harmful payload. Disassembling it with the standard library shows the opcodes a scanner has to reason about.

import pickle, pickletools

class Marker:
    def __reduce__(self):
        # a real attack would return (os.system, ("...",)); this one only prints
        return (print, ("pickle executed a call at load time",))

blob = pickle.dumps(Marker(), protocol=4)
pickletools.dis(blob)      # shows STACK_GLOBAL for builtins.print, then REDUCE
pickle.loads(blob)         # prints the message: the call happened during load

The important opcodes are GLOBAL and STACK_GLOBAL, which resolve a module attribute, and REDUCE, INST, OBJ and NEWOBJ, which call or construct. GLOBAL carries module and name inline; STACK_GLOBAL takes them from the stack, possibly via the memo, which is where naive scanners go wrong.

Scanning pickles without fooling yourself

A scanner should allowlist, not denylist. A denylist of os.system and subprocess.Popen misses the hundreds of other callables that reach the same place. An allowlist built from the globals your own known-good checkpoints use, typically PyTorch's tensor rebuild helpers, storage classes and collections.OrderedDict, rejects everything else. The exact set depends on the PyTorch version, so derive it from your own files rather than copying a list.

import pickletools, zipfile

ALLOWED = {("collections", "OrderedDict"), ("torch._utils", "_rebuild_tensor_v2"),
           ("torch._utils", "_rebuild_parameter"), ("torch", "FloatStorage"),
           ("torch", "BFloat16Storage"), ("torch", "HalfStorage")}   # derive yours

class Rejected(Exception):
    pass

STRING_OPS = {"SHORT_BINUNICODE", "BINUNICODE", "UNICODE", "BINUNICODE8"}

def scan_pickle(data: bytes):
    memo, strs, found, last = {}, [], set(), None
    try:
        for op, arg, _pos in pickletools.genops(data):
            name = op.name
            if name in STRING_OPS:
                strs.append(arg)
                last = arg
                continue
            if name == "MEMOIZE":
                memo[len(memo)] = last          # last is None if not a string
            elif name in ("PUT", "BINPUT", "LONG_BINPUT"):
                memo[arg] = last
            elif name in ("GET", "BINGET", "LONG_BINGET"):
                strs.append(memo.get(arg))      # None marks "not a known string"
            elif name == "GLOBAL":
                found.add(tuple(arg.split(" ", 1)))
            elif name == "STACK_GLOBAL":
                if len(strs) < 2 or None in strs[-2:]:
                    raise Rejected("STACK_GLOBAL with unresolvable names")
                found.add((strs[-2], strs[-1]))
            last = None
    except Rejected:
        raise
    except Exception as e:               # truncated or malformed stream
        raise Rejected(f"unparseable pickle: {e}")   # fail closed
    bad = found - ALLOWED
    if bad:
        raise Rejected(f"disallowed globals: {sorted(bad)}")

def scan_checkpoint(path):
    with zipfile.ZipFile(path) as z:                 # raises on non-zip: reject
        for info in z.infolist():
            if info.filename.endswith(".pkl"):
                scan_pickle(z.read(info))

This is a teaching sketch, not a full stack emulator: it tracks strings through the memo and rejects when a name cannot be resolved, rather than guessing. Maintained tools such as picklescan, modelscan and fickling go further; whichever you use, two design decisions matter more than the details. Fail closed on parse errors. In early 2025 researchers reported models on Hugging Face, dubbed nullifAI, whose PyTorch files were compressed with 7z instead of zip and whose pickle streams were deliberately broken after the payload. The payload sat at the start of the stream, so it ran when deserialized even though the stream later errored, while the scanner of the time reported an error rather than a detection. A scanner that treats an error as a pass is a scanner that passes attacks. Scan what you load. Scanning the repository's .safetensors file and then loading its .bin sibling defeats the point.

Safer formats and the other code paths

The stronger defence is not to execute pickles at all. safetensors stores a JSON header with tensor names, dtypes, shapes and byte offsets followed by raw tensor bytes; loading it parses the header and maps memory, with no code path that calls arbitrary functions. PyTorch's own torch.load gained a restricted unpickler, weights_only=True, that only allows tensor-rebuilding globals, and since PyTorch 2.6 that is the default. Code that passes weights_only=False to get an old checkpoint working has switched the protection off; grep for it.

Other formats have their own code paths:

  • Keras Lambda layers store Python bytecode. CVE-2024-3660 covered Keras versions before 2.13, which deserialized Lambda layers from legacy formats without checks; later versions load with safe_mode=True by default and refuse Lambda deserialization. Researchers have since published bypasses of safe mode in specific versions, so keep Keras patched and do not pass safe_mode=False for third-party files.
  • trust_remote_code=True in Hugging Face Transformers downloads and runs Python files from the model repository to define custom architectures. That is code execution by design. If you need it, pin revision= to a full commit hash you have reviewed, because a branch name can be moved to new code after your review.
  • Chat templates are Jinja templates stored in tokenizer configs or in GGUF metadata. CVE-2024-34359 in llama-cpp-python (versions 0.2.30 up to 0.2.72, where it was fixed) rendered the template from a GGUF file in an unsandboxed Jinja environment, so a crafted template achieved code execution through server-side template injection. Any runtime that renders model-supplied templates must use a sandboxed environment.

Getting you to download the wrong thing

Many attacks never touch the file format. They make you download the wrong file. Typosquatting publishes a model under a name one character away from a popular one. Namespace reuse exploits the fact that pipelines refer to models by owner/name: if the original owner deletes or renames their account and someone else registers the old name, every pipeline that still pulls owner/name now pulls the newcomer's artifacts. Revision drift is the quiet version: you reviewed main last month, and main now points somewhere else.

The defence is to stop resolving names at run time. Mirror approved artifacts into storage you control, address them by content hash, and record that hash in an AI bill of materials so an incident query can find every deployment that uses it; see AI bill of materials.

Poisoning datasets you never stored

Web-scale image-text datasets such as LAION-400M ship as URL lists, so the bytes you download are not necessarily the bytes the curators saw. Carlini and colleagues estimated in 2023 that buying expired domains from those lists, split-view poisoning, could have poisoned 0.01 percent of LAION-400M or COYO-700M for about 60 US dollars. The supply-chain lesson is the same as for models: record a content hash when the item is vetted and drop anything that no longer matches when it is fetched. Frontrunning snapshots, fetch-time verification and the rest of the poisoning picture are covered in data poisoning attacks; weight-level triggers in LLM backdoors.

Worked example: vetting a community fine-tune

An engineer wants to evaluate a community fine-tune published as someorg/llama-ft. The repository has a pytorch_model.bin, a config.json declaring a custom architecture with an auto_map entry, and a tokenizer config with a chat template. Walk it through the intake path:

  1. Resolve the current commit hash of the repository and record it; never fetch by branch again.
  2. Download into an isolated job with no cloud credentials and no network egress except to the mirror.
  3. The auto_map entry means loading needs trust_remote_code. Read the referenced Python files. If the architecture is a standard one in disguise, load it with the built-in class instead and drop the remote code.
  4. Run the opcode scanner on every .pkl member. It reports a STACK_GLOBAL resolving to a module outside the allowlist: reject, record the hash, report the repository.
  5. If the scan had passed, convert to safetensors inside the same sandbox, hash the output, and from then on load only the converted file with weights_only semantics.
  6. Render the chat template in a sandboxed Jinja environment against a fixed conversation and diff the output against the base model's template; unexpected filters or attribute access are a red flag.

Failure modes

  • Errors counted as clean. Scanners that time out, hit an unknown compression format or a malformed stream must reject, not skip.
  • Denylist scanning. Attackers pick callables the list does not name.
  • Conversion on a trusted host. Converting a pickle to safetensors requires loading the pickle; doing that on a CI runner with deploy keys moves the compromise rather than preventing it.
  • Pinned tags, not commits. Tags and branches move; only a commit hash or content hash pins bytes.
  • Scanning one file, loading another. Libraries may prefer a different file than the one you inspected, depending on what is present.
  • Trusting a clean scan as proof of a clean model. Opcode scanning finds load-time code; it says nothing about backdoored weights or poisoned training data.

Operating it

Operationally, give model loading the same treatment as running a downloaded binary. Loads of new third-party artifacts happen in a sandbox: a container with a read-only root, no credentials, no egress and a short lifetime. The sandbox's output is a converted, hashed artifact in a private registry, and production only ever loads from that registry. Container-level hardening for the serving side is in securing LLM containers.

Monitor for regressions in the policy itself: count loads with weights_only=False, trust_remote_code=True or safe_mode=False in code review and in runtime telemetry, and treat each one as an exception that needs an owner and an expiry.

Trade-offs

ControlStopsCost
safetensors onlyPickle and marshal code execution at loadSome architectures and old checkpoints need conversion
Allowlist opcode scanningUnknown globals in picklesAllowlist maintenance per framework version
Commit or content-hash pinningRevision drift, namespace reuseManual updates when you want a new version
Sandboxed loading and conversionDamage from anything the scanner missedExtra infrastructure and latency for intake
Hashed dataset indexesSplit-view poisoningDead links grow as content changes legitimately

What to do next

  1. Grep your code for torch.load, weights_only=False, trust_remote_code and safe_mode=False; list every hit with an owner.
  2. Upgrade to PyTorch 2.6 or later, current Keras and a patched llama-cpp-python, and record the versions.
  3. Build the allowlist from your known-good checkpoints and run the opcode scanner, failing closed, on every new pickle.
  4. Convert approved pickled checkpoints to safetensors in a sandbox and load only the converted files.
  5. Replace model names in pipelines with mirrored, content-hashed artifacts and record them in your AIBOM.
  6. For URL-list datasets, store a hash per item and drop mismatches at download time.
  7. Run the worked example end to end on one external model this week and write down where your intake path diverged from it.
Key takeaway: Treat every model artifact as code until proven otherwise. Prefer formats with no execution path, scan pickles against an allowlist and fail closed on errors, load and convert third-party artifacts only in a sandbox, pin by content hash instead of by name, and hash dataset items at index time so the bytes you train on are the bytes someone actually reviewed.