Running an open-weight model moves the whole security stack onto you. With a hosted API, the provider chooses the file format, runs the loader, patches the inference server, applies abuse monitoring and decides what the model refuses. Download the weights and every one of those becomes your job, and most teams discover this one incident at a time: a model file that executes code when loaded, a template that executes code when rendered, an inference port reachable from the internet, a fine-tune whose safety behaviour was quietly removed.

This article maps the trust boundaries of an open-weight deployment, then walks each one with the control that closes it: an intake pipeline that pins and hashes artifacts, sandboxed template rendering, signing, behavioural evaluation and a hardened serving endpoint. Malicious pickle files get one paragraph here because our ML supply chain attacks article covers them in full.

Trust boundaries

An open-weight deployment: every box is yours to securemodel hubuntrusted publisherintake pipelinepin, hash, scaninternal registrysigned, approvedeval gatesafety, backdoorsdownloadloaderweights, tokenizer, templateinference serverauth proxy in frontguardrailsinput and outputapproved artifactArtifact risks: code in pickles, remote code, templates. Runtime risks: open ports.Behaviour risks: no provider moderation, removable safety tuning, planted triggers.Nothing downstream of the hub runs anything the intake pipeline did not approve.
Trust boundaries. The hub is untrusted input; the registry is the only source the runtime may load from.

A model repository is not one artifact. It is weights, a tokenizer, a configuration file, often a chat template, sometimes Python code, and a model card with a licence. Each is parsed or executed by different software, and each has a different failure mode. The organising rule is that the hub is untrusted input, exactly like a package registry, and production loads only from an internal registry that the intake pipeline populated.

What each file can do

ComponentWho parses itWhat can go wrong
Pickle-based weights (.bin, .pt, .pkl)Python unpicklerArbitrary code execution at load
safetensors weightsA minimal header and tensor parserData only; malformed headers are the residual risk
GGUF weights and metadatallama.cpp and its bindingsMetadata strings, including the chat template, reach other code
config.json with custom architecturetransformers with trust_remote_codeArbitrary Python from the repository runs on import
Chat template (Jinja)A template engine in the loader or serverTemplate injection if rendered without a sandbox
Tokenizer filesTokenizer libraryAltered vocab or special tokens change behaviour silently

The first row is the famous one: a pickle is a program for rebuilding objects, and a malicious one can run any command during deserialisation. The fix is to accept safetensors only and to refuse repositories that ship nothing else. The fourth row is the one teams forget. Passing trust_remote_code=True to transformers imports Python files from the repository, so approving a model with custom code means approving its code, line by line, at a pinned revision.

Chat templates are code

Worked example: when a chat template is code. Chat templates are Jinja programs that turn a message list into the exact prompt string the model was trained on. GGUF files carry the template in their metadata, and loaders render it on every request. In 2024, CVE-2024-34359 showed what that means. llama-cpp-python versions 0.2.30 through 0.2.71 rendered the template from a GGUF file with an unsandboxed Jinja2 environment, so a crafted template could reach Python internals through attribute access and execute commands on the server the moment the model was loaded and used. Version 0.2.72 fixed it by rendering in a sandbox. Researchers at the time counted thousands of GGUF models on Hugging Face that depended on this path (sources: the NVD entry for CVE-2024-34359 and the JFrog research note on GGUF template injection).

The lesson generalises beyond one library: any field in a model file that some component evaluates is code. Pin a fixed loader version, and if you render templates yourself, use Jinja2's immutable sandbox and reject templates that reach for dunder attributes before you ever render them.

from jinja2.sandbox import ImmutableSandboxedEnvironment
from jinja2.exceptions import SecurityError
import re

SUSPICIOUS = re.compile(r"__\w+__|\bos\b|subprocess|import|popen|globals|builtins")

def render_chat(template_src, messages, **kw):
    if SUSPICIOUS.search(template_src):
        raise ValueError("chat template rejected: suspicious construct")
    env = ImmutableSandboxedEnvironment(trim_blocks=True, lstrip_blocks=True)
    try:
        return env.from_string(template_src).render(messages=messages, **kw)
    except SecurityError as e:
        raise ValueError(f"chat template tried unsafe access: {e}") from None

Run the static check at intake, store the approved template text in your registry next to the weights, and have the server load the stored template rather than the one embedded in the file. That also removes a quieter risk: a template edit that drops the system turn or changes special tokens degrades safety behaviour without touching a single weight.

An intake pipeline

The intake pipeline turns an untrusted repository into an approved, immutable artifact. Its job is mechanical: pin the exact revision, download only allowed file types, record a hash of every file, refuse remote code, and write a manifest that the runtime verifies before loading.

import hashlib, json, pathlib
from huggingface_hub import snapshot_download

ALLOWED = ["*.safetensors", "*.json", "*.jinja", "tokenizer.model", "*.txt", "*.md"]

def sha256_file(f):
    h = hashlib.sha256()
    with open(f, "rb") as fh:
        for block in iter(lambda: fh.read(1 << 20), b""):
            h.update(block)
    return h.hexdigest()

def intake(repo_id, revision_sha, dest):
    assert len(revision_sha) == 40, "pin a full commit hash, never a branch"
    path = pathlib.Path(snapshot_download(
        repo_id, revision=revision_sha, allow_patterns=ALLOWED, local_dir=dest))
    cfg = json.loads((path / "config.json").read_text())
    if "auto_map" in cfg:
        raise RuntimeError("custom code required; route to manual code review")
    if not list(path.glob("*.safetensors")):
        raise RuntimeError("no safetensors weights; refusing pickle-only repo")
    manifest = {"repo": repo_id, "revision": revision_sha, "files": {}}
    for f in sorted(path.rglob("*")):
        if f.is_file() and ".cache" not in f.parts:   # hub download metadata
            manifest["files"][str(f.relative_to(path))] = sha256_file(f)
    (path / "MANIFEST.json").write_text(json.dumps(manifest, indent=2))
    return manifest

An auto_map entry in config.json is how a repository points transformers at its own Python classes, so its presence is a reliable signal that loading needs remote code. At load time, the runtime recomputes hashes against the manifest and calls from_pretrained(path, trust_remote_code=False, use_safetensors=True) on the local directory, so a later push to the hub, or a renamed repository taken over by someone else, cannot change what production runs.

Signing. Hashes prove the files did not change after intake; signatures prove who approved them. The OpenSSF model-signing project, built on Sigstore, reached version 1.0 in April 2025 and signs a model directory as a unit, with model_signing sign and model_signing verify on the command line. Sign at the end of intake with your pipeline identity and verify in the serving container before load. If the publisher also signs, verify their signature at intake; most do not yet, which is exactly why your own signature matters.

Conversions, repacks and load-time checks

Most teams do not run the original release. They run a quantised GGUF, an AWQ or GPTQ build, or a merge, published by someone other than the lab that trained the base model. Each conversion is a new publisher in your supply chain, and its weights are not byte-comparable to the original, so a hash check against the base release proves nothing. Treat a repack as its own candidate: pin it, hash it, evaluate it against the base, and record both lineages in the manifest.

Where you can, convert it yourself. Quantising an approved safetensors release with a pinned version of the conversion tool gives you an artifact whose provenance you control, and it removes embedded metadata you did not write, including templates. When you must use a community build, compare its outputs with the original on a fixed prompt set at greedy decoding; quantisation shifts outputs slightly, but systematic divergence on particular prompts is a reason to look closer.

Finally, verify at load time, not only at intake. Storage gets overwritten, caches get poisoned and containers get rebuilt from the wrong path. The serving process should refuse to start on any mismatch.

import json, pathlib      # sha256_file as defined in the intake script

def verify_manifest(model_dir):
    root = pathlib.Path(model_dir)
    manifest = json.loads((root / "MANIFEST.json").read_text())
    for rel, expected in manifest["files"].items():
        if sha256_file(root / rel) != expected:
            raise SystemExit(f"refusing to serve: {rel} hash mismatch")
    extra = {str(f.relative_to(root)) for f in root.rglob("*")
             if f.is_file() and ".cache" not in f.parts}
    extra -= set(manifest["files"]) | {"MANIFEST.json"}
    if extra:
        raise SystemExit(f"refusing to serve: unexpected files {sorted(extra)}")

The unexpected-files check matters as much as the hashes: a dropped-in pickle or Python file next to approved weights is exactly what a loader with a permissive default will pick up.

Behaviour you have to verify

Clean files can still carry unsafe behaviour. Three properties of open weights matter here. There is no provider moderation layer: whatever filtering a hosted API applies to inputs and outputs is simply absent. Safety tuning is removable: anyone with the weights can fine-tune or edit refusal behaviour away, and community variants advertised as uncensored do exactly that, so a derivative of a well-behaved base model is not evidence of well-behaved weights. And weights can carry planted triggers that ordinary evaluation never exercises, covered in model backdoors and backdoor detection.

So every candidate runs an evaluation gate before it reaches the registry: a refusal regression set drawn from your own policy, compared with the base model it claims to derive from; a task suite that checks it is the model it claims to be; and trigger probes using rare tokens and odd formatting. Then put guardrails outside the model, an input classifier and an output filter, because the model's own refusals are the control you cannot fully trust.

def eval_gate(candidate, base, refusal_set, task_set, max_refusal_drop=0.05):
    def refusal_rate(model):
        return sum(model.refuses(p) for p in refusal_set) / len(refusal_set)
    r_cand, r_base = refusal_rate(candidate), refusal_rate(base)
    task = sum(candidate.solves(t) for t in task_set) / len(task_set)
    report = {"refusal_candidate": r_cand, "refusal_base": r_base, "task": task}
    if r_base - r_cand > max_refusal_drop:
        raise RuntimeError(f"safety regression vs base: {report}")
    return report

Serving without exposure

The most common open-model incident is not exotic at all: an inference server reachable from places it should not be. Ollama, for example, listens on 127.0.0.1 port 11434 by default and has no built-in authentication; setting OLLAMA_HOST=0.0.0.0 to reach it from another machine publishes an unauthenticated endpoint, and internet scans have repeatedly found very large numbers of exposed instances. Anyone who finds one can run inference on your hardware and, depending on version, use management endpoints to pull or delete models. Other servers differ in detail; check what yours binds to and whether it authenticates, rather than assuming.

  • Bind inference servers to loopback or a private interface and put an authenticating reverse proxy in front.
  • Block management endpoints, such as model pull and delete, at the proxy for every client except your deployment pipeline.
  • Run the server as a non-root user in a container with a read-only model mount and no outbound network, so even a loader bug cannot fetch a second stage.
  • Rate-limit per identity and alert on unusual token volume; stolen GPU time looks like a traffic spike.
  • Track loader and server versions like any dependency and patch on advisories, since the loader parses untrusted files.

Failure modes

  • Branch, not commit. Loading by repository name pulls whatever was pushed last. Pin a 40-character revision.
  • trust_remote_code by habit. Copied from a quick-start, it runs repository code on every load. Default to False and review exceptions.
  • Embedded templates trusted. The server renders whatever the file carries. Store and load the reviewed template.
  • Derivative assumed safe. A fine-tune of a safe base can have refusals removed. Gate on measured refusal rate.
  • Port opened for convenience. A test box with an open inference port becomes free compute for strangers. Authenticate everything.
  • Loader never patched. The component that parses untrusted files is the one most worth patching quickly.

Trade-offs

ChoiceGainCost
safetensors onlyNo code execution from weightsSome repositories unusable without conversion
No remote codeNo repository Python in productionNewest architectures wait for upstream support
Internal registry plus signingImmutable, attributable artifactsStorage and a pipeline to maintain
Eval gate before approvalCatches removed safety and some triggersHours of GPU time per candidate
External guardrailsIndependent of model behaviourLatency and false positives

What to do next

  1. List every open-weight model in use, where each was downloaded from and at which revision.
  2. Stand up an internal registry and make production load only from it.
  3. Implement the intake script: pinned revision, allow-listed files, auto_map rejection, a hash manifest.
  4. Sign approved models with model_signing and verify the signature in the serving container before load.
  5. Store reviewed chat templates beside the weights, render them in an immutable Jinja sandbox, and confirm your loader version is past known template fixes.
  6. Build a refusal regression set from your policy and gate every candidate against its base model.
  7. Audit every inference port: loopback or private bind, authenticating proxy, management endpoints blocked.
Key takeaway: Treat the model hub as untrusted input: accept safetensors, refuse remote code, pin and hash every file, sign what you approve and load only from your registry. Render chat templates in a sandbox, measure refusal behaviour against the base model, keep guardrails outside the model, and never expose an inference server without authentication.