Most advice about LLM supply chains focuses on what you download: model weights that run code when loaded, poisoned datasets, typosquatted packages. That matters, and this site covers it elsewhere. This article hardens the other half, the path by which your own ML code is built and released. Machine learning repositories are unusually attractive targets for that path. Their CI runners often hold GPU credentials, model hub write tokens and cloud keys; their dependency trees are deep and fast-moving; and their releases are installed into notebooks and training clusters that have broad access to data.
Two incidents show the shape of the risk. In December 2024 the Ultralytics package on PyPI shipped releases containing a cryptocurrency miner. A workflow triggered by pull requests ran attacker-controlled code with enough privilege to poison a GitHub Actions cache, the legitimate release job restored that cache and built the malicious wheel, and a later pair of releases was pushed with a stored token. In March 2025 the popular tj-actions/changed-files action was modified and its existing version tags were repointed to the malicious commit, which printed CI secrets into build logs across thousands of repositories (CVE-2025-30066). Neither attack touched a model file. Both went through the build.
Three trust zones
Think of CI as three trust zones. The untrusted zone runs code from forks and outside contributors; it gets a read-only token and no secrets. The trusted build zone runs on the default branch after review; it installs dependencies from a lock file and builds artifacts. The release zone publishes; it is gated by an environment with required reviewers and obtains short-lived credentials by OIDC rather than holding long-lived tokens. Data may flow from a lower zone to a higher one only through review, never through a side channel such as a shared cache or artifact.
Never run untrusted code with privileges
GitHub's pull_request trigger runs a fork's workflow with a read-only token and no secrets, which is safe. pull_request_target runs in the context of the base repository, with its secrets and a write-capable token, so it is meant for labelling and commenting, not for building the fork's code. The dangerous pattern is pull_request_target followed by a checkout of the PR head and a build or test step: that is attacker code running with your privileges.
The second trap is script injection. Expressions like ${{ github.head_ref }} or ${{ github.event.pull_request.title }} are substituted into the shell script before it runs, so a branch name containing shell syntax becomes code. Pass untrusted values through environment variables, where the shell treats them as data.
# Before: privileged trigger, attacker code checked out, injectable expression
on: pull_request_target
jobs:
test:
runs-on: [self-hosted, gpu]
steps:
- uses: actions/checkout@v4
with: { ref: "${{ github.event.pull_request.head.sha }}" }
- run: echo "testing ${{ github.head_ref }}" && pytest
# After: unprivileged trigger, no secrets, value passed as data
on: pull_request
permissions: { contents: read }
jobs:
test:
runs-on: ubuntu-latest # no GPU or cloud credentials for fork code
steps:
- uses: actions/checkout@<full-40-char-commit-sha> # v4.x
with: { persist-credentials: false }
- env: { BRANCH: "${{ github.head_ref }}" }
run: echo "testing $BRANCH" && pytestIf fork PRs genuinely need GPU tests, run them on ephemeral runners with no attached credentials, after a maintainer applies a label, and treat everything they produce as untrusted output.
Pin actions, scope permissions, distrust caches
A tag such as v4 is a mutable pointer. The tj-actions attack worked because every consumer referencing a tag received the new commit automatically. Pin third-party actions to a full commit SHA, keep the human-readable version in a comment, and let a dependency bot propose updates as reviewable diffs. Set a top-level permissions: contents: read in every workflow and grant more only per job, so a compromised step inherits as little as possible.
Caches deserve the same suspicion as dependencies, because they are inputs. Do not restore caches in release jobs: build releases from a clean environment with hash-checked dependencies, accepting the extra minutes. Where caching is necessary in trusted builds, key caches on the lock file hash and avoid any cache that untrusted workflows can write. Static analysers for workflows, such as the open source zizmor, flag many of these patterns; the short audit below catches the worst ones with no dependencies beyond PyYAML.
import re, sys, yaml, pathlib
SHA = re.compile(r"@[0-9a-f]{40}$")
INJECT = re.compile(r"\$\{\{\s*github\.(head_ref|event\.(pull_request|issue|comment))")
def audit(path):
wf, issues = yaml.safe_load(path.read_text()), []
triggers = wf.get(True, wf.get("on", {})) # PyYAML reads bare 'on' as True
if "pull_request_target" in str(triggers):
issues.append("pull_request_target: confirm it never checks out PR code")
if "permissions" not in wf:
issues.append("no top-level permissions block")
for name, job in (wf.get("jobs") or {}).items():
for step in job.get("steps", []):
uses = step.get("uses", "")
if uses and not uses.startswith("./") and not SHA.search(uses):
issues.append(f"{name}: unpinned action {uses}")
if INJECT.search(str(step.get("run", ""))):
issues.append(f"{name}: untrusted expression inside run")
if "cache" in uses and "release" in name:
issues.append(f"{name}: cache restored in a release job")
return issues
for f in pathlib.Path(".github/workflows").glob("*.y*ml"):
for i in audit(f):
print(f"{f.name}: {i}")
Tokens out, OIDC in
A long-lived PyPI or model hub token in repository secrets is a standing invitation: anything that can read secrets can publish. PyPI's trusted publishing replaces the token with an OIDC exchange: you register the repository, workflow file and environment on PyPI, the job requests id-token: write, and PyPI issues a short-lived upload credential only to that workflow. The PyPA publish action has generated PEP 740 attestations by default since version 1.11.0 when used this way, binding each uploaded file to the workflow identity that built it.
release:
needs: build
runs-on: ubuntu-latest
environment: pypi # required reviewers configured on the environment
permissions:
id-token: write # OIDC for trusted publishing
attestations: write
contents: read
steps:
- uses: actions/download-artifact@<sha>
with: { name: dist, path: dist }
- uses: actions/attest-build-provenance@<sha>
with: { subject-path: "dist/*" }
- uses: pypa/gh-action-pypi-publish@<sha> # no username or passwordFor model hubs, use fine-grained tokens scoped to a single repository and write permission only, store them in the release environment rather than the repository, and rotate them on a schedule. A training job that only needs to read a base model should never hold a write token.
Dependencies: locked, delayed, one index
Dependencies enter at install time, so make installs reproducible and slightly behind the frontier. Lock every transitive dependency with hashes, using pip install --require-hashes -r requirements.txt or a uv lock file, so a re-uploaded or swapped file fails the install. Add a cooldown: uv's exclude-newer setting resolves only versions published before a given date, which means a malicious release has days to be noticed by others before it can reach you. Many malicious versions are reported and pulled within days, as both Ultralytics pairs were, so the delay costs little.
Dependency confusion is the classic ML case. In December 2022 PyTorch nightly builds installed via pip pulled a malicious torchtriton package from PyPI because a public package with the same name as an internal dependency took precedence. Serve internal packages from a private index that proxies the public one, and never mix indexes with --extra-index-url. A newer variant comes from coding assistants that suggest plausible but non-existent package names; attackers register those names. Require that any new dependency in a pull request is reviewed against the package's history, maintainers and download record before merge.
Runners, secrets and egress
Assume that some day a step in your pipeline will be malicious, and limit what it can do. Both incidents above were, at heart, code running inside CI that reached secrets or the network. In the tj-actions case the payload read the runner process's memory and printed secrets into public build logs in an encoded form that slipped past log masking, so masking is not a control you can rely on.
Use ephemeral runners that are destroyed after every job, especially for self-hosted GPU machines, so nothing persists from an untrusted job into a trusted one. Restrict runner egress to the package index, the model hub and your artifact store; a step that tries to reach an unknown host should fail and alert. Keep secrets in environments that only release jobs can use, give each secret the narrowest scope the platform allows, and prefer cloud OIDC federation for training and evaluation jobs so there is no long-lived cloud key to steal. Pin the cooldown and index policy in the repository rather than in each developer's environment:
# pyproject.toml
[tool.uv]
exclude-newer = "2026-09-29T00:00:00Z" # bump deliberately; never resolve same-day releases
index-strategy = "first-index" # do not fall through to other indexes by name
[[tool.uv.index]]
name = "internal"
url = "https://pypi.internal.example/simple" # proxies public PyPI
default = trueReview the date bump like any other dependency change. The lock file diff it produces is the list of new code you are about to trust.
Provenance someone verifies
Provenance helps only if someone checks it. On the consuming side, verify release artifacts before deploying them with gh attestation verify dist/pkg.whl --owner your-org, which confirms the file was built by a workflow in your organisation. Pin model downloads to a commit revision rather than a branch, verify the file hash you recorded at intake, load only safetensors, and keep trust_remote_code=False unless the remote code has been reviewed and vendored. At runtime, serving containers should run with outbound network blocked and the hub client in offline mode, so a compromised dependency cannot fetch a second stage. Signing and loading are covered in depth in model signing and supply chain attacks on ML.
Worked example: hardening a fine-tuning library
A team maintains an open source fine-tuning library with GPU tests and releases to PyPI and a model hub. The audit script reports eleven unpinned actions, a pull_request_target workflow that checks out PR code to run GPU tests on a self-hosted runner, no permissions block, and a pip cache restored in the release job, which publishes with a PyPI token stored as a repository secret.
The fixes land in one week. GPU tests move to pull_request for maintainers' branches and to a label-gated ephemeral runner for forks. Actions are pinned by SHA with a dependency bot proposing updates. Every workflow gets a read-only default. The release job builds from scratch, attests the wheel and publishes by trusted publishing from a protected environment; the PyPI token is revoked. Dependencies move to a uv lock file with a seven-day cooldown. The hub token becomes a fine-grained write token in the release environment. Downstream, the deployment pipeline verifies the attestation before promoting a version. The remaining risk is a malicious maintainer or a compromised reviewer account, which is what branch protection with two reviewers addresses.
Failure modes
- Pinning your actions but not theirs. A composite action you pinned can itself reference a mutable tag; audit what pinned actions call.
- Trusted publishing with an unprotected environment. Without required reviewers, any push to the workflow file can publish.
- Attestations nobody verifies. Generating provenance is cheap theatre unless deployment fails closed when verification fails.
- Self-hosted runners reused across jobs. A fork job can leave a persistent process for the next privileged job; use ephemeral runners.
- Cooldown overrides that become the norm. Emergency security patches need a documented bypass, not a permanently disabled delay.
Trade-offs
SHA pinning adds update toil, which a dependency bot mostly absorbs. Clean release builds cost minutes. Cooldowns delay legitimate security fixes, so keep an explicit, reviewed bypass for advisories. Label-gated fork testing slows contributors and depends on maintainers reading the diff before applying the label. Against those costs, every control closes a path that has been used in a real incident against a popular project.
What to do next
- Run the audit script on every repository that builds ML code and fix
pull_request_targetand injection findings first. - Pin all actions by SHA and add a top-level read-only permissions block.
- Remove caches from release jobs and build releases clean.
- Move PyPI publishing to trusted publishing behind a protected environment and revoke the old token.
- Lock dependencies with hashes, add a cooldown and route installs through one private index.
- Attest artifacts and make deployment verify them; see the supply chain program for wiring this into intake, and container supply chain for the serving image.