Teams that pin every Python dependency to a hash will often train on s3://data/latest/. A dataset is the largest and least-checked input to an LLM. It passes through as many hands as a software package: a hub repository, a scraper, a filtering job, a tokenizer, a storage bucket and a mixing config. Any of them can change the bytes the model learns from, and almost none of them leave a record.
This article treats the dataset as a versioned artifact with a supply chain. It covers what can go wrong at each stage, content-addressed manifests, pinning hub datasets to commits, signed provenance for every transform, lineage from a shard back to its sources, gating refreshes on a diff, and a worked example. Two neighbouring topics are covered elsewhere: malicious model files in supply chain attacks on ML, and the poisoning content itself in data poisoning in depth.
The chain from source to shard
Draw the chain before defending it. Each stage consumes artifacts, runs code and emits new artifacts, so each one is a place where bytes can change without anyone deciding they should.
| Stage | What goes wrong | Control |
|---|---|---|
| Source | Hub repo force-pushed or account taken over; URL content changes after curation (split-view) | Pin commit SHA; store content hashes at curation |
| Acquire | Loader code runs on your machine with your credentials | Data-only formats (Parquet, JSONL); no remote code |
| Transform | A filter or dedup dependency is compromised or simply changes behaviour | Pinned container digest; output manifest |
| Tokenize | Tokenizer version drift changes every shard silently | Tokenizer hash in the manifest |
| Storage | Shards overwritten in place in a mutable bucket | Content-addressed paths; object lock |
| Mix | Weights or sources edited without review | Mix config in git, referenced by root digest |
| Licence | A source's terms change or a sub-source was never licensed | Licence recorded per source, carried to the output |
Acquisition runs code: loaders and archives
The acquisition step is the one that executes code. For years the Hugging Face datasets library let a repository ship a Python loading script that ran when you called load_dataset, gated in later releases behind trust_remote_code=True. Version 4.0.0 removed script support entirely, so loading a scripted dataset now fails with RuntimeError: Dataset scripts are no longer supported. That makes the safe path the default. The pressure it creates is the risk: teams pin datasets<4 to keep an old pipeline running, and the remote-code path quietly comes back. Treat that pin as a security exception with an owner and an expiry date, and convert the dataset to Parquet instead.
The rest of the chain runs code too. Decompression of untrusted archives (path traversal, zip bombs), custom deserializers for pickled feature columns, and notebooks that eval config strings all count. The rule is the same as for models: data in, data out, no code from the data's publisher.
Pin sources and hash what you hold
A branch name is a pointer the publisher can move, so a pin has to be a commit. Resolve the commit once, at review time, record it, and fetch by that SHA from then on.
from huggingface_hub import HfApi, snapshot_download
from datasets import load_dataset
REPO = "some-org/instruct-mix"
sha = HfApi().dataset_info(REPO).sha # resolve ONCE, during review; store in the manifest
local = snapshot_download(REPO, repo_type="dataset", revision=sha,
allow_patterns=["data/*.parquet", "README.md"])
ds = load_dataset("parquet", data_files=f"{local}/data/*.parquet", split="train")The commit pins what the hub serves, but it does not prove the bytes on your disk are the bytes you reviewed. Hash the files yourself and keep the result as a manifest. Make the dataset's identity a single root digest over the sorted (path, hash) list, so the training config can reference one value:
import hashlib, pathlib
def sha256_file(path, chunk=1 << 20):
h = hashlib.sha256()
with open(path, "rb") as f:
for block in iter(lambda: f.read(chunk), b""):
h.update(block)
return h.hexdigest()
def build_manifest(root, name, sources, code_ref, licence):
root = pathlib.Path(root)
files = [{"path": p.relative_to(root).as_posix(), "bytes": p.stat().st_size, "sha256": sha256_file(p)}
for p in sorted(root.rglob("*")) if p.is_file()]
leaves = "".join(f'{f["path"]}\0{f["sha256"]}\n' for f in files)
return {"name": name, "root_sha256": hashlib.sha256(leaves.encode()).hexdigest(),
"files": files, "sources": sources, "code_ref": code_ref, "licence": licence}
def verify(root, manifest):
root = pathlib.Path(root)
on_disk = {p.relative_to(root).as_posix() for p in root.rglob("*") if p.is_file()}
listed = {f["path"] for f in manifest["files"]}
if on_disk != listed: # extra files matter as much as changed ones
raise RuntimeError(f"file set differs: extra={on_disk - listed} missing={listed - on_disk}")
for f in manifest["files"]:
if sha256_file(root / f["path"]) != f["sha256"]:
raise RuntimeError(f"digest mismatch: {f['path']}")For URL-list sources the unit is the item, not the file: store a hash per item at curation and drop anything that no longer matches at fetch time. The data poisoning article has that code. MLCommons' Croissant metadata format also carries per-file checksums, so if you already publish Croissant, generate it from the same manifest rather than keeping two records.
Signed provenance for every transform
A manifest says what a dataset is. A provenance attestation says how it was made: which input roots, which code at which commit, in which container image, run by which identity. Use the in-toto Statement layout that SLSA build provenance uses, so existing verification tools apply. The subject is the output root, and the inputs are listed as materials:
{
"_type": "https://in-toto.io/Statement/v1",
"subject": [{"name": "support-sft-v7", "digest": {"sha256": "<root_sha256 of output>"}}],
"predicateType": "https://slsa.dev/provenance/v1",
"predicate": {
"buildDefinition": {
"buildType": "https://example.internal/dataset-transform/v1",
"externalParameters": {"pipeline": "git+https://git.internal/data/sft@<commit>", "config": "filters.yaml"},
"resolvedDependencies": [
{"name": "instruct-mix", "digest": {"sha256": "<root of M1>"}},
{"name": "tickets-2026-09", "digest": {"sha256": "<root of warehouse export>"}}
]
},
"runDetails": {"builder": {"id": "ci://data-builds"}, "metadata": {"invocationId": "<run id>"}}
}
}Sign the statement from CI with keyless Sigstore signing. For example, cosign sign-blob statement.json --bundle statement.bundle signs with the workflow's OIDC identity, and cosign verify-blob statement.json --bundle statement.bundle --certificate-identity <workflow> --certificate-oidc-issuer <issuer> checks it. The verifier at the start of training then walks the chain backwards: shard root, transform attestation, input roots, source pins. It refuses to start if any link is unsigned, signed by the wrong identity, or references a digest that does not match. That walk is also your lineage. When someone asks which models trained on ticket export 2026-09, the answer is a graph query, not an archaeology project.
Shards, tokenizers and the mix config
Lineage tends to break at the last two stages, because tokenized shards look nothing like their sources. A shard of packed token ids cannot be traced back to a document by inspecting it, so the trace has to be written down at the moment the shard is built. Three records make it work:
- A shard index. For each shard, list the source root and row ids that were packed into it, with token offsets. Store it next to the shard and include its hash in the shard manifest. Then a takedown request or a poisoned row maps to exact shards and, through the training logs, to exact steps.
- The tokenizer's identity. Record the hash of the tokenizer files, not just a version string. A changed merge table or a new special token re-tokenizes everything, and a shard built with the wrong tokenizer trains without any error.
- The mix config as a reviewed artifact. Sampling weights decide how often each source is seen, and doubling one source's weight does as much as adding data to it. Keep the mix in git, list each component by root digest, and have the training job log the resolved digest of the whole mix. A weight change then goes through code review like any other change to model behaviour.
With these in place, which runs consumed row 81,442 of tickets-2026-09? becomes a join over manifests. That turns a data-protection or poisoning incident from a full retrain into a scoped decision.
Gate every refresh on a diff
Most real compromises arrive in an update, not on day one. Every refresh should produce a diff against the last approved manifest, and the diff should be gated like a code review. Useful signals: the share of rows changed, rows added per contributor or domain, new near-duplicate clusters, and label-distribution shift. Thresholds are local, but they must exist:
def refresh_gate(old_rows, new_rows, key, contributor, max_changed=0.02, max_one_contributor=0.25):
old = {key(r): r for r in old_rows}
new = {key(r): r for r in new_rows}
changed = [k for k in new if k not in old or new[k] != old[k]]
share = len(changed) / max(len(new), 1)
by_who = {}
for k in changed:
by_who[contributor(new[k])] = by_who.get(contributor(new[k]), 0) + 1
top = max(by_who.values(), default=0) / max(len(changed), 1)
if share > max_changed or top > max_one_contributor:
raise RuntimeError(f"refresh needs review: {share:.1%} changed, top contributor {top:.0%} of changes")
Worked example: a three-source fine-tuning mix
An illustrative support-assistant fine-tune draws on three sources: a public instruction dataset on the hub (about 52,000 rows of Parquet), a warehouse export of 180,000 resolved tickets, and 12,000 product-doc pages from a URL list. The first build pins the hub dataset to commit a1b2c3..., hashes the warehouse export as M1b, stores per-page hashes for the docs, runs the transform in a pinned container, and publishes the shards under s3://training/sha256/<root>/ with object lock on.
A month later the refresh job runs. The docs fetch drops 3% of pages on hash mismatch. Almost all of those are on one subdomain whose registration lapsed, which is exactly the split-view pattern, so the domain is removed from the list. The hub dataset has a new commit changing 4,100 rows, 7.9% of the set, which is over the 2% gate. The diff shows 61% of the changes came from one new contributor account and add a recurring phrase to answers. The refresh is held, the previous root stays in the training config, and the source goes to review. The model never sees the change, and the audit trail shows which roots every past run used.
Failure modes
- Pinning the code but not the data. The training job is reproducible except for its biggest input. Reference datasets by root digest only.
- Hashing after transformation. A manifest taken after a compromised filter step certifies the compromise. Hash at every boundary.
- Mutable storage paths.
latest/or an overwritable prefix lets anyone with write access swap shards. Use content-addressed paths and object lock. - Verify-and-forget. A verifier that warns instead of failing gets ignored within a week. Make it a hard gate with an explicit, logged override.
- Licence loss in transit. Licences recorded at the source are dropped by the mixing step, so nobody can answer a takedown. Carry them in every manifest.
- Eval contamination. Benchmark items leak in through a refreshed source. Hash eval sets and check every new root against them.
Trade-offs
Hashing terabytes costs real I/O, so hash shards as they are written instead of re-reading them later. Attesting every transform step adds CI plumbing, but most of it is boilerplate once one pipeline does it. Strict refresh gates slow down data updates, so tier them: auto-approve small diffs from internal sources and require review for public, contributor-driven ones. Governance metadata such as owners, retention and consent belongs in data governance for AI; this chain is about integrity.
What to do next
- List every dataset your current models trained on and how each is referenced. Every branch name,
latestpath or unpinned URL is a finding. - Resolve hub datasets to commit SHAs and fetch only by SHA; remove any
datasets<4pin kept for loading scripts and convert those datasets to Parquet. - Add the manifest builder at each stage boundary and reference datasets in training configs by root digest.
- Move shards to content-addressed, object-locked storage.
- Emit and sign an in-toto provenance statement for each transform run; make the pre-training verifier a hard gate.
- Put a diff gate on every refresh, with thresholds per source tier.
- Fold all of this into your AI supply chain program so models, data and packages share one intake path, and apply the same pinning to retrieval corpora (RAG poisoning).