Every model debugging session eventually asks the same question: exactly which data trained this model? If the answer is "the training bucket, around March", you cannot reproduce the model, cannot tell whether a regression came from data or code, and cannot honour a request to show what personal data a model saw. Code has had version control for decades; data versioning gives training data the same property, an immutable identifier that always resolves to the same bytes.

This article explains the architecture from first principles: what a data version must capture, the three main ways systems implement it, how a training run should record the version it used, and the operational traps, especially retention, that quietly break reproducibility. Feature serving and model promotion are adjacent problems covered in feature store architecture and model registry.

Advertisement

What a data version has to pin

A dataset used for training is more than files. To rebuild a model, a version has to fix five things:

  • Bytes: the exact content of every record, not a path that might be overwritten.
  • Membership: which records are in the set, including the filter or query that selected them and the train, validation and test split.
  • Schema: column names, types and meanings, since a renamed or re-encoded column changes the data without changing the file count.
  • Labels: labels are often stored and revised separately from inputs, so they need their own version.
  • Transform code: the commit of the code that turned raw data into training examples.

The test of a good scheme is simple: given only the identifier recorded with a model, can you obtain identical training examples a year later? If any of the five can drift, the answer is no. Each architecture below pins bytes well; membership, labels and transform code depend on how you use it.

Three architectures

Systems give data an identity in three main ways. They differ in the unit of versioning, a file, a prefix of an object store, or a table, and they suit different data shapes.

Three ways to give a dataset an identity, and how a training run records itContent addressingDVC: file hash -> cache; .dvc / dvc.lock in GitObject-store branchinglakeFS: commits over S3 prefixesTable snapshotsDelta / Iceberg: version, snapshot id, tagRun manifestdata version + code commit + config + envmd5 / lockcommit idversion / tagTraining jobreads ONLY pinned versionsModel registrymodel version -> manifestRetention ties it together: a version you cannot read any more is not a version.
Each approach produces an identifier. The run manifest collects it with the code commit and configuration, and the registry links the model to the manifest.
ApproachUnitBest forWeak at
Content addressing (DVC and similar)Files and directoriesImages, audio, documents, model inputs as filesRow-level updates; very large numbers of small files
Object-store branching (lakeFS)Whole repository of objectsMixed data lakes, branch-per-experiment, atomic multi-file changesAnother service to run
Table snapshots (Delta Lake, Apache Iceberg)TablesTabular and event data, incremental updates, SQL accessUnstructured blobs; retention must be managed
Advertisement

Content addressing

Content addressing names data by a hash of its bytes. Change one byte and the name changes; keep the bytes and the name never does. DVC applies this to files tracked alongside a Git repository. dvc add data/images hashes the files, moves them into a local cache keyed by hash, and writes a small .dvc file containing the hash, which you commit to Git. dvc push copies the cache to remote storage, and dvc checkout at any Git commit restores exactly the files that commit referenced. Because storage is keyed by content, unchanged files are stored once however many versions refer to them.

Pipelines extend this to derived data. A dvc.yaml stage declares its command, inputs and outputs:

stages:
  prepare:
    cmd: python prepare.py --min-len 20
    deps:
      - prepare.py
      - data/raw
    outs:
      - data/prepared
  train:
    cmd: python train.py --config train.yaml
    deps:
      - train.py
      - train.yaml
      - data/prepared
    outs:
      - models/model.pt

dvc repro runs stages whose inputs changed and records the hash of every dependency and output in dvc.lock. Committing that lock file pins bytes, transform code and configuration together in one Git commit, which covers four of the five requirements. Labels are covered if they live in tracked files.

The limits are scale and granularity. Hashing and listing millions of small files is slow, and a single changed row in a large file creates a new version of the whole file. For tabular data that changes incrementally, table formats fit better.

Branching over an object store

lakeFS takes Git's model up a level: a repository maps onto object-store storage, and branches, commits and merges apply to the whole set of objects. Creating a branch is a metadata operation, so it is cheap even over very large data. Writers change objects on a branch in isolation; a commit records an immutable point-in-time view; a merge publishes many file changes atomically.

This fits two patterns well. Experiments can each get a branch, so a team can try a new cleaning rule without copying data or disturbing production. And ingestion can follow write-audit-publish: write a new batch to a branch, run validation on that branch, and merge only if checks pass, so consumers on the main branch never see partial or failed loads. The commit id becomes the dataset version recorded with each training run.

Table snapshots: Delta Lake and Iceberg

Open table formats version data as a by-product of how they commit writes. Every write produces a new table version (Delta) or snapshot (Iceberg) that lists the data files making up the table at that moment. Old files are not overwritten, so older versions remain readable. Internals are covered in Delta Lake architecture and Iceberg with Spark; here only the versioning interface matters.

-- Delta Lake: read the table as of a version or a time
SELECT * FROM events VERSION AS OF 1842;
SELECT * FROM events TIMESTAMP AS OF '2026-09-30 00:00:00';

-- Iceberg (Spark SQL extensions): name a snapshot and keep it for a year
ALTER TABLE prod.ml.events CREATE TAG train_2026_09 RETAIN 365 DAYS;
SELECT * FROM prod.ml.events VERSION AS OF 'train_2026_09';

Time travel is only as long as retention. Delta's VACUUM deletes data files no longer referenced by recent versions, with a default retention threshold of seven days; once it runs, older versions fail to read. Log retention separately bounds how far back time travel can reach. Iceberg's snapshot expiry does the same. A training job that records only "version 1842" will be unreproducible a week later unless that version is protected. Iceberg tags with a retention period exist for exactly this case; in Delta, the usual approach is to set longer retention on training tables, or to materialise the training set into its own table or files when the run starts.

Reading a pinned version from the training job

The most common way reproducibility fails is not storage but the reader: a training job that reads "the table" gets whatever is current when it starts. Make the version an explicit, required input to the job and resolve it once, at the start, to an immutable id:

# Delta Lake: read an exact table version
df = (spark.read.format("delta")
      .option("versionAsOf", cfg.data_version)      # required config, no default
      .load(cfg.table_path))

# Apache Iceberg: read an exact snapshot
df = (spark.read.format("iceberg")
      .option("snapshot-id", cfg.snapshot_id)
      .load("prod.ml.events"))

If the job is given a tag or a timestamp, resolve it to the concrete version or snapshot id first and record that id, because a tag can be moved and "as of" a timestamp depends on clock and commit timing. Fail the job when no version is supplied rather than falling back to the latest data. The same rule applies to DVC and lakeFS inputs: a Git commit or a lakeFS commit id, never a branch name, belongs in the manifest.

The run manifest

Whatever the storage layer, a training run should write a manifest that names everything needed to rebuild it, and the model registry should store a pointer to that manifest. Computing a fingerprint over the actual data read, not just recording the identifier requested, catches the case where a version was silently changed or a path resolved differently:

import hashlib, json, subprocess, sys, platform

def file_digest(path, chunk=1 << 20):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        while block := f.read(chunk):
            h.update(block)
    return h.hexdigest()

def dataset_fingerprint(paths):
    # Order-independent root: hash of sorted (path, digest) pairs.
    entries = sorted((p, file_digest(p)) for p in paths)
    root = hashlib.sha256("\n".join(f"{p} {d}" for p, d in entries).encode()).hexdigest()
    return root, len(entries)

def write_manifest(out_path, data_ref, paths, config):
    root, n = dataset_fingerprint(paths)
    manifest = {
        "data_ref": data_ref,          # e.g. {"kind": "iceberg_tag", "table": "...", "tag": "train_2026_09"}
        "data_sha256_root": root,
        "file_count": n,
        "code_commit": subprocess.check_output(["git", "rev-parse", "HEAD"], text=True).strip(),
        "config": config,
        "python": sys.version.split()[0],
        "platform": platform.platform(),
    }
    with open(out_path, "w") as f:
        json.dump(manifest, f, indent=2, sort_keys=True)
    return manifest

For table data, hash the materialised training files, or record the snapshot id and the list of data files it references. The manifest belongs in the same artifact store as the model and must be written before training starts, so a crashed run still leaves a record of what it was reading. Treat the manifest as part of the model's identity in your ML CI/CD pipeline: a model without one should not be promotable.

Labels, deletions and immutability

Labels change more often than inputs: relabelling campaigns, adjudicated disagreements, corrected classes. Store them as their own versioned table or file keyed by example id, and record the label version in the manifest separately. Joining inputs version X with labels version Y is then explicit and reproducible.

Deletion requests collide with immutability. If a person's records must be erased, keeping them in old versions defeats the purpose. Plan for it from the start: keep personal data in tables whose history you can rewrite or expire on a known schedule, record which model versions were trained on which data versions so affected models can be identified, and document the retention period so "reproducible" and "erasable" have a clear boundary. Content-addressed caches need the same care, because a deduplicated blob may be referenced by many versions.

A worked example

A churn model's recall drops four points between two monthly retrains. The team has run manifests for both.

Step one compares manifests: the code commit and configuration are identical, but the Iceberg tag differs and the label version moved from 14 to 15. Step two retrains the new code on the old data tag and old labels and recovers the old recall, which proves the change is in the data. Step three diffs label version 15 against 14 by example id and finds that a relabelling rule reassigned customers who paused their subscription from churned to retained. Step four reviews the rule with the business owner, who confirms paused customers should count as churned for this model. The rule is fixed, label version 16 is published, and the retrain recovers.

Without versioned labels and manifests, the same investigation is days of guessing. With them it is a few comparisons.

Failure modes

  • Version recorded, data expired: retention removed the files; the identifier now points at nothing. Protect training versions explicitly.
  • Reading latest by default: the training job reads the current table and logs nothing; reproduction is impossible.
  • Mutable paths: data at s3://bucket/train/ is overwritten in place; the manifest names a path, not content.
  • Split drift: test examples leak into training after a re-split because the split was computed at run time with a different seed or row order.
  • Unversioned labels: inputs pinned, labels pulled live from the labelling tool.
  • Non-deterministic transforms: the same inputs produce different examples because of sampling without a fixed seed or a dependency upgrade.

Trade-offs

DecisionOption AOption B
Unit of versioningFiles: simple, works with any dataTables: row-level updates, SQL time travel
Where versions liveGit plus a cache (DVC): no new serviceA versioning service (lakeFS): branches over the whole lake
Training inputRead a pinned version in place: no copyMaterialise a snapshot per run: copy cost, immune to retention
RetentionLong: reproducible, storage costShort: cheap, old models cannot be rebuilt

What to do next

  1. Pick one recent model and try to rebuild its training set; write down what was missing.
  2. Make every training job write a manifest with data reference, data fingerprint, code commit and configuration before it starts.
  3. Version labels separately and record the label version in the manifest.
  4. Protect training versions from cleanup: Iceberg tags with retention, longer Delta retention, or materialised copies.
  5. Use branches or write-audit-publish so failed ingestion never reaches the version training reads.
  6. Block model promotion when the manifest is missing or its data no longer resolves.
Key takeaway: A data version is an identifier that always resolves to the same training examples, which means pinning bytes, membership, schema, labels and transform code. Content addressing, object-store branching and table snapshots each provide the identifier; a run manifest with a data fingerprint ties it to the model; and retention policy decides whether it still works next year. Version the labels, protect the snapshots, and refuse to promote models that cannot say what they were trained on.