A model inherits whatever its training data contains: personal information nobody meant to collect, text under licences that forbid the use, labels written by people working under unknown conditions, benchmark answers that inflate evaluations, and poisoned examples planted on purpose. Most of these risks are invisible in the weights. They are only visible if someone wrote down how the data was gathered and what it contains.

Datasheets for Datasets, proposed by Timnit Gebru and colleagues (arXiv 2018, published in Communications of the ACM in 2021), is the standard way to write that down: a structured set of questions that a dataset's creators answer and its users read before training. This article treats the datasheet as a security artifact. It covers what the questions are for, how to make a datasheet machine-checkable, how to bind it to the bytes it describes, a worked example, and how datasheets go wrong.

The seven sections and their security reading

The original paper groups its questions into seven sections that follow a dataset's life. Each one has a security reading, and that reading tells you which answers a reviewer must not accept as vague.

SectionExample questionsSecurity question it answers
MotivationWhy was it created? Who funded it?Was it built for a purpose compatible with yours?
CompositionWhat are the instances? Is there personal or sensitive data? Is anything missing?What could leak through memorisation, and what attack surface is in the content?
Collection processHow was data acquired? Over what timeframe? Was consent obtained?Could an outsider have injected data, and is there a lawful basis?
Preprocessing, cleaning, labellingWhat was removed or transformed? Is the raw data kept?Did filtering actually remove PII and duplicates, and can it be audited?
UsesWhat has it been used for? What should it not be used for?Which downstream uses create unacceptable risk?
DistributionWho gets it? Under what licence or terms?Can you legally train on it, and can you share derivatives?
MaintenanceWho maintains it? How are errors and removal requests handled?Can a takedown or a poisoning report reach every copy?

The questions are deliberately open-ended. The paper's authors present them as a prompt for reflection by creators, not a checklist to automate, and the most useful answers often describe what is not known: time ranges with no data, sources that could not be verified, filters that were not applied.

Why undocumented data is a security problem

Several attacks and failures on LLM systems trace back to training data that nobody documented:

  • Poisoning. An attacker who can write to a source the scraper reads (a wiki, a code repository, an expired domain) can plant examples. The collection-process answers say which sources were open to writes and what snapshot was taken when. See data poisoning for the attacks.
  • Memorisation and leakage. If composition says the data contains support transcripts with email addresses, the model may reproduce them. The datasheet decides whether extraction testing is required.
  • Licence and consent. Distribution and collection answers are what legal review reads; without them, no one can say whether a takedown applies.
  • Evaluation contamination. If composition lists overlap with public benchmarks, scores on those benchmarks are not evidence of capability.

In each case the datasheet does not stop the attack. It makes the risk visible at the point where someone decides whether to train, and gives incident responders a record of what went into a model.

Where the datasheet sits in the pipeline

A datasheet in a shared wiki drifts from the data within weeks. To be a control, it has to sit in the pipeline: generated alongside a manifest of the exact files, checked in CI, and required by the training job.

The datasheet is bound to the exact bytes it describes and checked before trainingSourcesscrape, vendor, logsIngest + filterPII scrub, dedup, licenceShardssha256 per fileManifesthashes + versionDatasheetprose + machine fieldsmanifest hashCI gatecomplete? matches?Training jobrefuses unknown datapassModel cardcites datasheet hashA change to any shard changes the manifest hash, which invalidates the datasheet until someone reviews it again.
The manifest ties the datasheet to specific bytes; the gate blocks training on data whose datasheet is missing, incomplete or stale.

Machine-readable datasheets

Prose answers are for people. Gates need fields. Two formats are widely used. Hugging Face dataset cards are a README.md with a YAML header carrying fields such as license, language, task_categories and size_categories, followed by prose sections. Croissant, from MLCommons, is JSON-LD metadata that describes files, record structure and checksums, and its Responsible AI extension adds properties that line up closely with datasheet questions, including rai:dataCollection, rai:dataCollectionTimeframe, rai:personalSensitiveInformation, rai:dataBiases, rai:dataLimitations, rai:dataUseCases and rai:dataReleaseMaintenancePlan.

A trimmed Croissant record with RAI fields looks like this (check the current specification for required context and properties before validating):

{
  "@context": { "...": "see the Croissant 1.0 and RAI specifications" },
  "@type": "sc:Dataset",
  "name": "support-instruct-v3",
  "license": "internal-use-only",
  "rai:dataCollection": "Support tickets 2023-2025 from consenting enterprise tenants; vendor-written responses.",
  "rai:dataCollectionTimeframe": "2023-01-01/2025-06-30",
  "rai:personalSensitiveInformation": "Emails, phone numbers and order IDs removed by regex + NER; residual rate measured at 0.3 per 10k records.",
  "rai:dataLimitations": "English only; under-represents small-business tenants.",
  "rai:dataUseCases": "Fine-tuning support assistants. Not for training general chat models.",
  "rai:dataReleaseMaintenancePlan": "Quarterly refresh; erasure requests applied within 30 days to all shards.",
  "distribution": [
    { "@type": "cr:FileObject", "name": "shard-000.jsonl.zst",
      "contentUrl": "shards/shard-000.jsonl.zst", "encodingFormat": "application/zstd",
      "sha256": "9f2c...e41a" }
  ]
}

Keep the long prose datasheet as well. The machine fields let a gate check that answers exist and that hashes match; they cannot tell whether an answer is honest.

Binding the datasheet to the bytes

A datasheet is only evidence about the version it was written for. Bind them with a manifest hash, and have the gate fail when the data changes and the datasheet has not been reviewed since. The manifest lists every shard with its size and SHA-256; the datasheet records the manifest's hash and its reviewer.

import hashlib, json, sys
from pathlib import Path

REQUIRED = ["motivation", "composition", "collection", "preprocessing", "uses", "distribution", "maintenance"]
MUST_BE_SPECIFIC = {"composition.personal_data", "collection.consent", "distribution.license", "maintenance.erasure"}

def sha256(path, buf=1 << 20):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        while chunk := f.read(buf):
            h.update(chunk)
    return h.hexdigest()

def build_manifest(data_dir: Path) -> dict:
    files = sorted(p for p in data_dir.rglob("*") if p.is_file())
    entries = [{"path": str(p.relative_to(data_dir)), "bytes": p.stat().st_size, "sha256": sha256(p)} for p in files]
    blob = json.dumps(entries, sort_keys=True).encode()
    return {"files": entries, "manifest_sha256": hashlib.sha256(blob).hexdigest()}

def gate(data_dir: Path, sheet_path: Path) -> list[str]:
    sheet = json.loads(sheet_path.read_text())
    errors = [f"missing section: {s}" for s in REQUIRED if not sheet.get(s)]
    for key in MUST_BE_SPECIFIC:
        sec, field = key.split(".")
        val = str(sheet.get(sec, {}).get(field, "")).strip().lower()
        if val in ("", "n/a", "unknown", "tbd", "see above"):
            errors.append(f"{key} needs a specific answer, got {val!r}")
    if sheet.get("manifest_sha256") != build_manifest(data_dir)["manifest_sha256"]:
        errors.append("datasheet describes a different version of the data; re-review required")
    if not sheet.get("reviewed_by"):
        errors.append("no named reviewer")
    return errors

if __name__ == "__main__":
    errs = gate(Path(sys.argv[1]), Path(sys.argv[2]))
    print("\n".join(errs) or "datasheet gate: pass")
    sys.exit(1 if errs else 0)

Run the gate in CI for the dataset repository and again at the start of every training job, against the files the job will actually read. Then have the model card cite the manifest hash, so anyone holding a model can find the exact datasheet. The model cards article covers the model side of the same binding, and provenance covers signed manifests for content.

Worked example: a support-assistant dataset

Consider a hypothetical team assembling an instruction-tuning set for a support assistant from three sources: 180,000 historical support tickets, 20,000 responses written by a labelling vendor, and 50,000 question and answer pairs scraped from public product forums. Writing the datasheet surfaces problems the pipeline did not:

  1. Composition. Answering "does it contain personal data?" forces a measurement. Sampling 2,000 records after the regex scrubber finds customer names in free text that the regex never targeted. The team adds a named-entity pass and records the measured residual rate rather than writing "PII removed".
  2. Collection. Tickets come from tenants whose contracts allow model training; three tenants opted out. The datasheet records the opt-out list version, and the ingest job reads the same list.
  3. Collection, again. Forum posts are writable by anyone. The datasheet states the snapshot date and that posts from accounts younger than 30 days were dropped, which narrows the poisoning window and tells responders where to look if a backdoor appears.
  4. Preprocessing. Near-duplicate removal against the evaluation set finds 1,200 forum questions that also appear in the team's own benchmark. They are removed and the overlap check becomes part of the gate.
  5. Distribution and uses. The vendor contract forbids using the responses outside this product. The datasheet's uses section says so, and the dataset's access policy enforces it.
  6. Maintenance. An erasure request must reach the raw tickets, every shard, and any derived set. The datasheet names the owner and the procedure, and the manifest makes it checkable that the shard changed.

Datasheets and regulation

Datasheets also feed regulatory documentation, though they are not the same thing. The EU AI Act requires data governance for high-risk systems' training, validation and test data (Article 10), covering collection processes, data preparation, assumptions, availability and suitability, and examination for possible biases. Providers of general-purpose AI models must publish a sufficiently detailed summary of training content using a template from the AI Office (Article 53(1)(d)); the European Commission published that template on 24 July 2025. A maintained datasheet per dataset is the natural source for both, but check the official texts for what each document must contain rather than assuming a datasheet satisfies them.

Data inventories, retention and deletion across stores are covered in data governance for LLM systems; the datasheet is the per-dataset record those processes point to.

Failure modes

  • Written once. The data was refreshed five times; the datasheet describes the first version. The manifest hash check catches this.
  • Aspirational answers. "PII removed" with no measurement. Require numbers and the method used.
  • Only good news. Missing time ranges, unverified sources and known label noise are left out. Ask reviewers to look for what is absent.
  • The datasheet leaks. Internal bucket paths, annotator names or examples of sensitive records copied into a public card. Review the datasheet itself for disclosure before publishing it.
  • No owner. Maintenance names a team that was reorganised. Name a role with an on-call path.
  • Gate checks presence, not truth. Automation can confirm a field exists; a named reviewer has to confirm it is right, and spot-check claims against samples.

Trade-offs

Datasheets cost the most for the people who know the data best, at the moment they least want to stop. Keep the cost proportional: a full prose datasheet for any dataset that trains a shipped model or contains personal data, a short machine-readable record for internal experiments, and the same manifest binding for both. The payoff comes later, during a takedown, a leak investigation or an audit, when the alternative is reconstructing provenance from logs that were never kept.

Use the datasheet during incidents, not only before training. When a model emits a customer's phone number, the composition and preprocessing answers say which sources could hold it and which filter should have removed it. When a backdoor trigger is reported, the collection answers list the sources that were open to outside writes and the snapshot dates to search. When a licensor withdraws permission, the distribution answers and the manifest identify every shard and every model trained on them. Write each answer with that future reader in mind: someone under time pressure who needs a source name, a date, a method and an owner, not a reassurance.

What to do next

  1. List every dataset that trains or evaluates a model you ship, with an owner for each.
  2. Adopt the seven-section template and require specific, measured answers for personal data, consent, licence and erasure.
  3. Generate a SHA-256 manifest for each dataset version and record its hash in the datasheet.
  4. Add the datasheet gate to the dataset's CI and to the start of every training job.
  5. Publish machine-readable fields as Croissant RAI properties or a dataset-card YAML header.
  6. Cite the manifest hash from the model card, and rehearse an erasure request end to end.
Key takeaway: A datasheet answers seven groups of questions about why a dataset exists, what it contains, how it was collected and cleaned, how it may be used and shared, and who maintains it. Treated as a security control, it demands measured answers, carries machine-readable fields, is bound to a manifest hash, and is checked before every training run.