Most writing about Hugging Face Datasets is about consuming data: call load_dataset, map a tokenizer, train. This article is about the other side, the one that decides whether those training runs mean anything: producing a dataset that other people (or your own team, six months from now) can load, trust and reproduce. A dataset repository on the Hub is a git repository with a README and some data files, and almost every bad result traced to data comes from a decision made while building it: a column whose type silently changed, a test split that shares customers with the train split, a duplicate set that inflates scores, or a version that moved under an experiment.
We will build a dataset from raw records to a tagged, validated release, explaining each step from first principles, with a worked example you can adapt. If you want the library internals first (Arrow memory mapping, the fingerprint cache and streaming), read HuggingFace Datasets Architecture; for the Hub's repositories, cache and revisions in general, read the Hugging Face Hub in depth.
What a dataset repository is, after the 4.0 change
A dataset on the Hub is a repository of type dataset: versioned files, a README.md whose YAML front matter doubles as machine-readable metadata, and data files in a format the library can read without custom code, such as Parquet, CSV, JSON Lines, plain text, or image and audio files. When you call load_dataset("org/name"), the library resolves the repository at a revision, reads the YAML to find subsets (called configs) and splits, downloads the matching files and builds Arrow tables from them.
For years there was a second path: a Python loading script inside the repository that downloaded and parsed data itself. Datasets 4.0 removed it. Loading a repository that still relies on a script fails with RuntimeError: Dataset scripts are no longer supported, but found X.py, and the trust_remote_code escape hatch is gone with it. The practical consequence for producers is simple and good: the data files are the dataset. Whatever cleaning, joining and splitting you do happens before upload, in your own pipeline, and what lands on the Hub is static, inspectable and loadable by any Parquet reader, not only by Python.
If you maintain an old script-based dataset, convert it: load it once with a 3.x install that can still run the script, then push the resulting splits as Parquet with the code below, and delete the script in the same commit. Consumers who pinned the old revision keep working on 3.x; everyone else gets files.
The pipeline at a glance
Treat publishing as a small data pipeline with a gate at the end, not as a one-off upload from a notebook. Each stage below exists because skipping it produces a specific, recognisable failure later.
Schema first: declare Features, do not infer them
When you build a dataset from JSON or pandas, the library infers column types from the data. Inference is fine for exploration and a liability for a release. A column that is null in the first batch is inferred as null and then fails, or an ID that looks numeric in one shard is an integer there and a string in the next, and Parquet files with different schemas do not concatenate. Declare the schema with Features and cast to it explicitly, so a bad record fails loudly in your pipeline instead of quietly in someone else's training job.
from datasets import Dataset, Features, Value, ClassLabel
LABELS = ["billing", "bug", "feature_request", "account", "other"]
FEATURES = Features({
"ticket_id": Value("string"), # string even if it looks numeric
"customer_id": Value("string"), # the GROUP key used for splitting
"created_at": Value("timestamp[s]"),
"text": Value("string"),
"label": ClassLabel(names=LABELS), # stored as int, names travel with the data
})
rejects = []
def records():
for row in read_export("tickets_2026q3.jsonl"): # your own reader
if row.get("queue") not in LABELS:
rejects.append(row) # reviewed later, never silently dropped
continue
yield {
"ticket_id": str(row["id"]),
"customer_id": str(row["account"]),
"created_at": row["created"],
"text": row["subject"] + "\n\n" + row["body"],
"label": row["queue"], # ClassLabel encodes the name to its index
}
ds = Dataset.from_generator(records, features=FEATURES)
# from_generator caches by fingerprint: a cached rerun skips records(), so persist rejects on the first runClassLabel stores the label names with the data, so no consumer maps integer 2 to the wrong class from a list copied in a different order.
Deduplicate before you split, and check across splits after
Exact duplicates are common in exported data: retried writes, forwarded emails, templated messages. Remove them before splitting, using a normalised hash of the text, and then check that no hash appears in two splits. Near-duplicate detection (MinHash, embeddings) catches more at more cost; start exact.
import re
import hashlib
def norm_hash(text: str) -> str:
t = re.sub(r"\s+", " ", text.lower()).strip()
return hashlib.sha1(t.encode()).hexdigest()
ds = ds.map(lambda r: {"h": norm_hash(r["text"])})
seen = set()
def first_seen(r):
if r["h"] in seen:
return False
seen.add(r["h"])
return True
ds = ds.filter(first_seen) # single process: the closure holds state, so no num_proc hereThe comment matters: filter with num_proc runs the function in separate processes, each with its own seen set, and duplicates survive across workers. Stateful logic belongs in a single pass, or in a group-by over the hash column done in pandas or Polars before the data enters the library.
Splits that do not leak
The default train_test_split shuffles rows. That is correct only if rows are independent, and in real data they rarely are: the same customer files many tickets, the same document is chunked into many passages, the same patient has many scans. A row-level split puts near-identical examples on both sides, and the test score measures memorisation. The fix is to split by a group key, deterministically, so the assignment never changes when you add data or rerun the pipeline.
def bucket(key: str, salt: str = "tickets-v1") -> int:
h = hashlib.sha256(f"{salt}:{key}".encode()).hexdigest()
return int(h[:8], 16) % 100
def assign_split(row):
b = bucket(row["customer_id"])
row["split"] = "test" if b < 10 else "validation" if b < 20 else "train"
return row
ds = ds.map(assign_split)
splits = {s: ds.filter(lambda r, s=s: r["split"] == s).remove_columns("split")
for s in ("train", "validation", "test")}Hashing the group key gives a stable split: a customer who appears next quarter lands in the same split as before, so a v2 release does not leak v1's test set into v2's train set. Keep the salt fixed for the life of the dataset and record it in the card. If you need a time-based evaluation instead (train on the past, test on the future), split on created_at and say so; both are valid, but they answer different questions.
Worked example: publishing the ticket dataset
Suppose the export holds 120,000 tickets from 9,000 customers. After casting, 1,800 rows fail validation (unknown queue names) and are written to a reject file for review rather than dropped silently. Exact deduplication removes about 4 percent, mostly auto-replies. The hash split gives roughly 80/10/10 by customer, which is not exactly 80/10/10 by row, and that is expected: big customers move whole blocks of rows. Now push it.
from datasets import DatasetDict
from huggingface_hub import HfApi
dd = DatasetDict({k: v.remove_columns("h") for k, v in splits.items()})
dd.push_to_hub(
"acme-ml/support-tickets",
config_name="default",
private=True,
max_shard_size="500MB",
commit_message="v1.0: 2026Q3 export, customer-hash split (salt tickets-v1)",
)
api = HfApi()
api.create_tag("acme-ml/support-tickets", tag="v1.0", repo_type="dataset")push_to_hub writes each split as one or more Parquet files under data/, named by split and shard index, and writes the YAML that maps them, so the repository loads with no further configuration. Shards keep any single file a manageable size for upload, resumable download and parallel reading. The tag gives the release a name that consumers can pin; pushing again later creates new commits on the main branch but never moves v1.0.
Configs and splits in YAML
When you lay files out yourself rather than through push_to_hub, the README's YAML decides what loads. The configs field lists subsets, each with data_files mapping split names to paths or glob patterns, and config_name is required even when there is only one. Mark one config default: true so a bare load_dataset call works.
---
license: other
language:
- en
task_categories:
- text-classification
configs:
- config_name: default
default: true
data_files:
- split: train
path: "data/train-*.parquet"
- split: validation
path: "data/validation-*.parquet"
- split: test
path: "data/test-*.parquet"
- config_name: rejects
data_files:
- split: train
path: "rejects/*.parquet"
---Without YAML, the library infers splits from names: directories or files containing train or training, validation, valid, val or dev, and test, testing, eval or evaluation, delimited by non-word characters; if nothing matches, everything becomes one train split. That inference is convenient and is also a trap, covered under failure modes. For anything you publish deliberately, write the YAML.
The dataset card is part of the data
The README body is the dataset card, and for a training dataset it is the documentation that prevents misuse. Write it for a reader who will train on the data without asking you anything. The useful sections are concrete:
- Source and time range: where the records came from and which period they cover.
- Processing: the schema, filters, rejection rules, deduplication method and the split rule, including the group key and salt.
- Label definitions with examples of borderline cases, and how labels were produced (human, heuristic, model).
- Known gaps and biases: under-represented languages, customers or classes; anything the data is not suitable for.
- Personal data handling: what was removed or pseudonymised and how; who may access the repository.
- Licence and permitted use, and a changelog with one line per tagged version.
The YAML front matter carries machine-readable metadata such as license, language and task_categories, which the Hub uses for search and display. Keep it accurate; an incorrect licence tag is worse than none.
Versioning and reproducibility
Every push is a git commit, so every state of the dataset has an immutable commit hash. Tags give human names to commits. Consumers should load by tag or hash, never by the moving main branch, and log the resolved commit with the training run.
from datasets import load_dataset
from huggingface_hub import HfApi
REV = "v1.0"
dd = load_dataset("acme-ml/support-tickets", revision=REV)
sha = HfApi().dataset_info("acme-ml/support-tickets", revision=REV).sha
run.log_params({"dataset": "acme-ml/support-tickets", "dataset_rev": REV, "dataset_sha": sha})Adopt a versioning rule and write it in the card. A reasonable one: a new major version when the schema, label set or split rule changes (scores are not comparable across it); a minor version when rows are added under the same rules; and never rewrite a published tag. Withdrawn data means a new version and a changelog entry.
A validation gate that runs on what consumers will see
Validate twice: once on the in-memory dataset before pushing, and once on the dataset loaded back from the Hub at the new tag, because that is what consumers will actually get. The second run catches upload and YAML mistakes that the first cannot.
def validate(dd, features, min_rows):
problems = []
for split, d in dd.items():
if d.features != features:
problems.append(f"{split}: schema {d.features} != expected")
if d.num_rows < min_rows[split]:
problems.append(f"{split}: {d.num_rows} rows < {min_rows[split]}")
empty = sum(1 for t in d["text"] if not t.strip())
if empty:
problems.append(f"{split}: {empty} empty texts")
groups = {s: set(d["customer_id"]) for s, d in dd.items()}
for a in groups:
for b in groups:
if a < b and groups[a] & groups[b]:
problems.append(f"customer overlap between {a} and {b}")
if problems:
raise ValueError("; ".join(problems))
validate(load_dataset("acme-ml/support-tickets", revision="v1.0"), FEATURES,
{"train": 80_000, "validation": 9_000, "test": 9_000})
Failure modes seen in practice
- A stray file becomes a split. With no YAML, a file called
test_notes.csvor a folder ofevalscreenshots is picked up by name-based split detection. Fix: explicit YAML configs. - Shards with drifting schemas. One shard has an integer column, another a string; loading fails with a cast error, sometimes only for the consumer who touches that shard. Fix: cast every shard to one declared Features before writing.
- Moving main. A training job loads main, someone pushes a fix mid-sweep, and half the runs train on different data. Fix: pin a tag or hash and log the sha.
- Parallel dedup that does not dedup. Stateful filters under
num_prockeep duplicates across workers. Fix: single pass or group-by. - Evaluation contamination. Public benchmark items present in your training split. Fix: hash-match benchmark texts against your corpus before release and record the result in the card.
- Personal data in free text. Names, emails and account numbers survive in ticket bodies even after column-level removal. Fix: a scrubbing pass with audited recall, private repository and access controls until it is proven.
Trade-offs worth deciding explicitly
| Decision | Option A | Option B |
|---|---|---|
| File format | Parquet: typed, columnar, compressed, readable everywhere | JSON Lines: human-diffable, but types are inferred and files are larger |
| Split rule | Group hash: stable, leak-resistant | Time cut: measures drift to the future, needs a date you trust |
| Dedup depth | Exact hash: cheap, explainable | Near-duplicate: catches more, costs compute and tuning |
What to do next
- Pick one dataset your team trains on and find out which revision each recent run actually used; if you cannot, start logging the commit sha.
- Write an explicit Features schema for it and cast every shard; send rejected rows to a reviewed reject file.
- Replace row-level splits with a salted hash of the real group key, and record the key and salt in the card.
- Add exact deduplication before splitting and an overlap check across splits after.
- Publish with push_to_hub, write explicit YAML configs, tag the release, and never move a tag.
- Run the validation gate against the dataset loaded back from the Hub at the tag, in CI, before anyone trains on it.
- Convert any script-based dataset you own to Parquet. Then read Transformers tokenizers and the Trainer to see how the published splits flow into training.