Every model repository on the Hugging Face Hub has a README.md at its root, and the Hub renders that file as the model card. It is tempting to write it last, from memory. That is a mistake for two reasons. First, the card has two readers: a person deciding whether the model is safe and suitable to use, and the Hub itself, which parses the YAML header to drive search filters, the license badge, the task widget, the model tree of fine-tunes and quantisations, the evaluation panel and the access gate. Second, a wrong card is worse than a thin one. A card that claims the wrong base model, an evaluation run on a different split, or a license the weights do not carry will be copied into other people's decisions.
This page explains the model card as an engineering artifact. It covers where the idea came from, the metadata contract key by key, how evaluation results are recorded, how gated access is configured, which prose sections carry real information, and how to generate and check cards automatically so the card is produced by the same pipeline that produced the weights. Key names and API calls were checked against the Hugging Face documentation in October 2026.
Where model cards come from
The model card was proposed by Margaret Mitchell and colleagues in the paper Model Cards for Model Reporting (arXiv 1810.03993, presented at FAT* 2019). Their argument was that a trained model ships without the context needed to use it responsibly: what it was built for, what data it saw, how it performs across the groups it will be used on, and where it should not be used. The paper proposed a short, standard document travelling with the model, covering intended use, training and evaluation data, results across groups, and caveats.
Hugging Face gave the idea a concrete form: Markdown for people and a YAML header for machine-readable facts. Its documentation still points to the Mitchell paper for the intended-use and limitations content. What the Hub added is the second reader: because the header is parsed, errors in it propagate into search results, derived-model trees and leaderboards.
The anatomy of a card
A model card is one file with two parts. The YAML header sits between two lines of three dashes at the very top. Everything after it is Markdown. The huggingface_hub library exposes the split directly: card.data is a ModelCardData object holding the header, card.text is the body without the header, and card.content is both together.
The metadata contract, key by key
Here is a complete header for a fine-tuned ticket classifier. Each key is explained below with what it drives on the Hub.
---
language:
- en
license: apache-2.0
library_name: transformers # required for transformers repos created after Aug 2024
pipeline_tag: text-classification # picks the widget and the task filter
base_model: distilbert/distilbert-base-uncased
base_model_relation: finetune # adapter | merge | quantized | finetune
datasets:
- acme/support-tickets-2026q3 # Hub dataset ids; plain names do not link
tags:
- customer-support
- routing
model-index:
- name: ticket-router-v3
results:
- task:
type: text-classification
dataset:
name: Support tickets, held-out Sept 2026
type: acme/support-tickets-2026q3
split: test
metrics:
- type: accuracy
value: 0.912
- type: f1
name: Macro F1
value: 0.874
---| Key | What the Hub does with it | Common mistake |
|---|---|---|
license | License badge and the license filter. Use a known identifier, or other with license_name and license_link. | Copying the base model's license when your training data adds restrictions. |
library_name | Which library snippet and loader the page shows. Repos created after August 2024 are no longer assumed to be transformers just because a config.json exists. | Leaving it out and getting no usage snippet. |
pipeline_tag | The task filter and which inference widget to show. Inferred from config.json for transformers models, but overridable. | A generation model tagged as classification by an old config. |
base_model | The model tree: the parent links to its fine-tunes, adapters, quantisations and merges. A list means a merge. | Pointing at the model you downloaded, not the one it was trained from. |
base_model_relation | States the relation explicitly: adapter, merge, quantized or finetune. The Hub infers it otherwise. | Letting a quantised fine-tune be inferred as a plain fine-tune. |
datasets | Links to Hub datasets and the dataset filter. Must be Hub ids to link. | Free-text dataset names that link nowhere. |
new_version | Links the page to a newer model; chains are followed to the latest. | Forgetting it, so users keep downloading v2. |
Two rules make the header trustworthy. Use identifiers, not prose, wherever the Hub expects an id: dataset ids, model ids and license identifiers are what make the links and filters work. And never set a key you cannot back with evidence from the training run. If you do not know the base model revision or the exact training set, say so in the body rather than guessing in the header.
Recording evaluation results
The Hub supports two ways of recording scores. The established one is the model-index block in the header, based on the Papers with Code model-index specification and shown above. Each result names a task type, a dataset with an id and optionally a split, and a list of metrics with a type, an optional display name and a value. A source entry can attribute the result to where it was computed. The Hub renders these in an evaluation panel on the model page.
The newer mechanism keeps results out of the README altogether: YAML files in a .eval_results/ folder of the model repo, each entry naming a Hub dataset that has been registered as a benchmark, a task id defined in that dataset's eval.yaml, and a value, with optional date, source and notes. Hugging Face labels this feature as work in progress, and benchmark registration is beta with an allow-list, so treat it as an addition to model-index rather than a replacement for now.
Whichever you use, the discipline is the same. A score is only meaningful with the dataset revision, split, prompt or preprocessing, and the code that computed it. Record the split explicitly, keep the evaluation script in the repo, and never mix harnesses in one table silently. For the wider question of how a model gets promoted once it has numbers, see the model registry article.
Licenses and gated access
If the model's terms are not a standard license, set license: other and give license_name and license_link; the link can be a file in the repo. Gating is switched on in the repository settings, not in the header, and can grant access automatically or require manual approval. The header then customises the request form:
---
license: other
license_name: acme-research-license
license_link: LICENSE.md # a file in the repo is fine
extra_gated_prompt: "Access is for evaluation and non-commercial research."
extra_gated_fields:
Company: text
Country: country
Intended use:
type: select
options: [Research, Evaluation, Other]
I agree to the license terms: checkbox
---Field types are text, checkbox, date_picker, country and select. The heading, description and button text can be changed with extra_gated_heading, extra_gated_description and extra_gated_button_content. A gate is a distribution control, not a security boundary: anyone who is approved can copy the weights, so the license terms are what actually bind them.
The body: what people need to read
The prose is where the Mitchell outline earns its place. A useful body says what the model is, what it is for, what it must not be used for, what data trained it, how it was evaluated and where it is weak, and how to run it. Most cards skip out-of-scope use, the section that prevents misuse. Write it concretely: this ticket router was trained on English support tickets from one product line, so it should not be used for other languages, for legal or safety escalations, or as the only decision on refunds.
Limitations should be measured, not generic. A sentence such as macro F1 falls from 0.87 to 0.71 on tickets shorter than ten words is worth more than a paragraph about possible bias. If you ran sliced evaluations, put the slices in a table. If you did not, say that you did not. Document the training procedure well enough to reproduce it. If the model came out of a structured fine-tuning pipeline such as the one in the fine-tuning operations pipeline, most of this already exists in the run manifest and should be pulled from there.
Generating the card in CI
Cards rot when they are written by hand. The fix is to generate the header and the factual parts of the body from the training run's own outputs, and keep only judgement sections, intended use and limitations, as reviewed human text. The huggingface_hub library provides everything needed: ModelCardData for the header, EvalResult for model-index entries, and ModelCard.from_template, which renders a Jinja template (Jinja2 must be installed) with the card data and any extra variables.
import json
from huggingface_hub import ModelCard, ModelCardData, EvalResult
run = json.load(open("artifacts/run_summary.json")) # written by the training job
evals = [
EvalResult(
task_type="text-classification",
dataset_type=run["eval_dataset_id"],
dataset_name=run["eval_dataset_name"],
dataset_split="test",
metric_type=name,
metric_value=round(value, 4),
)
for name, value in run["metrics"].items() # {"accuracy": 0.912, "f1": 0.874}
]
data = ModelCardData(
model_name=run["model_name"], # required when eval_results is set
language="en",
license="apache-2.0",
library_name="transformers",
pipeline_tag="text-classification",
base_model=run["base_model"],
datasets=[run["train_dataset_id"]],
tags=["customer-support", "routing"],
eval_results=evals,
)
card = ModelCard.from_template(
data,
template_path="templates/model_card.md", # your Jinja template
intended_use=open("docs/intended_use.md").read(), # written by people, reviewed once
limitations=open("docs/limitations.md").read(),
training_commit=run["git_sha"],
train_rows=run["train_rows"],
)
card.validate() # Hub-side metadata checks
card.save("dist/README.md")Note that model_name is required when you pass eval_results, because it becomes the name of the model-index entry. Then lint the result before it is published. The library's validate call checks the header against the Hub's rules; a small lint of your own enforces the house rules the Hub does not know about:
from huggingface_hub import ModelCard
REQUIRED_KEYS = {"license", "library_name", "pipeline_tag", "base_model", "datasets"}
REQUIRED_HEADINGS = ["## Intended use", "## Out-of-scope use", "## Limitations",
"## Training data", "## Evaluation"]
def lint(path):
card = ModelCard.load(path) # also accepts a repo id
meta = card.data.to_dict()
problems = [f"missing metadata: {k}" for k in sorted(REQUIRED_KEYS - meta.keys())]
problems += [f"missing section: {h}" for h in REQUIRED_HEADINGS if h not in card.text]
if meta.get("license") == "other" and not meta.get("license_link"):
problems.append("license: other needs license_name and license_link")
if "[More Information Needed]" in card.text:
problems.append("template placeholder left in body")
return problems
if __name__ == "__main__":
import sys
issues = lint(sys.argv[1])
print("\n".join(issues) or "card ok")
sys.exit(1 if issues else 0)Finally, push the card as a pull request so a person reviews it like code, and use metadata_update for small header fixes later. Updating an existing key needs overwrite=True; without it the call refuses to change a key that is already set.
Worked example: shipping ticket-router-v3
A support team fine-tunes a DistilBERT checkpoint on 180,000 labelled tickets and wants to publish it to the organisation's private Hub namespace. The training job writes run_summary.json with the base model id, the training and evaluation dataset ids, the git commit, row counts and metrics: accuracy 0.912 and macro F1 0.874 on the September held-out split. CI renders the card from the template, and the lint fails twice: the template's limitations placeholder is still in the body, and the Out-of-scope use heading is missing because the template renamed it.
An engineer fills in limitations from the sliced evaluation, and restores the out-of-scope heading. The card passes, is pushed as a pull request, and the reviewer notices that base_model points at the team's earlier v2 checkpoint, which v3 was not trained from. Two months later a quantised build is published as a separate repo with base_model set to v3 and base_model_relation: quantized, and it appears under v3 in the model tree.
Failure modes
- License laundering. A fine-tune declares a permissive license while the base model or training data carries restrictions. Check both before choosing the license key.
- Phantom evaluation. Numbers copied from the base model's card or a paper appear in model-index as if measured on this model. Only record what your harness produced.
- Unpinned provenance.
base_modelnames a repo but not the revision trained from; the parent is later updated and the claim silently changes meaning. Record the revision in the body. - Template residue. Default template placeholders remain in the published card and readers learn to ignore the whole thing.
- Stale card, fresh weights. Weights are re-uploaded after a retrain while the README is untouched. Generating both in one job prevents this.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Generate the card in CI | Header and numbers always match the weights | Template and lint to maintain; judgement sections still need people |
| model-index in the header | Mature, widely parsed, rendered on the page | Bloats the README; one schema for all benchmarks |
.eval_results/ files | Benchmark leaderboards, verified and community badges | Work in progress; only registered benchmark datasets |
| Gated access | Know who downloads; collect terms acceptance | Friction for users; not a technical protection |
What to do next
- Open the cards of the models you have published and check license, library_name, pipeline_tag, base_model and datasets against what you actually did.
- Add base_model_relation to every adapter, merge and quantised build so the model tree is correct.
- Write a house template with intended use, out-of-scope use, limitations, training data, evaluation and how to use, and render it with ModelCard.from_template.
- Make the training job write a run summary, and build ModelCardData and EvalResult objects from it rather than by hand.
- Add the lint script to CI and fail the release on missing keys, missing sections or template residue.
- Publish cards through push_to_hub with create_pr=True and review them like code.
- Run at least one sliced evaluation per release and put the weakest slice in the limitations section.
- If you choose a model for production rather than publish one, read its card against the provider selection guide and test its claims on your own data.