Every model repository on the Hugging Face Hub has a README.md at its root, and the Hub renders that file as the model card. It is tempting to write it last, from memory. That is a mistake for two reasons. First, the card has two readers: a person deciding whether the model is safe and suitable to use, and the Hub itself, which parses the YAML header to drive search filters, the license badge, the task widget, the model tree of fine-tunes and quantisations, the evaluation panel and the access gate. Second, a wrong card is worse than a thin one. A card that claims the wrong base model, an evaluation run on a different split, or a license the weights do not carry will be copied into other people's decisions.

This page explains the model card as an engineering artifact. It covers where the idea came from, the metadata contract key by key, how evaluation results are recorded, how gated access is configured, which prose sections carry real information, and how to generate and check cards automatically so the card is produced by the same pipeline that produced the weights. Key names and API calls were checked against the Hugging Face documentation in October 2026.

Advertisement

Where model cards come from

The model card was proposed by Margaret Mitchell and colleagues in the paper Model Cards for Model Reporting (arXiv 1810.03993, presented at FAT* 2019). Their argument was that a trained model ships without the context needed to use it responsibly: what it was built for, what data it saw, how it performs across the groups it will be used on, and where it should not be used. The paper proposed a short, standard document travelling with the model, covering intended use, training and evaluation data, results across groups, and caveats.

Hugging Face gave the idea a concrete form: Markdown for people and a YAML header for machine-readable facts. Its documentation still points to the Mitchell paper for the intended-use and limitations content. What the Hub added is the second reader: because the header is parsed, errors in it propagate into search results, derived-model trees and leaderboards.

The anatomy of a card

A model card is one file with two parts. The YAML header sits between two lines of three dashes at the very top. Everything after it is Markdown. The huggingface_hub library exposes the split directly: card.data is a ModelCardData object holding the header, card.text is the body without the header, and card.content is both together.

One README.md, two readers: people and the HubYAML headerbetween --- lines at the topMarkdown bodyuses, limits, training, evalsSearch filterslicense, task, languageModel treebase_model, relationWidget and loaderspipeline_tag, library_nameEval results panelmodel-index or .eval_results/Access gate formextra_gated_* keysHuman readerdecides whether to use itread and judgedCI: build ModelCardData from the training run -> render template -> validate -> push_to_hub(create_pr=True)the card is generated from facts, not typed from memory
The header feeds Hub features; the body is read by people. Generating both from the training run keeps them consistent.
Advertisement

The metadata contract, key by key

Here is a complete header for a fine-tuned ticket classifier. Each key is explained below with what it drives on the Hub.

---
language:
  - en
license: apache-2.0
library_name: transformers          # required for transformers repos created after Aug 2024
pipeline_tag: text-classification   # picks the widget and the task filter
base_model: distilbert/distilbert-base-uncased
base_model_relation: finetune       # adapter | merge | quantized | finetune
datasets:
  - acme/support-tickets-2026q3     # Hub dataset ids; plain names do not link
tags:
  - customer-support
  - routing
model-index:
  - name: ticket-router-v3
    results:
      - task:
          type: text-classification
        dataset:
          name: Support tickets, held-out Sept 2026
          type: acme/support-tickets-2026q3
          split: test
        metrics:
          - type: accuracy
            value: 0.912
          - type: f1
            name: Macro F1
            value: 0.874
---
KeyWhat the Hub does with itCommon mistake
licenseLicense badge and the license filter. Use a known identifier, or other with license_name and license_link.Copying the base model's license when your training data adds restrictions.
library_nameWhich library snippet and loader the page shows. Repos created after August 2024 are no longer assumed to be transformers just because a config.json exists.Leaving it out and getting no usage snippet.
pipeline_tagThe task filter and which inference widget to show. Inferred from config.json for transformers models, but overridable.A generation model tagged as classification by an old config.
base_modelThe model tree: the parent links to its fine-tunes, adapters, quantisations and merges. A list means a merge.Pointing at the model you downloaded, not the one it was trained from.
base_model_relationStates the relation explicitly: adapter, merge, quantized or finetune. The Hub infers it otherwise.Letting a quantised fine-tune be inferred as a plain fine-tune.
datasetsLinks to Hub datasets and the dataset filter. Must be Hub ids to link.Free-text dataset names that link nowhere.
new_versionLinks the page to a newer model; chains are followed to the latest.Forgetting it, so users keep downloading v2.

Two rules make the header trustworthy. Use identifiers, not prose, wherever the Hub expects an id: dataset ids, model ids and license identifiers are what make the links and filters work. And never set a key you cannot back with evidence from the training run. If you do not know the base model revision or the exact training set, say so in the body rather than guessing in the header.

Recording evaluation results

The Hub supports two ways of recording scores. The established one is the model-index block in the header, based on the Papers with Code model-index specification and shown above. Each result names a task type, a dataset with an id and optionally a split, and a list of metrics with a type, an optional display name and a value. A source entry can attribute the result to where it was computed. The Hub renders these in an evaluation panel on the model page.

The newer mechanism keeps results out of the README altogether: YAML files in a .eval_results/ folder of the model repo, each entry naming a Hub dataset that has been registered as a benchmark, a task id defined in that dataset's eval.yaml, and a value, with optional date, source and notes. Hugging Face labels this feature as work in progress, and benchmark registration is beta with an allow-list, so treat it as an addition to model-index rather than a replacement for now.

Whichever you use, the discipline is the same. A score is only meaningful with the dataset revision, split, prompt or preprocessing, and the code that computed it. Record the split explicitly, keep the evaluation script in the repo, and never mix harnesses in one table silently. For the wider question of how a model gets promoted once it has numbers, see the model registry article.

Licenses and gated access

If the model's terms are not a standard license, set license: other and give license_name and license_link; the link can be a file in the repo. Gating is switched on in the repository settings, not in the header, and can grant access automatically or require manual approval. The header then customises the request form:

---
license: other
license_name: acme-research-license
license_link: LICENSE.md                 # a file in the repo is fine
extra_gated_prompt: "Access is for evaluation and non-commercial research."
extra_gated_fields:
  Company: text
  Country: country
  Intended use:
    type: select
    options: [Research, Evaluation, Other]
  I agree to the license terms: checkbox
---

Field types are text, checkbox, date_picker, country and select. The heading, description and button text can be changed with extra_gated_heading, extra_gated_description and extra_gated_button_content. A gate is a distribution control, not a security boundary: anyone who is approved can copy the weights, so the license terms are what actually bind them.

The body: what people need to read

The prose is where the Mitchell outline earns its place. A useful body says what the model is, what it is for, what it must not be used for, what data trained it, how it was evaluated and where it is weak, and how to run it. Most cards skip out-of-scope use, the section that prevents misuse. Write it concretely: this ticket router was trained on English support tickets from one product line, so it should not be used for other languages, for legal or safety escalations, or as the only decision on refunds.

Limitations should be measured, not generic. A sentence such as macro F1 falls from 0.87 to 0.71 on tickets shorter than ten words is worth more than a paragraph about possible bias. If you ran sliced evaluations, put the slices in a table. If you did not, say that you did not. Document the training procedure well enough to reproduce it. If the model came out of a structured fine-tuning pipeline such as the one in the fine-tuning operations pipeline, most of this already exists in the run manifest and should be pulled from there.

Generating the card in CI

Cards rot when they are written by hand. The fix is to generate the header and the factual parts of the body from the training run's own outputs, and keep only judgement sections, intended use and limitations, as reviewed human text. The huggingface_hub library provides everything needed: ModelCardData for the header, EvalResult for model-index entries, and ModelCard.from_template, which renders a Jinja template (Jinja2 must be installed) with the card data and any extra variables.

import json
from huggingface_hub import ModelCard, ModelCardData, EvalResult

run = json.load(open("artifacts/run_summary.json"))   # written by the training job

evals = [
    EvalResult(
        task_type="text-classification",
        dataset_type=run["eval_dataset_id"],
        dataset_name=run["eval_dataset_name"],
        dataset_split="test",
        metric_type=name,
        metric_value=round(value, 4),
    )
    for name, value in run["metrics"].items()          # {"accuracy": 0.912, "f1": 0.874}
]

data = ModelCardData(
    model_name=run["model_name"],                       # required when eval_results is set
    language="en",
    license="apache-2.0",
    library_name="transformers",
    pipeline_tag="text-classification",
    base_model=run["base_model"],
    datasets=[run["train_dataset_id"]],
    tags=["customer-support", "routing"],
    eval_results=evals,
)

card = ModelCard.from_template(
    data,
    template_path="templates/model_card.md",            # your Jinja template
    intended_use=open("docs/intended_use.md").read(),   # written by people, reviewed once
    limitations=open("docs/limitations.md").read(),
    training_commit=run["git_sha"],
    train_rows=run["train_rows"],
)
card.validate()                                         # Hub-side metadata checks
card.save("dist/README.md")

Note that model_name is required when you pass eval_results, because it becomes the name of the model-index entry. Then lint the result before it is published. The library's validate call checks the header against the Hub's rules; a small lint of your own enforces the house rules the Hub does not know about:

from huggingface_hub import ModelCard

REQUIRED_KEYS = {"license", "library_name", "pipeline_tag", "base_model", "datasets"}
REQUIRED_HEADINGS = ["## Intended use", "## Out-of-scope use", "## Limitations",
                     "## Training data", "## Evaluation"]

def lint(path):
    card = ModelCard.load(path)          # also accepts a repo id
    meta = card.data.to_dict()
    problems = [f"missing metadata: {k}" for k in sorted(REQUIRED_KEYS - meta.keys())]
    problems += [f"missing section: {h}" for h in REQUIRED_HEADINGS if h not in card.text]
    if meta.get("license") == "other" and not meta.get("license_link"):
        problems.append("license: other needs license_name and license_link")
    if "[More Information Needed]" in card.text:
        problems.append("template placeholder left in body")
    return problems

if __name__ == "__main__":
    import sys
    issues = lint(sys.argv[1])
    print("\n".join(issues) or "card ok")
    sys.exit(1 if issues else 0)

Finally, push the card as a pull request so a person reviews it like code, and use metadata_update for small header fixes later. Updating an existing key needs overwrite=True; without it the call refuses to change a key that is already set.

Worked example: shipping ticket-router-v3

A support team fine-tunes a DistilBERT checkpoint on 180,000 labelled tickets and wants to publish it to the organisation's private Hub namespace. The training job writes run_summary.json with the base model id, the training and evaluation dataset ids, the git commit, row counts and metrics: accuracy 0.912 and macro F1 0.874 on the September held-out split. CI renders the card from the template, and the lint fails twice: the template's limitations placeholder is still in the body, and the Out-of-scope use heading is missing because the template renamed it.

An engineer fills in limitations from the sliced evaluation, and restores the out-of-scope heading. The card passes, is pushed as a pull request, and the reviewer notices that base_model points at the team's earlier v2 checkpoint, which v3 was not trained from. Two months later a quantised build is published as a separate repo with base_model set to v3 and base_model_relation: quantized, and it appears under v3 in the model tree.

Failure modes

  • License laundering. A fine-tune declares a permissive license while the base model or training data carries restrictions. Check both before choosing the license key.
  • Phantom evaluation. Numbers copied from the base model's card or a paper appear in model-index as if measured on this model. Only record what your harness produced.
  • Unpinned provenance. base_model names a repo but not the revision trained from; the parent is later updated and the claim silently changes meaning. Record the revision in the body.
  • Template residue. Default template placeholders remain in the published card and readers learn to ignore the whole thing.
  • Stale card, fresh weights. Weights are re-uploaded after a retrain while the README is untouched. Generating both in one job prevents this.

Trade-offs

ChoiceGainCost
Generate the card in CIHeader and numbers always match the weightsTemplate and lint to maintain; judgement sections still need people
model-index in the headerMature, widely parsed, rendered on the pageBloats the README; one schema for all benchmarks
.eval_results/ filesBenchmark leaderboards, verified and community badgesWork in progress; only registered benchmark datasets
Gated accessKnow who downloads; collect terms acceptanceFriction for users; not a technical protection

What to do next

  1. Open the cards of the models you have published and check license, library_name, pipeline_tag, base_model and datasets against what you actually did.
  2. Add base_model_relation to every adapter, merge and quantised build so the model tree is correct.
  3. Write a house template with intended use, out-of-scope use, limitations, training data, evaluation and how to use, and render it with ModelCard.from_template.
  4. Make the training job write a run summary, and build ModelCardData and EvalResult objects from it rather than by hand.
  5. Add the lint script to CI and fail the release on missing keys, missing sections or template residue.
  6. Publish cards through push_to_hub with create_pr=True and review them like code.
  7. Run at least one sliced evaluation per release and put the weakest slice in the limitations section.
  8. If you choose a model for production rather than publish one, read its card against the provider selection guide and test its claims on your own data.
Key takeaway: A Hugging Face model card is the README.md at the root of a model repo: a YAML header the Hub parses and a Markdown body people read. The header drives the license badge, filters, usage snippet, widget, model tree, evaluation panel and access gate, so every key must be an identifier you can back with evidence. Record scores with their dataset, split and harness, in model-index or the newer work-in-progress .eval_results folder. Write concrete intended-use, out-of-scope and measured limitation sections. Generate the card from the training run with huggingface_hub, lint it in CI, and publish it as a reviewed pull request.