A model card is a short document that travels with a trained model and says what it is, what it was trained and evaluated on, what it is for, what it should not be used for, and where it fails. The idea was formalised by Mitchell and colleagues in "Model Cards for Model Reporting" (FAT* 2019), and it is now the default README format on the Hugging Face Hub, an input to AI bills of materials, and part of what regulators expect providers to hand to the people who build on their models.

Most writing treats model cards as an ethics or transparency exercise. This article treats them as a security artifact. A card is a set of claims that downstream teams act on: they pick a model because the card says it was evaluated against prompt injection, they ship it because the license field says commercial use is allowed, they trust it because the base model is a known one. If the card is wrong, stale or detached from the weights it describes, every one of those decisions is wrong too. The sections below cover what a card contains, the machine-readable format, how to bind a card to the exact bytes it describes, a CI gate with working code, a worked example, how cards mislead, and a checklist.

Training rundata manifest, configEvaluation suitequality + security evalsCard generatorYAML metadata + proseCI policy gatefields, hashes, formatsSign and registercard + weights + ML-BOMpassModel registryimmutable versionDeploy admissionverify signature + hashProductionintended use onlyIncidents, new evals, driftcard revision, new version, never silent editre-evaluateThe card is generated from the same pipeline that produced the weights, and is checked, signed and verified with them.A card written by hand after the fact describes what someone remembers, not what shipped.
A model card in the release pipeline: generated from run records, gated in CI, signed with the weights, verified at deploy, and revised rather than silently edited.

What a model card contains

The original paper proposed nine sections, and they are still a good skeleton: model details (who built it, version, type, license, contact), intended use and out-of-scope uses, factors (the groups, environments or input conditions performance may vary across), metrics, evaluation data, training data, quantitative analyses broken down by those factors, ethical considerations, and caveats and recommendations. The key design decision was disaggregation: a single headline accuracy hides the subgroup where the model fails, so the card reports results per factor.

For large language models the same structure needs a few additions that matter for security. A card should name the base model and the exact revision a fine-tune started from, because the base model's weaknesses are inherited. It should list the safety and security evaluations run, such as jailbreak resistance, prompt-injection behaviour when the model reads tool output, data-extraction probes and refusal rates, with the evaluation set version and date. It should say what serialisation format the weights ship in, what system prompt or guardrails the reported numbers assumed, and what the model was not tested for. Frontier labs now publish longer system cards that cover the model plus its deployment safeguards; the same principle applies at smaller scale to any model you release internally.

A related artifact, the datasheet for datasets (Gebru et al.), documents data the way a card documents a model. A good card links to the datasheets of its training and evaluation sets rather than re-describing them.

What a model card contains

The original paper proposed nine sections, and they are still a good skeleton: model details (who built it, version, type, license, contact), intended use and out-of-scope uses, factors (the groups, environments or input conditions performance may vary across), metrics, evaluation data, training data, quantitative analyses broken down by those factors, ethical considerations, and caveats and recommendations. The key design decision was disaggregation: a single headline accuracy hides the subgroup where the model fails, so the card reports results per factor.

For large language models the same structure needs a few additions that matter for security. A card should name the base model and the exact revision a fine-tune started from, because the base model's weaknesses are inherited. It should list the safety and security evaluations run, such as jailbreak resistance, prompt-injection behaviour when the model reads tool output, data-extraction probes and refusal rates, with the evaluation set version and date. It should say what serialisation format the weights ship in, what system prompt or guardrails the reported numbers assumed, and what the model was not tested for. Frontier labs now publish longer system cards that cover the model plus its deployment safeguards; the same principle applies at smaller scale to any model you release internally.

A related artifact, the datasheet for datasets (Gebru et al.), documents data the way a card documents a model. A good card links to the datasheets of its training and evaluation sets rather than re-describing them.

The machine-readable format

On the Hugging Face Hub a model card is the repository's README.md: a YAML metadata block between --- lines followed by Markdown prose. The metadata is what tools read. Fields such as license, language, datasets, base_model, pipeline_tag and tags drive search and filtering, and model-index carries structured evaluation results. The huggingface_hub library exposes this as ModelCard (with .data for metadata and .text for the body), ModelCardData and EvalResult, so cards can be generated and checked in code instead of edited by hand.

---
language: en
license: apache-2.0
base_model: example-org/base-7b          # exact repo id and revision you started from
datasets:
  - internal/support-tickets-2026q3      # pointer to a data card, not the data
pipeline_tag: text-generation
tags: [customer-support, lora-merged, internal]
weights_sha256:                          # house convention: one entry per shipped file
  model-00001-of-00002.safetensors: 3f1c...e9
  model-00002-of-00002.safetensors: a07d...41
model-index:
  - name: support-assistant-v3
    results:
      - task: {type: text-generation}
        dataset: {type: internal/support-eval-v4, name: Support eval v4}
        metrics:
          - {type: accuracy, value: 0.87}
---

The weights_sha256 key is not a Hub standard. It is a house convention, and unknown keys survive a load and save round trip through ModelCardData. It exists because nothing in the standard format ties the card to the bytes it describes: the same README can sit next to any weights. For supply-chain tooling, CycloneDX 1.5 added a machine-learning-model component type with a modelCard field, so the same information can be embedded in an ML bill of materials alongside dataset components and their hashes.

Binding the card to what it describes

Treat the card as a claim about specific bytes, and make the claim checkable. Three bindings matter.

  • Card to weights. Record a SHA-256 for every weight file in the card metadata and refuse to load or deploy if any file's hash differs. This catches the most common failure, a card copied forward from an earlier version, and also catches weights swapped in a registry or cache.
  • Card to evaluations. Every number in the card should come from an evaluation run with an ID, a dataset version and the hash of the weights it ran on. The generator copies numbers from the run record; humans write only the prose around them. A number with no run ID is an anecdote.
  • Card to signer. Sign the card, the weights manifest and the ML-BOM together with the same identity your pipeline uses for other release artifacts, and verify the signature at deploy admission. A signature on the weights alone lets an attacker change the card, which changes what users believe they deployed.

Version cards with the model. When a new evaluation reveals a weakness, publish a new card revision with a changelog entry rather than silently editing the old one; consumers who approved a model against the earlier card need to know the facts changed.

A CI gate for model cards

The script below is a CI gate for a model directory about to be registered. It loads the card with huggingface_hub, requires a minimal set of metadata keys and prose sections, requires machine-readable evaluation results, checks that every shipped .safetensors file matches a declared hash and that no declared file is missing, and rejects pickle-based weight formats, which can execute code on load.

import hashlib
from pathlib import Path
from huggingface_hub import ModelCard

REQUIRED_META = ["license", "base_model", "datasets"]
REQUIRED_SECTIONS = ["## Intended use", "## Out-of-scope use", "## Training data",
                     "## Evaluation", "## Security evaluation", "## Limitations"]
PICKLE_SUFFIXES = {".bin", ".pt", ".pth", ".pkl", ".ckpt"}

def sha256_file(path, chunk=1 << 20):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        for block in iter(lambda: f.read(chunk), b""):
            h.update(block)
    return h.hexdigest()

def check_model_dir(model_dir):
    model_dir = Path(model_dir)
    card = ModelCard.load(model_dir / "README.md")
    meta = card.data.to_dict()
    body = card.text.lower()
    problems = [f"missing metadata: {k}" for k in REQUIRED_META if not meta.get(k)]
    problems += [f"missing section: {s}" for s in REQUIRED_SECTIONS if s.lower() not in body]
    if not card.data.eval_results:
        problems.append("no machine-readable evaluation results (model-index)")

    declared = meta.get("weights_sha256") or {}
    shipped = {f.name: f for f in model_dir.glob("*.safetensors")}
    for name, f in shipped.items():
        if declared.get(name) != sha256_file(f):
            problems.append(f"weights not bound to card: {name}")
    for name in declared.keys() - shipped.keys():
        problems.append(f"card declares a file that is not shipped: {name}")
    for f in model_dir.iterdir():
        if f.suffix in PICKLE_SUFFIXES:
            problems.append(f"pickle-format weights present: {f.name}")
    return problems

if __name__ == "__main__":
    import sys
    issues = check_model_dir(sys.argv[1])
    for i in issues:
        print("FAIL", i)
    sys.exit(1 if issues else 0)

Run it after the card generator and before signing. Keep the required section list short and enforce it strictly; a gate that demands twenty sections teaches people to write "N/A" twenty times. The sections chosen here are the ones a reviewer cannot infer from anywhere else: what the model is for, what it must not be used for, where its data came from, how it was evaluated, what security testing found and what is known to be broken. For a model trained from scratch, drop base_model from the required list or require the literal value none.

Worked example: a support assistant

Consider a team that fine-tunes an open 7B base model on 40,000 resolved support tickets to draft replies for human agents. Their first card, written by hand, said: "Fine-tuned for customer support. Accuracy 87%. Apache 2.0." Walk it through the questions a security reviewer asks.

QuestionFirst cardRevised card
Which base, which revision?Not statedRepo id and commit; base model's own card linked
Intended useCustomer supportDrafts for human agents; never sent to customers without review
Out of scopeNot statedAutonomous replies, refunds or account changes, legal or medical advice
Training dataSupport ticketsDatasheet link; PII scrubbed with named tool and version; date range; opt-out handling
Evaluation87%87% reply acceptability on eval v4 (1,200 held-out tickets, run ID, weight hash)
Security evaluationNoneInjection set: share of tickets with embedded instructions that changed behaviour; extraction probes for memorised PII; both with run IDs
LicenseApache 2.0Base model license checked for fine-tune and commercial terms; data rights confirmed with legal

The revision exposed two real problems. The license line was the team's own choice, but the base model's license carried use restrictions that applied to derivatives, so the field misstated what consumers could do. And the extraction probe found that the model reproduced fragments of phone numbers from tickets the scrubber had missed, which changed the out-of-scope list and triggered a data fix and retrain. Neither issue was visible from the first card, and both would have been inherited by every team that picked the model up because "it has a card". Note that the revised card states the security results as measured rates on a named set; it does not claim the model is "safe against prompt injection", which no evaluation can show.

How model cards mislead

  • Drift. The card describes version 2 and the weights are version 3. Hash binding catches this mechanically; nothing else does reliably.
  • Unverifiable numbers. Benchmarks reported without the evaluation set version, prompt format or decoding settings cannot be reproduced. Contaminated benchmarks, where test items leaked into training data, produce honest-looking numbers that mean nothing. Prefer internal held-out sets and state how contamination was checked.
  • License laundering. A fine-tune declares a permissive license that the base model or training data does not allow. The card field is a claim, not a grant; check the lineage.
  • Impersonation on public hubs. Typosquatted repositories copy a well-known card verbatim and ship different, sometimes malicious, weights. A polished card proves nothing about the publisher. Pin repository IDs and revisions and verify hashes.
  • Cards as injection carriers. Agents and coding assistants increasingly read model cards to choose or configure models. A card is untrusted text; any tool that feeds it to an LLM must treat it as data, the same as a web page.
  • Selective omission. A card that lists only the evaluations the model passed is technically accurate and misleading. Require a security evaluation section even when the honest content is "not evaluated for X".

Operating model cards

Make the card part of the release, not documentation about it. Generate it in the training pipeline from run metadata, gate it in CI, store it in the registry with the weights, and expose it at the serving layer so an operator can ask a running endpoint which card it serves. Map card fields to the controls you already report on: the intended and out-of-scope uses feed the risk assessment, the security evaluation section feeds the red-team backlog, and the data lineage feeds privacy reviews.

For regulatory work, a card is useful evidence but rarely sufficient on its own. Under the EU AI Act, providers of general-purpose AI models have documentation obligations toward the AI Office and toward downstream providers, applicable since 2 August 2025, and high-risk systems need fuller technical documentation. A well-structured card covers much of the content, so keep the fields aligned and let the legal team map them, rather than maintaining two documents that drift apart. NIST's AI Risk Management Framework similarly treats documentation of intended use and known limitations as part of its Map function.

Finally, review cards the way you review code: a named owner, a reviewer from security for any model that touches customer data or takes actions, and a rule that changing a number in the card requires the run record that produced it.

What to do next

Related reading on this site: model and data provenance, data governance for AI systems, EU AI Act compliance, the NIST AI RMF and preparing for an AI audit.

  1. Inventory every model you serve and record which have a card, which card version, and whether anyone can show the card matches the deployed weights.
  2. Add the six required sections and the metadata keys above to a card template, with a security evaluation section that must say "not evaluated" explicitly when that is the truth.
  3. Generate cards from training and evaluation run records so that every number carries a run ID and weight hash.
  4. Put the gate script, adapted to your formats, in CI before model registration, and fail the build on pickle-format weights.
  5. Sign card, weights manifest and ML-BOM together, and verify the signature and hashes at deploy admission.
  6. For every third-party model you adopt, read its card critically: check the base model's license, look for evaluation set versions, and pin the repository revision you reviewed.
  7. Schedule a card review whenever a new red-team finding or incident touches a model, and publish the revision with a changelog.
Key takeaway: A model card is a set of claims that downstream teams act on, so treat it as a security artifact. Generate it from training and evaluation records, require intended use, out-of-scope use, data lineage, evaluation and security sections, bind it to weight hashes, and sign and verify it with the weights. Read third-party cards sceptically: check lineage and licenses, pin revisions, and treat card text as untrusted input.