An AI bill of materials (AIBOM, also called an ML-BOM) is a machine-readable inventory of what an AI system is made of: the models, with exact hashes of their weights; the datasets they were trained or tuned on; the base models they descend from; the software libraries they load; and the services that call them. It extends the software bill of materials (SBOM) idea, which lists packages and versions, to the parts that make AI systems different: artifacts that are large binaries with lineage, and data that carries licences and personal information.

The value of an AIBOM shows up on a bad day. A base model is found to be backdoored, a library has a deserialisation flaw, or a data subject asks you to delete their records, and someone asks which production services are affected. Without an inventory, that is a week of messages. With one, it is a graph query. This article covers the format field by field, generates a document in code, queries it, records not-affected decisions, and lists the ways AIBOMs quietly become wrong. It focuses on the document itself; the surrounding intake process is in AI Supply Chain Security Program.

What an AIBOM is for

Be clear about what the document is for, because scope creep kills these projects. An AIBOM is an inventory with identity and relationships. It answers what is in this system, exactly which bytes, where they came from, and what depends on what. It is not a model card, which explains intended use and evaluation to a human reader (see Model Cards, in depth). It is not a datasheet, which documents how a dataset was collected (see Datasheets for Datasets). And it is not a security proof: a complete AIBOM for a poisoned model is still a poisoned model.

Write down the questions your AIBOM must answer before choosing fields. A good starting set is:

  1. Which services run model X, or any model fine-tuned from X?
  2. Which models were trained on dataset D, and which versions of D?
  3. Does any production model load weights through a format that executes code, such as pickle?
  4. What licence applies to each model and dataset, and does it allow our use?
  5. Is the model running in production byte-for-byte the one that was evaluated and approved?

Every field you add should serve one of these questions. Fields that serve none are paperwork, and paperwork goes stale first.

Choosing a format and the fields that matter

Two open standards cover AI artifacts. CycloneDX added machine-learning support in version 1.5 (2023): a component type machine-learning-model, a component type data, and an embedded modelCard object. Version 1.7 was released in October 2025 and is described by the project as backward compatible with 1.4 to 1.6; the examples here use 1.6. SPDX 3.0 (2024) takes a different route, with separate AI and Dataset profiles layered on its core model. Both work. Pick the one your existing SBOM tooling already reads, because the AIBOM is far more useful when it sits in the same store and query path as your software SBOMs.

The CycloneDX 1.6 fields that matter for AI, checked against the published JSON schema:

FieldWhat it recordsQuestion it answers
components[].typemachine-learning-model, data, library, othersWhat kind of thing is this?
hashes[]Algorithm and digest, e.g. SHA-256, one per weight fileIs this the approved artifact?
pedigree.ancestorsComponents this one was derived fromWhich models descend from X?
modelCard.modelParametersapproach, task, architectureFamily, datasets, inputs, outputsWhat was it trained on and for what?
modelCard.quantitativeAnalysisperformanceMetrics with type, value, sliceWhat was measured at approval?
modelCard.considerationsuseCases, technicalLimitations, fairnessAssessments and moreWhere must it not be used?
data[]type (e.g. dataset), classification, governanceWho owns the data, how sensitive is it?
dependencies[]ref and dependsOn edgesWhat breaks if this breaks?
vulnerabilities[].analysisstate and justificationAre we affected, and why not?

Two details save pain later. First, modelParameters.datasets accepts either an inline data object or a reference by bom-ref; use references, so a dataset is described once and linked from every model trained on it. Second, a model you call through a hosted API has no weights you can hash. Record it in the top-level services array with its provider, endpoint and the model version string you pin, and treat silent provider-side upgrades as a known gap.

Where the document comes from

The architecture principle is that the pipeline writes the AIBOM, never a person. Each stage that creates an artifact emits its part: the data snapshot job records the dataset and its digest; the training job records the base model and the dataset references; packaging hashes the final weight files; and deployment adds the service and its dependency edges. The pieces are merged into one document per release and stored with the artifact.

The AIBOM is produced by the pipeline at each stage, stored next to the artifact, and queried during incidentsData snapshothash, owner, classTraining jobbase model, code, configPackagingweights hashed per fileDeployservice pins modeldata componenttype: datapedigree + modelCardancestors, datasetsmodel componentSHA-256 hashesdependenciesservice to model to dataemitemitemitemitBOM storeone signed CycloneDX document per release, keyed by serialNumberAdvisory or takedownlibrary CVE, bad base model, data deletionGraph queryreverse dependsOn edgesAnswer + VEXaffected services, not_affected with reason
Each pipeline stage emits the fields it alone can know. The store holds one document per release, and incident response queries the dependency graph in reverse.

Generating it in code

Here is a generator for one fine-tuned classifier, using only the standard library. It hashes the weight files from disk, links the base model through pedigree, references the training dataset by bom-ref, and records two sliced metrics. Every key was checked against the CycloneDX 1.6 schema.

import hashlib
from pathlib import Path

def sha256_file(path: Path) -> str:
    h = hashlib.sha256()
    with open(path, "rb") as f:
        for block in iter(lambda: f.read(1 << 20), b""):
            h.update(block)
    return h.hexdigest()

def model_component(name, version, weight_files, base_ref, dataset_refs, metrics):
    return {
        "type": "machine-learning-model",
        "bom-ref": f"model:{name}@{version}",
        "name": name,
        "version": version,
        "hashes": [{"alg": "SHA-256", "content": sha256_file(p)} for p in weight_files],
        "licenses": [{"license": {"id": "Apache-2.0"}}],
        "pedigree": {"ancestors": [{"type": "machine-learning-model",
                                    "bom-ref": base_ref, "name": base_ref.split(":", 1)[1]}]},
        "modelCard": {
            "modelParameters": {
                "approach": {"type": "supervised"},
                "task": "text-classification",
                "architectureFamily": "transformer",
                "datasets": [{"ref": r} for r in dataset_refs],
                "inputs": [{"format": "string"}],
                "outputs": [{"format": "string"}],
            },
            "quantitativeAnalysis": {"performanceMetrics": [
                {"type": m, "value": str(v), "slice": s} for m, v, s in metrics]},
            "considerations": {
                "technicalLimitations": ["English only; weaker on tickets under 10 words"]},
        },
        "properties": [{"name": "internal:weights-format", "value": "safetensors"}],
    }

The surrounding document adds bomFormat CycloneDX, specVersion 1.6, a random serialNumber URN, a metadata.component naming the service, a dataset component of type data with its governance.owners, the libraries, and the edges:

"dependencies": [
  {"ref": "svc:ticket-api@5.1.0",
   "dependsOn": ["model:ticket-router@3.2.0", "pkg:pypi/torch@2.4.1"]},
  {"ref": "model:ticket-router@3.2.0",
   "dependsOn": ["data:tickets-2026q2@7"]}
]

Use package URLs (purl) for libraries so scanners can match advisories automatically. For models pulled from a hub, the purl specification defines a huggingface type; pin it to a commit revision, never a branch name, so the identifier names fixed bytes.

Querying the graph during an incident

The dependency edges point from consumer to dependency. Incident response needs the reverse direction: given a compromised artifact, find everything above it. A short traversal is enough:

def blast_radius(bom, ref):
    parents = {}
    for dep in bom.get("dependencies", []):
        for child in dep.get("dependsOn", []):
            parents.setdefault(child, set()).add(dep["ref"])
    seen, stack = set(), [ref]
    while stack:
        for p in parents.get(stack.pop(), ()):
            if p not in seen:
                seen.add(p)
                stack.append(p)
    return seen

blast_radius(bom, "data:tickets-2026q2@7")
# {'model:ticket-router@3.2.0', 'svc:ticket-api@5.1.0'}

Worked example. Legal receives a deletion request covering records in tickets-2026q2 version 7. Run the query across every stored document, not just the latest: older model versions may still be serving in another region. The result lists the models trained on that snapshot and the services that load them. For pedigree, run a second pass over pedigree.ancestors, because a model fine-tuned from an affected model inherits the problem without depending on the dataset directly. The answer is a list of owners to page and models to retrain or retire, produced in minutes. In a real store, load documents into a graph database or a pair of SQL tables (nodes and edges) rather than scanning JSON on every query.

Recording not-affected decisions

Most advisories against your dependencies will not affect you, and the expensive part is proving that repeatedly. CycloneDX lets the document carry the decision, in the style of a VEX (vulnerability exploitability exchange) statement. The analysis.state values in 1.6 are resolved, resolved_with_pedigree, exploitable, in_triage, false_positive and not_affected; a not-affected decision should carry a justification such as code_not_reachable or protected_by_mitigating_control.

"vulnerabilities": [{
  "id": "INT-2026-014",
  "source": {"name": "internal-ml-advisories"},
  "affects": [{"ref": "pkg:pypi/torch@2.4.1"}],
  "analysis": {
    "state": "not_affected",
    "justification": "code_not_reachable",
    "detail": "Advisory concerns loading pickle checkpoints; this service loads safetensors only."
  }
}]

The detail text is the part auditors read, so make it a falsifiable claim. Then back it with a check: the admission controller should refuse pickle-format weights for that service, so the justification cannot silently become untrue.

Failure modes

AIBOMs fail by drifting away from reality while still looking complete. The common ways:

  • Hand-written documents. Anything typed by a person is wrong within a release or two. Generate every field from the artifact or the job that produced it.
  • Hashing a name instead of bytes. A digest of a repository name or a config file proves nothing about the weights. Hash each weight file, including every shard of a sharded checkpoint, and the tokenizer files.
  • Missing hidden models. Retrieval systems run an embedding model, often a reranker, and sometimes a guard or moderation classifier. Each is a model with its own lineage; leaving them out leaves out much of the system.
  • Mutable dataset references. A dataset named only by path or table changes underneath you. Record a snapshot version and a content digest or manifest hash.
  • Only the latest document is queried. Rollbacks, canaries and regional lag mean several versions serve at once. Keep every released document and record which one each deployment uses.
  • Sensitive detail in the BOM. Internal bucket paths, customer names in dataset names and personal contact details leak if the document is shared with customers. Keep an internal and an external rendering.
  • A document nobody verifies. If the runtime never checks that loaded weights match the recorded hashes, the AIBOM describes intent, not reality.

Operating it

Operating an AIBOM is mostly about making the right path the easy one.

  • Generate in CI at each stage and fail the build when a required field is missing: weight hashes, dataset references, licence and owner.
  • Validate against the official schema with a schema validator or the CycloneDX command-line tool before storing.
  • Sign the document and store it in the same registry as the model, keyed by serialNumber, so the two cannot be separated.
  • At load time, recompute weight hashes and compare with the document; refuse to serve on mismatch.
  • Diff the AIBOM between releases and show the diff in review: a new dataset or a changed base model deserves human attention even when tests pass.
  • Feed the store into your incident runbook (see LLM incident response) so the blast-radius query is a documented first step.

Trade-offs

Granularity is the central trade-off. Per-file hashes and per-snapshot datasets make queries exact but produce large documents and more pipeline work; a single hash per model and a dataset name are cheap but cannot prove anything. Choosing CycloneDX or SPDX matters less than generating automatically and verifying at load. Embedding the model card in the BOM keeps one artifact, but long human prose bloats a machine document; many teams keep short structured fields in the BOM and link the full card through externalReferences of type model-card. Finally, an AIBOM shared with customers builds trust but exposes your supply chain, so decide what the external version omits.

What to do next

  • List the five questions your AIBOM must answer and map each to the field that answers it.
  • Inventory every model in one production service, including embedding, reranking and guard models, and every hosted API model it calls.
  • Add a generator to the training and packaging jobs that hashes each weight and tokenizer file and references datasets by snapshot version.
  • Validate the output against the CycloneDX 1.6 schema in CI and fail on missing hashes, licences or owners.
  • Store one signed document per release with the artifact, and verify hashes at load time.
  • Run the blast-radius query in a tabletop exercise against a fake deletion request and time how long the answer takes.
Key takeaway: An AI bill of materials is an inventory with identity and relationships: hashed weight files, referenced dataset snapshots, base-model pedigree and service-to-model-to-data dependency edges. Its value is answering which services are affected in minutes when a model, library or dataset goes bad, and that only works if the pipeline generates it, the store keeps every release, and the runtime verifies the recorded hashes before loading.