An AI bill of materials (AIBOM) is a machine-readable inventory of what went into an AI system: the model files, the base model it was fine-tuned from, the datasets, the training and inference software, and the declared properties of each. Most writing about AIBOMs takes the producer's point of view and asks which fields to fill in. This article takes the other side of the transaction. You have just downloaded a model you did not train, it comes with an AIBOM, and you must decide whether it may load in production. What does the document actually prove, how do you check that it describes the bytes on your disk, and which claims can you turn into an automatic allow-or-reject decision?
The short answer is that an AIBOM proves very little on its own. It becomes evidence only when three things hold: every file it lists is pinned by a cryptographic digest, the document itself is signed by an identity you trust, and your admission gate recomputes the digests rather than believing them. The rest of this article builds that gate, maps the two main formats against each other, and walks through a worked rejection. For the producer side (generating a CycloneDX ML-BOM in CI, blast-radius queries and VEX), read AI Bill of Materials, in depth first; this page assumes you already know what an SBOM is, as covered in SBOMs and supply chain security.
What an AIBOM can prove
Start by separating the claims an AIBOM makes into three classes, because each needs a different kind of check.
| Claim class | Example | How you can check it |
|---|---|---|
| Identity of bytes | model.safetensors has SHA-256 9f2c... | Recompute locally; exact match or reject |
| Provenance | fine-tuned from base model X on dataset Y by org Z | Signature by Z; a signed build attestation; X and Y have their own BOMs |
| Declared properties | license Apache-2.0, no sensitive personal data, low safety risk | Cannot be verified from bytes; trust the signer, then apply policy |
Identity claims are the only ones you can verify mechanically, and they anchor everything else. Provenance claims are only as good as the signer and whatever build system produced the attestation. Declared properties, such as whether a training set contained personal information, are assertions by a party; the BOM makes them explicit, attributable and auditable, which is valuable, but it does not make them true. A gate that treats all three classes the same way either rejects everything or trusts everything.
SPDX 3.0 and CycloneDX, side by side
Two formats dominate. CycloneDX added a machine-learning BOM in version 1.5, with a machine-learning-model component type, a data component type and a model card object. SPDX 3.0, released in 2024, reorganised the specification into profiles and added an AI profile and a Dataset profile. The SPDX AI profile defines an AIPackage class whose optional properties include typeOfModel, domain, autonomyType, hyperparameter, informationAboutTraining, informationAboutApplication, limitation, metric, metricDecisionThreshold, modelDataPreprocessing, modelExplainability, safetyRiskAssessment, standardCompliance, useSensitivePersonalInformation and a group of energy-consumption fields. The Dataset profile adds a DatasetPackage with properties such as datasetType, dataCollectionProcess, intendedUse, knownBias and datasetSize.
| Question a consumer asks | SPDX 3.0 | CycloneDX 1.5+ |
|---|---|---|
| What kind of model is this? | AIPackage typeOfModel | modelCard modelParameters (approach, task) |
| What was it trained on? | relationships to DatasetPackage elements | data components and modelCard datasets |
| How risky is it? | safetyRiskAssessment (serious, high, medium, low) | modelCard considerations |
| Personal data involved? | useSensitivePersonalInformation; dataset-level properties | data component governance and considerations |
| How good is it? | metric, metricDecisionThreshold | modelCard quantitativeAnalysis |
| Which exact files? | Package/File elements with verifiedUsing hashes | component hashes |
Neither format is strictly better. SPDX models everything as a graph of typed elements and relationships, which suits lineage questions (this adapter was trained from that base on these datasets). CycloneDX is closer to the tooling most security teams already run. A consumer gate should accept both, normalise them into one internal record, and never let the format decide the policy. Treat any property you rely on as required in your policy even though both specifications make it optional; a field the producer left out must produce a decision, not silence.
Binding the BOM to the bytes
Binding is the step most AIBOM programmes skip. A model is usually a directory: weight shards, a tokenizer, a config, sometimes custom code. The producer computes a digest for every file, records those digests in the BOM, and signs. The OpenSSF model-signing project (pip install model-signing, built on Sigstore, version 1.0 released in 2025) signs a whole model directory by hashing each file into a manifest and signing the manifest, with either keyless Sigstore identities or ordinary key pairs. Its README shows commands of this form:
# keyless signing with a Sigstore OIDC identity
model_signing sign bert-base-uncased --signature model.sig
# consumer verification: the signer identity and issuer are pinned
model_signing verify bert-base-uncased \
--signature model.sig \
--identity "release-bot@example.org" \
--identity-provider "https://accounts.example.org"
# key-pair variant for air-gapped environments
model_signing verify key bert-base-uncased --signature model.sig --public-key key.pubSign the model and sign the BOM, or include the BOM file inside the signed directory so a single signature covers both. What matters is that a consumer can show three facts together: these bytes, this description, this signer.
An admission gate
The gate below runs before any loader touches the files. It normalises the BOM into a flat record (the normaliser per format is omitted for space), verifies the signature by calling the CLI, recomputes digests, and applies policy. Every check produces a reason string, so a rejection is explainable in a ticket.
import hashlib, subprocess
from pathlib import Path
BANNED_SUFFIXES = {".pkl", ".pickle", ".bin", ".pt"} # pickle-capable formats
ALLOWED_LICENSES = {"Apache-2.0", "MIT", "BSD-3-Clause"}
MAX_RISK = {"low", "medium"} # serious/high need a human
def sha256(path: Path) -> str:
h = hashlib.sha256()
with path.open("rb") as f:
for chunk in iter(lambda: f.read(1 << 20), b""):
h.update(chunk)
return h.hexdigest()
def admit(model_dir: Path, bom: dict, bom_path: Path, sig: Path,
identity: str, issuer: str):
reasons = []
r = subprocess.run(["model_signing", "verify", str(model_dir),
"--signature", str(sig), "--identity", identity,
"--identity-provider", issuer], capture_output=True)
if r.returncode != 0:
return False, ["signature verification failed"]
listed = {f["path"]: f["sha256"] for f in bom["files"]}
# the BOM cannot list its own digest, and the signature covers it instead
on_disk = {p.relative_to(model_dir).as_posix() for p in model_dir.rglob("*")
if p.is_file() and p not in (sig, bom_path)}
for extra in sorted(on_disk - set(listed)):
reasons.append(f"file not in BOM: {extra}")
for path, want in listed.items():
p = model_dir / path
if not p.is_file():
reasons.append(f"BOM lists missing file: {path}")
elif sha256(p) != want:
reasons.append(f"digest mismatch: {path}")
if p.suffix in BANNED_SUFFIXES:
reasons.append(f"pickle-capable format: {path}")
if bom.get("license") not in ALLOWED_LICENSES:
reasons.append(f"license not allowed: {bom.get('license')}")
if bom.get("safetyRiskAssessment") not in MAX_RISK:
reasons.append(f"risk level needs review: {bom.get('safetyRiskAssessment')}")
if bom.get("useSensitivePersonalInformation") != "no":
reasons.append("personal-data use is yes or undeclared")
if not bom.get("baseModel", {}).get("sha256"):
reasons.append("base model not pinned by digest")
return (not reasons), reasonsThree design choices matter more than the code. First, files on disk that the BOM does not list are a rejection, not a warning: a stray modeling_custom.py is exactly how remote code arrives. Second, an absent field fails closed; useSensitivePersonalInformation must say no, and noAssertion or a missing value goes to a human. Third, the signer identity is pinned per supplier in configuration, never read from the BOM itself, because a BOM signed by whoever produced it proves nothing about who that was.
Worked example: a rejected download
Suppose a team downloads a 7B instruction-tuned model from a public hub. The directory holds four .safetensors shards, a tokenizer, a config, an SPDX 3.0 AIBOM and a signature. The gate runs and reports:
signature verification: ok (identity release-bot@vendor.example)
digest mismatch: model-00003-of-00004.safetensors
file not in BOM: training_args.bin
risk level needs review: None
DECISION: REJECT (3 reasons)Each line has a distinct cause. The signature is valid, so the BOM is authentic, but shard 3 does not match it: either the hub served a different revision than the one signed, or something altered the file after signing. Re-downloading at the pinned revision fixed it in this case, which is why the gate records the hub revision alongside the decision. The training_args.bin file is a pickle written by a training framework; it is harmless metadata in the honest case and an arbitrary-code vector in the dishonest one, so the fix is to delete it before admission rather than to allow-list the suffix. The missing safetyRiskAssessment went to the model-risk reviewer, who recorded a decision that the gate now reads from a local override file keyed by the model digest, so the same question is never asked twice.
Failure modes
- Digests over the wrong unit. Hashing a tarball while the loader reads extracted files means the check passes and the loaded bytes are never verified. Hash what the loader opens.
- Time-of-check versus time-of-use. Verifying in one container and loading from a shared volume another job can write lets a file change between the two. Verify, then copy into read-only storage, then load from there.
- Lineage that stops one level up. A fine-tune's BOM names its base model by name only. Require the base model's digest and, where available, its own BOM; a poisoned base inherits into every adapter, as described in supply chain attacks on ML.
- Self-attested sensitive fields. Dataset properties such as known bias or personal-data use are declarations. Treat them as the supplier's legal statement, keep them with the decision record, and do not present them internally as verified facts.
- BOM drift after fine-tuning. Your own adapters are new artifacts. Emit a BOM for each, with a relationship back to the admitted base digest, or the next consumer inside your company faces the same blind download.
- Model card and BOM disagree. Reconcile them in the gate. When the human-readable model card states a different license or intended use from the BOM, reject until the supplier fixes one of them.
Operating it and the trade-offs
Run the gate as a single service that every loading path must call: notebooks, batch jobs, serving images and evaluation harnesses. Cache decisions by the digest of the BOM plus the digests of all files, so a re-pull of identical bytes is instant and any change forces a fresh decision. Store each decision with its reasons, the signer identity, the hub revision and the reviewer, since that record is what an auditor asks for under regimes such as the EU AI Act. Expect cost in the hashing step: SHA-256 runs at very roughly a gigabyte per second per core on modern CPUs, so a 15 GB checkpoint costs seconds to tens of seconds, which is acceptable at admission time and a reason not to re-hash on every pod start.
The trade-off throughout is strictness against adoption. A gate that rejects every unsigned model will block most of a public hub on day one. A workable rollout runs the gate in report-only mode for a few weeks, publishes the rejection reasons to the teams affected, then enforces the identity checks (signature, digests, no extra files) first and the declared-property checks later, once suppliers have caught up.
What to do next
- Pick one production model and list every file the loader actually opens; compare that list with what its BOM, if any, claims.
- Install
model-signing, sign one internal model with a key pair, and verify it on a second machine with the public key only. - Write the normaliser for both SPDX 3.0 and CycloneDX into one internal record, and add unit tests with a missing field in each.
- Deploy the admission gate in report-only mode in front of every loader, including notebooks.
- Pin the signer identity per supplier in configuration and reject anything signed by an identity you did not pin.
- Emit a BOM, with the base model digest, for every adapter or fine-tune you produce.
- After a few weeks of reports, switch the identity checks to enforce, then the policy checks.