When a security team says it tracks vulnerabilities, it usually means a scanner matches installed package versions against a database of known flaws. That model works for the code in an AI system. It covers only part of the rest. A prompt-injection weakness in a deployed assistant has no package version. A malicious model file on a public hub is an artifact, not a bug in anyone's code. A documented bias failure in a classifier belongs to one model under one evaluation, and may not reproduce on the next release.
So there is no single AI vulnerability database. There is a set of sources, each with a different unit of record: CVE, GitHub advisories and OSV for software flaws; CWE, MITRE ATLAS and the OWASP lists for classes of weakness and attack technique; and AVID and the AI Incident Database for model behaviour and real-world harm. This article explains what each one records and how its data is shaped. It then builds an internal register that joins them to your own inventory, with code you can run. Specific identifiers below were checked against public records; where a detail could not be confirmed, it is left out.
Three layers, three kinds of record
Sort every AI risk record by the layer it describes, because the layer decides whether matching can be automated.
- Code. Inference servers, training frameworks, tokenizers, model loaders, vector databases and agent frameworks are ordinary software. Their bugs get CVE IDs and affected version ranges, and a dependency scanner can match them.
- Artifacts. Model weights and serialized files can carry code. A pickle-based checkpoint can run arbitrary Python when loaded. A deliberately malicious upload is malware, not a vulnerability, so it is usually handled by hub scanning and takedown rather than a CVE.
- Behaviour. Jailbreaks, data leakage through outputs, harmful bias and tool misuse are properties of a model, a prompt and an integration together. They are measured, not version-matched.
The code layer: CVE, NVD, GHSA and OSV
CVE is the identifier system run by the CVE Program, with IDs assigned by CVE Numbering Authorities. NVD, run by NIST, enriches CVE records with severity scores and product identifiers (CPE). That enrichment has lagged publication since 2024, so do not make a CVSS score from NVD a precondition for triage. GitHub Security Advisories (GHSA) are curated for package ecosystems, and GitHub publishes its advisory database in the OSV format. OSV.dev aggregates GHSA, PyPI and many other sources, and its schema is built for open-source matching: each record names a package in an ecosystem and gives affected ranges as introduced and fixed events. A record's primary ID may be a GHSA ID, with the CVE in aliases, so deduplicate on aliases.
A concrete AI example shows why this layer matters. CVE-2025-32434 is a remote code execution flaw in PyTorch: torch.load with weights_only=True, a setting widely recommended as the safe way to load checkpoints, could still be made to execute code in versions up to 2.5.1. The fix shipped in 2.6.0, and the weakness class is CWE-502, deserialisation of untrusted data. A scanner that reads your lockfile finds this automatically. Nothing else in this article is that easy.
The query below asks OSV about one pinned package. The /v1/query endpoint accepts a package and version and returns full records. /v1/querybatch returns only IDs for many packages at once, and large results can be paginated with a next_page_token.
import json
import urllib.request
OSV_QUERY = "https://api.osv.dev/v1/query"
def osv_vulns(name, version, ecosystem="PyPI"):
"""Known vulnerabilities for one pinned package, from OSV.dev."""
body = json.dumps({"package": {"name": name, "ecosystem": ecosystem},
"version": version}).encode()
req = urllib.request.Request(OSV_QUERY, data=body,
headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=30) as resp:
data = json.load(resp)
found = []
for v in data.get("vulns", []):
fixed = sorted({ev["fixed"]
for aff in v.get("affected", [])
for rng in aff.get("ranges", [])
for ev in rng.get("events", []) if "fixed" in ev})
found.append({"id": v["id"], "aliases": v.get("aliases", []),
"summary": v.get("summary", ""), "fixed_in": fixed})
return found
for rec in osv_vulns("torch", "2.5.1"):
print(rec["id"], rec["aliases"], rec["fixed_in"], rec["summary"][:60])
Class catalogues: CWE-1427, ATLAS and OWASP
The second group of sources names kinds of flaw rather than instances. They have no version ranges, so they cannot be scanned for. Their job is to give each finding a shared label, so that a pen-test report, a red-team result and an incident can be counted together.
- CWE-1427, Improper Neutralization of Input Used for LLM Prompting, was added in CWE 4.16 in November 2024. It describes products that build prompts from external data in a way that lets the model confuse that data with developer instructions. It gives prompt injection a standard weakness ID next to CWE-502 and the rest.
- MITRE ATLAS is an ATT&CK-style matrix of adversary tactics and techniques against AI systems, backed by case studies and published as machine-readable data. Use it to describe how an attack proceeds; the ATLAS deep dive shows how to load and use it.
- The OWASP Top 10 for LLM Applications is a risk list for application builders. It is coarse by design and works best as a review checklist and a reporting vocabulary; see the OWASP LLM article.
A good record carries one label from each: a CWE for the weakness, an ATLAS technique for how it was exploited, and an OWASP category for the executive summary.
Behaviour and incidents: AVID and the AI Incident Database
The AI Vulnerability Database (AVID) is an open-source project of the AI Risk and Vulnerability Alliance, a nonprofit. It separates two record types. A vulnerability, with IDs such as AVID-2022-V001, is a general failure mode, and AVID compares it to a CVE. A report, with IDs such as AVID-2022-R0001, is one observed instance supported by a qualitative or quantitative evaluation. Its taxonomy has two views. The effect view groups risks into security, ethics and performance, with coded categories. The lifecycle view places each risk at a stage of the ML workflow. AVID also defines record classes for LLM evaluations, AI Incident Database incidents, CVE entries and third-party reports, so it can point at the other sources rather than replace them.
The AI Incident Database, at incidentdatabase.ai, collects reports of AI systems causing or nearly causing real-world harm. Its unit is an incident, which is an event, not a reproducible flaw. It answers a different question from a scanner: has this kind of system failed this way in the wild, and what did it cost? That makes it useful input for threat models and for persuading a product owner that a risk is not hypothetical.
Neither source can tell you that your deployment is affected. A report that one model leaked training data under a given prompt set is evidence about that model, version and harness. Treat behaviour records as test cases to run against your own system, not as matches.
Where the categories blur
The categories leak into each other, and the edge cases are where programs fail.
- AI features in hosted products get CVEs. EchoLeak, CVE-2025-32711, was a prompt injection against Microsoft 365 Copilot: a crafted email could lead the assistant to exfiltrate sensitive data. It has a CVE because it is a flaw in a vendor's product, fixed server-side. There was no package to upgrade, so no dependency scanner would ever have flagged it. Track vendor advisories for every hosted AI service you use as a separate feed.
- Behaviour flaws in your own product are your vulnerabilities. A CVE program exercise concluded there is no reason vulnerabilities in large language models cannot receive CVE IDs. In practice, most findings about your own assistant will live only in your tracker, which is why the register below matters.
- Malicious models are not CVEs. A poisoned checkpoint on a public hub is a supply-chain threat. Defend with format choice (safetensors rather than pickle), loader settings, scanning and provenance, as covered in the AI supply chain program article.
- Scanners for model files have bugs too. Pickle scanners have themselves received advisories for bypasses. They reduce risk; they do not make untrusted pickles safe.
An internal AI vulnerability register
Since no outside source maps findings to your deployments, build a register that does. Each entry joins an external or internal finding to an asset in your AI inventory. The record shape below is a suggested internal schema, not a standard.
{
"id": "AIV-2026-0042",
"layer": "code", # code | artifact | behaviour
"sources": ["GHSA-...", "CVE-2025-32434"], # external IDs, aliases kept
"classes": {"cwe": "CWE-502", "atlas": "<technique ID>", "owasp": "LLM03"},
"assets": ["svc-rag-api@prod", "train-cluster-a"],
"exposure": "loads third-party checkpoints from object storage",
"status": "open", # open | mitigated | accepted | fixed
"fix": "torch>=2.6.0; safetensors only for external weights",
"evidence": "lockfile scan 2026-10-05; loader audit",
"owner": "ml-platform",
"due": "2026-10-12"
}Three feeds fill it. The code feed runs an OSV query over every lockfile and container image in the AI inventory on each build and nightly, because new advisories arrive for old versions. The vendor feed watches advisories for each hosted model and AI SaaS product you depend on. The behaviour feed turns red-team results, AVID reports and incidents into entries against your own assets, and each of these needs a person to judge applicability. The OWASP value in the example uses the 2025 list's numbering; pin the edition you label against, because the categories have been renumbered between editions.
Worked example: a RAG assistant
Consider a retrieval-augmented support assistant with this inventory: a FastAPI service pinned to torch==2.5.1 that loads a reranker from a public hub, a hosted LLM API, a vector database and a tool that can send email. Walk it through each layer.
- Code. The OSV query returns CVE-2025-32434 for the pinned torch. The service loads an external checkpoint, so the flaw is exposed. Entry: open, fix by upgrading to 2.6.0 or later, owner platform, due within the critical SLA.
- Artifact. The reranker is distributed as a pickle checkpoint. No database has a record for it, and that is the point. Entry: convert to safetensors, pin by content hash, load with the restricted loader in the patched version.
- Vendor. The hosted LLM provider's advisories contain nothing relevant this month. Record the check, so absence is evidence rather than silence.
- Behaviour. The assistant reads retrieved documents and can send email, which is the same shape as EchoLeak. Label it CWE-1427 and run indirect-injection tests that try to make retrieved content trigger the email tool. If they succeed, the finding is your own vulnerability. Fix it by requiring confirmation for outbound email and stripping links and images from rendered output.
Only the first step was automatic. The other three needed an inventory that records how each model is loaded and what each tool can do, which is the real prerequisite for any AI vulnerability program.
Scoring, sources and operations
CVSS scores code flaws reasonably well and behaviour flaws badly, since a jailbreak's impact depends entirely on what the model can reach. Score behaviour findings on reachable impact: what data the model can read, what tools it can call, and whether a human approves actions. A prompt injection in a read-only chatbot and the same injection in an agent with email and file access are different severities.
| Source | Unit | Machine-matchable | Use it for |
|---|---|---|---|
| CVE / NVD | flaw in a product | partly (CPE matching is noisy) | canonical IDs, vendor products |
| GHSA / OSV | flaw in a package version range | yes | lockfile and image scanning |
| CWE-1427 and other CWEs | weakness class | no | labelling findings consistently |
| MITRE ATLAS | attack technique and case study | no | attack chains, coverage, red-team plans |
| AVID | failure mode and evaluation report | no | behaviour test ideas and taxonomy |
| AI Incident Database | real-world incident | no | threat modelling, business cases |
| Vendor advisories | flaw in a hosted service | no | AI SaaS and model API risk |
Operationally, give code-layer entries the same SLAs as any dependency, by exposure. Re-test behaviour entries on every model or prompt change, since a fix can silently regress. Feed confirmed incidents into your AI incident response process, and report suitable findings in third-party products through their disclosure or bounty programs.
Failure modes
- Scanner-only programs. Lockfile scanning covers the code layer and misses hosted services, artifacts and behaviour, which is where most AI-specific risk sits.
- Waiting for enrichment. Triage that requires an NVD CVSS score can stall for weeks. Use the CNA's or GHSA's score, or your own exposure-based one.
- Alias duplicates. The same flaw appears as a GHSA ID, a CVE and a PyPI advisory. Deduplicate on aliases or your counts inflate.
- Copying behaviour reports as facts. A report about another model and harness is a hypothesis about yours. Reproduce it before you open an entry or close one.
- No inventory of loaders and tools. Without knowing which services load external weights or hold dangerous tools, no source can tell you whether you are exposed.
- Unpinned taxonomies. OWASP and ATLAS change between releases. Record the version you labelled against.
What to do next
- List every AI service you run or buy, with how it loads models and which tools it can call.
- Run the OSV query above against each AI service's lockfile, and add it to CI.
- Check every PyTorch install is 2.6.0 or later, and move external weights to safetensors.
- Subscribe to advisories from each hosted model and AI SaaS vendor and log each review.
- Create the register schema, then label existing findings with a CWE, an ATLAS technique and an OWASP category.
- Turn three AVID reports or incidents relevant to your systems into tests, and run them against your own deployment.