AVID, the AI Vulnerability Database, is an open, community-maintained catalogue of failures in AI systems: security flaws, harmful bias, toxic output, privacy leaks and performance problems. It is run by the AI Risk and Vulnerability Alliance, a nonprofit, at avidml.org. It comes with an open data model and a Python library, avidtools, so that evaluation tools can write records in a shared format. The garak LLM scanner, for example, can convert its run reports into AVID records with one command.

The overview of AI vulnerability databases places AVID beside CVE, OSV, ATLAS and incident databases. This article goes one level deeper. It covers AVID's two record types field by field, its taxonomy and codes, how to write a valid report in code, how to turn scanner output into records, and how to run an internal intake pipeline on the same schema. Field names and enum values below come from the avidtools source. The schema is still evolving, so pin the version you build against.

What AVID records, and what it does not

A CVE names a flaw in a specific product version that a patch fixes. Most AI failures do not fit that shape. A jailbreak success rate, a demographic disparity in a classifier, or a model that repeats memorised personal data are measurements of a model, under a test harness, at a point in time. AVID's design accepts that. Its basic unit is evidence about an artifact, not a version range to match.

That sets clear expectations for what you can do with it.

  • It is not a scanner feed. There is no package-and-version matching. An AVID record about a model is a hypothesis to test against your deployment, not a verdict on it.
  • It is a shared vocabulary. Its taxonomy gives security, ethics and performance issues comparable codes, which helps when one register has to hold all three.
  • It links rather than duplicates. Record classes cover AI Incident Database incidents, ATLAS case studies, CVE entries, LLM evaluations and third-party reports. The impact block can carry ATLAS, CVSS and CWE mappings next to AVID's own codes, so one record can point to all the other systems.

Reports and vulnerabilities

AVID's two record types and what flows between themevaluation or findinggarak run, red team, paperReportone observed instance + metricsVulnerabilitygeneral failure modewriteingestImpactavid, atlas, cvss, cwe, odinreports[]summaries of linked reportsAVID taxonomyrisk domain, SEP, lifecycleA report says: this artifact, measured this way, showed this problem.A vulnerability says: this kind of problem exists, and here are the reports that show it.
Reports carry evidence; vulnerabilities generalise across reports and list the reports that support them.

AVID keeps two record types. A report (IDs like AVID-2022-R0001) describes one observed instance: an artifact, a detection method, the measured result and the references. A vulnerability (IDs like AVID-2022-V001) describes a general failure mode and carries a list of report summaries, each with a report ID, type and name. In avidtools, Vulnerability.ingest(report) copies a report's affects, problem type, description, references, impact and credit into a vulnerability and stamps the published and modified dates. It also drops the report-level vulnerability ID from the taxonomy block.

The split mirrors how evidence accumulates. A dozen scanner runs against different models produce a dozen reports. If they all show the same prompt-injection pattern, one vulnerability record ties them together. Your own register should copy this: keep raw findings and generalised issues as separate objects with a link between them, or every rerun becomes a new issue.

Anatomy of a report

These are the top-level fields of Report and what each holds.

FieldTypeWhat goes in it
data_typestrNamespace, AVID by default
data_versionstrSchema or data version
metadata.report_idstrThe report ID once assigned
affectsAffectsDeveloper list, deployer list, artifacts with type Model, Dataset or System and a name
problemtypeProblemtypeclassof (record class), optional type (Issue, Advisory, Measurement, Detection), description
metricslist[Metric]Metric name, detection method (Significance Test or Static Threshold) and a free-form results dict
referenceslist[Reference]Label and URL of evidence
descriptionLangValueLanguage code and text
impactImpactavid taxonomy; optionally atlas, cvss, cwe and odin mappings
creditlist[LangValue]Who found it
reported_datedateWhen it was reported

Two design choices matter in practice. First, affects separates the developer of a model from the deployer who put it in a product. A finding against a hosted assistant may belong to either, and remediation differs: the deployer adds a guardrail, the developer retrains. Second, metrics.results is an untyped dictionary. That makes it flexible enough for any evaluation, but it also means two tools can report the same measurement under different keys. Fix your own keys before you aggregate.

The taxonomy: SEP and lifecycle views

The AVID taxonomy has two views, and a record can carry several codes from each.

The effect view, called SEP, groups issues into three risk domains, each with numbered categories and subcategories:

DomainTop-level codesExamples of subcategories
Security (S)S0100 Software Vulnerability, S0200 Supply Chain Compromise, S0300 Over-permissive API, S0400 Model Bypass, S0500 Exfiltration, S0600 Data PoisoningS0301 Information Leak, S0403 Adversarial Example, S0502 Model theft, S0601 Ingest Poisoning
Ethics (E)E0100 Bias/Discrimination, E0200 Explainability, E0300 User actions, E0400 MisinformationE0101 Group fairness, E0301 Toxicity, E0402 Generative Misinformation
Performance (P)P0100 Data issues, P0200 Model issues, P0300 Privacy, P0400 SafetyP0102 Concept drift, P0204 Accuracy, P0301 Anonymization, P0401 Psychological Safety

The lifecycle view places the issue in the ML workflow: L01 Business Understanding, L02 Data Understanding, L03 Data Preparation, L04 Model Development, L05 Evaluation and L06 Deployment. Most findings from testing a finished model land in L05. The lifecycle code exists to answer a different question, which is where the fix belongs. A biased training set is an L03 problem even if L05 testing found it.

The codes are coarse on purpose. Classify with the most specific code you can defend, list several when a finding spans domains, and keep detail such as the attack technique in the ATLAS mapping or the description rather than inventing codes. Pin the taxonomy_version string in your records. garak currently writes an empty string there, so do not depend on that field being populated in records you import.

Worked example: writing a report

Here is a worked example. An internal evaluation ran 400 indirect prompt-injection attempts against a support assistant built on a third-party model. In 37 of them, the assistant followed instructions hidden in a retrieved document. The threshold agreed with the product owner is 2%, so this is a finding. The code below writes it as an AVID report.

from datetime import date
from avidtools.datamodels.report import Report
from avidtools.datamodels.components import (
    Affects, Artifact, AvidTaxonomy, Detection, Impact, LangValue,
    Metric, Problemtype, Reference,
)
from avidtools.datamodels.enums import (
    ArtifactTypeEnum, ClassEnum, LifecycleEnum, MethodEnum, SepEnum, TypeEnum,
)

hits, attempts = 37, 400
report = Report(
    affects=Affects(
        developer=["Example Model Co"],
        deployer=["Acme support assistant team"],
        artifacts=[Artifact(type=ArtifactTypeEnum.system, name="acme-support-rag v2.3"),
                   Artifact(type=ArtifactTypeEnum.model, name="example-chat-8b 2026-08")],
    ),
    problemtype=Problemtype(
        classof=ClassEnum.llm, type=TypeEnum.measurement,
        description=LangValue(lang="eng", value="Indirect prompt injection via retrieved documents"),
    ),
    metrics=[Metric(
        name="attack success rate",
        detection_method=Detection(type=MethodEnum.thres, name="ASR above 0.02"),
        results={"attempts": attempts, "successes": hits, "asr": hits / attempts,
                 "harness": "inj-suite 1.4", "temperature": 0.7},
    )],
    references=[Reference(label="internal eval run 8812", url="https://evals.example.internal/runs/8812")],
    description=LangValue(lang="eng", value=(
        "Instructions embedded in knowledge-base articles were followed in 37 of 400 trials, "
        "including 9 that called the ticket-update tool. Payloads withheld.")),
    impact=Impact(avid=AvidTaxonomy(
        risk_domain=["Security"],
        sep_view=[SepEnum.S0400, SepEnum.S0301],
        lifecycle_view=[LifecycleEnum.L05, LifecycleEnum.L06],
        taxonomy_version="0.2",
    )),
    credit=[LangValue(lang="eng", value="Acme AI red team")],
    reported_date=date(2026, 10, 2),
)
report.save("findings/acme-support-inj.avid.json")

Several choices here are deliberate. Both the system and the model are listed as artifacts, because the result depends on the retrieval integration as much as the model. The results dict records the harness version and sampling temperature, because an attack success rate without its conditions cannot be reproduced. The description says the payloads were withheld: a record can be useful without being an exploit kit. The taxonomy version string is your choice and should match the codes you validated against.

From garak runs to AVID records

garak, the open-source LLM vulnerability scanner, writes a JSONL report per run, named like garak.<uuid>.report.jsonl. Its reporting option converts that into AVID records:

python -m garak -r garak.6f1c.report.jsonl
# writes garak.6f1c.avid.jsonl, one AVID report per line

The conversion makes one report per probe. Each report records the target as a Model artifact and the target type as deployer, with class LLM Evaluation and type Measurement. Its metrics hold per-detector pass counts and scores, its risk domain and SEP codes are derived from the probe's tags, and its lifecycle is L05. Because each probe becomes its own record, a run with forty probes produces forty records, most of which will show no problem. Filter before anything reaches a human. Here is a minimal triage pass:

import json
from avidtools.datamodels.report import Report

def failing_reports(path, max_fail_rate=0.02):
    out = []
    with open(path, encoding="utf-8") as fh:
        for line in fh:
            r = Report.model_validate_json(line)       # rejects unknown enum codes
            for m in r.metrics or []:
                rows = m.results                        # column -> {row_index: value}
                for i, det in rows.get("detector", {}).items():
                    total = rows["total_evaluated"][i]
                    failed = total - rows["passed"][i]
                    if total and failed / total > max_fail_rate:
                        out.append((r.problemtype.description.value, det, failed, total))
    return out

The results layout above follows how garak serialises a pandas frame into the dict, with column names mapping to row-indexed values. Inspect one line of your own output before relying on it, because it is a convention of the producer, not the schema.

An internal intake pipeline

An intake pipeline that keeps AVID records useful inside your organisationscannersgarak -r outputred teammanual findingspublic AVIDexternal recordsvalidatepydantic modelstriagethresholds, dedupinternal registerasset, owner, statusredact and publishoptional, after fixValidation rejects records with unknown taxonomy codes before they reach a human.
Internal intake: every source is parsed into the same models, validated, triaged, and only then joined to assets.

Use the AVID schema as an interchange format, not as your whole register. AVID records say nothing about which of your services use an artifact, who owns the fix, or whether it is done. Keep those in an internal register keyed to your inventory, as described in the overview article, and store the AVID report as attached evidence.

Validate on entry. Parse every record with the pydantic models. A typo in a SEP code then fails at intake rather than surfacing months later as an empty dashboard bucket.

Deduplicate by artifact, problem and harness. Reruns of the same probe against the same model version are new measurements of one finding. Append them as metrics history instead of opening new findings.

Map to classes. Fill the ATLAS mapping where a technique applies (see MITRE ATLAS) and CWE where a code-level weakness exists. Mappings make the register searchable by the frameworks auditors ask about.

Pull public records as test cases. When AVID publishes a report against a model family you use, turn it into a test in your evaluation suite and run it against your deployment. Record the outcome as your own report.

Failure modes

  • Treating records as matches. A public report against a model does not mean your deployment, with its system prompt, filters and tools, is affected. Retest.
  • Results without conditions. A score with no harness version, sampling settings or prompt set cannot be compared over time. Make those keys mandatory in your own results dicts.
  • Publishing exploit payloads. A behaviour record can work as a ready-made attack. Describe the class and the rate, and hold payloads until the developer and deployer have fixed or mitigated. Disclosure norms covers timelines.
  • Record floods. Per-probe conversion of every scanner run buries real findings. Triage by threshold and deduplicate before human review.
  • Schema drift. The impact block has gained fields over time, including CVSS, CWE and 0DIN mappings. Pin avidtools and test your parser against new releases before upgrading.
  • Stale model identity. Hosted models change behind a stable name. Record the exact model version or snapshot date in the artifact name, or the record cannot be tied to what was tested.

Trade-offs

OptionGood forLimitation
AVID schema internallyShared vocabulary across security, ethics and performance; tool supportNo ownership, asset or status fields; evolving schema
CVE for AI product flawsVendor-fixable flaws in a product; scanner matchingPoor fit for measured model behaviour
Incident databasesEvidence of real-world harmEvents, not reproducible tests
Bespoke register onlyFits your process exactlyNothing interoperates; mapping work repeats

Most teams end up with the hybrid shown above: a bespoke register for ownership and status, AVID-shaped evidence attached to each finding, and class mappings for reporting. For how findings flow into response, see AI incident databases.

What to do next

  1. Install a pinned version of avidtools and read the report, components and enums modules. The schema is short.
  2. Run garak against one internal model, convert the run with -r, and inspect a record by hand.
  3. Write the triage filter, agree a failure threshold per probe family with the product owner, and route only failures to humans.
  4. Add mandatory result keys (harness version, sampling settings, prompt set hash) to your own reports.
  5. Store AVID reports as evidence on internal register entries keyed by asset and owner, with ATLAS and CWE mappings filled.
  6. Each month, search public AVID records for model families you use and turn relevant ones into regression tests.
  7. Decide your publication policy: what you share, after which fix milestone, and with payloads withheld.
Key takeaway: AVID records AI failures as evidence: reports that measure one artifact under one harness, and vulnerabilities that generalise across reports, coded with SEP risk categories and lifecycle stages. Use its avidtools models to validate and exchange findings, convert scanner runs with garak -r, triage before humans see them, and keep ownership and status in your own register with AVID records attached.