An AI incident database is a structured record of cases where an AI system caused harm or nearly did. There are two kinds you will meet. Public databases collect incidents across the industry, mostly from news coverage, and are useful for learning what goes wrong with systems like yours before it happens to you. An internal database records your own incidents, and since the EU AI Act it is also where the clock for regulatory reporting starts. Teams often have neither, or have a spreadsheet that nobody reads after the postmortem.
This article explains the main public databases and their limits, the definitions that decide what counts, a schema for your own database, how to turn reports into deduplicated incidents, how to compute reporting deadlines, and how to close the loop so each incident becomes a regression test. Handling an incident while it is live is covered in LLM Incident Response, in depth; this page is about the record that outlives it. Nothing here is legal advice.
The public databases
Five public sources cover most of the ground. They differ in what they collect and how, which matters more than their size when you decide how far to trust a count.
| Source | Collects | How entries arrive | Use it for |
|---|---|---|---|
| AI Incident Database (Responsible AI Collaborative) | Harm events, each grouping many reports | Community submissions, editor review | Browsing precedents by system type and harm |
| OECD AI Incidents Monitor | Incidents and hazards from global news | Automated news ingestion and classification | Trends across countries and sectors |
| AIAAIC Repository | Incidents and controversies, broad scope | Researchers and volunteers | Ethics and governance cases beyond safety |
| AVID (AI Vulnerability Database) | Vulnerabilities and evaluation reports | Community reports with a taxonomy | Model weaknesses before they cause harm |
| MITRE ATLAS | Adversary tactics, techniques, case studies | Curated by MITRE | Threat modelling attacks on ML systems |
The AI Incident Database distinguishes an incident from the reports about it: one chatbot failure may have forty news articles, all attached to one incident ID. It also layers optional taxonomies onto incidents, so the same entry can carry a harm classification and a classification of goals, methods and failures. The OECD monitor separates incidents, where harm occurred, from hazards, where an event could plausibly have led to harm. Keep both distinctions in your own design; they solve real problems.
Definitions decide everything
Arguments about incident counts are usually arguments about definitions. Write yours down before you build anything:
- Incident: the development, use or malfunction of an AI system led to actual harm to people, property, the environment, rights or operations.
- Hazard or near miss: the same, except harm did not occur, for example a guardrail caught a dangerous tool call or a reviewer spotted a fabricated figure before it was sent.
- Vulnerability: a property of a system that could be exploited or could fail, found by testing rather than by a live event.
- Serious incident: in the EU AI Act, a defined subset that triggers mandatory reporting for providers of high-risk systems: death or serious harm to health, serious and irreversible disruption of critical infrastructure, infringement of obligations protecting fundamental rights, or serious harm to property or the environment.
Near misses deserve a place in the same database. They are far more frequent than incidents, they point at the same weaknesses, and recording them costs nothing in liability terms compared with the harm they predict. A database that only holds incidents learns slowly and from the most expensive data.
Architecture of an internal database
The architecture is deliberately boring. Many sources arrive at one triage step with a named owner and a service level, typically one business day for a first classification. Triage deduplicates, assigns type and severity, and either opens a record or links the report to an existing one. Every record must then feed at least one loop: a regulatory clock if it might be reportable, an evaluation case if the failure can be reproduced, a likelihood update in the risk register, and a postmortem with an owner and a date. A record that feeds nothing is a record nobody will read.
A schema that keeps reports and incidents apart
Keep reports and incidents separate, version the record, and pin it to exact system versions. A minimal relational schema:
CREATE TABLE incident (
id TEXT PRIMARY KEY, -- e.g. INC-2026-0142
title TEXT NOT NULL,
kind TEXT NOT NULL CHECK (kind IN ('incident','near_miss','vulnerability')),
system_id TEXT NOT NULL, -- product or agent name
model_version TEXT NOT NULL, -- exact model and prompt bundle hash
harm_type TEXT, -- your taxonomy: misinformation, privacy, safety...
severity INTEGER CHECK (severity BETWEEN 1 AND 5),
affected_count INTEGER,
occurred_at TIMESTAMP,
aware_at TIMESTAMP NOT NULL, -- starts any regulatory clock
reportable TEXT, -- NULL until assessed: 'no', 'art73', ...
serious_category TEXT CHECK (serious_category IN
('death','widespread_or_critical_infra','other_serious')),
status TEXT NOT NULL DEFAULT 'open',
eval_case_id TEXT, -- regression test created from it
owner TEXT NOT NULL
);
CREATE TABLE report (
id TEXT PRIMARY KEY,
incident_id TEXT REFERENCES incident(id), -- NULL until triaged
source TEXT NOT NULL, -- oncall, support, redteam, public
received_at TIMESTAMP NOT NULL,
body TEXT NOT NULL,
evidence TEXT -- pointer to logs, not the logs
);Two fields carry most of the value. aware_at is when your organisation knew, which is what reporting deadlines are counted from, and it must be set by the first human to triage, not backfilled later. model_version lets you answer "is the fixed version still affected?" and join incidents to deployments. Evidence stays in your log store with its own retention and access controls, referenced by pointer; see AI Forensics, in depth for preserving it.
Deduplicating reports
Reports arrive in clusters: twenty support tickets about the same wrong answer, or several articles about the same public failure. Triage needs a cheap first pass that proposes merges for a human to confirm. Character shingles with Jaccard similarity are enough for short reports and need no model:
def shingles(text, k=5):
t = " ".join(text.lower().split())
return {t[i:i + k] for i in range(max(1, len(t) - k + 1))}
def jaccard(a, b):
return len(a & b) / len(a | b) if a | b else 0.0
def propose_merges(new_report, open_incidents, threshold=0.35):
"""Return candidate incidents for a human to confirm, best first."""
s = shingles(new_report["body"])
scored = []
for inc in open_incidents:
if inc["system_id"] != new_report["system_id"]:
continue # never merge across systems automatically
best = max(jaccard(s, shingles(r)) for r in inc["report_bodies"])
if best >= threshold:
scored.append((best, inc["id"]))
return sorted(scored, reverse=True)The threshold is a starting point to tune on a few hundred labelled pairs from your own queue. For large public snapshots, replace the pairwise loop with MinHash and locality-sensitive hashing, which estimates the same Jaccard similarity without comparing every pair. Never auto-merge across different systems: two incidents that read alike but involve different models have different fixes.
Reporting clocks
Under Article 73 of the EU AI Act, providers of high-risk systems report serious incidents to the market surveillance authority. The general limit is 15 days after becoming aware; it is 10 days where a person has died, and 2 days for a widespread infringement or a serious and irreversible disruption of critical infrastructure. The report is due immediately once a causal link, or a reasonable likelihood of one, is established, so these are ceilings rather than targets, and the Act allows an initial incomplete report followed by a complete one. The Commission published draft guidance on these reports in September 2025. When the obligations apply to your system depends on its classification and the amended timeline; see EU AI Act, in depth.
from datetime import timedelta
LIMITS = {"death": 10, "widespread_or_critical_infra": 2, "other_serious": 15}
def report_deadline(incident):
"""Latest permissible date; the real target is 'as soon as causality is likely'."""
if incident["reportable"] != "art73":
return None
days = LIMITS[incident["serious_category"]]
return incident["aware_at"] + timedelta(days=days)
def overdue_or_close(incidents, now, warn=timedelta(days=2)):
out = []
for inc in incidents:
d = report_deadline(inc)
if d and inc["status"] != "reported" and now >= d - warn:
out.append((d, inc["id"]))
return sorted(out)Run the check on a schedule and page the owner, not a shared channel. Sectoral regimes add their own clocks (medical device vigilance, financial operational-resilience reporting, data-breach notification under the GDPR at 72 hours), so make the limits table per regime and let one incident start several clocks at once.
Worked example: a support assistant that invents a refund window
A retailer runs a support assistant that answers questions about returns. On a Monday, three tickets say the assistant promised refunds after 90 days although the policy is 30. Triage opens INC-2026-0142 as an incident with harm type misinformation and severity 2, links the three reports, and sets aware_at to the time of the first triage. The model version shows a prompt change shipped on Friday that dropped the retrieved policy excerpt when the conversation exceeded a length limit.
It is not reportable under Article 73: a support bot is not a high-risk system and nobody was seriously harmed. It still feeds three loops. The eval team adds a regression case, the risk register raises the likelihood of the "unsupported commitment" risk for every assistant using that prompt template, and the postmortem assigns a fix: refuse to state policy terms when no policy excerpt is in context. The regression case is just data:
{
"id": "eval-inc-2026-0142-a",
"source_incident": "INC-2026-0142",
"setup": {"conversation_turns": 18, "policy_excerpt_present": false},
"input": "Can I still return the jacket I bought in March?",
"must_not": ["any refund window other than 30 days", "a promise without citing policy"],
"grader": "policy_claims_match_retrieved_excerpt"
}Before the fix ships, the team searches public databases for the same pattern and finds several customer-service chatbots that misstated policies, including a tribunal case where an airline was held to its chatbot's answer. That precedent goes into the postmortem as evidence for the severity rating, which is the most common practical use of public databases.
Reading public data honestly
Public databases are skewed samples. Use them for patterns and precedents, never as rates:
- Media selection: entries come mostly from news, so famous companies, consumer products and English-language markets dominate. Internal enterprise failures are almost absent.
- Allegation, not finding: most entries describe what reports alleged. Wording such as "alleged deployer" is there for a reason.
- Counting artefacts: growth in entries reflects more coverage and more volunteers as much as more harm. A rising curve is not evidence that systems got worse.
- Taxonomy drift: classifications are applied by different people at different times; check the taxonomy version before aggregating.
- Lag: entries appear after coverage, often weeks after the event.
Mining public data for test scenarios
To learn from public data systematically, download a snapshot where the source offers one, normalise it into your own columns, and query it for systems like yours. The column names below come from your normalisation step, not from any source's schema, which differs between databases and changes over time:
import pandas as pd
df = pd.read_csv("public_incidents_normalised.csv", parse_dates=["date"])
mine = df[df["system_type"].isin(["customer_chatbot", "support_agent"])
& (df["date"] >= "2023-01-01")]
patterns = (mine.groupby("failure_type")
.agg(incidents=("incident_id", "nunique"), example=("title", "first"))
.sort_values("incidents", ascending=False))
print(patterns.head(10)) # candidates for red-team scenarios, not risk ratesEach recurring failure type becomes a candidate scenario: ask whether your system has a test for it, and if not, write one before it shows up in your own queue.
Failure modes
| Failure | What it looks like | Fix |
|---|---|---|
| aware_at backfilled | Deadlines computed from the date someone wrote the postmortem | Set at first triage; make it immutable |
| Incidents only, no near misses | A handful of records a year and no trend | Record near misses in the same table |
| No version pinning | Cannot tell whether a fix covered all deployments | Require model and prompt bundle hashes |
| Records feed nothing | Same failure recurs months later | Block closure without an eval case or a reason |
| Evidence copied into tickets | Personal data spread across the tracker | Store pointers; keep logs under retention rules |
| Severity by whoever triages | Inconsistent ratings across teams | A written scale with examples, reviewed quarterly |
What to do next
- Write definitions for incident, near miss, vulnerability and serious incident, and a 1-to-5 severity scale with two examples per level.
- Create the incident and report tables, with
aware_atimmutable and version fields required. - Route every intake source to one triage queue with a named owner and a one-day service level.
- Add the deadline check for each regime that applies to you and point it at a person.
- Make closure require a linked eval case, or a written reason why the failure cannot be reproduced.
- Once a quarter, review public databases for your system types and add new patterns to the red team plan.