A risk matrix is the grid most organisations use to rank risks: likelihood on one axis, severity on the other, and a colour in each cell. For AI systems it is usually the first artefact a governance board asks for, and usually the weakest one in the pack. The grid is not the problem. The problem is that axes labelled rare and major mean different things to every person who scores, the colours are picked by eye, and the cell number is computed by multiplying two ranks as if they were quantities.
This article treats the matrix as a measuring instrument you design, test and calibrate. You will anchor each axis to numbers, choose a scoring rule that respects those numbers, test the colour scheme in code, see a worked example with four LLM risks, and learn where the instrument stops being trustworthy. For the quantitative analysis that should take over at that point, see AI risk assessment with loss scenarios and Monte Carlo; for what happens to a scored risk afterwards, see the AI risk register.
What a risk matrix is, and what it is not
A risk matrix maps two ordinal ratings to a priority. Ordinal means the ranks are ordered but the gaps between them are not defined: a likelihood of 4 is more than 3, but nothing says it is a third more. That is fine for sorting, and wrong for arithmetic, which is where most matrices go astray.
Used well, the matrix does three jobs. It triages a long list of risks quickly, so attention goes to the right handful. It gives non-specialists a common language, so a product manager and a security engineer can disagree about the same cell rather than about the meaning of words. And it routes decisions: a red cell needs an executive owner and a treatment plan, a green one can be accepted at team level. It does not estimate how much risk you carry, it cannot add risks together, and it says nothing about uncertainty.
Step 1: anchor the likelihood axis
Anchor likelihood to a frequency of the harmful event at your real exposure, not to adjectives. Decade bands work well because estimates of rare events are rarely better than a factor of three, and decade bands make the scoring rule simple, as the next step shows.
AI risks need one conversion step that traditional registers rarely do: per-request rates must become events per year. A jailbreak that succeeds on 0.2% of attempts in your red-team suite is not rare if adversaries try it 4,000 times a year; that is about eight successes a year, band L4. Write the conversion into the risk record so a reviewer can check it and so it is redone when traffic or the model changes.
| Band | Events per year | Plain meaning | AI example |
|---|---|---|---|
| L1 | 0.001 to 0.01 | once in 100 to 1,000 years | training-data poisoning by a nation-state actor at a small firm |
| L2 | 0.01 to 0.1 | once in 10 to 100 years | model weights exfiltrated from a well-run registry |
| L3 | 0.1 to 1 | once in 1 to 10 years | indirect prompt injection that leaks a document |
| L4 | 1 to 10 | several times a year | PII written into application logs |
| L5 | 10 to 100 | every month or more | a toxic answer reaching a user |
Anything above the top band or below the bottom band needs explicit handling: events more frequent than 100 a year are operational quality issues for a dashboard, not matrix entries, and events rarer than one in a thousand years go to the override rules below if their severity is extreme.
Step 2: anchor the severity axis
Severity is the consequence of one occurrence. AI harms are multi-dimensional, so score each dimension against its own anchors and take the worst one, never the average: an incident that is cheap but injures someone is a serious incident. Financial anchors in decades line up with the likelihood bands.
| Band | Financial | People and safety | Privacy | Legal and trust |
|---|---|---|---|---|
| S1 | $1k to 10k | inconvenience | one record, low sensitivity | internal note |
| S2 | $10k to 100k | distress, quickly remedied | tens of records | customer complaints |
| S3 | $100k to 1M | material harm to a few people | hundreds of records or one sensitive | regulator inquiry |
| S4 | $1M to 10M | serious harm to one person | thousands or special category | enforcement, press |
| S5 | $10M to 100M | life-threatening harm | mass exposure | licence or market at risk |
Non-financial columns are aligned to the financial one by agreement, not by pretending a privacy breach has a precise price. The alignment is a policy choice the board should sign, and the risk appetite is the natural place to record it.
Step 3: the scoring rule, L + S rather than L x S
Most matrices compute a score as L times S. With decade bands that is the wrong operation. The midpoint expected loss of cell (L, S) is about 10 to the power (L + S - 1) dollars a year under the bands above, so the quantity that tracks expected loss is the sum L + S, and each step of the sum is a factor of ten.
The product misranks risks that the sum orders correctly. Take risk A at (L1, S5), rare but catastrophic, and risk B at (L3, S2), occasional and moderate. By product, A scores 5 and B scores 6, so B ranks higher. By sum, A scores 6 and B scores 5, and that is right: A has a midpoint expected loss of about $100k a year against B at about $10k. Multiplying ranks systematically understates rare, severe risks, which are exactly the ones AI governance worries about most.
Step 4: test the colour thresholds in code
Colour boundaries should be tested, not drawn. A minimum requirement, which Cox called weak consistency, is that no risk in a red cell can be quantitatively smaller than a risk in a green cell. Because each cell is a rectangle of frequencies and losses, you can compute its lowest and highest possible expected loss and check every pair of cells:
L_EDGES = [0.001, 0.01, 0.1, 1.0, 10.0, 100.0] # events per year
S_EDGES = [1e3, 1e4, 1e5, 1e6, 1e7, 1e8] # dollars per event
RANK = {"green": 0, "amber": 1, "red": 2}
def band(x, edges):
for i in range(len(edges) - 1):
if edges[i] <= x < edges[i + 1]:
return i + 1
raise ValueError(f"{x} is outside the matrix: handle it explicitly")
def colour(l, s):
k = l + s
return "red" if k >= 8 else "amber" if k >= 6 else "green"
def cell_range(l, s):
"""Lowest and highest annual expected loss a risk in this cell can have."""
return L_EDGES[l - 1] * S_EDGES[s - 1], L_EDGES[l] * S_EDGES[s]
def overlaps():
cells = [(l, s) for l in range(1, 6) for s in range(1, 6)]
out = []
for a in cells:
for b in cells:
if RANK[colour(*a)] > RANK[colour(*b)]:
if cell_range(*a)[0] < cell_range(*b)[1]:
out.append((a, colour(*a), b, colour(*b)))
return out
bad = overlaps()
print(len(bad), sum(1 for x in bad if x[1] == "red" and x[3] == "green"))With these bands and thresholds the script prints 32 and 0. No red cell can hold a smaller risk than a green cell, so the scheme is weakly consistent. The 32 overlaps are all between neighbouring colours on adjacent anti-diagonals: each cell spans two decades of expected loss, so an amber risk can be smaller than a green one by up to a factor of ten. That is the resolution limit of a 5x5 grid, and it is why risks near a boundary are escalated rather than trusted. Run the same check on your own matrix; hand-drawn colour schemes, especially ones with a red cell in a low-likelihood corner, often fail the red-versus-green test outright. Note the explicit error for values outside the bands: the grid has edges, and silently clamping a $500M loss into S5 is how tail risks disappear.
Worked example: four LLM risks
Four risks from an internal assistant with email access, scored by the band and colour functions above:
| Risk | Events per year | Loss per event | Cell | L x S | L + S | Colour | Expected loss |
|---|---|---|---|---|---|---|---|
| Indirect prompt injection exfiltrates a mailbox | 0.3 | $400k | L3 S3 | 9 | 6 | amber | $120k |
| Hallucinated dosage reaches a clinician | 0.02 | $30M | L2 S5 | 10 | 7 | amber | $600k |
| PII written to application logs | 4 | $60k | L4 S2 | 8 | 6 | amber | $240k |
| Toxic answer reaches a user | 40 | $2k | L5 S1 | 5 | 6 | amber | $80k |
All four land in amber, while their expected losses differ by a factor of 7.5. The product column would put the toxic answer last by a wide margin and the dosage risk first; the sum puts three at 6 and the dosage risk at 7. The honest output of the matrix is that these four share a priority tier, and the next decision needs better information than the grid holds.
Two things should happen next. The dosage risk carries a life-threatening severity, which the override rule below turns red regardless of likelihood. The other three are close enough that the team should estimate ranges and compare treatment costs, not debate cell borders.
Overrides for tails, ties and correlated risks
- Severity floor. Any S5 harm to people's safety or any irreversible harm is red whatever the likelihood. Expected loss is the wrong lens when the outcome is unrecoverable.
- Out-of-range values. Losses above the top severity edge or frequencies above the top likelihood edge are rejected at entry, forcing a human decision rather than a clamp.
- Ties and boundaries. A risk whose plausible range crosses a threshold takes the hotter colour until someone quantifies it. Record the range, not just the point.
- Correlated risks. Ten amber risks that share one cause, such as a single retrieval index with no access control, are one red risk. Group by root cause before scoring, because the matrix cannot add.
- Unknown likelihood. New capabilities, such as an agent gaining a payment tool, have no history. Score from red-team attack success rates times exposure, and mark the score provisional with a re-score date.
Overrides are where the matrix meets your threat-model scoring: DREAD-style factors can feed the likelihood estimate but never replace the anchored bands.
Operating the matrix: calibration and re-scoring
A matrix is only as consistent as the people scoring it. Three practices make it repeatable. First, calibrate: give five scorers the same ten written scenarios, have them score independently, and measure agreement with a weighted kappa. Low agreement on likelihood usually means the conversion from per-request rate to events per year is missing; on severity it usually means people are averaging dimensions. Second, store the inputs, not just the cell: frequency estimate, exposure, loss estimate, dimension scored, and the evidence. A cell with no inputs cannot be re-scored when traffic doubles. Third, define re-score triggers: a model upgrade, a new tool or data source, a traffic change of more than a band's width, an incident or near miss, or a new public attack technique against your model family.
risk = {
"id": "R-114",
"scenario": "indirect prompt injection in inbound email exfiltrates a mailbox",
"likelihood": {"per_attempt": 0.003, "attempts_per_year": 100,
"events_per_year": 0.3, "evidence": "redteam-run-2026-09-30"},
"severity": {"financial": 4e5, "privacy": "S3", "worst": "S3"},
"cell": [3, 3], "colour": "amber", "overrides": [],
"rescore_on": ["model_change", "new_tool", "traffic_x10", "incident"],
}
Failure modes
- Range compression. Very different risks share a cell. Mitigate with decade bands and by escalating anything in the top tiers to quantitative analysis.
- Centring. Scorers avoid the extremes, so everything drifts to 3 by 3. Anchors with examples, and calibration sessions, counter it.
- Gaming. Owners nudge a risk one band down to escape red-cell governance. Requiring the inputs and evidence makes nudges visible in review.
- Stale scores. AI systems change monthly. Without re-score triggers the matrix describes last year's model.
- Mitigation double counting. Score inherent and residual separately and name the control that moves the cell; a residual score with no named control is a wish.
- Aggregation. Counting red cells as a portfolio metric treats them as additive. Report the list, not the count.
Trade-offs: matrix versus quantitative analysis
The matrix trades accuracy for speed and shared language. It is the right tool when you have dozens or hundreds of risks, little data and a mixed audience. Quantitative analysis, with loss distributions and simulation, is the right tool for the top ten, for anything near a decision threshold, and for comparing treatments with real costs. Five bands per axis is the usual compromise: three bands lose too much resolution, seven invite false precision that scorers cannot deliver. Whatever you choose, use one matrix across the organisation, because risks scored on different grids cannot be compared.
What to do next
- Rewrite both axes as numeric bands: events per year and loss per event, in decades.
- Add the per-request to per-year conversion to every AI risk record.
- Score severity per dimension and take the worst, never the average.
- Replace L x S with L + S and set colour thresholds on the sum.
- Run the consistency check on your thresholds and fix any red-versus-green overlap.
- Write the override rules: safety floor, out-of-range rejection, ties go hot, group by root cause.
- Hold a calibration session with ten scenarios and measure inter-rater agreement.
- Send every red cell and every boundary case to quantitative analysis.