A risk matrix is the grid most organisations use to rank risks: likelihood on one axis, severity on the other, and a colour in each cell. For AI systems it is usually the first artefact a governance board asks for, and usually the weakest one in the pack. The grid is not the problem. The problem is that axes labelled rare and major mean different things to every person who scores, the colours are picked by eye, and the cell number is computed by multiplying two ranks as if they were quantities.

This article treats the matrix as a measuring instrument you design, test and calibrate. You will anchor each axis to numbers, choose a scoring rule that respects those numbers, test the colour scheme in code, see a worked example with four LLM risks, and learn where the instrument stops being trustworthy. For the quantitative analysis that should take over at that point, see AI risk assessment with loss scenarios and Monte Carlo; for what happens to a scored risk afterwards, see the AI risk register.

What a risk matrix is, and what it is not

A risk matrix maps two ordinal ratings to a priority. Ordinal means the ranks are ordered but the gaps between them are not defined: a likelihood of 4 is more than 3, but nothing says it is a third more. That is fine for sorting, and wrong for arithmetic, which is where most matrices go astray.

Used well, the matrix does three jobs. It triages a long list of risks quickly, so attention goes to the right handful. It gives non-specialists a common language, so a product manager and a security engineer can disagree about the same cell rather than about the meaning of words. And it routes decisions: a red cell needs an executive owner and a treatment plan, a green one can be accepted at team level. It does not estimate how much risk you carry, it cannot add risks together, and it says nothing about uncertainty.

Where the matrix sits: an ordinal front end to the register, not a substitute for analysisLoss scenariowho, what, how oftenAnchored scoringL band, S band per dimensionMatrix ruleL+S, thresholds, overridesRegister + ownertreatment, review dateQuantify insteadtop cells, ties, tails, close callsescalateThe matrix triages hundreds of risks quickly; the handful that matter get numbers.
The matrix is the cheap, fast front end. Anything near a threshold, in the top cells or with a long tail goes to quantitative analysis.

Step 1: anchor the likelihood axis

Anchor likelihood to a frequency of the harmful event at your real exposure, not to adjectives. Decade bands work well because estimates of rare events are rarely better than a factor of three, and decade bands make the scoring rule simple, as the next step shows.

AI risks need one conversion step that traditional registers rarely do: per-request rates must become events per year. A jailbreak that succeeds on 0.2% of attempts in your red-team suite is not rare if adversaries try it 4,000 times a year; that is about eight successes a year, band L4. Write the conversion into the risk record so a reviewer can check it and so it is redone when traffic or the model changes.

BandEvents per yearPlain meaningAI example
L10.001 to 0.01once in 100 to 1,000 yearstraining-data poisoning by a nation-state actor at a small firm
L20.01 to 0.1once in 10 to 100 yearsmodel weights exfiltrated from a well-run registry
L30.1 to 1once in 1 to 10 yearsindirect prompt injection that leaks a document
L41 to 10several times a yearPII written into application logs
L510 to 100every month or morea toxic answer reaching a user

Anything above the top band or below the bottom band needs explicit handling: events more frequent than 100 a year are operational quality issues for a dashboard, not matrix entries, and events rarer than one in a thousand years go to the override rules below if their severity is extreme.

Step 2: anchor the severity axis

Severity is the consequence of one occurrence. AI harms are multi-dimensional, so score each dimension against its own anchors and take the worst one, never the average: an incident that is cheap but injures someone is a serious incident. Financial anchors in decades line up with the likelihood bands.

BandFinancialPeople and safetyPrivacyLegal and trust
S1$1k to 10kinconvenienceone record, low sensitivityinternal note
S2$10k to 100kdistress, quickly remediedtens of recordscustomer complaints
S3$100k to 1Mmaterial harm to a few peoplehundreds of records or one sensitiveregulator inquiry
S4$1M to 10Mserious harm to one personthousands or special categoryenforcement, press
S5$10M to 100Mlife-threatening harmmass exposurelicence or market at risk

Non-financial columns are aligned to the financial one by agreement, not by pretending a privacy breach has a precise price. The alignment is a policy choice the board should sign, and the risk appetite is the natural place to record it.

Step 3: the scoring rule, L + S rather than L x S

Most matrices compute a score as L times S. With decade bands that is the wrong operation. The midpoint expected loss of cell (L, S) is about 10 to the power (L + S - 1) dollars a year under the bands above, so the quantity that tracks expected loss is the sum L + S, and each step of the sum is a factor of ten.

The product misranks risks that the sum orders correctly. Take risk A at (L1, S5), rare but catastrophic, and risk B at (L3, S2), occasional and moderate. By product, A scores 5 and B scores 6, so B ranks higher. By sum, A scores 6 and B scores 5, and that is right: A has a midpoint expected loss of about $100k a year against B at about $10k. Multiplying ranks systematically understates rare, severe risks, which are exactly the ones AI governance worries about most.

A 5x5 matrix with decade bands: colour follows L + S, which tracks log10(expected loss)L+S=2L+S=3L+S=4L+S=5L+S=6L+S=3L+S=4L+S=5L+S=6L+S=7L+S=4L+S=5L+S=6L+S=7L+S=8L+S=5L+S=6L+S=7L+S=8L+S=9L+S=6L+S=7L+S=8L+S=9L+S=10$1k-10k$10k-100k$100k-1M$1M-10M$10M-100ML1: 1 per 100-1000 yrL2: 1 per 10-100 yrL3: 1 per 1-10 yrL4: 1-10 per yrL5: 10-100 per yrSeverity S (loss per event, worst dimension)green: L+S at most 5 amber: 6-7 red: 8 or more (plus overrides for safety-critical severity)
Colour thresholds on L + S. Each anti-diagonal is one decade of expected loss.

Step 4: test the colour thresholds in code

Colour boundaries should be tested, not drawn. A minimum requirement, which Cox called weak consistency, is that no risk in a red cell can be quantitatively smaller than a risk in a green cell. Because each cell is a rectangle of frequencies and losses, you can compute its lowest and highest possible expected loss and check every pair of cells:

L_EDGES = [0.001, 0.01, 0.1, 1.0, 10.0, 100.0]   # events per year
S_EDGES = [1e3, 1e4, 1e5, 1e6, 1e7, 1e8]          # dollars per event
RANK = {"green": 0, "amber": 1, "red": 2}

def band(x, edges):
    for i in range(len(edges) - 1):
        if edges[i] <= x < edges[i + 1]:
            return i + 1
    raise ValueError(f"{x} is outside the matrix: handle it explicitly")

def colour(l, s):
    k = l + s
    return "red" if k >= 8 else "amber" if k >= 6 else "green"

def cell_range(l, s):
    """Lowest and highest annual expected loss a risk in this cell can have."""
    return L_EDGES[l - 1] * S_EDGES[s - 1], L_EDGES[l] * S_EDGES[s]

def overlaps():
    cells = [(l, s) for l in range(1, 6) for s in range(1, 6)]
    out = []
    for a in cells:
        for b in cells:
            if RANK[colour(*a)] > RANK[colour(*b)]:
                if cell_range(*a)[0] < cell_range(*b)[1]:
                    out.append((a, colour(*a), b, colour(*b)))
    return out

bad = overlaps()
print(len(bad), sum(1 for x in bad if x[1] == "red" and x[3] == "green"))

With these bands and thresholds the script prints 32 and 0. No red cell can hold a smaller risk than a green cell, so the scheme is weakly consistent. The 32 overlaps are all between neighbouring colours on adjacent anti-diagonals: each cell spans two decades of expected loss, so an amber risk can be smaller than a green one by up to a factor of ten. That is the resolution limit of a 5x5 grid, and it is why risks near a boundary are escalated rather than trusted. Run the same check on your own matrix; hand-drawn colour schemes, especially ones with a red cell in a low-likelihood corner, often fail the red-versus-green test outright. Note the explicit error for values outside the bands: the grid has edges, and silently clamping a $500M loss into S5 is how tail risks disappear.

Worked example: four LLM risks

Four risks from an internal assistant with email access, scored by the band and colour functions above:

RiskEvents per yearLoss per eventCellL x SL + SColourExpected loss
Indirect prompt injection exfiltrates a mailbox0.3$400kL3 S396amber$120k
Hallucinated dosage reaches a clinician0.02$30ML2 S5107amber$600k
PII written to application logs4$60kL4 S286amber$240k
Toxic answer reaches a user40$2kL5 S156amber$80k

All four land in amber, while their expected losses differ by a factor of 7.5. The product column would put the toxic answer last by a wide margin and the dosage risk first; the sum puts three at 6 and the dosage risk at 7. The honest output of the matrix is that these four share a priority tier, and the next decision needs better information than the grid holds.

Two things should happen next. The dosage risk carries a life-threatening severity, which the override rule below turns red regardless of likelihood. The other three are close enough that the team should estimate ranges and compare treatment costs, not debate cell borders.

Overrides for tails, ties and correlated risks

  • Severity floor. Any S5 harm to people's safety or any irreversible harm is red whatever the likelihood. Expected loss is the wrong lens when the outcome is unrecoverable.
  • Out-of-range values. Losses above the top severity edge or frequencies above the top likelihood edge are rejected at entry, forcing a human decision rather than a clamp.
  • Ties and boundaries. A risk whose plausible range crosses a threshold takes the hotter colour until someone quantifies it. Record the range, not just the point.
  • Correlated risks. Ten amber risks that share one cause, such as a single retrieval index with no access control, are one red risk. Group by root cause before scoring, because the matrix cannot add.
  • Unknown likelihood. New capabilities, such as an agent gaining a payment tool, have no history. Score from red-team attack success rates times exposure, and mark the score provisional with a re-score date.

Overrides are where the matrix meets your threat-model scoring: DREAD-style factors can feed the likelihood estimate but never replace the anchored bands.

Operating the matrix: calibration and re-scoring

A matrix is only as consistent as the people scoring it. Three practices make it repeatable. First, calibrate: give five scorers the same ten written scenarios, have them score independently, and measure agreement with a weighted kappa. Low agreement on likelihood usually means the conversion from per-request rate to events per year is missing; on severity it usually means people are averaging dimensions. Second, store the inputs, not just the cell: frequency estimate, exposure, loss estimate, dimension scored, and the evidence. A cell with no inputs cannot be re-scored when traffic doubles. Third, define re-score triggers: a model upgrade, a new tool or data source, a traffic change of more than a band's width, an incident or near miss, or a new public attack technique against your model family.

risk = {
    "id": "R-114",
    "scenario": "indirect prompt injection in inbound email exfiltrates a mailbox",
    "likelihood": {"per_attempt": 0.003, "attempts_per_year": 100,
                   "events_per_year": 0.3, "evidence": "redteam-run-2026-09-30"},
    "severity": {"financial": 4e5, "privacy": "S3", "worst": "S3"},
    "cell": [3, 3], "colour": "amber", "overrides": [],
    "rescore_on": ["model_change", "new_tool", "traffic_x10", "incident"],
}

Failure modes

  • Range compression. Very different risks share a cell. Mitigate with decade bands and by escalating anything in the top tiers to quantitative analysis.
  • Centring. Scorers avoid the extremes, so everything drifts to 3 by 3. Anchors with examples, and calibration sessions, counter it.
  • Gaming. Owners nudge a risk one band down to escape red-cell governance. Requiring the inputs and evidence makes nudges visible in review.
  • Stale scores. AI systems change monthly. Without re-score triggers the matrix describes last year's model.
  • Mitigation double counting. Score inherent and residual separately and name the control that moves the cell; a residual score with no named control is a wish.
  • Aggregation. Counting red cells as a portfolio metric treats them as additive. Report the list, not the count.

Trade-offs: matrix versus quantitative analysis

The matrix trades accuracy for speed and shared language. It is the right tool when you have dozens or hundreds of risks, little data and a mixed audience. Quantitative analysis, with loss distributions and simulation, is the right tool for the top ten, for anything near a decision threshold, and for comparing treatments with real costs. Five bands per axis is the usual compromise: three bands lose too much resolution, seven invite false precision that scorers cannot deliver. Whatever you choose, use one matrix across the organisation, because risks scored on different grids cannot be compared.

What to do next

  1. Rewrite both axes as numeric bands: events per year and loss per event, in decades.
  2. Add the per-request to per-year conversion to every AI risk record.
  3. Score severity per dimension and take the worst, never the average.
  4. Replace L x S with L + S and set colour thresholds on the sum.
  5. Run the consistency check on your thresholds and fix any red-versus-green overlap.
  6. Write the override rules: safety floor, out-of-range rejection, ties go hot, group by root cause.
  7. Hold a calibration session with ten scenarios and measure inter-rater agreement.
  8. Send every red cell and every boundary case to quantitative analysis.
Key takeaway: A risk matrix is an ordinal triage instrument. Anchor both axes to decade bands of real numbers, score with L + S so each step is a factor of ten in expected loss, test that no red cell can hold a smaller risk than a green one, and add overrides for safety-critical harms, ties and shared causes. Then send the top cells to quantitative analysis, because the grid ranks risks but cannot measure them.