An AI risk assessment answers one practical question: should this specific AI system be deployed in this specific way, and if so, with which controls? It is a decision procedure, not a document genre. A good assessment ends with a named owner signing off on a residual exposure they can state in plain numbers, plus a list of events that force the assessment to be repeated. A bad one ends with a spreadsheet of red, amber and green cells that nobody can act on.

This article walks the procedure end to end for a language-model system: scoping, characterising the system, identifying concrete loss scenarios, estimating them with calibrated ranges instead of ordinal scores, simulating the annual loss distribution in Python, deciding against a stated appetite, and setting reassessment triggers. The outputs land in a risk register, and the procedure maps cleanly onto the Map and Measure functions of the NIST AI RMF. Threat discovery is covered in depth in threat modelling for LLM systems; here we focus on turning threats into numbers and numbers into a decision.

What an assessment is for

Three activities are often confused. Threat modelling asks what could go wrong and how, and its output is a list of attack paths. A risk register is a living inventory of accepted, treated and open risks across a portfolio. A risk assessment sits between them: for one system, it takes the threats and hazards, estimates how often each would cause loss and how much, and compares the result with what the organisation has agreed to tolerate.

For AI systems the hazard list is wider than for conventional software. Besides attackers, you must consider the model being confidently wrong, the model being right on average but systematically wrong for one group of users, sensitive data surfacing in outputs, and users over-trusting outputs they cannot check. Standards reflect this. ISO/IEC 23894 gives guidance on applying general risk management to AI, ISO/IEC 42001 makes assessment a requirement of an AI management system, and NIST AI 600-1, the generative AI profile, lists risk areas specific to generative models such as confabulation and information integrity. None of them prescribes a scoring formula, which is why the estimation method below is your choice and your responsibility.

The assessment pipeline

An AI risk assessment is a pipeline: each stage produces an artefact the next one consumes1. Scopeuse case, owner, decision2. Characterisedata flow, trust boundaries3. Identifyscenarios per harm type4. Estimatecalibrated 90% ranges5. SimulateMonte Carlo, exceedance6. Decideaccept, treat, avoid7. Treatcontrols, re-estimate8. Recordregister entries, sign-offReassessment triggersmodel swap, new tool, new data class, incident, drift in a key indicatormonitorre-openStages 4 to 7 loop until residual exposure sits inside the stated appetite, or the use case is changed or dropped.
The assessment pipeline. Stages 1 to 3 are qualitative; 4 to 6 are quantitative; the triggers turn a one-off exercise into a maintained control.

Scoping comes first and is the step most often skipped. Write down the use case in one sentence, the decision the system influences, who is affected, and who owns the outcome. A support assistant that drafts replies for a human agent and one that issues refunds on its own are different systems even if they share a model and a prompt. Scope also fixes what is out of scope, such as the upstream model vendor's training practices, so that the assessment does not silently absorb risks it cannot estimate.

Characterisation produces a data-flow diagram with trust boundaries: where user input enters, which retrieval sources feed the context window, which tools the model can call, with what credentials, and where outputs go. Every arrow that crosses a boundary is a place where a scenario can start. Record the model version, the system prompt hash and the tool list, because a change in any of them is later a reassessment trigger.

From threats to loss scenarios

Identification turns threats into loss scenarios. A scenario names an actor or cause, an event, an asset and a loss. Write each as one sentence that a sceptic could test, for example: an external user embeds instructions in a support ticket attachment, the assistant calls the refund tool for an order the user does not own, and the company pays out money it does not recover. Vague entries such as model misuse cannot be estimated and should be split until they can.

Use a fixed set of harm categories to make sure coverage is not shaped by whoever happens to attend the workshop:

Harm categoryTypical LLM scenarioEvidence you can collect
SecurityPrompt injection drives a tool call with the agent credentialsRed-team success rate per tool
PrivacyRetrieval returns another customer record into a replyCanary-document leak tests
IntegrityConfabulated policy or price stated to a customerGrounded-answer eval error rate
FairnessLower answer quality for one language or dialectSliced eval scores by cohort
SafetyHarmful instructions emitted despite filtersJailbreak benchmark pass rate
LegalStatement treated as a binding commitmentCount of outputs containing commitments

The last row is not hypothetical. In Moffatt v. Air Canada (2024), a Canadian tribunal held the airline responsible for a refund policy its website chatbot had invented. Integrity failures become legal ones as soon as a customer relies on them.

Each scenario should link back to evidence. The red-team programme supplies attack success rates, the evaluation suite supplies error rates, and production logs supply volumes. Without evidence, an estimate is a guess with a decimal point.

Estimating with calibrated ranges

Ordinal scales, likelihood 1 to 5 times impact 1 to 5, are popular and misleading. Multiplying ranks is not arithmetic on quantities, very different risks collapse to the same cell, and the scale cannot express uncertainty. Replace them with two estimates per scenario: how many times per year the loss event happens, and how much a single event costs, each given as a 90 percent confidence interval.

Frequency is built from the evidence. Suppose the assistant handles 400,000 conversations a year, about one in 2,000 contains an injection attempt, the red team saw 3 percent of attempts reach the refund tool, and an amount check stops half of those. The point estimate is 400,000 / 2,000 x 0.03 x 0.5 = 3 events a year. Because every factor is uncertain, state a range such as 1 to 8 rather than the point.

Magnitude covers direct payout, response effort, regulatory exposure and customer churn. Elicit the low and high bounds from people who would actually pay the bill, and train estimators first: ask ten trivia questions with 90 percent ranges and check that about nine contain the truth. Uncalibrated experts are systematically overconfident, and narrow ranges understate the tail that matters most.

Simulating annual loss

With ranges in hand, a short Monte Carlo simulation produces the quantity decision makers actually need: the probability that total annual loss exceeds a given amount. The script below treats frequency as uniform within its range, draws event counts from a Poisson process and draws each event cost from a lognormal fitted to the 90 percent interval. Lognormal suits costs because they are positive and right-skewed.

import math, random

Z90 = 1.6449  # a 90% interval spans +-1.645 standard deviations

def lognormal_from_ci(lo, hi):
    mu = (math.log(lo) + math.log(hi)) / 2
    sigma = (math.log(hi) - math.log(lo)) / (2 * Z90)
    return mu, sigma

def poisson(lam, rng):
    # count exponential inter-arrival times that fit inside one year
    n, t = 0, rng.expovariate(lam)
    while t < 1.0:
        n, t = n + 1, t + rng.expovariate(lam)
    return n

def simulate(scenarios, years=20000, seed=7):
    rng = random.Random(seed)
    totals = []
    for _ in range(years):
        loss = 0.0
        for s in scenarios:
            lam = rng.uniform(*s["freq"])          # uncertain rate
            mu, sigma = lognormal_from_ci(*s["cost"])
            for _ in range(poisson(lam, rng)):
                loss += rng.lognormvariate(mu, sigma)
        totals.append(loss)
    return sorted(totals)

def exceedance(totals, threshold):
    return sum(t > threshold for t in totals) / len(totals)

scenarios = [
    {"id": "inject-refund", "freq": (1, 8),     "cost": (2_000, 60_000)},
    {"id": "pii-leak",      "freq": (0.05, 0.5), "cost": (50_000, 2_000_000)},
    {"id": "false-policy",  "freq": (2, 20),    "cost": (300, 25_000)},
]
totals = simulate(scenarios)
print("mean annual loss", round(sum(totals) / len(totals)))
for t in (100_000, 500_000, 1_000_000):
    print(f"P(loss > {t:,}) = {exceedance(totals, t):.3f}")

Reading the output is the point of the exercise. The mean annual loss is useful for budgeting controls, but the exceedance probabilities are what you compare with appetite. In runs of this example the privacy scenario fires less than once every three years on average, yet it supplies about half the mean and almost all of the chance of losing more than a million, because its cost tail is wide. A heat map would have rated it low likelihood and moved on.

Keep the model honest. Run it with several seeds and confirm the exceedance figures are stable to two decimal places; if not, increase the year count. Change each input bound by a factor of two and see which scenario moves the answer most; that is where better evidence is worth buying.

Deciding against appetite

The decision needs an appetite statement written before the numbers arrive, otherwise the threshold drifts to fit the result. A usable form is: we accept at most a 5 percent chance per year of AI-attributable losses above 500,000, and no scenario with a credible path to regulated personal data leaving the tenant. The first clause is checked against the exceedance curve, the second is a hard constraint that no amount of averaging can satisfy.

There are four outcomes per scenario. Accept when the exposure is inside appetite and record who accepted it. Treat by adding controls, then re-estimate frequency or magnitude with fresh evidence, never by applying a guessed percentage reduction. Transfer through insurance or contract where that genuinely moves the loss. Avoid by changing the design, for example removing the refund tool and having the model draft a refund request a human approves. Avoidance is underused; it is often cheaper than any control.

For the running example, treatment might be: scope the refund tool to orders owned by the authenticated user, which collapses the injection path rather than filtering it; add per-tenant filters to retrieval; and ground policy answers in a retrieved policy document with a citation check. Re-running the red team and the grounded-answer evaluation after each change supplies the new ranges. Where a law imposes duties by risk tier, as the EU AI Act does, record the tier classification as a separate constraint alongside the quantitative result.

Reassessment triggers

An assessment describes one configuration at one point in time. Write explicit triggers into the sign-off so that it is reopened when that configuration changes:

  • A model version change, including a vendor silently updating a model behind the same name.
  • A new tool, a broader credential on an existing tool, or a new retrieval source.
  • A new class of data entering the context window, such as health or payment data.
  • An incident or near miss attributed to the system.
  • A key indicator crossing its threshold, for example red-team success above 5 percent.
  • A scheduled date, typically every six or twelve months, as a backstop.

Automate the cheap ones. Store the model identifier, prompt hash and tool manifest beside the assessment record and have the deployment pipeline fail when they differ from the assessed values without a linked reassessment ticket.

Failure modes

The common failure modes are predictable. Assessments done once at launch and never reopened describe a system that no longer exists. Scenarios written as categories rather than events cannot be estimated, so they get a default medium. Estimates come from the team that wants to ship, without calibration or challenge. Controls are credited with effectiveness nobody measured. The vendor model is assumed safe because a vendor assessment exists, although that assessment covered a different use. And the appetite is set after the results are known.

A subtler failure is aggregation. Ten scenarios each comfortably inside appetite can together exceed it, and several AI systems sharing one model or one vector store share failure causes. Simulate correlated scenarios together and assess shared components once, at portfolio level.

Trade-offs

Quantitative assessment costs more up front: calibration training, evidence collection and a modelling habit. In exchange it makes disagreements discussable, because people argue about a range and a source rather than whether something is amber. For low-stakes internal tools a short qualitative screen is proportionate; reserve the full procedure for systems that act on behalf of users, touch regulated data or make decisions about people. Precision is also a trap in reverse: a simulation output to the nearest dollar suggests certainty the inputs do not have. Report rounded figures and the ranges that drove them.

What to do next

  1. Write a one-sentence scope, owner and decision for your highest-stakes AI system.
  2. Draw its data flow with trust boundaries and record model version, prompt hash and tool list.
  3. Write five testable loss scenarios across at least three harm categories.
  4. Run a calibration exercise with your estimators before collecting any ranges.
  5. Collect evidence per scenario: red-team rates, eval error rates, traffic volumes.
  6. Run the simulation, check seed stability and do a sensitivity pass on each bound.
  7. Write the appetite statement, decide per scenario, and push results into the register.
  8. Wire the configuration hash check into your deployment pipeline as a reassessment trigger.
Key takeaway: Treat AI risk assessment as a decision procedure for one system in one configuration. Turn threats into testable loss scenarios, estimate frequency and cost as calibrated ranges backed by evidence, simulate the annual loss distribution, and compare its tail with an appetite written in advance. Prefer design changes that remove a path over filters that thin it, and make configuration changes reopen the assessment automatically.