Every organisation that ships an LLM feature ends up with a list of things that could go wrong. Usually it lives in a slide deck from the launch review, and by the third model upgrade nobody can say which items are still true, who agreed to live with them, or what evidence showed the mitigations worked. An AI risk register is that list turned into an operating record: one entry per risk, each with an owner, a score before and after controls, the evidence that the controls are real, an explicit decision about the remainder, and the conditions that force a second look.
This article covers testable risk statements, honest residual scores, key risk indicators, expiring acceptances and change triggers, with a worked coding-agent example run through a small register engine.
What a register is for
A register answers four questions for each risk: what could happen, how bad is it now, who decided that this level is tolerable, and how would we know if that stopped being true. Threat models, red-team reports and frameworks feed it; they are not substitutes for it. A threat model enumerates attack paths for one design. A framework such as the NIST AI RMF or ISO/IEC 42001 tells you which activities a management system must contain. The register is where the outputs of all of them meet a decision, and it is the document an auditor, a regulator or your own incident review will ask for first.
A useful register keeps history, so you can reconstruct what was known when an incident happened, and it is wired to the system: entries reference real controls, tests and metrics, so it goes stale loudly, not silently.
Anatomy of an entry
The fields below are the minimum that makes an entry actionable. Anything less and the entry cannot be scored, owned or reopened; much more and teams stop filling it in.
| Field | What goes in it | Why it matters |
|---|---|---|
| Statement | Cause, event and consequence in one sentence | Makes the risk testable and stops duplicates |
| System and asset | Which application, model, tool or data store | Lets a change to that component find the entry |
| Taxonomy tags | OWASP LLM Top 10 item, MITRE ATLAS technique, internal category | Aggregation and coverage reporting |
| Owner | One named team or person | An entry owned by a committee is owned by nobody |
| Inherent likelihood and impact | 1 to 5 each, with written anchors | The score with no controls, used to rank effort |
| Controls | Name, which axis it reduces, strength, last test date, evidence link | The only route from inherent to residual |
| Residual score and band | Computed, never typed | Compared with appetite to force a decision |
| Treatment and acceptance | Mitigate, avoid, transfer or accept; approver; expiry date | Records who owns the remainder and until when |
| Key risk indicators | Metric, threshold, data source | Early warning that the score is wrong |
| Triggers | Change types that reopen the entry | Ties the register to change management |
Give every score level a written anchor. Impact 5 might mean cross-tenant data exposure or an irreversible production action; likelihood 5 might mean a public technique works against this design today. Without anchors, two reviewers scoring the same risk routinely land points apart.
Writing risk statements you can test
Most weak registers fail at the sentence level. Entries such as "prompt injection" or "hallucination" name a category, not a risk. They cannot be scored because the impact depends entirely on what the model is connected to, and they cannot be closed because the category never goes away. Write the statement as cause, event and consequence for a specific system.
| Weak | Testable |
|---|---|
| Prompt injection | Instructions in a web page fetched by the coding agent cause it to push a malicious commit to a shared repository |
| Data leakage | Repository secrets returned by a tool call are repeated in a model response and leave the trust boundary |
| Model drift | A provider-side model update weakens refusals for destructive shell commands without a deploy on our side |
| Cost | A tool-calling loop in one session consumes the monthly inference budget within a day |
The testable form tells you what to red-team, which control applies to which half of the sentence, and what a key risk indicator should count.
Inherent and residual scores without self-deception
Inherent risk is the score with no controls in place; residual risk is what remains after the controls you can prove. The word prove carries the weight. The commonest failure in AI registers is residual scores that assume every listed control works, including the guardrail model nobody has evaluated since launch and the "human in the loop" who approves 400 actions an hour.
Three rules keep residual scores honest. First, every control says which axis it reduces. A pre-commit review gate does not make injection less likely; it limits the damage, so it lowers impact. An allowlist on tool calls makes the dangerous action less reachable, so it lowers likelihood. Second, a control earns credit only if it has been tested recently, with the test linked as evidence; 90 days is a reasonable default. Third, the residual is computed from the inherent score and the credited controls, never typed by hand. A typed residual is an opinion; a computed one can be audited.
Scores are ordinal: 12 is not twice as bad as 6. Use the product for ranking and banding, keep both axes visible, and let appetite rules refer to impact directly, for example requiring executive acceptance for any residual impact of 5.
The lifecycle and appetite
Appetite is the threshold above which a residual risk may not simply be lived with. A common structure is three bands: low risks are accepted by the owner, medium risks need a named director-level approver and an expiry no more than two quarters out, and high risks block release unless a senior executive signs an exception with a short expiry and a treatment plan.
A register engine in Python
Keep the register as data in version control with a small engine beside it, so changes are reviewed and checks run in CI. This engine computes residuals from freshly tested controls, compares KRIs with thresholds, flags missing or expired acceptances above the low band, and reopens entries whose triggers match this release.
from dataclasses import dataclass, field
from datetime import date, timedelta
FRESH = timedelta(days=90) # a control test older than this earns no credit
APPETITE = {"high": 15, "medium": 8} # residual score bands
@dataclass
class Control:
name: str
reduces: str # "likelihood" or "impact"
strength: int # points removed from that axis when the control is proven
last_tested: date | None
@dataclass
class Risk:
id: str
statement: str
owner: str
likelihood: int # inherent, 1-5
impact: int # inherent, 1-5
controls: list = field(default_factory=list)
kris: dict = field(default_factory=dict) # name -> (threshold, observed)
accepted_until: date | None = None
accepted_by: str | None = None
triggers: set = field(default_factory=set) # change types that reopen it
def residual(r, today):
lik, imp, credited = r.likelihood, r.impact, []
for c in r.controls:
if c.last_tested is None or today - c.last_tested > FRESH:
continue
if c.reduces == "likelihood":
lik -= c.strength
else:
imp -= c.strength
credited.append(c.name)
lik, imp = max(lik, 1), max(imp, 1)
return lik * imp, credited
def band(score):
if score >= APPETITE["high"]:
return "HIGH"
return "MEDIUM" if score >= APPETITE["medium"] else "LOW"
def review(register, today, changes):
for r in register:
score, credited = residual(r, today)
notes = []
for kri, (limit, seen) in r.kris.items():
if seen > limit:
notes.append(f"KRI {kri} {seen} > {limit}")
if band(score) != "LOW" and (r.accepted_until is None or r.accepted_until < today):
notes.append("acceptance missing or expired")
hit = r.triggers & changes
if hit:
notes.append("reopen: " + ",".join(sorted(hit)))
inherent = r.likelihood * r.impact
print(f"{r.id} inherent {inherent:>2} residual {score:>2} {band(score):<6} "
f"credit={len(credited)}/{len(r.controls)} " + ("; ".join(notes) or "ok"))In production, observed KRI values come from your metrics store and the change set from a diff of the release manifest.
Worked example: a coding agent on a model upgrade
Take a coding agent that reads issues, browses documentation on the web, edits files and pushes branches to a shared repository. It uses a hosted model, and this release moves to a newer model version. The four testable statements from earlier become four entries, reviewed on 4 October 2026 with the change set {"model_change"}.
T = date(2026, 10, 4)
REGISTER = [
Risk("R-01", "web page injection -> malicious push",
"platform-sec", 4, 5,
[Control("protected-branch review", "impact", 2, T - timedelta(days=20)),
Control("tool-call allowlist on push", "likelihood", 1, T - timedelta(days=130))],
kris={"injection_canary_hits_per_wk": (0, 3)},
accepted_until=date(2026, 12, 31), accepted_by="CISO",
triggers={"new_tool", "model_change"}),
Risk("R-02", "secrets in tool output leak via response",
"appsec", 3, 4,
[Control("secret scanner on tool output", "likelihood", 2, T - timedelta(days=10))],
kris={"secret_detections_per_wk": (2, 1)}),
Risk("R-03", "model update weakens shell refusals",
"ml-platform", 3, 4,
[Control("pinned model snapshot", "likelihood", 1, T - timedelta(days=40))],
kris={"refusal_eval_pass_rate_drop_pct": (2, 0)},
accepted_until=date(2026, 9, 30), accepted_by="VP Eng",
triggers={"model_change"}),
Risk("R-04", "agent loop exhausts budget",
"finops", 2, 3,
[Control("per-session token cap", "impact", 1, T - timedelta(days=5))],
kris={"sessions_over_cap_per_day": (5, 2)}),
]
review(REGISTER, T, changes={"model_change"})
# Output:
# R-01 inherent 20 residual 12 MEDIUM credit=1/2 KRI injection_canary_hits_per_wk 3 > 0; reopen: model_change
# R-02 inherent 12 residual 4 LOW credit=1/1 ok
# R-03 inherent 12 residual 8 MEDIUM credit=1/1 acceptance missing or expired; reopen: model_change
# R-04 inherent 6 residual 4 LOW credit=1/1 okRead each line against its inputs. R-01 starts at 4 by 5 = 20. The branch-protection review was tested 20 days ago and lowers impact by 2, so impact falls to 3. The push allowlist would lower likelihood, but its last test was 130 days ago, so it earns nothing and likelihood stays at 4: residual 12, medium. Its acceptance runs to December, so that check passes, but two other things fire. The canary KRI saw 3 hits against a threshold of 0, meaning planted injection strings in test pages were acted on this week, and the model change is one of its triggers. The fix is to retest the allowlist and rerun the injection suite against the new model, not to edit the score.
R-02 is quiet: one fresh control, residual 1 by 4 = 4. R-03 is the interesting one. Its residual of 2 by 4 = 8 is only medium, but its acceptance expired on 30 September and the release changes exactly the thing it is about. A pinned model snapshot is the control, and this release unpins it. The entry must be reassessed with the new model's refusal evaluation attached, and re-accepted or the release held. R-04 is low and within appetite; no action.
Key risk indicators that actually warn
A key risk indicator is a metric that moves before the risk materialises, or at least before it becomes an incident. Good indicators for LLM systems are cheap to collect and specific to one entry: canary injection strings acted on per week, secret-scanner hits on tool output, refusal-evaluation pass rate after each model change, share of agent actions approved in under two seconds (a sign the human reviewer has become a rubber stamp), sessions hitting the token cap, and retrieval of documents outside a user's permission set in shadow checks. Each needs a threshold chosen in advance and a data source a reviewer can open.
Avoid indicators that measure activity, such as red-team tests run: they always go up, so they never warn.
Running it: cadence, CI and reporting
Run the engine in CI on every change to the register and on every release, so a release that trips a trigger cannot ship without the owner's sign-off. Review the full register monthly with owners present and quarterly with whoever holds appetite. Report upwards as a short table: top residual risks, entries that changed band, expired acceptances and KRI breaches. Closing an entry requires the reason (component removed, risk transferred by contract, superseded by a merged entry) and is itself a dated, reviewed change.
Map entries to external taxonomies through tags, not by structuring the register around them. Tags to the OWASP LLM Top 10 or MITRE ATLAS show coverage, and the same entries feed ISO/IEC 42001 risk treatment records.
Failure modes
- Category entries. "Hallucination" cannot be scored or closed. Rewrite as cause, event, consequence.
- Typed residual scores. Without computed residuals, every control is assumed to work forever.
- Acceptance without expiry. A risk accepted for the launch model is silently accepted for every later model.
- No change triggers. A new tool or model version changes impact or likelihood overnight; the register should notice.
- Register as spreadsheet of record only. If it is not in review and CI, it is a snapshot of a past opinion.
Related reading
For framework alignment see the NIST AI Risk Management Framework and the OWASP Top 10 for LLM applications. To feed the register, use threat modelling for LLM applications and red teaming LLM systems; when a KRI turns into an incident, follow incident response for LLM systems.
What to do next
- Write score anchors for likelihood and impact, one sentence per level, and get them agreed before scoring anything.
- Rewrite every existing risk as cause, event and consequence for a named system; merge duplicates.
- List the controls for each entry with the axis they reduce, their strength and a link to their last test.
- Put the register in version control with an engine like the one above, and fail CI on expired acceptances.
- Define one KRI per medium or high entry, with a threshold and a data source, and wire it into review.
- Add trigger tags (new tool, model change, new data source) and generate the change set from your release manifest.