A risk assessment estimates how likely and how bad each failure is before launch. A risk register records those estimates, owners and controls. Risk monitoring is the part that tells you, after launch, when those estimates have become wrong. Without it the register describes the system you planned rather than the one you run. Model upgrades, prompt edits, new user populations and new tools all change risk without anyone updating a spreadsheet.

This article is about the runtime pipeline: which signals to collect, how to measure rates of harmful behaviour you cannot label exhaustively, how to detect a change quickly without paging people every night, and how alerts flow back into the register and into incident response. It assumes you already have a register; AI Risk Register, in depth explains how to build one with testable risk statements and key risk indicators, and AI Risk Assessment, in depth covers the estimates it starts from.

What monitoring must prove

Monitoring is also an obligation in the main frameworks. The NIST AI Risk Management Framework's MANAGE function calls for post-deployment monitoring plans, including mechanisms for capturing user input, appeal, override, incident response and change management. The EU AI Act's Article 72 requires providers of high-risk AI systems to run a post-market monitoring system, based on a monitoring plan, that actively collects and analyses performance data across the system's lifetime. ISO/IEC 42001 expects monitoring, measurement, analysis and evaluation of the management system and its controls. None of them prescribes a technique. They require that you can show what you watch, why, how often, and what happened when it moved.

That suggests the test of a monitoring design: for every risk in the register with material residual exposure, can you name the signal that would move if the risk materialised, how quickly it would move, and who would be told? Risks without an answer are unmonitored, whatever the dashboard looks like.

Architecture of a monitoring pipeline

A runtime risk-monitoring pipeline: from raw events to decisions, and back to the registerGateway logsprompts, outputs, costGuardrail verdictsblocks, flagsTool-call tracesargs, resultsUser feedbackreports, thumbsChange eventsmodel, prompt, dataEvent streamrisk-tagged, sampledRule detectorscounts, thresholdsSampled judge + human calibrationrates with boundsChange detectorsCUSUM, distribution shiftDetector canariesmonitor the monitorsAlert routerpage, ticket, digestRisk registerrescore, ownerIncidentresponse planEval suitenew casesEvery alert names a register entry and an owner.
Signals are tagged with the risk they inform, fed to detectors suited to their statistics, and routed to an owner. Every alert updates the register, opens an incident or adds an evaluation case.

The pipeline has four stages. Collection captures events from the model gateway, guardrails, tool executions, user feedback and change management. Tag each event with the system, model version, prompt version and the register entries it can inform at write time, because reconstructing that later is unreliable. Detection turns events into measurements. Routing decides who is told and how urgently. Feedback changes something: a register score, an incident, a new evaluation case or a control. Reuse your audit trail for collection where you can; the integrity properties of that store are covered in LLM audit logging architecture.

Choosing signals

Signals differ in latency, cost and noise, and a good design mixes them deliberately.

RiskSignalHow measuredLatency
Unsupported or false claimsRate of answers judged unsupportedSampled LLM judge, human-calibratedHours
Sensitive data in outputsPII detector hits per 10,000 outputsRule detector on every outputSeconds
Prompt injectionInjection classifier flags; tool calls after untrusted contentClassifier plus trace rulesSeconds
Tool misuseDenied or unusual tool calls per sessionPolicy engine logsSeconds
Abuse by usersAccounts crossing abuse score thresholdsEntity scoringMinutes
Cost runawayTokens per session, loops per requestGateway countersSeconds
Silent changeModel, prompt or retrieval index version changesChange eventsImmediate
Population driftShift in topic or language mix of inputsDistribution distance on clustersDays

Cheap, exhaustive rules catch what can be written as a rule. Everything that needs judgement, such as whether an answer is supported by the retrieved documents, has to be sampled, and that is where most monitoring systems quietly go wrong. User-abuse signals have their own architecture, described in LLM Abuse Detection Architecture.

Measuring rates you cannot label exhaustively

Suppose you sample 300 conversations a day and an LLM judge flags 14 as making unsupported claims. Two corrections are needed before that number means anything. First, 300 is a small sample, so report an interval, not a point. The Wilson score interval behaves well at small rates. Second, the judge itself makes mistakes. Measure its true-positive rate and false-positive rate on a few hundred human-labelled examples, then correct the observed rate. The Rogan-Gladen estimator does this: true rate equals observed minus false-positive rate, divided by true-positive rate minus false-positive rate.

import math

def wilson(k, n, z=1.96):
    # 95% Wilson score interval for k flagged out of n sampled
    if n == 0:
        return (0.0, 1.0)
    ph = k / n
    d = 1 + z * z / n
    centre = (ph + z * z / (2 * n)) / d
    half = z * math.sqrt(ph * (1 - ph) / n + z * z / (4 * n * n)) / d
    return (max(0.0, centre - half), min(1.0, centre + half))

def corrected(q, tpr, fpr):
    # Rogan-Gladen: true prevalence from an observed flag rate and judge error rates
    return min(1.0, max(0.0, (q - fpr) / (tpr - fpr)))

lo, hi = wilson(14, 300)                 # 0.028 .. 0.077 observed
print(corrected(lo, 0.90, 0.02),         # about 0.009
      corrected(hi, 0.90, 0.02))         # about 0.065

The result is sobering. Fourteen flags in 300 is consistent with a true rate anywhere from about 0.9 to 6.5 percent. A judge with a 2 percent false-positive rate generates most of the flags when the true rate is around 1 percent, so a raw flag rate overstates the problem and hides changes in it. Recalibrate the judge whenever you change the judge model or prompt, and stratify samples by tenant or feature so a small population with a large problem is not averaged away.

Detecting change quickly

A fixed daily threshold is the usual first design, and it is a poor one: set it tight and it fires on noise, set it loose and a real doubling takes weeks to cross it. Sequential change detection uses every sample as it arrives. The Bernoulli CUSUM accumulates evidence that the flag rate has moved from an in-control rate p0 to an out-of-control rate p1, and alarms when the evidence passes a threshold h.

import math

class BernoulliCusum:
    def __init__(self, p0, p1, h):
        self.up = math.log(p1 / p0)                 # added for a flagged sample
        self.down = math.log((1 - p1) / (1 - p0))   # added (negative) for a clean sample
        self.h, self.s = h, 0.0

    def update(self, flagged):
        self.s = max(0.0, self.s + (self.up if flagged else self.down))
        if self.s > self.h:
            self.s = 0.0                            # reset after alarming
            return True
        return False

Choose h by simulation against your own rates, because it trades false alarms against detection delay. The average run length in control, the expected samples between false alarms, should be long relative to your sample rate; the run length out of control is your detection delay. Run the detector on the judge's observed flags and set p0 and p1 in observed terms, using the judge's error rates.

Not every change is a rate. For population drift, cluster input embeddings into topics on a reference window and compare each week's topic mix to it with a distance such as the population stability index. Drift is not harm by itself, but it says your evaluation set no longer represents your traffic.

Worked example: a support assistant after a model upgrade

A support assistant's register entry says: the assistant states refund or warranty terms not present in the policy documents in no more than 1.5 percent of conversations. Baseline measurement put the true rate at 1.2 percent. The judge has a true-positive rate of 0.90 and a false-positive rate of 0.02, so the expected observed flag rate is 1.2 percent times 0.90 plus 98.8 percent times 0.02, about 3.06 percent. The team wants to detect a rise to 3 percent true, which is about 4.64 percent observed.

  • With p0 = 0.0306, p1 = 0.0464 and h = 4, 400 simulated runs gave a mean in-control run length of about 19,000 samples and a mean detection delay of about 920 samples.
  • At 300 samples a day, that is a false alarm about every two months and detection in about three days.
  • Tripling the sample to 1,000 a day also triples false alarms per calendar day at the same threshold, to one every 19 days or so. Raising h to 5 restores the spacing: the simulation gave a mean in-control run of about 53,500 samples, roughly 53 days, and detection after about 1,200 samples, roughly 1.2 days.
  • A daily threshold on 300 samples cannot do this well: the interval computed above spans the whole range between the healthy and the degraded rate.

Then a model upgrade ships. The change event alone opens a ticket asking the owner to confirm the register entry still holds, and the team doubles the sample rate for a week, the cheapest way to shrink detection delay exactly when risk is highest. On day two the CUSUM alarms. The router opens a ticket rather than a page, because the register marks the entry moderate severity; the owner confirms on 50 human-labelled conversations that the new model paraphrases policy more freely, rolls the prompt back, adds the failing conversations to the evaluation suite and rescores the register entry with the measured rate.

Routing, ownership and monitoring the monitors

Monitoring fails socially more often than technically. Each alert must name the register entry, the owner, the evidence and the runbook step. Route by register severity: page only for entries where harm accrues within hours, such as data leakage or unsafe tool actions; ticket for rate changes; digest for drift. Track each alert's outcome and review the precision of every detector monthly. A detector whose alerts are mostly dismissed is retuned or removed, because ignored alerts train people to ignore the real one. When an alert does indicate harm, hand over to LLM incident response rather than improvising.

Monitor the monitors. Inject synthetic canaries, known-bad conversations tagged so they never reach users, and alarm if a detector fails to flag them. Alarm on silence too: a guardrail emitting zero events for an hour is more often broken than perfect. Report coverage as the share of high-residual register entries with at least one live detector.

Failure modes

Failure modeConsequenceCountermeasure
Raw judge rates reportedProblem overstated, changes hiddenCalibrate the judge; report corrected intervals
Fixed daily thresholdsNoise pages or slow detectionSequential detectors tuned by simulation
Unversioned eventsCannot attribute a change to a releaseTag model, prompt and index versions at write time
Averages across tenantsSmall severe pockets invisibleStratified sampling and per-segment detectors
Silent detector failureFalse confidenceCanaries and silence alarms
Alerts without ownersNobody actsEvery detector maps to a register entry and owner
Sampling only easy trafficLong or multilingual cases unmonitoredSample by risk-weighted strata

Trade-offs

  • Sample size versus cost. Detection delay falls roughly in proportion to samples per day, and judge calls are the main cost. Spend more right after changes and less in quiet periods.
  • Sensitivity versus fatigue. A lower h detects sooner and alarms falsely more often. Set it per risk from severity, not one global value.
  • LLM judges versus humans. Judges scale; humans are needed to calibrate them and to confirm alarms. Budget a fixed weekly human sample.
  • Privacy versus visibility. Monitoring needs content, but stored content is itself a risk. Redact before storage, keep sampled transcripts for a short retention period, and restrict access.

What to do next

  1. List every register entry with material residual risk and write down the signal that would move if it materialised.
  2. Tag gateway, guardrail and tool events with system, model, prompt and index versions at write time.
  3. Label a few hundred examples by hand and measure your judge's true-positive and false-positive rates.
  4. Report sampled rates as corrected Wilson intervals, stratified by tenant or feature.
  5. Replace fixed thresholds with a CUSUM per rate-based risk, choosing h by simulation.
  6. Route alerts by register severity, attach owner and runbook, and review detector precision monthly.
  7. Add canaries and silence alarms, and raise the sample rate automatically after every model or prompt change.
Key takeaway: Risk monitoring keeps the register honest after launch. Tie every signal to a register entry and owner, correct sampled judge rates for judge error and report them as intervals, detect change sequentially rather than with fixed thresholds, raise sampling after every change, and test that your detectors still fire.