Ask a compliance team how their AI program is doing and you will usually get a slide with green circles. Ask what each circle measures, over which population, as of which date, and the answers get vague. That gap is the subject of this article. A compliance metric is only useful if it can be recomputed by someone who does not trust you, from data you did not curate by hand, and if moving it in the right direction actually reduces risk.
AI systems make this harder than ordinary IT compliance. The population changes weekly as teams wire new models into products; the controls include things with no binary outcome, such as evaluation quality or human-oversight effectiveness; and the evidence goes stale fast, because a model upgrade or prompt change can invalidate last month's red-team result. This article builds a measurement system from first principles: what a metric must declare, the three families worth tracking, how to compute them from control-test results with SQL, which metrics are specific to LLM systems, how to set thresholds, and the ways metrics get gamed. The program those metrics describe is covered in AI governance programs, in depth; here the focus is the measurement itself.
Why most compliance dashboards mislead
Most bad compliance metrics fail in one of four ways. The first is a missing denominator: "38 systems have model cards" says nothing unless you know whether there are 40 systems or 400. The second is an unstated population: does the count include vendor models embedded in SaaS tools, internal prototypes, or only systems in production? The third is stale evidence counted as current: a bias test run before the last retraining is evidence about a model that no longer exists. The fourth is measuring activity instead of outcome: "120 people completed AI training" is activity; "share of high-risk launches that passed the pre-launch gate on the first attempt" is closer to an outcome.
All four are fixed by the same discipline. Every metric is a declared query over system-of-record data, with a named population, a numerator, a denominator, a freshness rule and an owner. Nothing on the dashboard is typed in. If a value cannot be traced back to rows in the inventory, the control-test log and the evidence store, it does not go on the dashboard.
Three metric families: coverage, effectiveness, timeliness
Three families cover almost everything a board, regulator or auditor will ask about.
| Family | Question | Example metric | Leading or lagging |
|---|---|---|---|
| Coverage | Is every system that should be governed actually inside the program? | Share of production AI systems with an assigned risk tier | Leading |
| Effectiveness | When controls are tested, do they work? | Share of in-scope control tests passing in the last cycle | Mixed |
| Timeliness | Are problems found and fixed fast enough? | Median days from finding to closure, by severity | Lagging |
Coverage comes first because the other two are meaningless without it. A 100 percent pass rate over the 10 systems you know about tells you nothing about the 30 shadow deployments you do not. Effectiveness is where most of the risk signal sits. Timeliness shows whether the organisation reacts once it knows about a problem.
Leading indicators move before an incident: inventory coverage, evaluation freshness, open high-severity red-team findings. Lagging indicators move after: incidents per quarter, regulator inquiries, customer complaints tied to AI output. You need both. Leading indicators alone can look healthy while incidents climb, which usually means a control is being tested in a way that does not reflect real use. Lagging indicators alone tell you about damage after it has happened.
Metric definitions as code
Write metric definitions as data, check them into the same repository as the control library, and review changes to them the way you review code. A changed denominator is a changed metric, and the history must show when and why it changed. A minimal schema looks like this:
# metrics/eval_freshness.yaml
id: AIC-M-007
name: Evaluation freshness, high-risk systems
family: timeliness
owner: ai-risk@corp
population: >
inventory.systems WHERE tier = 'high' AND lifecycle = 'production'
numerator: >
systems whose latest passing eval run is newer than both
(a) 90 days and (b) the latest model or prompt version deploy
denominator: population
freshness: computed daily from eval_runs and deploy_events
thresholds: {green: ">= 0.95", amber: ">= 0.85", red: "< 0.85"}
excludes: systems with an approved exception (exception_id required)
version: 3
changelog:
- v3 2026-09: deploy-relative freshness added; v2 used calendar age onlyTwo fields deserve emphasis. excludes forces you to say what you leave out and to point every exclusion at an approved, expiring exception. Without it, people quietly drop the systems that would fail. The deploy-relative freshness rule captures the fact that evidence about a model is invalidated by changing the model, not by the calendar. A 20-day-old evaluation is stale if the system prompt changed yesterday.
The measurement pipeline
The pipeline has four inputs, all of which should already exist if the program is healthy. The inventory lists every AI system with its tier, lifecycle state, owner and jurisdictions. The control library lists the controls that apply to each tier. The control-test log records each test run with its outcome, tester and timestamp. The evidence store holds the artefacts with content hashes, so a passing test can point at exactly what it examined. Getting those artefacts collected is covered in AI audit preparation.
The metric engine is a scheduled job, typically SQL in the warehouse, that evaluates each definition and writes one row per metric per day into an append-only snapshot table. The dashboard reads only the snapshot table, which means any number shown for any date can be re-derived from that date's snapshot. When an auditor asks what you reported to the risk committee on 30 June, you can show the exact value and the query that produced it.
Computing effectiveness from control-test results
Here is the effectiveness metric written against a simple schema. Note that it reports the denominator alongside the ratio, and that it counts a control as failing if its latest test is older than the test frequency allows. That is the single most important rule for stopping stale passes from inflating the number.
-- Control effectiveness for production high-risk systems, as of :as_of
WITH scope AS (
SELECT s.system_id, cl.control_id, cl.test_every_days
FROM inventory_systems s
JOIN control_applicability cl ON cl.tier = s.tier
WHERE s.lifecycle = 'production' AND s.tier = 'high'
AND NOT EXISTS (SELECT 1 FROM exceptions e
WHERE e.system_id = s.system_id
AND e.control_id = cl.control_id
AND e.approved AND e.expires_on >= :as_of)
),
latest AS (
SELECT system_id, control_id, outcome, tested_at,
ROW_NUMBER() OVER (PARTITION BY system_id, control_id
ORDER BY tested_at DESC) AS rn
FROM control_tests WHERE tested_at <= :as_of
)
SELECT
COUNT(*) AS denominator,
SUM(CASE WHEN l.outcome = 'pass'
AND l.tested_at >= :as_of - sc.test_every_days
THEN 1 ELSE 0 END) AS numerator,
SUM(CASE WHEN l.system_id IS NULL THEN 1 ELSE 0 END) AS never_tested,
SUM(CASE WHEN l.outcome = 'pass'
AND l.tested_at < :as_of - sc.test_every_days
THEN 1 ELSE 0 END) AS stale_pass
FROM scope sc
LEFT JOIN latest l
ON l.system_id = sc.system_id AND l.control_id = sc.control_id AND l.rn = 1;Publishing never_tested and stale_pass next to the ratio matters. A drop in effectiveness that comes from stale passes means the testing calendar slipped; a drop from real failures means a control is broken. These need different fixes, and a single percentage hides which one you have.
Metrics specific to LLM systems
Generic IT-control metrics carry over, but LLM systems need several of their own. Each below is defined so that it can be computed from logs or test records, not estimated.
| Metric | Definition | Why it matters |
|---|---|---|
| Evaluation freshness | Share of systems whose latest passing eval postdates the latest model or prompt deploy | Model and prompt changes silently invalidate prior results |
| Red-team closure time | Median days from a high-severity red-team finding to verified fix | Shows whether adversarial testing changes anything |
| Guardrail drift | Week-over-week change in the input or output block rate, per system | A sudden fall often means a broken filter, not safer users |
| Human override rate | Share of AI recommendations that reviewers change, for human-in-the-loop systems | Near 0 percent suggests rubber-stamping; very high suggests a poor model |
| Disclosure coverage | Share of user-facing chat surfaces that show the required AI disclosure, checked by a synthetic probe | Transparency duties apply per surface, not per model |
| Incident reporting timeliness | Share of reportable incidents notified within the applicable window | Regulatory windows are short and start at awareness |
The human override rate needs care. A reviewer who accepts 99.8 percent of a model's suggestions may be reviewing nothing. Pair the metric with periodic seeded checks: insert known-bad recommendations into the review queue and measure how many are caught. That seeded catch rate measures whether oversight works, which the override rate alone cannot. The design of the review step itself is covered in human-in-the-loop controls.
Worked example: an insurer's September snapshot
Take a mid-sized insurer with 52 AI systems in the inventory: 9 high-risk (claims triage, fraud scoring, an underwriting assistant), 21 limited-risk (customer chat, document summarisation) and 22 minimal-risk internal tools. High-risk systems have 14 applicable controls each, tested quarterly or monthly.
In the September snapshot the effectiveness query returns a denominator of 9 x 14 = 126 control instances, minus 4 covered by approved exceptions, giving 122. Of those, 101 have a passing test within frequency, 7 are stale passes, 9 are failing and 5 have never been tested. Effectiveness is 101 / 122 = 82.8 percent, which is red against an 85 percent amber floor.
The breakdown changes the response. The 5 never-tested instances all belong to the underwriting assistant, launched in August without anyone adding it to the testing calendar; that is a process gap, and fixing it means making calendar enrolment part of the launch gate. The 7 stale passes are monthly prompt-injection tests that slipped because the red team was moved to another project; that is a capacity problem. The 9 failures are concentrated in two controls, logging retention and disclosure text, and both have a clear owner. The headline number fell, but the report now says exactly what to do, which a green circle never would.
Coverage is a separate check. A quarterly scan of cloud billing for AI API spend, plus egress logs for calls to model providers, turned up three unregistered systems. Coverage fell from 100 percent of known systems to 52 / 55 = 94.5 percent of discovered systems. Report the discovered figure. The known-systems figure is 100 percent by construction and therefore says nothing.
Thresholds, bands and trends
Thresholds turn a number into a decision, so derive them from risk appetite rather than from what the current value happens to be. A defensible approach has three steps. First, set the red line at the point where the risk committee would act: pause launches, escalate, or fund remediation. Second, set amber far enough above red to leave time to react, given how fast the metric usually moves. Third, show the trend over at least six snapshots, because a metric sitting at 86 percent and falling three points a month is a bigger problem than one sitting at 84 percent and rising.
Do not average across tiers. A 97 percent effectiveness figure that blends 22 minimal-risk tools with 9 high-risk systems is dominated by the systems that matter least. Report per tier, and if you must give one number, weight by tier or report the worst tier. The same logic applies to service levels in operations, covered in SLIs, SLOs and error budgets; a compliance threshold is effectively an SLO for risk controls.
Failure modes
- Denominator gaming. Systems get reclassified from high to limited risk just before the quarterly report. Detect it by reporting tier changes per period and requiring a recorded rationale for every downgrade.
- Exception inflation. Failing controls are moved into exceptions so they leave the denominator. Track the exception count and its age as a metric of its own, and make exceptions expire.
- Test-to-pass design. A control test that checks only that a document exists will pass when the document is empty. Sample a share of passing tests each quarter for quality review.
- Stale passes. Without the frequency rule in the query, an old pass counts forever. Always compare test age against the control's required frequency.
- Unreproducible history. Recomputing past months from current tables rewrites history when rows are corrected. Report from the immutable snapshot table.
- Too many metrics. Forty indicators means none is read. Keep about eight at committee level.
Trade-offs
| Choice | Benefit | Cost |
|---|---|---|
| Automated control tests | Daily freshness, cheap to rerun | Can only check what is machine-observable |
| Manual testing with sampling | Judges quality, not just existence | Slow, expensive, quarterly at best |
| Deploy-relative freshness | Catches evidence invalidated by changes | Teams that ship often must retest often |
| Discovered-population coverage | Honest about shadow AI | Needs billing and egress data access |
| Per-tier reporting | Risk is not diluted | More numbers for the committee |
| Strict exception expiry | No permanent carve-outs | Renewal workload for owners |
What to do next
- List the metrics you currently report and, for each, write down the population, numerator, denominator and data source. Retire any you cannot fill in.
- Put metric definitions in version control using a schema like the one above, with an owner and a changelog.
- Build the snapshot table and compute coverage, effectiveness and timeliness per tier from systems-of-record data only.
- Add the stale-pass rule and publish never-tested and stale-pass counts next to every effectiveness ratio.
- Run a discovery scan of billing and egress logs and report coverage against discovered systems, not known ones.
- Add two LLM-specific metrics this quarter, starting with evaluation freshness and seeded catch rate for human review.
- Agree red and amber thresholds with the risk committee and show a six-snapshot trend line beside every value. When you map metrics to framework outcomes, use the measurement function in the NIST AI RMF and the auditor's view in AI compliance audits.