Ask a compliance team how their AI program is doing and you will usually get a slide with green circles. Ask what each circle measures, over which population, as of which date, and the answers get vague. That gap is the subject of this article. A compliance metric is only useful if it can be recomputed by someone who does not trust you, from data you did not curate by hand, and if moving it in the right direction actually reduces risk.

AI systems make this harder than ordinary IT compliance. The population changes weekly as teams wire new models into products; the controls include things with no binary outcome, such as evaluation quality or human-oversight effectiveness; and the evidence goes stale fast, because a model upgrade or prompt change can invalidate last month's red-team result. This article builds a measurement system from first principles: what a metric must declare, the three families worth tracking, how to compute them from control-test results with SQL, which metrics are specific to LLM systems, how to set thresholds, and the ways metrics get gamed. The program those metrics describe is covered in AI governance programs, in depth; here the focus is the measurement itself.

Why most compliance dashboards mislead

Most bad compliance metrics fail in one of four ways. The first is a missing denominator: "38 systems have model cards" says nothing unless you know whether there are 40 systems or 400. The second is an unstated population: does the count include vendor models embedded in SaaS tools, internal prototypes, or only systems in production? The third is stale evidence counted as current: a bias test run before the last retraining is evidence about a model that no longer exists. The fourth is measuring activity instead of outcome: "120 people completed AI training" is activity; "share of high-risk launches that passed the pre-launch gate on the first attempt" is closer to an outcome.

All four are fixed by the same discipline. Every metric is a declared query over system-of-record data, with a named population, a numerator, a denominator, a freshness rule and an owner. Nothing on the dashboard is typed in. If a value cannot be traced back to rows in the inventory, the control-test log and the evidence store, it does not go on the dashboard.

Three metric families: coverage, effectiveness, timeliness

Three families cover almost everything a board, regulator or auditor will ask about.

FamilyQuestionExample metricLeading or lagging
CoverageIs every system that should be governed actually inside the program?Share of production AI systems with an assigned risk tierLeading
EffectivenessWhen controls are tested, do they work?Share of in-scope control tests passing in the last cycleMixed
TimelinessAre problems found and fixed fast enough?Median days from finding to closure, by severityLagging

Coverage comes first because the other two are meaningless without it. A 100 percent pass rate over the 10 systems you know about tells you nothing about the 30 shadow deployments you do not. Effectiveness is where most of the risk signal sits. Timeliness shows whether the organisation reacts once it knows about a problem.

Leading indicators move before an incident: inventory coverage, evaluation freshness, open high-severity red-team findings. Lagging indicators move after: incidents per quarter, regulator inquiries, customer complaints tied to AI output. You need both. Leading indicators alone can look healthy while incidents climb, which usually means a control is being tested in a way that does not reflect real use. Lagging indicators alone tell you about damage after it has happened.

Metric definitions as code

Write metric definitions as data, check them into the same repository as the control library, and review changes to them the way you review code. A changed denominator is a changed metric, and the history must show when and why it changed. A minimal schema looks like this:

# metrics/eval_freshness.yaml
id: AIC-M-007
name: Evaluation freshness, high-risk systems
family: timeliness
owner: ai-risk@corp
population: >
  inventory.systems WHERE tier = 'high' AND lifecycle = 'production'
numerator: >
  systems whose latest passing eval run is newer than both
  (a) 90 days and (b) the latest model or prompt version deploy
denominator: population
freshness: computed daily from eval_runs and deploy_events
thresholds: {green: ">= 0.95", amber: ">= 0.85", red: "< 0.85"}
excludes: systems with an approved exception (exception_id required)
version: 3
changelog:
  - v3 2026-09: deploy-relative freshness added; v2 used calendar age only

Two fields deserve emphasis. excludes forces you to say what you leave out and to point every exclusion at an approved, expiring exception. Without it, people quietly drop the systems that would fail. The deploy-relative freshness rule captures the fact that evidence about a model is invalidated by changing the model, not by the calendar. A 20-day-old evaluation is stale if the system prompt changed yesterday.

The measurement pipeline

A compliance metric is a query over control-test results, not a number someone types inAI system inventorythe denominatorControl librarywhat must be trueControl testsautomated + manualEvidence storehashes, timestampsMetric engineversioned definitionsnumerator / denominator / freshnessCoverageshare of systems in scopeEffectivenessshare of tests passingTimelinessevidence age, time to closeSnapshot archivepoint-in-time, immutableEvery number on the dashboard can be re-derived from the archive for any past date.If it cannot be re-derived, it is an opinion, and an auditor will treat it as one.
Inputs on the left are systems of record; the metric engine only reads them. Daily snapshots make every historical value reproducible.

The pipeline has four inputs, all of which should already exist if the program is healthy. The inventory lists every AI system with its tier, lifecycle state, owner and jurisdictions. The control library lists the controls that apply to each tier. The control-test log records each test run with its outcome, tester and timestamp. The evidence store holds the artefacts with content hashes, so a passing test can point at exactly what it examined. Getting those artefacts collected is covered in AI audit preparation.

The metric engine is a scheduled job, typically SQL in the warehouse, that evaluates each definition and writes one row per metric per day into an append-only snapshot table. The dashboard reads only the snapshot table, which means any number shown for any date can be re-derived from that date's snapshot. When an auditor asks what you reported to the risk committee on 30 June, you can show the exact value and the query that produced it.

Computing effectiveness from control-test results

Here is the effectiveness metric written against a simple schema. Note that it reports the denominator alongside the ratio, and that it counts a control as failing if its latest test is older than the test frequency allows. That is the single most important rule for stopping stale passes from inflating the number.

-- Control effectiveness for production high-risk systems, as of :as_of
WITH scope AS (
  SELECT s.system_id, cl.control_id, cl.test_every_days
  FROM inventory_systems s
  JOIN control_applicability cl ON cl.tier = s.tier
  WHERE s.lifecycle = 'production' AND s.tier = 'high'
    AND NOT EXISTS (SELECT 1 FROM exceptions e
                    WHERE e.system_id = s.system_id
                      AND e.control_id = cl.control_id
                      AND e.approved AND e.expires_on >= :as_of)
),
latest AS (
  SELECT system_id, control_id, outcome, tested_at,
         ROW_NUMBER() OVER (PARTITION BY system_id, control_id
                            ORDER BY tested_at DESC) AS rn
  FROM control_tests WHERE tested_at <= :as_of
)
SELECT
  COUNT(*)                                                    AS denominator,
  SUM(CASE WHEN l.outcome = 'pass'
            AND l.tested_at >= :as_of - sc.test_every_days
           THEN 1 ELSE 0 END)                                  AS numerator,
  SUM(CASE WHEN l.system_id IS NULL THEN 1 ELSE 0 END)        AS never_tested,
  SUM(CASE WHEN l.outcome = 'pass'
            AND l.tested_at < :as_of - sc.test_every_days
           THEN 1 ELSE 0 END)                                  AS stale_pass
FROM scope sc
LEFT JOIN latest l
  ON l.system_id = sc.system_id AND l.control_id = sc.control_id AND l.rn = 1;

Publishing never_tested and stale_pass next to the ratio matters. A drop in effectiveness that comes from stale passes means the testing calendar slipped; a drop from real failures means a control is broken. These need different fixes, and a single percentage hides which one you have.

Metrics specific to LLM systems

Generic IT-control metrics carry over, but LLM systems need several of their own. Each below is defined so that it can be computed from logs or test records, not estimated.

MetricDefinitionWhy it matters
Evaluation freshnessShare of systems whose latest passing eval postdates the latest model or prompt deployModel and prompt changes silently invalidate prior results
Red-team closure timeMedian days from a high-severity red-team finding to verified fixShows whether adversarial testing changes anything
Guardrail driftWeek-over-week change in the input or output block rate, per systemA sudden fall often means a broken filter, not safer users
Human override rateShare of AI recommendations that reviewers change, for human-in-the-loop systemsNear 0 percent suggests rubber-stamping; very high suggests a poor model
Disclosure coverageShare of user-facing chat surfaces that show the required AI disclosure, checked by a synthetic probeTransparency duties apply per surface, not per model
Incident reporting timelinessShare of reportable incidents notified within the applicable windowRegulatory windows are short and start at awareness

The human override rate needs care. A reviewer who accepts 99.8 percent of a model's suggestions may be reviewing nothing. Pair the metric with periodic seeded checks: insert known-bad recommendations into the review queue and measure how many are caught. That seeded catch rate measures whether oversight works, which the override rate alone cannot. The design of the review step itself is covered in human-in-the-loop controls.

Worked example: an insurer&#x27;s September snapshot

Take a mid-sized insurer with 52 AI systems in the inventory: 9 high-risk (claims triage, fraud scoring, an underwriting assistant), 21 limited-risk (customer chat, document summarisation) and 22 minimal-risk internal tools. High-risk systems have 14 applicable controls each, tested quarterly or monthly.

In the September snapshot the effectiveness query returns a denominator of 9 x 14 = 126 control instances, minus 4 covered by approved exceptions, giving 122. Of those, 101 have a passing test within frequency, 7 are stale passes, 9 are failing and 5 have never been tested. Effectiveness is 101 / 122 = 82.8 percent, which is red against an 85 percent amber floor.

The breakdown changes the response. The 5 never-tested instances all belong to the underwriting assistant, launched in August without anyone adding it to the testing calendar; that is a process gap, and fixing it means making calendar enrolment part of the launch gate. The 7 stale passes are monthly prompt-injection tests that slipped because the red team was moved to another project; that is a capacity problem. The 9 failures are concentrated in two controls, logging retention and disclosure text, and both have a clear owner. The headline number fell, but the report now says exactly what to do, which a green circle never would.

Coverage is a separate check. A quarterly scan of cloud billing for AI API spend, plus egress logs for calls to model providers, turned up three unregistered systems. Coverage fell from 100 percent of known systems to 52 / 55 = 94.5 percent of discovered systems. Report the discovered figure. The known-systems figure is 100 percent by construction and therefore says nothing.

Thresholds, bands and trends

Thresholds turn a number into a decision, so derive them from risk appetite rather than from what the current value happens to be. A defensible approach has three steps. First, set the red line at the point where the risk committee would act: pause launches, escalate, or fund remediation. Second, set amber far enough above red to leave time to react, given how fast the metric usually moves. Third, show the trend over at least six snapshots, because a metric sitting at 86 percent and falling three points a month is a bigger problem than one sitting at 84 percent and rising.

Do not average across tiers. A 97 percent effectiveness figure that blends 22 minimal-risk tools with 9 high-risk systems is dominated by the systems that matter least. Report per tier, and if you must give one number, weight by tier or report the worst tier. The same logic applies to service levels in operations, covered in SLIs, SLOs and error budgets; a compliance threshold is effectively an SLO for risk controls.

Failure modes

  • Denominator gaming. Systems get reclassified from high to limited risk just before the quarterly report. Detect it by reporting tier changes per period and requiring a recorded rationale for every downgrade.
  • Exception inflation. Failing controls are moved into exceptions so they leave the denominator. Track the exception count and its age as a metric of its own, and make exceptions expire.
  • Test-to-pass design. A control test that checks only that a document exists will pass when the document is empty. Sample a share of passing tests each quarter for quality review.
  • Stale passes. Without the frequency rule in the query, an old pass counts forever. Always compare test age against the control's required frequency.
  • Unreproducible history. Recomputing past months from current tables rewrites history when rows are corrected. Report from the immutable snapshot table.
  • Too many metrics. Forty indicators means none is read. Keep about eight at committee level.

Trade-offs

ChoiceBenefitCost
Automated control testsDaily freshness, cheap to rerunCan only check what is machine-observable
Manual testing with samplingJudges quality, not just existenceSlow, expensive, quarterly at best
Deploy-relative freshnessCatches evidence invalidated by changesTeams that ship often must retest often
Discovered-population coverageHonest about shadow AINeeds billing and egress data access
Per-tier reportingRisk is not dilutedMore numbers for the committee
Strict exception expiryNo permanent carve-outsRenewal workload for owners

What to do next

  1. List the metrics you currently report and, for each, write down the population, numerator, denominator and data source. Retire any you cannot fill in.
  2. Put metric definitions in version control using a schema like the one above, with an owner and a changelog.
  3. Build the snapshot table and compute coverage, effectiveness and timeliness per tier from systems-of-record data only.
  4. Add the stale-pass rule and publish never-tested and stale-pass counts next to every effectiveness ratio.
  5. Run a discovery scan of billing and egress logs and report coverage against discovered systems, not known ones.
  6. Add two LLM-specific metrics this quarter, starting with evaluation freshness and seeded catch rate for human review.
  7. Agree red and amber thresholds with the risk committee and show a six-snapshot trend line beside every value. When you map metrics to framework outcomes, use the measurement function in the NIST AI RMF and the auditor's view in AI compliance audits.
Key takeaway: A compliance metric is a declared query over the inventory, control tests and evidence, with a named population, an explicit denominator and a freshness rule. Measure coverage against discovered systems, count stale passes as failures, report per tier from immutable snapshots, and expect honest measurement to look worse before it looks better.