Sooner or later someone asks how mature your organisation is against the NIST AI Risk Management Framework, and wants a number. The awkward fact is that the AI RMF does not define maturity levels. It defines outcomes, and it offers profiles as a way to describe where you are and where you want to be. Any maturity score is therefore something you construct, and a badly constructed one is worse than none: it lets a single excellent policy hide five systems that nobody has assessed.
This article builds an evidence-based maturity scale over the RMF's outcome statements, shows how to score and aggregate it in code without averaging away gaps, works through an example organisation, and lists the ways these programs go wrong. It assumes you know the four functions; if not, start with NIST AI Risk Management Framework, in depth.
What the RMF says about maturity, and what it does not
AI RMF 1.0, published by NIST in January 2023 as NIST AI 100-1, is voluntary and outcome-based. Its core has four functions, GOVERN, MAP, MEASURE and MANAGE, broken into categories and then subcategories, which are the individual outcome statements you can assess. The generative AI profile, NIST AI 600-1, published in July 2024, adds risks and suggested actions specific to generative systems but still no levels. Profiles are the RMF's own mechanism for progress: a current profile describes outcomes achieved now, a target profile describes outcomes you intend to achieve, and the gap between them is your roadmap.
People often borrow the four Tiers from the NIST Cybersecurity Framework 2.0, which are Partial, Risk Informed, Repeatable and Adaptive. Use them with care. CSF Tiers characterise the rigour of an organisation's cybersecurity risk governance and management practices overall; NIST does not present them as a maturity model for scoring each outcome. NIST has also published a preliminary draft Cybersecurity Framework Profile for Artificial Intelligence, NIST IR 8596, which maps AI concerns onto the CSF 2.0 functions; its comment period closed on 30 January 2026, so check NIST for its current status before relying on its contents.
The shape of the core
The unit of assessment is the subcategory. The table gives the shape of the core; the subcategory counts are what make a spreadsheet approach feasible and a meeting-based approach painful.
| Function | Categories | Subcategories | What it asks for |
|---|---|---|---|
| GOVERN | 6 | 19 | Policies, accountability, culture, workforce, stakeholder engagement, third-party risk |
| MAP | 5 | 18 | Context, intended use, categorisation, capabilities, impacts on people |
| MEASURE | 4 | 22 | Methods and metrics, trustworthiness evaluation, tracking risks over time, feedback on measurement |
| MANAGE | 4 | 13 | Prioritised response, benefit maximisation, third-party risk, monitoring and communication |
| Total | 19 | 72 |
GOVERN is organisation-wide and cross-cutting. MAP, MEASURE and MANAGE apply per AI system, which is the first design decision for your scale: the organisational score for MAP is only meaningful as a roll-up of per-system scores.
An evidence-defined scale
A useful scale has levels defined by the evidence that would prove them, so two assessors looking at the same artifacts reach the same level. Here is a five-level scale that works across all 72 subcategories.
| Level | Name | Evidence required |
|---|---|---|
| 0 | Absent | Nothing, or only intent |
| 1 | Ad hoc | Some instances exist, owner-dependent, no written method |
| 2 | Defined | Written method or standard, named owner, approved |
| 3 | Operating | Records showing the method applied to every in-scope system in the last review period |
| 4 | Measured | Metrics on the activity itself, reviewed, with at least one documented improvement |
Three rules keep it honest. First, a level requires all lower levels: an excellent dashboard over an undocumented process is still level 1. Second, evidence expires; a record older than the review period does not count toward level 3. Third, for per-system subcategories, level 3 requires coverage of every in-scope system at the relevant risk tier, not a sample of the well-run ones.
What counts as evidence
Evidence is where most assessments become arguments. Agree in advance what counts for each kind of outcome, so that the assessor's job is checking artifacts, not judging intent. The table gives typical evidence at levels 2 and 3 for a few representative outcomes in each function.
| Function and outcome | Level 2 evidence | Level 3 evidence |
|---|---|---|
| GOVERN: roles and accountability | Approved RACI naming an accountable owner per system tier | Current owner recorded for every inventoried system, reviewed this period |
| GOVERN: third-party AI risk | Vendor AI review standard in procurement policy | Completed reviews for every AI vendor contract signed this period |
| MAP: intended use and context | Intended-use template with required fields | Completed, signed statements for every in-scope system |
| MEASURE: bias and fairness | Written test method with metrics and thresholds | Test reports per high-impact system, dated within the period |
| MEASURE: security and resilience | Red-team and evaluation plan | Findings and retest records bound to the deployed model version |
| MANAGE: incident response | AI-specific incident procedure | Exercise or real-incident records with follow-up actions closed |
Store each piece of evidence as a record with a subcategory reference, a kind, a date, the system it covers, and a link to the artifact. A record without a retrievable artifact is a claim, not evidence. Have the second line, risk or internal audit, re-perform a sample each cycle: pick a few level 3 claims at random and ask for the artifacts. The sampled pass rate is itself a useful health metric for the assessment process.
Overlaying the generative AI profile
If some systems are generative, overlay the generative AI profile on their scope. NIST AI 600-1 names twelve risks that are unique to or worsened by generative AI: CBRN information or capabilities, confabulation, dangerous, violent or hateful content, data privacy, environmental impacts, harmful bias and homogenisation, human-AI configuration, information integrity, information security, intellectual property, obscene, degrading or abusive content, and value chain and component integration. For each generative system, MAP should record which of these apply and why, and MEASURE should hold an evaluation record for each applicable risk. A system that maps confabulation as a relevant risk but has no measurement of it cannot claim level 3 for its measurement outcomes, however good the general evaluation method is. Treat the overlay as extra rows in the per-system table rather than a separate assessment.
Scoring as code
Scoring by hand drifts. Keep evidence as records and compute the levels, so a score can always be traced back to artifacts. The sketch below takes evidence records and an inventory and returns per-subcategory levels and a function roll-up.
from datetime import date, timedelta
PERIOD = timedelta(days=180)
def level(sub, evidence, systems, today):
ev = [e for e in evidence if e["sub"] == sub]
fresh = [e for e in ev if today - e["date"] <= PERIOD]
lvl = 0
if ev:
lvl = 1
if lvl == 1 and any(e["kind"] == "method" and e.get("approved") for e in ev):
lvl = 2
if lvl == 2:
if sub.startswith("GOVERN"):
covered = any(e["kind"] == "record" for e in fresh)
else:
done = {e["system"] for e in fresh if e["kind"] == "record"}
covered = systems <= done # every in-scope system
if covered:
lvl = 3
if lvl == 3 and any(e["kind"] == "improvement" for e in fresh):
lvl = 4
return lvl
def rollup(levels):
by_fn = {}
for sub, lvl in levels.items():
by_fn.setdefault(sub.split(" ")[0], []).append((lvl, sub))
return {fn: {"floor": min(v)[0], "weakest": min(v)[1],
"mean": round(sum(l for l, _ in v) / len(v), 2)}
for fn, v in by_fn.items()}The roll-up reports both the floor and the mean, deliberately. The mean shows breadth of effort; the floor is the level you can actually claim for that function, because an assessor or regulator will find the weakest outcome. Present them side by side and never a mean alone.
Worked example: fourteen systems
Take a company with 14 AI systems in its inventory, four of them rated high impact. GOVERN looks strong at first: an approved AI policy, a named accountable executive and a governance council, all level 2 or 3. But the subcategory on third-party AI risk is level 1, because vendor models are reviewed only when a security questionnaire happens to ask. GOVERN's floor is therefore 1, not the 2.6 mean the slide deck showed.
MAP is uneven. Intended-use statements exist for 11 of 14 systems; because three are missing, the subcategory on documenting intended purpose cannot reach level 3 even though the template is good. MEASURE has a defined evaluation method, but bias testing records exist for only two of the four high-impact systems. MANAGE has no written incident response procedure specific to AI, so the subcategories on responding to and recovering from AI incidents sit at 0.
The target profile the board approved is level 3 for every outcome within two cycles. Sorting gaps by size and dependency gives a short, defensible plan: write the AI incident procedure (MANAGE 0 to 2, then 3 after an exercise), finish intended-use statements for three systems (MAP 2 to 3), run bias testing on the two remaining high-impact systems (MEASURE 2 to 3), and add AI vendor review to procurement (GOVERN 1 to 2, then 3 once reviews are on record). Four work items move every function floor; polishing the policy would move none.
When reporting to the board, show three things on one page: the floor and mean per function against target, the four work items with owners and dates, and the evidence sample pass rate. Resist a single headline number. A board that sees MANAGE at 0 with a dated plan to reach 2 has learned something it can act on; a board that sees an overall 2.1 has not, and will be surprised by the first AI incident.
Sequencing the roadmap
Sequence matters because outcomes depend on each other. You cannot map a system you have not inventoried, measure risks you have not mapped, or manage risks you have not measured. A practical order:
- Inventory and tiering, the GOVERN outcomes that make the scope countable.
- Context and intended use for high-impact systems first, in MAP.
- Evaluation methods tied to the risks MAP surfaced, in MEASURE.
- Response, monitoring and decommissioning procedures, in MANAGE.
- Only then, measurement of the program itself, which is what level 4 asks for.
Reassess on a fixed cycle, such as every six months, and report the change in floors. Programs that already hold an ISO/IEC 42001 AI management system certification can reuse much of its evidence, since an audited management system produces exactly the method and record artifacts that levels 2 and 3 need; map controls once and tag evidence with both references. For governance structure, see AI Governance Council, in depth and AI Governance Program Structure.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Averaging | Score of 2.7 while one function has zero incident capability | Report floors per function |
| Policy scoring | Levels assigned because a document exists | Level 3 needs operating records per system |
| Self-assessment inflation | Every team rates itself 3 | Evidence-defined levels, sampled second-line review |
| Stale evidence | Last bias test was 18 months ago | Evidence expiry built into the scorer |
| Invented tiers | Claims of being at an official NIST AI level | State that the scale is internal and how it is defined |
| Scope games | Hard systems left out of the inventory | Inventory completeness is itself a scored outcome |
Trade-offs
Granularity costs effort. Scoring 19 organisation-wide GOVERN outcomes plus 53 per-system outcomes across 14 systems is about 760 judgements per cycle, which is why evidence must be collected as a by-product of work rather than gathered for the assessment. Some organisations score at category level to save time; that is acceptable for a first baseline but hides exactly the weak subcategory that matters. A strict floor can feel punitive and discourage teams; pairing it with the mean shows progress while keeping the claim honest. An external assessor adds credibility and cost; use one for the high-impact scope and self-assess the rest. For how the security leader owns this reporting, see CISO Role in AI Security.
What to do next
- Write down, in one paragraph, that your maturity scale is internal, built on AI RMF subcategories, and how levels are defined.
- Load the 72 subcategories into a table with a column per in-scope system for MAP, MEASURE and MANAGE.
- Define evidence kinds (method, record, improvement) and an expiry period.
- Score the current profile from evidence only, and compute floors and means per function.
- Agree a target profile by risk tier with the accountable executive.
- Turn the largest floor gaps into no more than five work items with owners and dates.
- Reassess on a fixed cycle and report floor movement, not just mean movement.