The NIST AI Risk Management Framework is the reference most US organisations reach for when someone asks how they manage AI risk. It is also widely misunderstood: people treat it as a checklist, a certification or a set of technical controls, and it is none of those. It is a structured vocabulary of outcomes that an organisation should achieve, organised so that different teams can agree on who does what.

This article reads the framework the way an engineer would: what each part says, how the parts fit together, and how to turn them into artefacts a team can maintain, namely an inventory, a risk register held as data, an evidence store and a release gate. Comparisons with ISO/IEC 42001 and the EU AI Act are covered in AI Safety Frameworks, in depth; this page stays inside the RMF and its generative AI profile, and works one example through end to end.

Advertisement

What the RMF is, and what it is not

NIST published AI RMF 1.0 as NIST AI 100-1 on 26 January 2023. It is voluntary, it is not a regulation, and there is no RMF certification: nobody can audit you and declare you compliant with it in the way an ISO management-system standard allows. It is also deliberately technology-neutral and sector-neutral, which is why it describes outcomes such as "legal and regulatory requirements involving AI are understood, managed, and documented" rather than prescribing specific tests.

In July 2024 NIST added NIST AI 600-1, the Generative AI Profile, which applies the framework to generative models. As of this writing the core document is still version 1.0; NIST has signalled that revisions and further profiles are coming, so check the NIST AI RMF page before citing version details in a policy. Around the documents sits the AI RMF Playbook, which lists suggested actions, references and documentation for each subcategory. None of the Playbook is mandatory either.

Why use a voluntary framework at all? Because it gives you shared names. A procurement questionnaire, an internal auditor, a regulator's guidance and your own engineering backlog can all point at "MEASURE 2.7" and mean the same thing. US federal policy has also leaned on it at various points, which the US AI executive orders article traces.

Trustworthiness: the seven characteristics

Part 1 of the framework defines what it is trying to protect. A trustworthy AI system is described by seven characteristics: valid and reliable; safe; secure and resilient; accountable and transparent; explainable and interpretable; privacy-enhanced; and fair with harmful bias managed. Valid and reliable is treated as the base the others rest on, and accountable and transparent as cutting across all of them.

The characteristics trade off against one another, and the framework says so. Stronger privacy protection can reduce accuracy; more interpretability can mean a simpler, less capable model; tighter safety filters can harm some users' access. The RMF does not resolve these tensions. It asks you to make the trade-off explicit, decide it at the right level of the organisation, and record the decision. For engineers this is the main practical point: when two metrics conflict, the answer is a documented decision with an owner, not a silent default in a config file.

Advertisement

The core: four functions and nineteen categories

Part 2 is the core. It has four functions. GOVERN is cross-cutting and applies throughout; MAP, MEASURE and MANAGE are applied to specific systems and repeat over the lifecycle. Each function is divided into categories, and each category into subcategories, which are the individual outcome statements.

FunctionCategoriesWhat the categories cover
GOVERN6Policies and practices in place; accountability structures; workforce diversity, equity, inclusion and accessibility; a culture that considers and communicates risk; engagement with relevant AI actors; third-party software, data and supply chain
MAP5Context established; system categorised; capabilities, goals, benefits and costs understood; risks mapped for all components including third parties; impacts on individuals, groups, communities, organisations and society characterised
MEASURE4Methods and metrics identified and applied; systems evaluated for trustworthy characteristics; mechanisms to track risks over time; feedback on whether measurement works
MANAGE4Risks prioritised and responded to; strategies to maximise benefit and minimise harm; third-party risks managed; treatments, response, recovery and communication documented and monitored
The AI RMF core as an engineering loopGOVERN (cross-cutting): policy, inventory, accountability, culture, third partiesMAPcontext, purpose, impactsMEASUREtests, metrics, trackingMANAGEprioritise, treat, monitorrisksresultsAI system inventoryGOVERN 1.6Risk registerrisks as dataEvidence storeevals, red team, runbooksscopewriteattachincidents and monitoring feed back into MAPProfiles (use-case, current and target) select which outcomes apply; the Playbook suggests actions for each.NIST AI 600-1 adds generative-AI risks and actions keyed to the same subcategories.
GOVERN wraps the system-level loop. MAP scopes risks for each inventoried system, MEASURE produces evidence, MANAGE decides and monitors, and incidents feed back into MAP. The three artefacts at the bottom are what an engineering team actually maintains.

The order MAP, MEASURE, MANAGE is a dependency, not a waterfall. You cannot measure a risk you have not mapped, and you cannot manage one you have not measured, but in practice the loop runs continuously: a production incident reopens MAP for that system.

Reading subcategories

Subcategories are the level at which work is assigned. Each has an identifier such as MAP 1.1 and a sentence describing an outcome. A few examples, quoted from the core, show the range:

  • GOVERN 1.6: "Mechanisms are in place to inventory AI systems and are resourced according to organizational risk priorities." This is the inventory, and most programmes should start here.
  • MAP 4.1: approaches for mapping technology and legal risks of components, including third-party data or software, are in place, followed and documented, including infringement risks.
  • MEASURE 2.7: "AI system security and resilience – as identified in the map function – are evaluated and documented." Red-team results and penetration tests are evidence here.
  • MEASURE 2.11: "Fairness and bias – as identified in the map function – are evaluated and results are documented."
  • MANAGE 2.4: mechanisms and assigned responsibilities exist "to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use." In engineering terms: a tested kill switch and a named owner.
  • MANAGE 4.1: post-deployment monitoring plans are implemented, including user feedback, appeal and override, decommissioning, incident response, recovery and change management.

Notice the phrase "as identified in the map function" in the MEASURE items. Measurement is scoped by mapping: you evaluate the security properties that matter for this system's context, not every property for every system. Notice too that the outcomes say what must be true, never how. Translating each into a control, an owner and a type of evidence is your job, and that translation is what the rest of this article does.

Profiles and the Playbook

A profile is a selection and tailoring of the core for a particular setting. NIST describes use-case profiles, for a sector or application, and temporal profiles: a current profile describing the outcomes you achieve today and a target profile describing where you want to be. The gap between the two is your roadmap. Building both is a good first exercise because it forces a decision on which subcategories apply.

The Playbook is the companion resource for each subcategory: suggested actions, transparency and documentation questions, and references. Use it as a source of ideas when you are deciding how to meet an outcome, not as a list to complete. Teams that try to perform every suggested action for every system tend to produce paperwork and little risk reduction.

The generative AI profile: twelve risks

NIST AI 600-1 names twelve risks that are unique to or exacerbated by generative AI, and proposes actions mapped to the same GOVERN, MAP, MEASURE and MANAGE subcategories:

RiskEngineering reading
CBRN Information or CapabilitiesUplift toward chemical, biological, radiological or nuclear harm; refusal and capability evaluations
ConfabulationConfident false output; grounding checks, citations, abstention
Dangerous, Violent, or Hateful ContentOutput filtering and policy evaluation
Data PrivacyTraining-data leakage, memorisation, inference about individuals
Environmental ImpactsEnergy and resource cost of training and inference
Harmful Bias and HomogenizationSkewed outputs; reduced diversity of outputs across users
Human-AI ConfigurationOver-reliance, automation bias, anthropomorphism
Information IntegrityMisinformation and synthetic content at scale; provenance
Information SecurityPrompt injection, data poisoning, model-assisted attacks
Intellectual PropertyReproduction of protected material
Obscene, Degrading, and/or Abusive ContentIncluding synthetic sexual imagery; filters and provenance
Value Chain and Component IntegrationThird-party models, datasets and plugins whose changes you do not control

The engineering reading column is our interpretation, not NIST's text. The profile's value is that each suggested action already carries a subcategory identifier, so generative AI risks slot into the same register as everything else instead of living in a separate document.

Worked example: a register as code

Take a customer-support assistant that can look up orders and issue refunds through a tool. MAP produces three risks. R1: the assistant states the wrong refund policy (confabulation). R2: text in a customer message makes the model call the refund tool (information security). R3: the vendor updates the hosted model without notice (value chain). Each risk is tagged with the subcategories it must satisfy, a likelihood and impact on a 1 to 5 scale, an owner, a review date and links to evidence.

from datetime import date

RISKS = [
    dict(id="R1", title="Assistant states wrong refund policy", gai="Confabulation",
         likelihood=4, impact=3, subcats=["MAP 1.1", "MEASURE 3.1"],
         evidence={"MEASURE 3.1": "eval/grounding_weekly.json"}, owner="support-eng",
         review_by=date(2026, 11, 1)),
    dict(id="R2", title="Prompt injection triggers refund tool", gai="Information Security",
         likelihood=3, impact=5, subcats=["MEASURE 2.7", "MANAGE 2.4"],
         evidence={"MEASURE 2.7": "redteam/2026-09.md", "MANAGE 2.4": "runbooks/kill_switch.md"},
         owner="appsec", review_by=date(2026, 10, 15)),
    dict(id="R3", title="Vendor model silently updated", gai="Value Chain and Component Integration",
         likelihood=3, impact=3, subcats=["MAP 4.1", "MANAGE 4.1"], evidence={},
         owner=None, review_by=date(2026, 9, 1)),
]

def check(register, today, appetite=12):
    findings = []
    for r in register:
        score = r["likelihood"] * r["impact"]
        missing = [s for s in r["subcats"] if s not in r["evidence"]]
        if missing:
            findings.append(f"{r['id']} score {score}: no evidence for {', '.join(missing)}")
        if r["owner"] is None:
            findings.append(f"{r['id']}: no owner (GOVERN 2)")
        if r["review_by"] < today:
            findings.append(f"{r['id']}: review overdue since {r['review_by']}")
        if score > appetite and missing:
            findings.append(f"{r['id']}: above appetite {appetite} with open gaps -> block release")
    return findings

for f in check(RISKS, date(2026, 10, 2)):
    print(f)

# Output:
# R1 score 12: no evidence for MAP 1.1
# R3 score 9: no evidence for MAP 4.1, MANAGE 4.1
# R3: no owner (GOVERN 2)
# R3: review overdue since 2026-09-01

Read the output against the inputs. R1 scores 4 times 3 = 12 and has evidence for MEASURE 3.1 (a weekly grounding evaluation) but not MAP 1.1, because nobody has documented the intended use and its limits. R2 has the highest score, 15, but both its subcategories have evidence, a red-team report and a kill-switch runbook, so it produces no finding. R3 has no evidence, no owner and a review date in the past. The release rule blocks only when a risk is above the appetite of 12 and still has gaps; none does today, but if R1's likelihood rose to 5 its score of 15 with a missing MAP 1.1 would block the release.

This is deliberately small, but it shows the shape: risks are data in version control, evidence is a path to an artefact produced by a pipeline, and the check runs in CI. The same structure scales to hundreds of systems and is what an auditor will sample; LLM Audit Preparation covers producing those populations.

Failure modes

  • Checklist theatre. Marking every subcategory done with a policy document and no system-level evidence. Fix: require a link to an artefact produced by a running process for each MEASURE and MANAGE item.
  • No inventory. Without GOVERN 1.6 you cannot know what you are managing; shadow uses of hosted models are the usual gap. Fix: discover systems from cloud bills, API keys and network egress, not only from surveys.
  • Mapping once. Context changes when a feature adds a tool, a new user group or a new jurisdiction. Fix: make significant changes trigger a MAP review in the same pipeline that ships them.
  • Unowned trade-offs. Accuracy versus privacy, or helpfulness versus safety, decided by whoever wrote the config. Fix: record the decision, the owner and the metric thresholds in the register.
  • Untested kill switches. MANAGE 2.4 satisfied on paper by a runbook nobody has run. Fix: exercise it on a schedule and attach the result as evidence.
  • Treating the RMF as legal compliance. It is not a substitute for binding obligations. Map those separately; see AI Regulation Deep Dive and EU AI Act, in depth.

Trade-offs in adopting it

The RMF's breadth is both strength and cost. Applying every subcategory to every system wastes effort on low-risk tools, so tier your systems and apply the full core only to the higher tiers. Its outcome language gives flexibility but leaves room for weak interpretations, so write down your own control for each outcome you adopt. And because it is voluntary and not certifiable, it works best as the internal backbone that other obligations map onto, rather than as the thing you show externally.

What to do next

  1. Build an inventory of every AI system, including hosted model APIs used by scripts and internal tools, and assign each an owner and a risk tier.
  2. For one high-tier system, write a current profile listing which subcategories you meet today with evidence, and a target profile.
  3. Run MAP for that system: document intended use, users, context and impacts, and list risks using the twelve generative AI risks as prompts.
  4. Put the risks in a register stored as data, tag each with subcategories, and add a CI check like the one above.
  5. Connect evidence to pipelines: evaluation results, red-team reports and incident records written automatically to the evidence store.
  6. Test the deactivation path for that system and record the result against MANAGE 2.4.
  7. Schedule a quarterly review and a trigger that reopens MAP whenever the system gains a new tool, data source or user group.
Key takeaway: The NIST AI RMF is a voluntary, outcome-based vocabulary, not a certification or a control catalogue. Its core has four functions and nineteen categories: GOVERN sets policy, inventory and accountability across the organisation; MAP, MEASURE and MANAGE scope, evidence and treat risks system by system in a continuous loop. Profiles tailor it, the Playbook suggests actions, and NIST AI 600-1 adds twelve generative AI risks keyed to the same subcategories. Get value from it by turning outcomes into an inventory, a risk register held as data, evidence produced by pipelines, and a release gate that blocks when high risks lack evidence.