Most AI security guidance tells you what can go wrong: prompt injection, data poisoning, leaked system prompts. A maturity model answers a different question. How reliably does this organisation do the things that prevent those failures, and what should it improve next? A red-team finding is a snapshot of one system. A maturity score describes the process that will produce the next system.
The OWASP AI Maturity Assessment (AIMA) is OWASP's answer for AI. Version 1.0 is dated 11 August 2025. It ships with documentation and an Excel toolkit (v1.0.1 at the time of writing) and is listed as an OWASP Incubator project, so expect revisions. AIMA adapts the structure of OWASP SAMM, the Software Assurance Maturity Model, to the AI lifecycle. This page explains that structure, how to run an evidence-based assessment, scoring code, a worked example and common failures. Rubrics and scoring rules marked as this page's own are not OWASP text.
What AIMA is: domains, streams and levels
AIMA defines eight domains. Each has two streams. Stream A, Create & Promote, is about embedding a practice into teams and workflows. Stream B, Measure & Improve, is about monitoring that practice, measuring it and refining it. Each stream has three maturity levels, from ad hoc or reactive at Level 1 to continuous, automated and auditable at Level 3. That gives a grid of 8 x 2 x 3 = 48 cells. An assessment places the organisation, or one AI system, somewhere in each stream.
AIMA offers two ways to fill the grid. The lightweight assessment is a yes/no checklist. It is quick and good for a first baseline. The detailed assessment is evidence-backed. An assessor asks for artefacts and judges them, which is slower but defensible to an auditor or a board.
| Domain | AIMA's stated focus | Example evidence (this page's suggestion) |
|---|---|---|
| Responsible AI | ethical values, transparency, fairness, societal impact | impact assessments, fairness test results |
| Governance | strategy, policies, metrics, cross-role training | owned AI policy, AI system inventory |
| Data Management | data quality, governance, accountability, training practices | dataset lineage, poisoning screens |
| Privacy | minimisation, privacy-by-design, user control, transparency | data-flow maps for prompts and logs |
| Design | threat assessment, secure architecture, security requirements | threat models covering retrieval and tools |
| Implementation | secure build, deployment, defect management | pinned, scanned model artefacts |
| Verification | security testing, requirement-based validation, architecture review | red-team reports, regression evals in CI |
| Operations | incident response, event monitoring, operational stability | AI incident runbook, tool-call alerting |
Inherited from SAMM, and what is new
If you know SAMM, the shape is familiar. SAMM v2 has five business functions: Governance, Design, Implementation, Verification and Operations. Each has three security practices, each practice has two streams, and each stream has three maturity levels. Five of AIMA's eight domain names are exactly those business functions. The other three, Responsible AI, Data Management and Privacy, are where AI differs most from conventional software. In an AI system, data is part of the program. Model behaviour raises fairness and transparency questions that conventional software rarely does. Prompts and logs collect personal data in ways a typical web form never did.
An organisation already running SAMM can reuse its interviews and evidence store and add three domains. But inherited domains need AI-specific evidence: a threat model that ignores retrieval, tool calls and model supply chain is SAMM evidence, not AIMA evidence.
AIMA is not a risk list and not a certification. No body issues AIMA levels, and a self-assessed score means only as much as the evidence behind it.
Scoping: organisation or system
Decide the unit of assessment before the first interview. There are two common choices, and mixing them produces meaningless averages.
- Organisation-wide. One grid for the whole company. Useful for a board report and for Governance, where the policy really is shared. Misleading for Verification and Operations, where the flagship chatbot may be well tested while a team's internal summariser is not.
- Per AI system. One grid per inventoried system, with shared domains such as Governance scored once and inherited. More work, but it finds the weakest system, which is usually where the incident comes from.
A workable compromise is to score Governance, Responsible AI policy and Privacy policy once, and score Data Management, Design, Implementation, Verification and Operations per system for the systems that touch customers, personal data or tools with side effects. That makes the AI inventory a prerequisite. If you cannot list your AI systems with owners, data sources and tool permissions, that is itself a Governance Level 1 finding, and the inventory is your first deliverable. Interview the people who run each practice, not only its policy owners. When they disagree, the practitioner is usually right.
Evidence rules that make a level mean something
The detailed assessment stands or falls on evidence rules. AIMA describes levels qualitatively, so you need a written rubric that turns them into decisions an independent assessor would repeat. The rubric below is this page's convention, built to match AIMA's level descriptions. It is not OWASP text.
| Level | AIMA's description | Evidence test (this page's rubric) |
|---|---|---|
| 1 | ad hoc, reactive | The practice has happened at least once, for at least one system, and an artefact proves it. |
| 2 | not quoted; this page reads it as defined and repeatable | A written standard exists. Artefacts exist for most in-scope systems from the last two quarters, produced by more than one person. |
| 3 | continuous, automated, auditable | The practice runs automatically or on a schedule, coverage is tracked against the inventory, and exceptions are visible in a report someone reviews. |
Two evidence rules keep scores honest. Recency: an artefact older than the last major model or prompt change does not count for that system, because the system it described no longer exists. Sampling: the assessor picks which systems to sample from the inventory. The team being assessed does not choose, otherwise every sample is the showcase project. For Verification Stream B Level 3, a CI injection-regression suite with per-system pass rates and a blocked release is evidence. A slide saying red-teaming happens quarterly is not.
Scoring as data
Keep the answers as data, in version control, next to the evidence links. The scorer below uses SAMM's answer scale of 0, 0.25, 0.5 and 1 per question, a convention borrowed from SAMM rather than prescribed by AIMA. It adds one rule that spreadsheets usually skip: a level counts only if every level below it is also met, so a team cannot claim automation of a practice it has not defined.
from dataclasses import dataclass
from statistics import mean
DOMAINS = ["Responsible AI", "Governance", "Data Management", "Privacy",
"Design", "Implementation", "Verification", "Operations"]
SCALE = {"no": 0.0, "some": 0.25, "half": 0.5, "yes": 1.0} # SAMM-style answer scale
MET = 0.75 # this page's rule: a level is met when its questions average >= 0.75
@dataclass
class Answer:
domain: str
stream: str # "A" (Create & Promote) or "B" (Measure & Improve)
level: int # 1, 2 or 3
value: str # key into SCALE
evidence: list # links; empty evidence downgrades the answer to "no"
def stream_level(answers, domain, stream):
"""Highest level L such that levels 1..L are all met. 0 means not even Level 1."""
achieved = 0
for level in (1, 2, 3):
qs = [a for a in answers if (a.domain, a.stream, a.level) == (domain, stream, level)]
if not qs:
break
score = mean(SCALE[a.value] if a.evidence else 0.0 for a in qs)
if score < MET:
break
achieved = level
return achieved
def scorecard(answers, targets):
rows = []
for d in DOMAINS:
a, b = stream_level(answers, d, "A"), stream_level(answers, d, "B")
rows.append({"domain": d, "A": a, "B": b,
"floor": min(a, b), "target": targets.get(d, 2),
"gap": max(0, targets.get(d, 2) - min(a, b))})
return sorted(rows, key=lambda r: (-r["gap"], r["floor"]))The scorer reports the floor of the two streams, not their average. A domain where teams build threat models but nobody measures coverage is un-measured, not halfway mature. Sorting by gap yields the roadmap, and answers without evidence links score zero.
Worked example: a support assistant
A payments company runs three AI systems: a customer-facing support assistant with retrieval over help articles, an internal code-review assistant, and a fraud-scoring model trained on transaction data. It sets a target of Level 2 everywhere and Level 3 for Verification and Operations on the customer-facing assistant. Interviews and evidence produce this per-system result for the support assistant:
| Domain | A | B | Floor | Target | What the evidence showed |
|---|---|---|---|---|---|
| Verification | 2 | 0 | 0 | 3 | Red-team report exists, but no regression suite and no pass-rate tracking |
| Data Management | 1 | 0 | 0 | 2 | Help articles are indexed with no review of who can edit them |
| Operations | 2 | 1 | 1 | 3 | Runbook exists; tool calls are logged but nobody alerts on them |
| Privacy | 1 | 1 | 1 | 2 | Chat logs are kept 400 days, against a 90-day policy |
| Design | 2 | 1 | 1 | 2 | Threat model omits the retrieval index as an injection source |
| Governance | 2 | 2 | 2 | 2 | Inventory and policy in place, reviewed this year |
Sorted by gap, the top of the roadmap is Verification and Data Management. They connect. The retrieval index is editable by any support agent, and nothing tests whether a poisoned help article can steer the assistant. One quarter's plan follows directly from the evidence. Add edit approval and change logging to the help-article store (Data Management A to Level 2). Add an indirect-injection regression set that plants instructions in a test article and fails the build if the assistant follows them, tracked per release (Verification B to Level 2). Extend the threat model to the index (Design B to Level 2). Fix log retention, a configuration change with an owner and a date. The scorecard after the quarter is the acceptance test for the plan. An average would have shown Verification at 1.0; the floor of 0 shows nobody can tell whether defences regress.
Crosswalk to other frameworks
Few organisations adopt one framework alone. AIMA sits beside regulation and management-system standards, and a crosswalk lets one piece of evidence serve several of them.
| Framework | What it is | How AIMA relates |
|---|---|---|
| NIST AI RMF 1.0 | Risk framework with Govern, Map, Measure, Manage functions; no maturity levels | AIMA levels can serve as the scale the RMF lacks; Governance and Responsible AI evidence maps largely to Govern and Map |
| ISO/IEC 42001:2023 | Certifiable AI management system standard | Certification checks that the system exists; AIMA levels describe how well its processes run |
| OWASP Top 10 for LLM Applications | Catalogue of LLM application risks | Supplies the test content for Verification and the threat list for Design |
| OWASP SAMM v2 | Software assurance maturity model | AIMA's parent; reuse SAMM evidence where the AI parts are covered |
Map evidence artefacts to clauses, so one regression dashboard serves AIMA, the RMF and an ISO 42001 audit at once.
How maturity programmes fail
- Checkbox theatre. The lightweight yes/no form is answered by the policy owner in an hour, and every answer is yes. Fix: require an evidence link for every yes before it scores.
- Averaging hides zeros. A mean across domains or streams of 1.6 conceals an unmeasured Verification practice. Report floors and per-stream levels, and show zero cells in red.
- Showcase sampling. The flagship is assessed; shadow AI on a personal API key is not. Sample from the inventory and scan egress logs for unlisted AI services.
- Stale evidence. Invalidate evidence when a system's model, prompt or tools change.
- Score as goal. Pair every target level with an outcome metric, such as regressions caught before release.
- Framework drift. AIMA is an incubator project. Record the AIMA version in every scorecard, so a re-baseline after a revision is not mistaken for progress or regression.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Lightweight checklist | A baseline in days; good for awareness | Self-reported, optimistic, weak for auditors |
| Detailed, evidence-backed | Defensible, repeatable, finds the real gaps | Weeks of assessor time; needs an evidence store |
| Organisation-wide unit | Simple board story | Hides the weakest system |
| Per-system unit | Finds where incidents will come from | Scales with inventory size; needs inheritance rules |
| Level 3 targets everywhere | Ambitious | Automation cost spent where risk is low; target by system risk instead |
What to do next
- Build or update the AI system inventory with owners, data sources and tool permissions. Without it, sampling is impossible.
- Download the AIMA documentation and toolkit from the OWASP project page, and record the version you assess against.
- Run the lightweight checklist as a baseline, then convert it to a detailed assessment one domain at a time, starting with Verification and Data Management.
- Write down your evidence rubric, recency rule and sampling rule before the first interview.
- Store answers as data with evidence links, and score with stream floors and gaps, not averages.
- Set targets per system by risk, and re-assess each quarter against the same rubric.
- Keep learning: an evidence-based maturity scale for the NIST AI RMF, the OWASP Top 10 for LLM Applications, building an AI red-team programme, running an AI governance programme and an AI supply-chain security programme.