Most AI security guidance tells you what can go wrong: prompt injection, data poisoning, leaked system prompts. A maturity model answers a different question. How reliably does this organisation do the things that prevent those failures, and what should it improve next? A red-team finding is a snapshot of one system. A maturity score describes the process that will produce the next system.

The OWASP AI Maturity Assessment (AIMA) is OWASP's answer for AI. Version 1.0 is dated 11 August 2025. It ships with documentation and an Excel toolkit (v1.0.1 at the time of writing) and is listed as an OWASP Incubator project, so expect revisions. AIMA adapts the structure of OWASP SAMM, the Software Assurance Maturity Model, to the AI lifecycle. This page explains that structure, how to run an evidence-based assessment, scoring code, a worked example and common failures. Rubrics and scoring rules marked as this page's own are not OWASP text.

What AIMA is: domains, streams and levels

AIMA defines eight domains. Each has two streams. Stream A, Create & Promote, is about embedding a practice into teams and workflows. Stream B, Measure & Improve, is about monitoring that practice, measuring it and refining it. Each stream has three maturity levels, from ad hoc or reactive at Level 1 to continuous, automated and auditable at Level 3. That gives a grid of 8 x 2 x 3 = 48 cells. An assessment places the organisation, or one AI system, somewhere in each stream.

AIMA offers two ways to fill the grid. The lightweight assessment is a yes/no checklist. It is quick and good for a first baseline. The detailed assessment is evidence-backed. An assessor asks for artefacts and judges them, which is slower but defensible to an auditor or a board.

OWASP AIMA v1.0: eight domains, two streams each, three levels per streamResponsible AIStream ACreate & PromoteStream BMeasure & ImproveL1 / L2 / L3 in each streamGovernanceStream ACreate & PromoteStream BMeasure & ImproveL1 / L2 / L3 in each streamData ManagementStream ACreate & PromoteStream BMeasure & ImproveL1 / L2 / L3 in each streamPrivacyStream ACreate & PromoteStream BMeasure & ImproveL1 / L2 / L3 in each streamDesignStream ACreate & PromoteStream BMeasure & ImproveL1 / L2 / L3 in each streamImplementationStream ACreate & PromoteStream BMeasure & ImproveL1 / L2 / L3 in each streamVerificationStream ACreate & PromoteStream BMeasure & ImproveL1 / L2 / L3 in each streamOperationsStream ACreate & PromoteStream BMeasure & ImproveL1 / L2 / L3 in each streamPink, amber, violet: the three domains AIMA adds. Green: the five SAMM v2 business-function names.8 domains x 2 streams x 3 levels = 48 cells. An assessment answers each cell with evidence, then scores it.lightweightyes/no checklistdetailedevidence-backed reviewroadmaptarget level per domain
Figure: the AIMA grid. Each domain splits into a build-it stream and a measure-it stream, and each stream climbs three levels. The domain names are AIMA's. The colour grouping is this page's reading of where they came from.
DomainAIMA's stated focusExample evidence (this page's suggestion)
Responsible AIethical values, transparency, fairness, societal impactimpact assessments, fairness test results
Governancestrategy, policies, metrics, cross-role trainingowned AI policy, AI system inventory
Data Managementdata quality, governance, accountability, training practicesdataset lineage, poisoning screens
Privacyminimisation, privacy-by-design, user control, transparencydata-flow maps for prompts and logs
Designthreat assessment, secure architecture, security requirementsthreat models covering retrieval and tools
Implementationsecure build, deployment, defect managementpinned, scanned model artefacts
Verificationsecurity testing, requirement-based validation, architecture reviewred-team reports, regression evals in CI
Operationsincident response, event monitoring, operational stabilityAI incident runbook, tool-call alerting

Inherited from SAMM, and what is new

If you know SAMM, the shape is familiar. SAMM v2 has five business functions: Governance, Design, Implementation, Verification and Operations. Each has three security practices, each practice has two streams, and each stream has three maturity levels. Five of AIMA's eight domain names are exactly those business functions. The other three, Responsible AI, Data Management and Privacy, are where AI differs most from conventional software. In an AI system, data is part of the program. Model behaviour raises fairness and transparency questions that conventional software rarely does. Prompts and logs collect personal data in ways a typical web form never did.

An organisation already running SAMM can reuse its interviews and evidence store and add three domains. But inherited domains need AI-specific evidence: a threat model that ignores retrieval, tool calls and model supply chain is SAMM evidence, not AIMA evidence.

AIMA is not a risk list and not a certification. No body issues AIMA levels, and a self-assessed score means only as much as the evidence behind it.

Scoping: organisation or system

Decide the unit of assessment before the first interview. There are two common choices, and mixing them produces meaningless averages.

  • Organisation-wide. One grid for the whole company. Useful for a board report and for Governance, where the policy really is shared. Misleading for Verification and Operations, where the flagship chatbot may be well tested while a team's internal summariser is not.
  • Per AI system. One grid per inventoried system, with shared domains such as Governance scored once and inherited. More work, but it finds the weakest system, which is usually where the incident comes from.

A workable compromise is to score Governance, Responsible AI policy and Privacy policy once, and score Data Management, Design, Implementation, Verification and Operations per system for the systems that touch customers, personal data or tools with side effects. That makes the AI inventory a prerequisite. If you cannot list your AI systems with owners, data sources and tool permissions, that is itself a Governance Level 1 finding, and the inventory is your first deliverable. Interview the people who run each practice, not only its policy owners. When they disagree, the practitioner is usually right.

Evidence rules that make a level mean something

The detailed assessment stands or falls on evidence rules. AIMA describes levels qualitatively, so you need a written rubric that turns them into decisions an independent assessor would repeat. The rubric below is this page's convention, built to match AIMA's level descriptions. It is not OWASP text.

LevelAIMA's descriptionEvidence test (this page's rubric)
1ad hoc, reactiveThe practice has happened at least once, for at least one system, and an artefact proves it.
2not quoted; this page reads it as defined and repeatableA written standard exists. Artefacts exist for most in-scope systems from the last two quarters, produced by more than one person.
3continuous, automated, auditableThe practice runs automatically or on a schedule, coverage is tracked against the inventory, and exceptions are visible in a report someone reviews.

Two evidence rules keep scores honest. Recency: an artefact older than the last major model or prompt change does not count for that system, because the system it described no longer exists. Sampling: the assessor picks which systems to sample from the inventory. The team being assessed does not choose, otherwise every sample is the showcase project. For Verification Stream B Level 3, a CI injection-regression suite with per-system pass rates and a blocked release is evidence. A slide saying red-teaming happens quarterly is not.

Scoring as data

Keep the answers as data, in version control, next to the evidence links. The scorer below uses SAMM's answer scale of 0, 0.25, 0.5 and 1 per question, a convention borrowed from SAMM rather than prescribed by AIMA. It adds one rule that spreadsheets usually skip: a level counts only if every level below it is also met, so a team cannot claim automation of a practice it has not defined.

from dataclasses import dataclass
from statistics import mean

DOMAINS = ["Responsible AI", "Governance", "Data Management", "Privacy",
           "Design", "Implementation", "Verification", "Operations"]
SCALE = {"no": 0.0, "some": 0.25, "half": 0.5, "yes": 1.0}   # SAMM-style answer scale
MET = 0.75   # this page's rule: a level is met when its questions average >= 0.75

@dataclass
class Answer:
    domain: str
    stream: str          # "A" (Create & Promote) or "B" (Measure & Improve)
    level: int           # 1, 2 or 3
    value: str           # key into SCALE
    evidence: list       # links; empty evidence downgrades the answer to "no"

def stream_level(answers, domain, stream):
    """Highest level L such that levels 1..L are all met. 0 means not even Level 1."""
    achieved = 0
    for level in (1, 2, 3):
        qs = [a for a in answers if (a.domain, a.stream, a.level) == (domain, stream, level)]
        if not qs:
            break
        score = mean(SCALE[a.value] if a.evidence else 0.0 for a in qs)
        if score < MET:
            break
        achieved = level
    return achieved

def scorecard(answers, targets):
    rows = []
    for d in DOMAINS:
        a, b = stream_level(answers, d, "A"), stream_level(answers, d, "B")
        rows.append({"domain": d, "A": a, "B": b,
                     "floor": min(a, b), "target": targets.get(d, 2),
                     "gap": max(0, targets.get(d, 2) - min(a, b))})
    return sorted(rows, key=lambda r: (-r["gap"], r["floor"]))

The scorer reports the floor of the two streams, not their average. A domain where teams build threat models but nobody measures coverage is un-measured, not halfway mature. Sorting by gap yields the roadmap, and answers without evidence links score zero.

Worked example: a support assistant

A payments company runs three AI systems: a customer-facing support assistant with retrieval over help articles, an internal code-review assistant, and a fraud-scoring model trained on transaction data. It sets a target of Level 2 everywhere and Level 3 for Verification and Operations on the customer-facing assistant. Interviews and evidence produce this per-system result for the support assistant:

DomainABFloorTargetWhat the evidence showed
Verification2003Red-team report exists, but no regression suite and no pass-rate tracking
Data Management1002Help articles are indexed with no review of who can edit them
Operations2113Runbook exists; tool calls are logged but nobody alerts on them
Privacy1112Chat logs are kept 400 days, against a 90-day policy
Design2112Threat model omits the retrieval index as an injection source
Governance2222Inventory and policy in place, reviewed this year

Sorted by gap, the top of the roadmap is Verification and Data Management. They connect. The retrieval index is editable by any support agent, and nothing tests whether a poisoned help article can steer the assistant. One quarter's plan follows directly from the evidence. Add edit approval and change logging to the help-article store (Data Management A to Level 2). Add an indirect-injection regression set that plants instructions in a test article and fails the build if the assistant follows them, tracked per release (Verification B to Level 2). Extend the threat model to the index (Design B to Level 2). Fix log retention, a configuration change with an owner and a date. The scorecard after the quarter is the acceptance test for the plan. An average would have shown Verification at 1.0; the floor of 0 shows nobody can tell whether defences regress.

Crosswalk to other frameworks

Few organisations adopt one framework alone. AIMA sits beside regulation and management-system standards, and a crosswalk lets one piece of evidence serve several of them.

FrameworkWhat it isHow AIMA relates
NIST AI RMF 1.0Risk framework with Govern, Map, Measure, Manage functions; no maturity levelsAIMA levels can serve as the scale the RMF lacks; Governance and Responsible AI evidence maps largely to Govern and Map
ISO/IEC 42001:2023Certifiable AI management system standardCertification checks that the system exists; AIMA levels describe how well its processes run
OWASP Top 10 for LLM ApplicationsCatalogue of LLM application risksSupplies the test content for Verification and the threat list for Design
OWASP SAMM v2Software assurance maturity modelAIMA's parent; reuse SAMM evidence where the AI parts are covered

Map evidence artefacts to clauses, so one regression dashboard serves AIMA, the RMF and an ISO 42001 audit at once.

How maturity programmes fail

  • Checkbox theatre. The lightweight yes/no form is answered by the policy owner in an hour, and every answer is yes. Fix: require an evidence link for every yes before it scores.
  • Averaging hides zeros. A mean across domains or streams of 1.6 conceals an unmeasured Verification practice. Report floors and per-stream levels, and show zero cells in red.
  • Showcase sampling. The flagship is assessed; shadow AI on a personal API key is not. Sample from the inventory and scan egress logs for unlisted AI services.
  • Stale evidence. Invalidate evidence when a system's model, prompt or tools change.
  • Score as goal. Pair every target level with an outcome metric, such as regressions caught before release.
  • Framework drift. AIMA is an incubator project. Record the AIMA version in every scorecard, so a re-baseline after a revision is not mistaken for progress or regression.

Trade-offs

ChoiceGainCost
Lightweight checklistA baseline in days; good for awarenessSelf-reported, optimistic, weak for auditors
Detailed, evidence-backedDefensible, repeatable, finds the real gapsWeeks of assessor time; needs an evidence store
Organisation-wide unitSimple board storyHides the weakest system
Per-system unitFinds where incidents will come fromScales with inventory size; needs inheritance rules
Level 3 targets everywhereAmbitiousAutomation cost spent where risk is low; target by system risk instead

What to do next

  1. Build or update the AI system inventory with owners, data sources and tool permissions. Without it, sampling is impossible.
  2. Download the AIMA documentation and toolkit from the OWASP project page, and record the version you assess against.
  3. Run the lightweight checklist as a baseline, then convert it to a detailed assessment one domain at a time, starting with Verification and Data Management.
  4. Write down your evidence rubric, recency rule and sampling rule before the first interview.
  5. Store answers as data with evidence links, and score with stream floors and gaps, not averages.
  6. Set targets per system by risk, and re-assess each quarter against the same rubric.
  7. Keep learning: an evidence-based maturity scale for the NIST AI RMF, the OWASP Top 10 for LLM Applications, building an AI red-team programme, running an AI governance programme and an AI supply-chain security programme.
Key takeaway: The OWASP AI Maturity Assessment measures whether AI security and responsible-AI practices happen reliably, not whether one system is currently safe. Its grid of eight domains, two streams and three levels comes from SAMM, with Responsible AI, Data Management and Privacy added for AI. Its value depends on the evidence rules you bring. Assess per system, require recent artefacts, score stream floors rather than averages, and turn the largest gaps into a quarterly roadmap.