A maturity model is a ladder you can put an organisation on: a set of capabilities, a small number of levels for each, and rules for deciding which rung you stand on. The idea comes from the Capability Maturity Model developed at Carnegie Mellon's Software Engineering Institute for software processes, whose five levels (initial, repeatable, defined, managed, optimising) and their successor CMMI still shape most models in use. Applied to AI compliance, a maturity model answers a question executives ask and auditors cannot: not whether a given control passed this quarter, but how reliably the organisation can keep meeting its obligations as systems, laws and vendors change.

One clarification first. There is no regulator-endorsed AI compliance maturity model. The EU AI Act, ISO/IEC 42001 and the NIST AI RMF define requirements and outcomes, not levels; any maturity scale is something you or a consultancy built on top. That is fine, but it means the scale is only as meaningful as its evidence rules. This site already covers security-oriented scales in NIST AI RMF maturity and the OWASP AI maturity model. This page builds a compliance-oriented model and adds three things those lack: dependencies between capabilities, calibrated assessors, and a roadmap computed from the dependency graph.

Eleven compliance capabilities

Compliance capability means the machinery that turns obligations into evidence. Eleven capabilities cover it for most organisations shipping AI, regardless of which regimes apply:

CapabilityWhat it doesDepends on
System inventoryKnows every AI system, its owner, model and purposenone
Obligations registerLists each applicable requirement in testable formnone
ClassificationAssigns each system its risk class per regimeinventory, obligations
Impact assessmentAssesses effects on people for in-scope systemsclassification
Control libraryMaps obligations to controls with ownersobligations, classification
Control testingTests controls on a schedule and on eventscontrol library
MonitoringWatches systems in production against thresholdsinventory, control library
Incident reportingDetects, triages and reports incidents on timemonitoring
Third-party AIAssesses vendors and models you did not buildinventory, obligations
Change intakeTurns new laws and guidance into register changesobligations
Regulator interfaceProduces evidence packs and answers inquiriescontrol testing, incident reporting

The dependency column is the important addition. Classification without an inventory classifies only the systems someone remembered. Incident reporting without monitoring reports only what customers complain about. A model that scores capabilities independently will happily report a level 4 regulator interface sitting on level 1 control testing, which describes an organisation that writes excellent letters about controls nobody checks.

Levels defined by evidence

Levels must be defined by observable evidence, never by self-description. The same five levels apply to every capability:

LevelNameEvidence required
1Ad hocActivity happens through individuals; no records survive their absence
2RepeatableWritten procedure; records exist for the most recent cycle for some systems
3DefinedOrganisation-wide standard; records for every in-scope system; named owner; evidence younger than one cycle
4MeasuredEvidence produced by tooling; metrics with thresholds; sampled tests pass at a stated rate
5OptimisingMetrics drive change; regulatory changes and incidents feed back with measured lead times

Two rules make the table hold. Evidence must be sampled, not presented: the assessor picks systems from the inventory at random and asks for their records, rather than accepting the examples the capability owner chooses. And evidence decays. A level 3 impact assessment capability requires assessments younger than one cycle for every in-scope system; if a third of them are two years old, the capability is at level 2 regardless of how good the template is. The metrics needed for level 4 are covered in AI compliance metrics.

Dependencies cap what you can claim

Raw scores come from the assessment. Effective scores apply one rule: a capability can be at most one level above the weakest capability it depends on. One level of headroom, rather than none, is deliberate; it lets a team build ahead of a foundation that is improving, but not far ahead.

Compliance capabilities as a dependency graphSystem inventoryObligations registerClassificationControl libraryChange intakeImpact assessmentMonitoringControl testingThird-party AIIncident reportingRegulator interfaceAn arrow means: the target cannot be more than one level above the source.Inventory also feeds classification, monitoring and third-party AI (edges omitted for clarity).
Capability dependencies. Foundations on the left cap what can be claimed on the right.
CAPS = {
    "inventory": [], "obligations": [],
    "classification": ["inventory", "obligations"],
    "impact_assessment": ["classification"],
    "controls": ["obligations", "classification"],
    "control_testing": ["controls"],
    "monitoring": ["inventory", "controls"],
    "incident_reporting": ["monitoring"],
    "third_party": ["inventory", "obligations"],
    "change_intake": ["obligations"],
    "regulator_interface": ["control_testing", "incident_reporting"],
}

def effective_levels(raw):
    eff = {}
    def level(cap):
        if cap not in eff:
            caps = [level(d) + 1 for d in CAPS[cap]]
            eff[cap] = min([raw[cap]] + caps)
        return eff[cap]
    for cap in CAPS:
        level(cap)
    return eff

Report raw and effective levels side by side. The difference is itself a finding: a large gap shows effort spent on visible capabilities while foundations lag, which is exactly the pattern that fails under audit.

Calibrating the assessors

Maturity scores are judgements, and two assessors reading the same evidence often disagree. If you never measure that, year-on-year movement may just be a change of assessor. The fix is borrowed from annotation work: have two assessors independently score a shared sample of capabilities, then compute agreement corrected for chance.

from collections import Counter

def weighted_kappa(a, b, levels=(1, 2, 3, 4, 5)):
    # Quadratic-weighted Cohen's kappa: disagreeing by 2 levels costs 4x disagreeing by 1.
    n, k = len(a), len(levels)
    w = lambda i, j: ((i - j) ** 2) / ((k - 1) ** 2)
    ca, cb = Counter(a), Counter(b)
    observed = sum(w(x, y) for x, y in zip(a, b)) / n
    expected = sum(ca[i] * cb[j] * w(i, j) for i in levels for j in levels) / (n * n)
    return 1.0 if expected == 0 else 1 - observed / expected

Plain kappa treats a 2 against a 3 the same as a 1 against a 5; the quadratic weighting reflects that levels are ordered. A common rule of thumb treats values above about 0.8 as strong agreement, but use it as a trend, not a pass mark. When agreement is low, the remedy is almost always sharper evidence definitions for the disputed level, not more assessor training. Rotate one external assessor into the pair every year to keep the internal pair from converging on shared blind spots.

Targets come from exposure

Not every capability needs level 5, and not every organisation needs the same targets. Set targets from exposure. A company with high-risk systems under Annex III of the EU AI Act needs level 3 or better in impact assessment and incident reporting, because those obligations carry deadlines and documentation duties; a company whose AI use is internal productivity tooling may rationally hold those at level 2. Write the rationale for each target into the model, so the board approving the roadmap is approving a risk decision, not a wish list. The AI compliance program page describes the obligations register that most targets trace back to.

Targets also need a horizon. A jump of two levels in one capability usually takes a year or more, because level 3 requires coverage of every in-scope system and level 4 requires tooling that produces evidence on its own. Record each target with the date by which it should be reached and the regulatory date that motivates it, if any. For a high-risk system under the EU AI Act, that date is the one in the current timeline after the Digital Omnibus changes, not the original text, so link the target to the obligations register entry rather than copying a date that may move again.

Finally, keep the assessment itself cheap enough to repeat. A workable cycle is two weeks: one to collect sampled evidence against the level definitions, one for the two assessors to score independently, reconcile disagreements and write findings. Anything longer tends to be skipped the following year, and a maturity model assessed once is a snapshot, not a measurement of progress.

A roadmap scheduled over the graph

With raw levels, effective levels, targets and weights, the roadmap is a scheduling problem over the graph. Each quarter, a capability can rise one level if its dependencies already allow the new level. Among the eligible moves, pick those with the largest weighted gap, up to the number of improvement efforts the organisation can staff.

def roadmap(raw, target, weight, per_quarter=3, quarters=6):
    raw, plan = dict(raw), []
    for q in range(1, quarters + 1):
        cur = effective_levels(raw)
        # only uncapped capabilities need effort; capped ones recover with their foundations
        ready = [c for c in CAPS
                 if cur[c] < target[c] and raw[c] == cur[c]
                 and all(cur[d] >= cur[c] for d in CAPS[c])]
        ready.sort(key=lambda c: (-weight[c] * (target[c] - cur[c]), c))
        moves = ready[:per_quarter]
        for c in moves:
            raw[c] += 1
        after = effective_levels(raw)
        plan.append((q, [(c, after[c]) for c in moves],
                     sorted(c for c in CAPS if after[c] > cur[c] and c not in moves)))
    return plan

A capability held down by the cap gets no effort slot of its own: it recovers when its foundation does, and the third element of each quarter lists those free recoveries. The greedy order is not optimal in every case, but it has the property that matters: it never schedules work that the foundations cannot support, so it does not produce the headline-capability roadmap that looks good in a board pack and collapses at audit.

Worked example: a lender

An illustrative lender running a credit-decision assistant and a customer chatbot scores its capabilities. Raw levels: inventory 3, obligations 2, classification 3, impact assessment 3, control library 2, control testing 1, monitoring 4, incident reporting 3, third-party 1, change intake 2, regulator interface 3.

Applying the cap changes two scores. Monitoring falls from 4 to 3, because the control library it monitors against is at 2. Regulator interface falls from 3 to 2, because control testing is at 1: the evidence packs are well formatted but contain little tested evidence. Everything else is unchanged. That second correction is the most useful line in the report, because it predicts what an examiner will find.

Targets, driven by the credit system being high-risk: 3 for everything except monitoring and control testing at 4. Control testing and third-party AI get weight 3, since the credit model is a vendor model; the rest get 1. With three efforts per quarter, the code raises control testing and third-party to 2 and change intake to 3 in the first quarter, and the regulator interface recovers to 3 for free as its cap lifts. The second quarter takes control testing, third-party and the control library to 3, which frees monitoring back to 4. The third finishes control testing at 4 and obligations at 3, and every target is met. Leadership had planned to invest in a regulator reporting portal first; the graph shows the regulator interface is fixed by funding control testing, while the portal would only have presented untested controls more attractively.

Failure modes

  • Self-assessment without sampling. Owners present their best examples. Sample from the inventory instead.
  • Goodhart drift. Once levels feed bonuses, evidence is manufactured to meet the definition. Keep maturity scores out of individual objectives; tie objectives to the underlying metrics with audits.
  • Averaging. A single organisational score hides a level 1 foundation behind level 4 showpieces. Publish the per-capability profile.
  • Ignoring decay. Evidence from last year scores this year. Require freshness inside each level definition.
  • Framework sprawl. Separate maturity models for each regime duplicate assessments. Score capabilities once and map them to regimes, as ISO/IEC 42001 clauses map to controls.
  • Level 5 as a goal. Optimising everything is waste. Targets come from exposure.

Trade-offs

Granularity trades against cost. Eleven capabilities with five levels is assessable in a few weeks; sixty sub-practices give finer signal but turn the assessment into a project nobody repeats. Assessing per business unit rather than per organisation finds the weak unit but multiplies effort; a common compromise is organisation-wide for foundations and per unit for monitoring and incident reporting. External assessment brings credibility and fresh eyes but costs money and tends to import the assessor's own model; internal assessment is cheap and repeatable but drifts without calibration. Finally, the dependency cap makes scores lower and less flattering, which is uncomfortable to report upward. That discomfort is the point.

What to do next

  1. Adopt the eleven capabilities, or your own list, and write down the dependencies between them.
  2. Define each level by evidence, including a freshness rule, before anyone scores anything.
  3. Run a first assessment with two assessors on a random sample, and compute weighted kappa.
  4. Compute effective levels and report them beside raw levels.
  5. Set targets from regulatory exposure, with a written rationale per capability.
  6. Generate a roadmap from the graph and staff the first quarter's moves.
  7. Reassess yearly, rotate in an external assessor, and track the raw-to-effective gap over time.
Key takeaway: An AI compliance maturity model is useful only when its levels are defined by sampled, fresh evidence and its scores respect dependencies: no capability can sit more than a level above the foundations it relies on. Calibrate assessors with weighted kappa, set targets from regulatory exposure, and let the dependency graph order the roadmap so investment goes where an audit would actually find the gaps.