ISO/IEC 42001:2023 is the first international standard for an AI management system, usually shortened to AIMS. Unlike a technical standard, it does not tell you how to make a model robust or which guardrail to deploy. It specifies how an organisation decides what AI risks it has, which controls it will run against them, who is accountable, how it checks that the controls work and how it improves. Because it is certifiable by accredited bodies, a certificate has become a common procurement question for anyone selling AI-backed products to enterprises.

This article is for the engineers and security leads who will actually build the system behind the certificate. It goes clause by clause through what must exist, explains Annex A and the Statement of Applicability, shows how to keep both as code with an evidence pipeline, works an example scope for a company running an LLM support product, and covers the failure modes that turn certification into a paperwork exercise. It deliberately stays at clause and domain level; use your licensed copy of the standard for exact control text.

What the standard is, and what it is not

The standard follows the harmonised structure shared by ISO management system standards, so clauses 4 to 10 line up with ISO/IEC 27001 and ISO 9001. That is deliberate: most organisations integrate an AIMS with an existing information security management system rather than running a second bureaucracy.

Four annexes sit behind the clauses. Annex A lists the reference control objectives and controls: 38 controls grouped into nine domains numbered A.2 to A.10, covering AI policies, internal organisation, resources for AI systems, assessing impacts, the AI system life cycle, data for AI systems, information for interested parties, use of AI systems, and third-party and customer relationships. Annex B gives implementation guidance for those controls. Annex C lists potential AI-related organisational objectives and risk sources, useful as a prompt list during risk assessment. Annex D discusses use of the management system across domains and sectors.

Two companion standards published in 2025 matter in practice. ISO/IEC 42005 gives guidance on AI system impact assessment, the most novel requirement in 42001. ISO/IEC 42006 sets additional requirements for the certification bodies that audit an AIMS, building on ISO/IEC 17021-1, including the competence auditors need to assess AI systems. ISO/IEC 23894 gives guidance on AI risk management and pairs naturally with clause 6.

What 42001 is not: it is not a statement that a model is safe, and certification on its own does not give presumption of conformity with the EU AI Act, whose harmonised standards are being developed separately by CEN-CENELEC JTC 21. It is evidence that you manage AI risk systematically, which is useful input to AI Act work but not a substitute for it.

Clauses 4 to 10, as things you must build

ISO/IEC 42001 is a management system: clauses 4 to 10 form a loop4 contextscope, roles, parties5 leadershipAI policy, accountability6 planningrisk, impact, SoA7 supportcompetence, documents8 operationrun 8.2, 8.3, 8.49 evaluationmonitor, audit, review10 improvementnonconformity, fixesupdate plansAnnex A: 38 controls in 9 domains (A.2 to A.10)selected through risk treatment, justified in the Statement of ApplicabilityAuditors test the loop: did risk drive the controls, and did evidence drive the changes?
Clauses 4 to 10 of ISO/IEC 42001 as a plan-do-check-act loop, with Annex A controls selected through planning.

Clause 4, context. Determine internal and external issues, the interested parties and their requirements, and the scope of the AIMS. The standard asks you to determine your role with respect to AI systems, for example AI provider, AI producer, AI customer or AI user, and the same organisation is often several. A company that calls a hosted model API and sells a product built on it is a customer of the model vendor and a provider to its own users. Scope is the single most consequential decision: too broad and the audit is unmanageable, too narrow and customers notice the product they buy is outside it.

Clause 5, leadership. Top management must establish an AI policy, assign roles and responsibilities and show commitment. In audit terms this means a signed policy, named owners, and minutes showing management actually reviewed AI risk.

Clause 6, planning. This is the engine. Clause 6.1.2 requires an AI risk assessment process with defined criteria. Clause 6.1.3 requires risk treatment: choose controls, compare them against Annex A so nothing necessary is omitted, produce a Statement of Applicability, and get risk owners to approve the treatment plan and residual risk. Clause 6.1.4 requires a process for AI system impact assessment covering consequences for individuals, groups and society. Clause 6.2 sets measurable AI objectives.

Clause 7, support. Resources, competence, awareness, communication and documented information. Expect auditors to ask how the engineers who tune prompts and tools know the AI policy and the risks of their system.

Clause 8, operation. Run what clause 6 planned: perform risk assessments at planned intervals and on significant change (8.2), implement the treatment plan (8.3) and perform impact assessments (8.4). "Significant change" needs a definition you can apply, such as a new model version, a new tool, a new data source or a new user population.

Clauses 9 and 10, evaluation and improvement. Monitor and measure, run internal audits, hold management reviews, and handle nonconformities with root cause analysis and corrective action. A first audit often finds more in clause 9 than anywhere else, because internal audit is the step teams skip.

Risk assessment versus impact assessment

Risk assessment and impact assessment answer different questions and are easy to blur. The risk assessment asks what could go wrong for the organisation and its objectives: a data leak through the assistant, a regulatory finding, an outage of the model vendor. The impact assessment asks what the AI system could do to people outside the organisation: a customer denied a refund by a wrong answer, a group systematically served worse, misuse of generated content. A single incident can appear in both, seen from two directions.

Make both repeatable by fixing the inputs. A risk entry needs a scenario specific enough to test ("a planted instruction in a support ticket causes the assistant to reveal another customer's order history"), likelihood and impact on defined scales, the controls that treat it, and an owner. Annex C's list of risk sources is a good completeness check. An impact assessment record needs the intended use, foreseeable misuse, affected groups, the severity and reversibility of harm, and the human oversight available; ISO/IEC 42005 gives a fuller structure.

The trigger matters as much as the template. Tie both assessments to your change process so that adding a tool to an agent or swapping the model version opens a reassessment ticket automatically. That is how clause 8 stays alive between audits.

The Statement of Applicability as code

The Statement of Applicability lists every Annex A control, says whether it is applied, and justifies both inclusions and exclusions. Teams that keep it as a spreadsheet find it is wrong within a quarter. Keep it as data next to the systems it describes, with a pointer from each control to the evidence that proves it.

# soa.yaml (one row per Annex A control; ids follow your licensed copy)
- control: A.6-lifecycle-verification     # internal key mapped to the standard's number
  domain: A.6
  applicable: true
  justification: "risk R-014, R-022: model and prompt changes can regress safety"
  owner: ml-platform-lead
  implementation: "release gate in ci/llm_release.yml runs eval suite and red team regression"
  evidence:
    query: "ci_runs where pipeline='llm_release' and status in ('pass','blocked')"
    max_age_days: 30
- control: A.7-data-provenance
  domain: A.7
  applicable: true
  justification: "R-031: RAG corpus includes customer-authored tickets"
  owner: data-governance
  implementation: "source manifest with origin and licence per corpus, checked at ingest"
  evidence:
    query: "ingest_log where manifest_check='fail'"
    max_age_days: 7

A small checker turns the file into an enforced artefact:

import yaml, datetime as dt

EXPECTED = 38                          # Annex A control count
rows = yaml.safe_load(open("soa.yaml"))
store = load_evidence_index()          # control -> latest evidence timestamp
problems = []
if len(rows) != EXPECTED:
    problems.append(f"SoA has {len(rows)} rows, expected {EXPECTED}")
for r in rows:
    if not r.get("justification"):
        problems.append(f"{r['control']}: no justification")
    if r["applicable"]:
        seen = store.get(r["control"])
        limit = dt.timedelta(days=r["evidence"]["max_age_days"])
        if seen is None or dt.datetime.utcnow() - seen > limit:
            problems.append(f"{r['control']}: evidence missing or stale")
for msg in problems:
    open_nonconformity(msg)            # ticket with owner, due date, root cause field
Evidence as a pipeline: controls point at queries, not at screenshotssoa.yaml38 rows, ownersevidence queriesCI, tickets, logsevidence storedated, immutableaudit packper samplefreshness checkstale or missing evidence opens a nonconformityowner fixesInternal audit samples from the same store an external auditor will use.
Evidence pipeline: each SoA row names a query; stale evidence becomes a tracked nonconformity before an auditor finds it.

Exclusions deserve the same rigour. "Not applicable" is acceptable when the risk treatment shows the control addresses a situation you do not have, for example a control aimed at training data collection when you only consume a third-party model, but then the data domain still applies to your retrieval corpus and fine-tuning sets. Auditors reject exclusions that are really "not done yet".

Worked example: scoping an LLM support product

A 300-person company sells a customer-support platform. Its AI features are an LLM assistant that drafts replies for human agents using a hosted model API and retrieval over the customer's help centre, and an in-house classifier that routes tickets by urgency.

Scope: the design, development, operation and support of both features in the production platform, including the vendor relationship with the model provider. Out of scope: internal employee use of general chat tools, which a separate usage policy covers. Roles: AI customer of the model vendor; AI provider and developer for the routing classifier; AI provider for the assistant to its own customers.

Top risks: cross-tenant leakage through retrieval; prompt injection in incoming tickets steering drafts; vendor model change degrading quality without notice; routing classifier under-prioritising tickets written in some languages. Impact assessment: the assistant is advisory with a human sending every reply, which lowers severity; the router has no human in the loop and can delay urgent help, so its impact assessment drives a per-language recall objective under clause 6.2 and a monitoring control.

Controls selected: tenant-scoped retrieval, provenance labels on ticket text, a release gate running evaluations on every model or prompt change, contractual change notification from the vendor plus a canary evaluation that detects silent changes, per-language recall dashboards, and disclosure to customers that drafts are AI-generated. Every one maps to a risk ID and an SoA row, and every row has an evidence query.

Path to certification: gap assessment, then several months of operating the system to generate evidence, one full internal audit and one management review, then the certification body's stage 1 audit (documentation and readiness) and stage 2 audit (implementation). Certificates under the usual 17021-1 model run three years with surveillance audits in between. Timelines vary widely with scope and existing ISO 27001 maturity, so treat any fixed figure with suspicion. The audit cycle itself is covered in LLM security certifications and AI compliance audit.

Failure modes and trade-offs

Paper AIMS. Policies written by a consultant, never read by the engineers who change prompts daily. The stage 2 auditor interviews an engineer and the gap is obvious. Fix it by putting controls in the tools engineers already use: CI gates, pull request templates, change tickets.

Static risk register. Assessed once for the audit and never again, so the agent that gained a write tool last month is assessed as read-only. Tie reassessment to change events; see AI risk register for testable risk statements and computed residual scores.

Scope gaming. Certifying a narrow internal tool while marketing the certificate against the flagship product. Customers increasingly read scope statements.

Copying 27001 controls. Information security controls cover confidentiality well and say little about impact on individuals, bias, transparency or human oversight. Integration should share processes, not substitute content.

Evidence by screenshot. Manually gathered evidence is stale by the time it is filed and impossible to sample. Query-based evidence costs more to set up and far less every year after.

The main trade-off is overhead against assurance. A management system adds meetings, records and audits; for a startup with one AI feature, aligning to the standard without certifying may be the right first step. For a company selling to regulated enterprises, the certificate often pays for itself in shortened security questionnaires. Integration with an existing management system cuts cost but can hide AI-specific gaps inside generic processes. For how 42001 relates to other frameworks, read AI safety frameworks; for the operating model around it, AI governance program structure.

Key takeaway: <p>ISO/IEC 42001 certifies how you manage AI risk, not how safe a model is. The work that makes it real is a defensible scope, risk and impact assessments triggered by change, an SoA kept as data with evidence queries, and internal audit that samples the same evidence an external auditor will.</p><ul><li>Buy the standard and read clauses 4 to 10 and Annex A end to end before planning.</li><li>Write a one-paragraph scope and list your role for each AI system: provider, producer, customer, user.</li><li>Define "significant change" and wire it to automatic reassessment tickets.</li><li>Draft risk entries as testable scenarios and an impact assessment per system using ISO/IEC 42005.</li><li>Create soa.yaml with all 38 controls, owners, justifications and evidence queries; run the checker in CI.</li><li>Integrate with your ISO 27001 processes where they exist, and add the AI-specific content they lack.</li><li>Run one full internal audit and one management review before booking a stage 1 audit.</li></ul>