An AI compliance audit is an independent examination of whether an organisation's AI systems and the processes around them meet a defined set of criteria: an internal policy, a management system standard such as ISO/IEC 42001, or obligations drawn from law. It is not a model evaluation and not a red-team exercise, although it may reuse both. The question it answers is narrower and harder: did the controls the organisation says it runs actually exist, fit the risk, and operate throughout the period?

This article is written from the auditor's side of the table, for internal auditors, second-line risk teams and engineers asked to run an audit of their own AI estate. Preparing to be audited is the mirror image and is covered in LLM Audit Preparation. Here you will learn how to choose criteria, plan a risk-based scope, test design and operating effectiveness of AI controls, re-perform an evaluation gate, size a sample, and write findings that get fixed.

Audit, attestation or certification

Start by knowing what kind of conclusion you are producing, because it decides the rigour required and who may rely on it.

EngagementWho performs itWhat it produces
Internal auditThe organisation's own independent audit functionA report to management and the audit committee with rated findings
SOC 2 examinationA licensed CPA firm under AICPA standardsAn attestation report on controls; not a certification
ISO/IEC 42001 certificationAn accredited certification bodyA certificate for the AI management system, maintained by surveillance audits
Regulatory conformityDepends on the regime and systemEvidence of conformity; see the regime's own rules

Two of these are often confused. A SOC 2 report is an attestation: the auditor expresses an opinion on controls, and the reader decides what that means for them; mapping it to LLM systems is covered in SOC 2 for LLM Applications. ISO/IEC 42001, published in 2023, is a management system standard against which an organisation can be certified. Certification bodies follow ISO/IEC 17021-1, and ISO/IEC 42006, published in July 2025, adds requirements specific to bodies that audit and certify AI management systems, including auditor competence. Certification follows the usual management system pattern: a Stage 1 audit of documentation and readiness, a Stage 2 audit of implementation, then surveillance audits during a three-year cycle before recertification.

Criteria and the AI control universe

An audit without written criteria is an opinion poll. Assemble the criteria before fieldwork and get management to agree them. For an AI audit they usually come from three sources: the organisation's own AI policy and standards, an external framework such as ISO/IEC 42001 or the NIST AI RMF, and legal obligations identified by counsel. The AI governance program should already maintain a control library mapping these to concrete controls; if it does not, that absence is itself a finding.

Typical AI controls in scope include:

  • An inventory of AI systems with risk tiers, kept current.
  • Pre-deployment impact or risk assessment for higher tiers, with sign-off.
  • Evaluation gates: a model, prompt or retrieval change cannot reach production unless defined evaluations pass.
  • Change management covering prompts, model versions and tool definitions, not just code.
  • Data controls: training and retrieval data provenance, consent and retention.
  • Human oversight where the policy requires it, and logging that proves it happened.
  • Monitoring, incident handling and user complaint channels.
  • Third-party and model-provider due diligence.

Planning a risk-based scope

You cannot test every control on every system. Plan by risk: rank systems by tier and exposure, then pick the controls whose failure would hurt most. A model that drafts internal meeting notes and one that decides credit limits do not deserve equal hours.

An AI compliance audit engagement: from criteria to an opinion, with evidence flowing back to the auditorCriteria42001, policy, lawRisk-based plansystems, controls, scopeDesign testdoes the control fitOperating testdid it run all periodInspectionrecords, configsRe-performancererun evals, gatesSamplingchanges, incidentsFindingscondition, criteria, cause, effect, recommendationReport and opinionrated findingsRemediation follow-upretest before closingThe auditor tests the control, not the model: evidence must show the control ran, for every item in the population or a valid sample.
Criteria drive a risk-based plan; each control gets a design test and an operating test using inspection, re-performance and sampling; exceptions become findings, which are reported and retested.

For each selected control, write the test step before fieldwork: what evidence you will request, what population it comes from, and what would count as an exception. Writing the exception definition in advance is what stops an audit from drifting into negotiation.

Testing design and operation

Audit testing has two levels. Design effectiveness asks whether the control, if it operated as described, would address the risk. An evaluation gate that tests only accuracy on a benchmark does not address a policy requirement about harmful outputs, however reliably it runs. Operating effectiveness asks whether it actually ran, every time it should have, across the audit period.

The techniques are the classic ones, applied to new objects.

TechniqueApplied to AI controlsStrength
InquiryInterview the team about how evaluation gates workWeak alone; never sufficient
InspectionRead deployment records, evaluation reports, risk assessments, configsModerate; shows a record exists
ObservationWatch a release go through the gateModerate; only covers what you saw
Re-performanceRerun the evaluation on the released artefact and compareStrong; independent of the team's own result
Data analyticsCompare the full deployment log with the gate logStrong; covers the whole population

Re-performance is the technique most AI audits underuse. If a control says 'no model version is deployed unless the safety evaluation pass rate is at least 98 percent', the auditor can take the deployed version and the pinned evaluation set and run it again.

import hashlib, json

def reperform_gate(deploy_record, eval_runner, threshold=0.98, tolerance=0.01):
    # 1. The evaluation set used must be the controlled, versioned one.
    data = open(deploy_record["eval_set_path"], "rb").read()
    if hashlib.sha256(data).hexdigest() != deploy_record["eval_set_sha256"]:
        return "exception: evaluation set differs from the controlled version"
    # 2. Rerun against the exact artefact that was deployed.
    result = eval_runner(model=deploy_record["model_version"],
                         prompt=deploy_record["prompt_version"],
                         cases=json.loads(data))
    claimed = deploy_record["reported_pass_rate"]
    if result.pass_rate < threshold:
        return f"exception: reperformed {result.pass_rate:.3f} below {threshold}"
    if abs(result.pass_rate - claimed) > tolerance:
        return f"exception: reported {claimed:.3f} not reproducible ({result.pass_rate:.3f})"
    return "no exception"

The tolerance exists because model outputs and model-graded evaluations vary between runs. Agree it with management before testing, and note it in the workpapers.

Populations and sampling

Operating effectiveness is usually tested on a sample of the population: every production change to a model, prompt or tool in the period. Get the population from a system of record, not from the team, and check it for completeness by reconciling it with an independent source such as the deployment platform's own log. A population assembled by hand will omit exactly the hotfix you most need to see.

For attribute testing where you expect no exceptions, a simple calculation gives the sample size: if the true exception rate were at the tolerable level, how many items must you draw so that seeing zero exceptions would be unlikely?

import math, random

def zero_exception_sample_size(tolerable_rate=0.10, confidence=0.95):
    # smallest n with (1 - rate) ** n <= 1 - confidence
    return math.ceil(math.log(1 - confidence) / math.log(1 - tolerable_rate))

def draw(population_ids, n, seed):
    rng = random.Random(seed)            # record the seed in the workpapers
    return sorted(rng.sample(population_ids, min(n, len(population_ids))))

print(zero_exception_sample_size())          # 29
print(zero_exception_sample_size(0.05))      # 59

With a 10 percent tolerable rate and 95 percent confidence you need 29 items with no exceptions. When the population is small or fully logged, skip sampling and test all of it with analytics: joining the deployment log against the gate log is cheap and leaves no sampling risk.

Writing findings that get fixed

A finding that does not get fixed was badly written. The standard structure has five parts.

  • Condition: what you found, with numbers. '3 of 29 sampled prompt changes were deployed without an evaluation run.'
  • Criteria: the requirement it breaches, quoted. 'AI Standard 4.2: every production change to prompts passes the tier-appropriate evaluation gate.'
  • Cause: why it happened. 'Prompts are stored in a configuration service whose deploy path bypasses the CI pipeline that hosts the gate.'
  • Effect: the risk in business terms. 'Unevaluated prompt changes can reach customers in the claims assistant, a tier 1 system.'
  • Recommendation: addressing the cause, not the symptom. 'Route configuration deploys through the gate, or block configuration writes to production prompts outside it.'

Rate findings on an agreed scale, agree an owner and date with management, and close them only after retesting. Closing on the strength of a team's assurance repeats the original control failure.

Keep workpapers that another auditor could follow without asking you anything: the test step, the population source and its reconciliation, the sample seed, exported evidence with hashes, and the reasoning behind each exception. Independence matters as much as rigour. An engineer from the team that built the evaluation gate should not be the one re-performing it, however convenient that is.

Worked example: auditing customer-facing AI at a bank

An internal audit of a bank's customer-facing AI estate scopes three tier 1 systems: a dispute-summary assistant, a complaint classifier and a chatbot. Criteria are the bank's AI standard plus ISO/IEC 42001 clauses the bank has adopted.

  1. Inventory test: the auditor compares the AI inventory with the model-provider billing records and finds two production uses of an external model API not in the inventory. Finding: inventory incomplete.
  2. Design test of the evaluation gate: the gate checks accuracy and toxicity but not the policy requirement that the chatbot must not give individual financial advice. Finding: gate design gap.
  3. Operating test: the deployment population has 212 changes; analytics join it to the gate log and find 7 prompt changes with no gate record. Cause traced to the configuration service bypass.
  4. Re-performance: for 5 model upgrades, rerunning the safety evaluation reproduces the reported pass rates within tolerance. No exception.
  5. Human oversight: the complaint classifier policy requires review of low-confidence labels. A sample of 29 shows review timestamps for all. No exception.

The report carries three findings, one rated high, with the bypass fixed and retested within the quarter. For EU-facing systems, obligations depend on the system's classification and the current timeline; see EU AI Act rather than hard-coding dates into audit criteria.

Failure modes

  • Auditing the model instead of the control. Spending fieldwork on benchmark scores while never checking whether the gate ran for every deployment.
  • Team-supplied populations. Accepting a spreadsheet of changes without reconciling it to a system of record.
  • Prompts out of scope. Treating prompts and tool definitions as content rather than production configuration.
  • Unreproducible evidence. Screenshots of dashboards instead of exported records with hashes and timestamps.
  • Certificate as assurance. Treating an ISO/IEC 42001 certificate or a vendor's SOC 2 report as proof that a specific system is safe; read the scope.
  • Auditor competence gaps. Auditors who cannot read an evaluation report will accept a weak one; pair them with an engineer who is independent of the audited team.

Trade-offs

ChoiceBenefitCost
Periodic auditDeep, independent, well understoodPoint-in-time; problems surface months late
Continuous control monitoringFull populations, fast detectionEngineering investment; monitors themselves need review
SamplingCheap for manual evidenceSampling risk; misses rare failures
Full-population analyticsNo sampling riskNeeds reliable logs joined across systems
External certificationRecognised by customersCost; scope may not cover the systems you care about

A mature programme uses continuous monitoring for high-volume controls such as gates and logging, and periodic audit for judgement-heavy controls such as risk assessments and oversight design.

What to do next

  1. Write the audit criteria first and get management to agree them before fieldwork starts.
  2. Rank AI systems by tier and exposure and scope the top few deeply rather than all of them thinly.
  3. For each control, define the evidence, the population source and the exception before testing.
  4. Reconcile every population to an independent system of record.
  5. Re-perform at least one evaluation gate end to end on a deployed artefact.
  6. Use analytics on full populations wherever logs allow; sample only where they do not.
  7. Write findings with condition, criteria, cause, effect and recommendation, and retest before closing.
  8. Feed recurring causes back to the governance programme as control design changes.
Key takeaway: An AI compliance audit tests controls, not models: did the inventory, assessments, evaluation gates, change management and oversight exist, fit the risk and operate for the whole period? Agree criteria first, scope by risk, reconcile populations to systems of record, re-perform gates on deployed artefacts, prefer full-population analytics to sampling, and write five-part findings that you retest before closing.