An AI impact assessment asks a narrow question that a security review or a model evaluation does not: if this system is used as intended, and also as it will realistically be misused, what happens to the people on the other end of its decisions? It looks outward, at individuals, groups and society, rather than inward at the organisation's own losses. That outward view is why regulators ask for it, and why it is so often done badly: the team that built the system is the team least able to imagine how it fails for someone else.

This article treats the assessment as an engineering artefact. You will see which regimes ask for one, how to run a single versioned record that produces each of their views, how Canada's published scoring rules turn answers into an impact level, a worked example for a public-benefits triage assistant, the triggers that force reassessment, and the failure modes that make assessments worthless. Loss-focused risk scoring is a different exercise, covered in AI Risk Assessment, in depth.

What an impact assessment is

Three properties separate an impact assessment from neighbouring practices. Its subject is the affected person, not the asset: the question is who is harmed and how badly, not what it costs the company. Its timing is before deployment and at every material change, not once a year. Its output is a decision with conditions: proceed, proceed with named mitigations, or stop. A document that ends without a decision is a description, not an assessment.

It also differs from a data protection impact assessment (DPIA). A DPIA under GDPR Article 35 is about risks to people arising from the processing of personal data. Many AI harms involve no personal data at all, such as a model that systematically misreads a dialect in anonymous text, and many involve personal data but harm people through the decision rather than the processing. An AI impact assessment covers the decision, the people affected, the humans in the loop and the redress path, and it can reuse a DPIA for the data part. The GDPR side is covered in GDPR for LLM Applications.

The regimes that ask for one

Four sources shape practice today. Read the primary texts before relying on any summary, including this one.

SourceWho it bindsWhat it asks for
ISO/IEC 42005:2025Voluntary guidanceHow and when to assess impacts on individuals and societies across the life cycle, how to document it, and how to integrate it with risk management and an AI management system
ISO/IEC 42001Certified organisationsAn AI system impact assessment process inside the management system, with evidence an auditor can sample
EU AI Act, Article 27Certain deployers of Annex III high-risk systemsA fundamental rights impact assessment (FRIA) before first use, results notified to the market surveillance authority
Canada, Directive on Automated Decision-MakingFederal institutionsThe Algorithmic Impact Assessment questionnaire, an impact level I to IV, and published results

Article 27 applies to deployers that are bodies governed by public law, private entities providing public services, and deployers using high-risk systems for creditworthiness or credit scoring and for life and health insurance pricing. The assessment must describe (a) the deployer's processes in which the system is used, (b) the period and frequency of use, (c) the categories of people and groups likely to be affected, (d) the specific risks of harm to them, (e) how human oversight is implemented, and (f) the measures to take if risks materialise, including internal governance and complaint arrangements. Where a DPIA exists, the FRIA complements it. Under the amended timeline described in EU AI Act, in depth, Annex III obligations now apply from 2 December 2027; check the consolidated text before you plan against that date.

ISO/IEC 42005 was published in May 2025. It is guidance, not a certifiable standard, and it is most useful as the process skeleton that ISO/IEC 42001 auditors expect to see behind the assessment records, described in ISO/IEC 42001, in depth.

Architecture: one record, many views

One assessment record, several regulatory viewsChange intakenew use, model, dataScreeningwho, what decisionAssessment recordversioned, in gitSign-offowner + reviewerin scopeFRIA viewArt. 27 (a)-(f)DPIA viewGDPR Art. 35AIA viewCanada level I-IVInternal view42001 AIMS evidenceMonitoringKRIs, complaintsTriggersre-open the recordthreshold crossedThe record is the source of truth; each regime gets a generated view, never a separate document.
Every regime gets a view generated from one record. Monitoring feeds triggers, and a trigger re-opens the record rather than starting a fresh document.

The commonest structural mistake is to write one document per regime. Each drifts, the FRIA says one thing about human oversight and the DPIA another, and nobody knows which is current. Instead keep one structured record per system and use, in version control, and generate each regime's view from it. The record holds the facts once: purpose, decision supported, affected groups, data, model and version, oversight design, redress, risks with likelihood and severity, mitigations with owners, and the decision. A pull request is the change log, and the reviewer's approval is the sign-off.

Running the assessment

A workable process has seven steps, sized so the screening is cheap and the full assessment is reserved for systems that need it.

  1. Screen. Five questions: does the output feed a decision about a person; is that decision about access to something important (money, work, housing, benefits, health, education, liberty); is the person able to opt out; are vulnerable groups likely affected; is the system in an Annex III category. Any yes means a full assessment.
  2. Describe the use, not the model. Write the deployer's process step by step, marking where the AI output enters and who acts on it. This is element (a) and most of (b).
  3. Map affected people. List direct subjects, indirect parties (family, co-applicants), and groups whose error rates may differ. Talk to at least one person from the affected population or an advocate for it.
  4. Enumerate harms by mechanism. Wrong output, right output used wrongly, automation bias, unequal error rates, loss of contestability, chilling effects. For each, estimate severity and reversibility, not just likelihood.
  5. Test the claims. Measure error rates per group on held-out data, and measure oversight: how often do reviewers overturn the model? The metrics are covered in AI Fairness, in depth.
  6. Decide with conditions. Proceed, proceed with mitigations that have owners and dates, or stop. Record who decided.
  7. Wire the triggers. Turn every assumption into a monitored quantity with a threshold that re-opens the record.

Scoring with Canada's AIA

Canada's Algorithmic Impact Assessment is the most concrete public scoring scheme, and it is worth implementing even outside Canadian government because it forces explicit answers. It has 65 risk questions with a maximum raw impact score of 169 and 41 mitigation questions with a maximum of 77. The published rule is: if the mitigation score is less than 80% of the maximum, the current score equals the raw impact score; if it is 80% or more, "15% is deducted from the raw impact score". This code reads that as multiplying by 0.85. The level bands are 0 to 25%, 26 to 50%, 51 to 75% and 76 to 100% of the maximum; because the published bands leave gaps for fractional percentages, the code uses explicit upper bounds so 25.6% falls in Level II. Confirm both readings against the current questionnaire before relying on them.

RAW_MAX, MIT_MAX = 169, 77

def aia_level(raw: int, mitigation: int) -> tuple[float, int]:
    """Return (current score as percent of RAW_MAX, impact level 1-4)."""
    if not (0 <= raw <= RAW_MAX and 0 <= mitigation <= MIT_MAX):
        raise ValueError("score out of range")
    current = raw * 0.85 if mitigation >= 0.8 * MIT_MAX else raw
    pct = 100.0 * current / RAW_MAX
    for upper, level in ((25.0, 1), (50.0, 2), (75.0, 3)):
        if pct <= upper:
            return pct, level
    return pct, 4

assert aia_level(80, 50) == (100 * 80 / 169, 2)      # 47.3%, no deduction
assert aia_level(80, 62)[1] == 2                      # 62 >= 61.6: 40.2%
assert aia_level(130, 70)[1] == 3                     # 110.5 -> 65.4%

The level then selects obligations: peer review, notice to affected people, the depth of human involvement in decisions, explanation and training requirements all scale with it. Keeping the scoring in code means a change to one answer in the record re-computes the level in CI, and a level change fails the build until someone signs it off.

Worked example: housing triage

Consider a city agency that deploys an LLM assistant to triage applications for emergency housing assistance. It reads the free-text application and supporting documents and proposes a priority band; a caseworker confirms or changes it. As a public body using AI to evaluate eligibility for public assistance, the agency falls under Article 27 once Annex III obligations apply, and it also processes special category data, so a DPIA is required too.

Use. About 1,200 applications a month, every application scored, bands drive the order in which caseworkers open files. Affected people. Applicants, their children, and applicants writing in a second language, whose text the model may read as less urgent. Harms. The serious one is delay: a family wrongly placed in the lowest band waits weeks, which is severe and partly irreversible. Automation bias compounds it, because caseworkers under load accept the band.

Testing. On 2,000 historical applications with known outcomes, the agency compares the miss rate (truly urgent cases placed in the lowest band) for applications written in English and in other languages, and audits a sample of caseworker decisions to measure override rates. Suppose the miss rate is 3% for English and 9% for other languages, and caseworkers overturn the band in 4% of files. Those numbers are the finding: a group-dependent error with weak human correction.

Decision. Proceed with conditions: translate applications before scoring and re-test; never assign the lowest band without a caseworker reading the file; show the extracted urgency evidence beside the band; give applicants a phone line to request review. The post-translation re-test must bring the miss-rate ratio to 1.5x or better before go-live. Triggers. Monthly miss-rate ratio above 1.5x, an override rate below 2% or above 8% in a random audit sample, and any complaint alleging delay.

Reassessment triggers as code

An assessment is only true for the system and use it describes. Write reassessment triggers as code alongside the record so they cannot be forgotten.

TRIGGERS = {
    "model_change":   lambda ev: ev["model_id"] != RECORD["model_id"],
    "new_population": lambda ev: ev["new_language_share"] > 0.05,
    "parity_breach":  lambda ev: ev["miss_rate_ratio"] > 1.5,
    "weak_oversight": lambda ev: ev["override_rate_audit"] < 0.02,
    "high_override":  lambda ev: ev["override_rate_audit"] > 0.08,
    "purpose_drift":  lambda ev: ev["new_decision_type"] is not None,
}

def check(ev: dict) -> list[str]:
    fired = [name for name, rule in TRIGGERS.items() if rule(ev)]
    if fired:
        open_reassessment(RECORD["id"], reasons=fired)   # creates a ticket + PR
    return fired

Note the weak-oversight trigger. An override rate near zero is not proof that the model is right; it is often proof that nobody is checking. Tie triggers into the risk register described in AI Risk Register, in depth so a fired trigger is visible to the people who own the risk.

Failure modes

  • Assessing the model instead of the use. The same model is harmless drafting emails and harmful ranking tenants. Every record names one use.
  • Written after launch. An assessment dated after go-live is evidence that the process is ceremonial. Gate deployment on a signed record.
  • No affected-person input. Teams consistently miss harms that are obvious to the people subject to the system, such as the cost of a phone-only appeal.
  • Averages only. An overall error rate hides group differences; report per group, with confidence intervals for small groups.
  • Mitigations without owners. "Monitor for bias" is not a mitigation; a named metric, threshold, owner and review date is.
  • Document drift. Separate FRIA, DPIA and internal documents disagree within a quarter. One record, generated views.

Trade-offs

Depth costs time, and a heavy process applied to everything teaches teams to game the screening. Keep screening light and honest, and spend effort only where people can be hurt. Quantitative scoring like the AIA is reproducible and auditable but rewards answering the questionnaire rather than thinking; pair it with the narrative harm analysis. Independent review catches blind spots but slows releases; use it for the highest levels. Publishing results builds trust and invites scrutiny of weak mitigations, which is the point. Finally, a regime-neutral record adds a small mapping cost up front and removes a large reconciliation cost later.

What to do next

  1. List every AI use that feeds a decision about a person, and run the five-question screen on each this week.
  2. Create one structured, version-controlled assessment record per in-scope use, with fields covering Article 27 elements (a) to (f) and your DPIA.
  3. Implement the scoring and trigger code in CI so a changed answer recomputes the level and a fired trigger opens a reassessment.
  4. Measure error rates per affected group and reviewer override rates before sign-off.
  5. Talk to at least one affected person or advocate, and record what changed as a result.
  6. Gate deployment on a signed record dated before go-live, and map open Annex III uses to the 2 December 2027 date.
Key takeaway: An AI impact assessment looks outward at the people a system's decisions touch, happens before deployment, and ends in a decision with owned conditions. Keep one versioned record per use and generate the FRIA, DPIA and internal views from it, score it reproducibly where a scheme like Canada's AIA exists, test error rates per group and the strength of human oversight, and wire every assumption to a trigger that re-opens the record.