Most organisations that use AI already run some compliance training: a slide deck, a quiz, a completion percentage in a dashboard. Most of it fails the two tests that matter. It does not change what people do with AI tools, and when an auditor or regulator asks who was trained on which version of which policy before a given incident, nobody can answer from records.
This article treats AI compliance training as an engineered control. It covers why training is now an explicit obligation, a role-based curriculum kept as data and tied to the obligations register, assignment triggers driven by identity and policy changes, training as a precondition for access, assessments that measure behaviour instead of attendance, and evidence an auditor can sample. It sits alongside the compliance program that owns the register and the employee AI usage policy whose rules the training teaches.
Why training is now an obligation
Three frameworks make AI training an explicit requirement rather than good practice.
| Source | What it asks for | What evidence looks like |
|---|---|---|
| EU AI Act, Article 4 (AI literacy) | Providers and deployers take measures on the AI literacy of staff and other persons operating or using AI systems on their behalf, considering their knowledge, experience and the context of use | A documented programme, role-based content and records of who received it |
| ISO/IEC 42001, clauses 7.2 and 7.3 | Competence for people whose work affects AI performance, and awareness of the AI policy and their contribution | Competence requirements per role, training records, effectiveness evaluation |
| NIST AI RMF, GOVERN 2.2 | Personnel and partners receive AI risk management training to perform their duties consistent with policies | Curriculum mapped to roles and policies; completion and assessment data |
Article 4 has applied since 2 February 2025. The AI Digital Omnibus amendments adopted in 2026 recast it from ensuring a sufficient level of AI literacy towards taking measures that support its development, an obligation of effort rather than of result; confirm the exact wording against the consolidated text, as the EU AI Act article advises. The engineering consequence is unchanged: you still need to show that measures exist, that they fit each role and context, and that they reached the people concerned, including contractors acting on your behalf.
A role-based curriculum kept as data
Start from roles, not from content. Each role touches AI differently, and the obligations that apply follow from that. A workable starting matrix:
| Role | What they do with AI | Core module | Depth |
|---|---|---|---|
| All staff and contractors | Use approved assistants | Approved tools, data tiers, reporting problems | 45 min, annual |
| Engineers building AI features | Ship prompts, retrieval, agents | Prompt injection, data handling, evaluation gates, logging | 3 h plus a lab |
| Human overseers and reviewers | Approve or correct model outputs | Automation bias, seeded-error practice, escalation | 4 h plus calibration |
| Product owners | Decide use cases and risk tier | Risk classification, transparency duties, change control | 2 h |
| Procurement | Buy AI products | Training-on-inputs terms, model change notices | 1.5 h |
| Executives | Accept residual risk | Accountability, incident escalation, metrics | 1 h |
Keep the curriculum as data in version control, next to the obligations register, so that every module states which obligations and policy clauses it covers, who must take it, what triggers it and what it unlocks:
modules:
- id: AIL-101
title: Using approved AI tools safely
version: "2026.09"
material_change: true # policy POL-AIUSE changed what users may paste
audience: [all_staff, contractor]
covers: [EUAIA-ART4, ISO42001-7.3, POL-AIUSE-4.2]
triggers: [hire, role_change, annual]
due_days: 30
unlocks: [tool_tier_1]
assessment: {kind: scenario, items: 8, pass: 0.75}
- id: AIL-310
title: Human oversight of model decisions
version: "2026.07"
material_change: false
audience: [credit_reviewer, claims_reviewer]
covers: [EUAIA-ART26, ISO42001-7.2, POL-OVERSIGHT-2]
triggers: [hire, role_change, annual]
due_days: 14
unlocks: [review_queue]
assessment: {kind: seeded_errors, cases: 40, pass_catch_rate: 0.8}The covers field closes the loop with the register: a nightly job can list every obligation that requires awareness or competence and fail if no module covers it, the same traceability check the program applies to technical controls.
Event-driven assignment
Assignments come from events, not from an annual spreadsheet. Five triggers cover almost everything: a hire, a role change, gaining access to a new tool tier, a material policy change and the annual refresh. An incident lesson can add a sixth, targeted at the teams involved. The engine joins identity data with the curriculum and with completion records:
from dataclasses import dataclass
from datetime import date, timedelta
@dataclass(frozen=True)
class Assignment:
person: str
module: str
version: str
due: date
reason: str
def assign(people, modules, events, done, today):
# people: id -> set(roles); events: set of (id, trigger); done: (id, module) -> (version, date)
out = []
for pid, roles in people.items():
for m in modules:
if not roles & set(m["audience"]):
continue
prior = done.get((pid, m["id"]))
reason = None
if prior is None:
reason = "not yet completed"
elif prior[0] != m["version"] and m["material_change"]:
reason = f"material change {prior[0]} -> {m['version']}"
elif (today - prior[1]).days > 365:
reason = "annual refresh"
elif any((pid, t) in events for t in m["triggers"]):
reason = "trigger event"
if reason:
out.append(Assignment(pid, m["id"], m["version"],
today + timedelta(days=m["due_days"]), reason))
return outTwo details matter. Only a material change forces re-training; editorial fixes bump the version without reassigning, or people stop reading notifications. And contractors and agency staff must be in people: they are often missing from the HR feed, yet Article 4 explicitly covers persons acting on your behalf.
Training as a precondition for access
Training that has no consequence gets deferred forever. The strongest design ties completion to access: the identity provider or AI gateway grants a tool tier only when the modules it requires are complete at their current material version. Joiners and people facing a new material version get a grace period ending at the due date, so work is not blocked, and access lapses automatically if the date passes.
Gating works because it changes the default. Without it, completion rates sit around whatever reminders can achieve; with it, the people who use a tool most are, by construction, trained on its current rules. It also produces evidence for free: every grant records the evidence id that justified it.
Assessments that measure behaviour
Completion proves attendance. Effectiveness, which ISO/IEC 42001 asks you to evaluate, needs assessments that resemble the work. Three formats measure behaviour well:
- Scenario items for general staff: a realistic request ("summarise this customer complaint, which includes an account number, using the public chatbot") with choices that reflect real options. Rotate items from a bank so answers cannot be memorised.
- Labs for engineers: a small retrieval app with a planted indirect prompt injection and an over-permissive tool; the pass condition is a finding and a fix, graded against a rubric.
- Seeded-error calibration for human overseers: a queue of real-looking cases where a known share of model outputs is wrong. Measure the catch rate and the false-alarm rate. This is the only format that directly tests resistance to automation bias.
Then look for effects outside the course. Gateway telemetry shows whether blocked attempts to send restricted data fall after the module ships; incident reports show whether staff report AI problems sooner. These signals belong in the compliance metrics as effectiveness measures, next to coverage and timeliness.
Evidence an auditor can sample
An auditor will sample. Expect questions like: show me that this reviewer was trained on the oversight policy in force on 3 March, and show me what the training said. The completion record has to answer both:
{
"evidence_id": "trn-2026-0412-88f1",
"person": "u-20817",
"roles_at_completion": ["credit_reviewer"],
"module": "AIL-310",
"module_version": "2026.07",
"content_sha256": "9c4e...e1",
"covers": ["EUAIA-ART26", "ISO42001-7.2", "POL-OVERSIGHT-2"],
"assigned_reason": "hire",
"completed_at": "2026-04-12T10:41:07Z",
"assessment": {"kind": "seeded_errors", "catch_rate": 0.85, "false_alarm_rate": 0.05, "passed": true}
}The content hash pins the exact material, so later edits do not rewrite history. Archive every module version, keep records for at least as long as the systems they relate to are in service plus your retention policy, and make the store append-only. That is what turns a training programme into an audit population; audit preparation covers sampling it.
Worked example: a lending rollout
A lender with 1,200 employees and 150 contractors deploys an internal assistant and an LLM that drafts credit-decision summaries for 60 reviewers. The curriculum assigns AIL-101 to all 1,350 people, a 3-hour engineering module to 180 engineers, AIL-310 to the 60 reviewers, a procurement module to 25 buyers and a briefing to 12 executives. Total effort is 1,350 x 0.75 + 180 x 3 + 60 x 4 + 25 x 1.5 + 12 x 1, about 1,840 person-hours a year.
In this example, after 30 days AIL-101 coverage is 94 percent of employees but only 71 percent of contractors, because the agency feed was added late; that gap is the first finding. The tool tier 1 gate is switched on at day 31, and contractor completion reaches 97 percent within two weeks. For reviewers, a baseline seeded-error run before training catches 52 percent of planted errors; after AIL-310 the catch rate is 81 percent with a false-alarm rate of 6 percent. Two reviewers below 60 percent get coaching and a re-test before keeping queue access. These figures are illustrative, but the shape is typical: gating moves coverage, and seeded errors reveal differences that completion data cannot.
Failure modes
- Checkbox training. A generic video and a five-question quiz with obvious answers. Completion is high and behaviour does not change; the first incident shows it.
- Stale content. The policy changed in June and the module still teaches the old data rules. Tie module versions to policy versions and flag material changes.
- Contractors missing. The HR feed covers employees only. Add agency and vendor identities, and check that people with tool access all appear in the training population.
- Reassignment fatigue. Every typo fix re-assigns 1,350 people. Separate editorial from material versions.
- Training instead of controls. A module tells staff not to paste secrets while the gateway lets them. Training complements technical controls; it does not replace them.
- Evidence that cannot be reconstructed. Records without version and content hash cannot answer what someone was taught at a given date.
Trade-offs
Gating access gives the best coverage but creates friction and an operational dependency: if the learning platform is down, access requests stall, so the gate needs a fail-open grace rule with logging. Scenario and lab assessments cost more to build than quizzes and need item rotation, but they are the only way to evaluate effectiveness. Short, frequent, role-specific modules beat one long annual course for retention, at the price of more versions to maintain. Size the programme to risk: general staff need awareness; reviewers of high-risk decisions need competence that is measured.
What to do next
- List every role that uses, builds, oversees, buys or approves AI, including contractors, and map each to the obligations it touches.
- Put the curriculum in version control with covers, audience, triggers, unlocks and material-change flags.
- Add a nightly check that every awareness or competence obligation in the register is covered by a module.
- Drive assignments from identity events and policy versions, not annual lists.
- Gate the highest-risk tool tiers and review queues on current completion, with a grace period and logged break-glass.
- Replace quizzes with scenario items, engineering labs and seeded-error calibration for overseers, and report catch rates.
- Store completion records with module version and content hash in an append-only store, and run a mock audit sample.