An AI compliance program answers one question continuously: for every legal, regulatory and contractual duty that applies to our AI systems, can we show that it is met? That sounds like governance, and the two overlap, but they are different machines. Governance decides who may approve a launch, which risks are acceptable, and how disputes are escalated; that is covered in AI governance programs and the AI governance council. The compliance program is the plumbing underneath: it turns external text into atomic obligations, maps them to controls, schedules tests, collects evidence, manages exceptions and handles regulator contact.
This article treats that plumbing as an engineering system. You will see how to take in regulatory change without drowning, how to write an obligations register that machines can check, a traceability job that finds gaps every night, how applicability is computed from the inventory, the exception lifecycle, the testing calendar, how to deal with regulators and auditors, a worked example, and the failure modes that sink programs. It is not legal advice; counsel decides what a rule means, and the program makes sure the organisation keeps doing what counsel decided.
Anatomy of the program
The program has a spine and some supporting parts. The spine is a chain of four record types: obligations (what an external source requires), controls (what we do to meet it), tests (how we check the control works) and evidence (what the test examined). Supporting them are the AI system inventory, which determines which obligations apply where; the exception register, for duties you knowingly do not meet yet; and the external interface to regulators and auditors.
The design rule is many-to-many with no orphans. One control often satisfies several obligations: a single logging control can meet the EU AI Act's record-keeping duty for high-risk systems, an internal security standard and a customer contract clause. But every obligation must reach at least one tested control, and every control must exist because of at least one obligation or an explicit internal decision. Orphan controls waste testing effort; orphan obligations are compliance gaps.
Regulatory change intake
Regulatory change arrives constantly: new laws, amendments, regulator guidance, enforcement decisions, standards updates and customer contract clauses. The EU AI Act alone changed its high-risk timeline in 2026 through the AI Digital Omnibus, moving Annex III high-risk duties to 2 December 2027 and Annex I product duties to 2 August 2028 (details in the EU AI Act, in depth). A program that relies on someone reading the news will miss things.
Run intake as a queue with service levels. Each item gets a source link, a received date, a triage outcome and an owner. Triage has three outcomes: not applicable (with a one-line reason, so the decision can be audited), monitor (proposed text, track until final), or impact assessment. An impact assessment names the affected systems from the inventory, the new or changed obligations, the controls that need to change, and the date the duty applies. Working back from that date gives the engineering deadline, which is usually months earlier once you count build, test and evidence collection.
An obligations register machines can check
The register is where most programs go wrong, because they store obligations as paragraphs copied from the law. A paragraph cannot be tested. Break each source provision into atomic obligations, each with one actor, one action and one condition, and keep a pointer to the exact source text.
# obligations/eu_ai_act_art50_chat.yaml
id: OBL-EUAIA-050-01
source: Regulation (EU) 2024/1689, Article 50(1)
source_version: consolidated text as amended 2026
actor: provider
applies_when:
system.interacts_with_natural_persons: true
system.jurisdictions: [EU]
requirement: >
Design the system so people are informed they are interacting
with an AI system, unless obvious from context.
interpretation_ref: legal-memo-2026-031 # counsel's reading, dated
controls: [CTL-DISC-01] # chat surfaces show disclosure
effective_from: 2026-08-02
owner: compliance-eu@corpTwo fields carry most of the weight. applies_when is a predicate over inventory attributes, so applicability is computed rather than argued system by system. interpretation_ref points to counsel's dated memo. When the interpretation changes, you can find every obligation that depended on it. Check the effective date against the current consolidated text; dates in this example follow the sibling EU AI Act article and should be re-verified whenever the law moves.
Nightly traceability
Once the register, control library, test log and inventory are data, gaps become a query. Run this nightly and treat any output as a defect with an owner, not as a report to read later.
from datetime import date, timedelta
def applies(obl, system):
for field, want in obl["applies_when"].items():
have = system.get(field)
if isinstance(want, list): # any overlap, e.g. jurisdictions
have = have if isinstance(have, list) else [have]
if not set(want) & set(have):
return False
elif have != want:
return False
return True
def trace_gaps(obligations, controls, tests, systems, exceptions, today=None):
today = today or date.today()
waived = {(e["system"], e["obligation"]) for e in exceptions
if e["approved"] and e["expires_on"] >= today}
gaps = []
for obl in obligations:
if obl["effective_from"] > today + timedelta(days=180):
continue # tracked by intake, not yet live
if not obl["controls"]:
gaps.append(("NO_CONTROL", obl["id"], None))
continue
for s in systems:
if not applies(obl, s) or (s["id"], obl["id"]) in waived:
continue
for ctl in obl["controls"]:
last = tests.get((s["id"], ctl)) # latest test record or None
freq = controls[ctl]["test_every_days"]
if last is None:
gaps.append(("NEVER_TESTED", obl["id"], s["id"]))
elif last["outcome"] != "pass":
gaps.append(("FAILING", obl["id"], s["id"]))
elif last["tested_at"] < today - timedelta(days=freq):
gaps.append(("STALE", obl["id"], s["id"]))
return gapsThe 180-day look-ahead is deliberate: obligations that take effect within six months show up as gaps now, which is when engineering can still act. Obligations further out stay in the intake queue. The same job should also report controls that no obligation references, which are candidates for removal.
Exceptions that expire
Some obligations will not be met on the effective date, and pretending otherwise pushes the gap underground. An exception is a recorded, approved decision to accept a known gap for a limited time. Every exception needs the system, the obligation, the reason, the compensating control (what reduces the risk in the meantime), the remediation plan, the risk owner who accepts it, the approver, and an expiry date no more than two quarters out.
Expiry is the important part. When an exception expires, the traceability job puts the gap back in the defect list automatically. Renewal requires the same approval as the original, plus evidence of progress on the remediation plan. Watch the exception count and median age as metrics of their own; a growing pile of renewed exceptions is a program failing slowly. Who is allowed to approve an exception at each risk tier is a governance decision, and should be written into the council's charter.
The testing calendar and event triggers
Tests run on a calendar set by control frequency and system tier, with event triggers on top. A typical split: automated technical controls (logging enabled, disclosure banner present, retention policy applied) run daily or on every deploy; evaluation-based controls (bias, robustness, prompt-injection resistance) run monthly and on every model or prompt change; process controls (risk assessment reviewed, human oversight staffed) are sampled quarterly by the second line. Internal audit, the third line, tests the program itself annually.
Event triggers matter more for AI than for most systems. A model upgrade, a new tool given to an agent, a new jurisdiction or a new user population should each re-open the affected tests, regardless of the calendar. Wire this into the deploy pipeline: when a deployment changes a field in the inventory record, the tests linked to obligations whose applies_when reads that field are marked due.
Working with regulators and auditors
The program is also the organisation's interface to supervisors and auditors, and that interface needs runbooks. Three kinds of contact are common. Information requests ask for documentation within a deadline; with the spine in place, an evidence pack for a system is a query plus an export, and the job is mostly reviewing it before it leaves. Incident notifications have fixed windows: for serious incidents involving high-risk systems, Article 73 of the EU AI Act sets a general limit of 15 days from awareness, with shorter limits for deaths and widespread incidents, and other regimes such as data protection breach notification run on their own clocks. The incident process in LLM incident response should call the compliance runbook at triage, not after containment. Audits and inspections test the spine directly; see AI compliance audits for what the auditor's sampling will look like.
Keep a correspondence log: every submission, its content hash, who approved it and when. Statements to a regulator are commitments; the log lets you check that later statements are consistent with earlier ones.
Worked example: a lender's account assistant
A consumer lender adds an LLM assistant that answers account questions and drafts hardship-plan proposals for human agents to approve. It serves customers in the EU and the United States. Intake had already registered the relevant obligations; the launch triggers applicability.
The inventory record says: interacts with natural persons, EU and US, processes personal data, influences decisions about credit terms, human approval required. Applicability returns 23 obligations: chatbot disclosure under Article 50, data protection duties including the GDPR rules on automated decisions (Article 22, which the human-approval design is meant to keep the system outside of), consumer credit fair-treatment rules from the lender's existing compliance universe, and 9 internal policy obligations. Counsel's memo concludes the assistant is not an Annex III creditworthiness system because it does not evaluate creditworthiness or set scores, but that conclusion depends on the human-approval step. It is recorded as an interpretation with that dependency named.
The traceability job shows 19 obligations already met by existing controls, such as logging, data retention and access control, once the new system is enrolled in their tests. Four are gaps: disclosure text in the mobile app, an override-rate monitor to prove human approval is real, a complaint-handling route for AI-drafted proposals, and a Spanish-language disclosure. Three are closed before launch. The Spanish disclosure gets a 60-day exception, with English-plus-icon disclosure as the compensating control and the product owner as risk owner. Six months later, an agent-efficiency project proposes auto-approving small hardship plans. The inventory change re-runs applicability, the Annex III interpretation's dependency flags it, and the change goes to counsel before it ships rather than after.
Failure modes
- Paragraph registers. Obligations stored as copied law cannot be mapped or tested. Split them into atomic, predicate-driven records.
- Spreadsheet spine. Mappings in a spreadsheet drift from reality within a quarter. Keep them as reviewed data in version control, checked by a job.
- Interpretation drift. Counsel's reading changes but nobody finds the obligations that relied on it. Reference memos by ID, with dates.
- Late intake. A duty is noticed when it takes effect. Use look-ahead in the traceability job and plan back from effective dates.
- Permanent exceptions. Renewals without progress turn gaps into policy. Enforce expiry and report exception age.
- Inventory blind spots. Applicability is only as good as the inventory. Reconcile it against billing and network egress regularly.
Trade-offs
| Choice | Benefit | Cost |
|---|---|---|
| Atomic obligations | Testable, mappable, diffable | Upfront legal and analyst effort |
| Controls shared across frameworks | Test once, satisfy many | One failure hits several regimes |
| Predicate applicability | Consistent, automatic scoping | Inventory fields must be accurate |
| Commercial GRC platform | Workflow and reporting built in | Licence cost, data model lock-in |
| Repository-as-register | Code review, history, CI checks | Needs engineers to maintain tooling |
| Short exception expiry | Gaps stay visible | More approval work |
Many teams run a hybrid: the register and mappings live in a repository with CI checks, and a GRC tool consumes them for workflow, attestations and committee reporting. Whichever you choose, the repository or the tool must be the single source of truth, never both.
What to do next
- Pick your three most important sources and decompose them into atomic obligations with an applicability predicate and source pointer each.
- Add the inventory fields those predicates need, and backfill them for every production system.
- Map obligations to existing controls; write down every obligation with no control as a gap.
- Stand up the nightly traceability job with a 180-day look-ahead and route its output to owners as defects.
- Create the exception register with mandatory compensating control, risk owner and expiry.
- Wire deploy events to re-open affected tests when inventory fields change.
- Write runbooks for information requests and incident notifications, and rehearse one of each this quarter.