An AI compliance program answers one question continuously: for every legal, regulatory and contractual duty that applies to our AI systems, can we show that it is met? That sounds like governance, and the two overlap, but they are different machines. Governance decides who may approve a launch, which risks are acceptable, and how disputes are escalated; that is covered in AI governance programs and the AI governance council. The compliance program is the plumbing underneath: it turns external text into atomic obligations, maps them to controls, schedules tests, collects evidence, manages exceptions and handles regulator contact.

This article treats that plumbing as an engineering system. You will see how to take in regulatory change without drowning, how to write an obligations register that machines can check, a traceability job that finds gaps every night, how applicability is computed from the inventory, the exception lifecycle, the testing calendar, how to deal with regulators and auditors, a worked example, and the failure modes that sink programs. It is not legal advice; counsel decides what a rule means, and the program makes sure the organisation keeps doing what counsel decided.

Anatomy of the program

Obligations engineering: every rule traces to a control, a test and evidenceRegulatory sourceslaws, guidance, contractsChange intaketriage + impactObligations registeratomic, sourcedControl libraryone control, many dutiesTest calendarfrequency per tierEvidence storehashed, retainedAI inventorytier, jurisdictionappliesExceptionsowner, expiryRegulators, auditorsrequests, reportsevidence packsTraceability checknightly, fails loudlyGovernance decides who may approve what. The compliance program proves the rules are met.
The spine runs top to bottom: obligation, control, test, evidence. The traceability check walks it nightly and reports every broken link.

The program has a spine and some supporting parts. The spine is a chain of four record types: obligations (what an external source requires), controls (what we do to meet it), tests (how we check the control works) and evidence (what the test examined). Supporting them are the AI system inventory, which determines which obligations apply where; the exception register, for duties you knowingly do not meet yet; and the external interface to regulators and auditors.

The design rule is many-to-many with no orphans. One control often satisfies several obligations: a single logging control can meet the EU AI Act's record-keeping duty for high-risk systems, an internal security standard and a customer contract clause. But every obligation must reach at least one tested control, and every control must exist because of at least one obligation or an explicit internal decision. Orphan controls waste testing effort; orphan obligations are compliance gaps.

Regulatory change intake

Regulatory change arrives constantly: new laws, amendments, regulator guidance, enforcement decisions, standards updates and customer contract clauses. The EU AI Act alone changed its high-risk timeline in 2026 through the AI Digital Omnibus, moving Annex III high-risk duties to 2 December 2027 and Annex I product duties to 2 August 2028 (details in the EU AI Act, in depth). A program that relies on someone reading the news will miss things.

Run intake as a queue with service levels. Each item gets a source link, a received date, a triage outcome and an owner. Triage has three outcomes: not applicable (with a one-line reason, so the decision can be audited), monitor (proposed text, track until final), or impact assessment. An impact assessment names the affected systems from the inventory, the new or changed obligations, the controls that need to change, and the date the duty applies. Working back from that date gives the engineering deadline, which is usually months earlier once you count build, test and evidence collection.

An obligations register machines can check

The register is where most programs go wrong, because they store obligations as paragraphs copied from the law. A paragraph cannot be tested. Break each source provision into atomic obligations, each with one actor, one action and one condition, and keep a pointer to the exact source text.

# obligations/eu_ai_act_art50_chat.yaml
id: OBL-EUAIA-050-01
source: Regulation (EU) 2024/1689, Article 50(1)
source_version: consolidated text as amended 2026
actor: provider
applies_when:
  system.interacts_with_natural_persons: true
  system.jurisdictions: [EU]
requirement: >
  Design the system so people are informed they are interacting
  with an AI system, unless obvious from context.
interpretation_ref: legal-memo-2026-031      # counsel's reading, dated
controls: [CTL-DISC-01]                       # chat surfaces show disclosure
effective_from: 2026-08-02
owner: compliance-eu@corp

Two fields carry most of the weight. applies_when is a predicate over inventory attributes, so applicability is computed rather than argued system by system. interpretation_ref points to counsel's dated memo. When the interpretation changes, you can find every obligation that depended on it. Check the effective date against the current consolidated text; dates in this example follow the sibling EU AI Act article and should be re-verified whenever the law moves.

Nightly traceability

Once the register, control library, test log and inventory are data, gaps become a query. Run this nightly and treat any output as a defect with an owner, not as a report to read later.

from datetime import date, timedelta

def applies(obl, system):
    for field, want in obl["applies_when"].items():
        have = system.get(field)
        if isinstance(want, list):              # any overlap, e.g. jurisdictions
            have = have if isinstance(have, list) else [have]
            if not set(want) & set(have):
                return False
        elif have != want:
            return False
    return True

def trace_gaps(obligations, controls, tests, systems, exceptions, today=None):
    today = today or date.today()
    waived = {(e["system"], e["obligation"]) for e in exceptions
              if e["approved"] and e["expires_on"] >= today}
    gaps = []
    for obl in obligations:
        if obl["effective_from"] > today + timedelta(days=180):
            continue                                   # tracked by intake, not yet live
        if not obl["controls"]:
            gaps.append(("NO_CONTROL", obl["id"], None))
            continue
        for s in systems:
            if not applies(obl, s) or (s["id"], obl["id"]) in waived:
                continue
            for ctl in obl["controls"]:
                last = tests.get((s["id"], ctl))       # latest test record or None
                freq = controls[ctl]["test_every_days"]
                if last is None:
                    gaps.append(("NEVER_TESTED", obl["id"], s["id"]))
                elif last["outcome"] != "pass":
                    gaps.append(("FAILING", obl["id"], s["id"]))
                elif last["tested_at"] < today - timedelta(days=freq):
                    gaps.append(("STALE", obl["id"], s["id"]))
    return gaps

The 180-day look-ahead is deliberate: obligations that take effect within six months show up as gaps now, which is when engineering can still act. Obligations further out stay in the intake queue. The same job should also report controls that no obligation references, which are candidates for removal.

Exceptions that expire

Some obligations will not be met on the effective date, and pretending otherwise pushes the gap underground. An exception is a recorded, approved decision to accept a known gap for a limited time. Every exception needs the system, the obligation, the reason, the compensating control (what reduces the risk in the meantime), the remediation plan, the risk owner who accepts it, the approver, and an expiry date no more than two quarters out.

Expiry is the important part. When an exception expires, the traceability job puts the gap back in the defect list automatically. Renewal requires the same approval as the original, plus evidence of progress on the remediation plan. Watch the exception count and median age as metrics of their own; a growing pile of renewed exceptions is a program failing slowly. Who is allowed to approve an exception at each risk tier is a governance decision, and should be written into the council's charter.

The testing calendar and event triggers

Tests run on a calendar set by control frequency and system tier, with event triggers on top. A typical split: automated technical controls (logging enabled, disclosure banner present, retention policy applied) run daily or on every deploy; evaluation-based controls (bias, robustness, prompt-injection resistance) run monthly and on every model or prompt change; process controls (risk assessment reviewed, human oversight staffed) are sampled quarterly by the second line. Internal audit, the third line, tests the program itself annually.

Event triggers matter more for AI than for most systems. A model upgrade, a new tool given to an agent, a new jurisdiction or a new user population should each re-open the affected tests, regardless of the calendar. Wire this into the deploy pipeline: when a deployment changes a field in the inventory record, the tests linked to obligations whose applies_when reads that field are marked due.

Working with regulators and auditors

The program is also the organisation's interface to supervisors and auditors, and that interface needs runbooks. Three kinds of contact are common. Information requests ask for documentation within a deadline; with the spine in place, an evidence pack for a system is a query plus an export, and the job is mostly reviewing it before it leaves. Incident notifications have fixed windows: for serious incidents involving high-risk systems, Article 73 of the EU AI Act sets a general limit of 15 days from awareness, with shorter limits for deaths and widespread incidents, and other regimes such as data protection breach notification run on their own clocks. The incident process in LLM incident response should call the compliance runbook at triage, not after containment. Audits and inspections test the spine directly; see AI compliance audits for what the auditor's sampling will look like.

Keep a correspondence log: every submission, its content hash, who approved it and when. Statements to a regulator are commitments; the log lets you check that later statements are consistent with earlier ones.

Worked example: a lender&#x27;s account assistant

A consumer lender adds an LLM assistant that answers account questions and drafts hardship-plan proposals for human agents to approve. It serves customers in the EU and the United States. Intake had already registered the relevant obligations; the launch triggers applicability.

The inventory record says: interacts with natural persons, EU and US, processes personal data, influences decisions about credit terms, human approval required. Applicability returns 23 obligations: chatbot disclosure under Article 50, data protection duties including the GDPR rules on automated decisions (Article 22, which the human-approval design is meant to keep the system outside of), consumer credit fair-treatment rules from the lender's existing compliance universe, and 9 internal policy obligations. Counsel's memo concludes the assistant is not an Annex III creditworthiness system because it does not evaluate creditworthiness or set scores, but that conclusion depends on the human-approval step. It is recorded as an interpretation with that dependency named.

The traceability job shows 19 obligations already met by existing controls, such as logging, data retention and access control, once the new system is enrolled in their tests. Four are gaps: disclosure text in the mobile app, an override-rate monitor to prove human approval is real, a complaint-handling route for AI-drafted proposals, and a Spanish-language disclosure. Three are closed before launch. The Spanish disclosure gets a 60-day exception, with English-plus-icon disclosure as the compensating control and the product owner as risk owner. Six months later, an agent-efficiency project proposes auto-approving small hardship plans. The inventory change re-runs applicability, the Annex III interpretation's dependency flags it, and the change goes to counsel before it ships rather than after.

Failure modes

  • Paragraph registers. Obligations stored as copied law cannot be mapped or tested. Split them into atomic, predicate-driven records.
  • Spreadsheet spine. Mappings in a spreadsheet drift from reality within a quarter. Keep them as reviewed data in version control, checked by a job.
  • Interpretation drift. Counsel's reading changes but nobody finds the obligations that relied on it. Reference memos by ID, with dates.
  • Late intake. A duty is noticed when it takes effect. Use look-ahead in the traceability job and plan back from effective dates.
  • Permanent exceptions. Renewals without progress turn gaps into policy. Enforce expiry and report exception age.
  • Inventory blind spots. Applicability is only as good as the inventory. Reconcile it against billing and network egress regularly.

Trade-offs

ChoiceBenefitCost
Atomic obligationsTestable, mappable, diffableUpfront legal and analyst effort
Controls shared across frameworksTest once, satisfy manyOne failure hits several regimes
Predicate applicabilityConsistent, automatic scopingInventory fields must be accurate
Commercial GRC platformWorkflow and reporting built inLicence cost, data model lock-in
Repository-as-registerCode review, history, CI checksNeeds engineers to maintain tooling
Short exception expiryGaps stay visibleMore approval work

Many teams run a hybrid: the register and mappings live in a repository with CI checks, and a GRC tool consumes them for workflow, attestations and committee reporting. Whichever you choose, the repository or the tool must be the single source of truth, never both.

What to do next

  1. Pick your three most important sources and decompose them into atomic obligations with an applicability predicate and source pointer each.
  2. Add the inventory fields those predicates need, and backfill them for every production system.
  3. Map obligations to existing controls; write down every obligation with no control as a gap.
  4. Stand up the nightly traceability job with a 180-day look-ahead and route its output to owners as defects.
  5. Create the exception register with mandatory compensating control, risk owner and expiry.
  6. Wire deploy events to re-open affected tests when inventory fields change.
  7. Write runbooks for information requests and incident notifications, and rehearse one of each this quarter.
Key takeaway: A compliance program is a traceable chain from obligation to control to test to evidence, computed against the inventory rather than argued. Decompose law into atomic predicate-driven obligations, check the chain nightly with look-ahead, make exceptions expire, re-test on change, and treat regulator contact as a rehearsed runbook.