An AI governance program is the set of roles, rules, records and routines that lets an organisation say, for every AI system it builds, buys or lets staff use, who owns it, how risky it is, which controls apply, and what evidence shows those controls ran. Most organisations start with a policy document and a committee. Neither works on its own. A policy nobody can apply to a real launch decision becomes shelfware. A committee with no inventory reviews only the projects that happen to be brought to it.
This article covers the structure of the whole program: operating model, artefacts, lifecycle gates and framework mapping. The decision-making body itself is covered in depth in the governance council article. Here it appears as one component among several. By the end you should be able to draft an operating model, stand up an inventory with automated tiering, and plan the first two quarters of a program that produces evidence rather than slides.
What the program governs and what it produces
Start by fixing scope, because scope fights come up early and keep coming back. A workable definition covers three populations. Built systems: models you train or fine-tune, and applications that call a model API, including retrieval pipelines and agents. Bought systems: vendor products with AI features, from a CRM's lead scoring to a support desk's reply drafting. Used systems: general-purpose assistants employees use for their own work, governed mainly through an acceptable-use policy and approved tooling.
The program's outputs are concrete. It makes decisions (approve, approve with conditions, reject, retire), requires controls (an evaluation suite, a human approval step, a data-retention limit), and keeps evidence (the eval report, the approval record, the monitoring dashboard). If an activity produces none of these three, question whether it belongs in the program. Security asks whether an attacker can misuse a system, and model risk asks whether it is accurate enough. Governance asks whether someone accountable made a recorded, reviewable decision to accept the risk that remains.
What the program governs and what it produces
Start by fixing scope, because scope fights come up early and keep coming back. A workable definition covers three populations. Built systems: models you train or fine-tune, and applications that call a model API, including retrieval pipelines and agents. Bought systems: vendor products with AI features, from a CRM's lead scoring to a support desk's reply drafting. Used systems: general-purpose assistants employees use for their own work, governed mainly through an acceptable-use policy and approved tooling.
The program's outputs are concrete. It makes decisions (approve, approve with conditions, reject, retire), requires controls (an evaluation suite, a human approval step, a data-retention limit), and keeps evidence (the eval report, the approval record, the monitoring dashboard). If an activity produces none of these three, question whether it belongs in the program. Security asks whether an attacker can misuse a system, and model risk asks whether it is accurate enough. Governance asks whether someone accountable made a recorded, reviewable decision to accept the risk that remains.
The operating model: three lines, hub and spokes
The structure that holds up in practice combines the Institute of Internal Auditors' Three Lines Model with a hub-and-spoke delivery pattern. The first line is the teams that build and run AI systems. They own the risk, because they are the only people who can change the system. The second line sets standards and challenges the first: risk, legal, privacy, compliance and security. The third line, internal audit, gives independent assurance that the first two are doing what they claim.
The hub is a small AI governance office, often two to six people. It owns the policy, inventory, tiering, control library and tooling, but it does not review every system. The spokes are champions in each product area. They run low-tier reviews locally and escalate the rest, so the hub never becomes a queue.
| Role | Owns | Typical failure if missing |
|---|---|---|
| Executive sponsor | Risk appetite, budget, final escalation | Program has authority on paper only |
| Governance office (hub) | Policy, inventory, tiering, control library, metrics | Every team invents its own process |
| Council | Tier-3 approvals, exceptions, policy changes | Hard calls get made in hallways, unrecorded |
| Champions (spokes) | Tier-1/2 reviews, local inventory hygiene | Hub becomes the bottleneck |
| System owner | One named person per inventory record | Incidents with no one to page |
| Second line | Standards, sign-off on their domain | Legal or privacy finds out after launch |
| Internal audit | Sampling evidence against controls | Controls drift, nobody notices |
Write decision rights down per gate, not per policy: the argument is always about who can stop a launch. The AI security leader is usually accountable for the security controls, not the whole program. Placing the whole program under security tends to shrink it into threat review.
The five artefacts the program runs on
Five artefacts carry the program. Get these right and most process questions answer themselves.
- A four-layer policy stack. A two-page policy mandates the program. Standards make it testable, for example "tier-3 systems need a documented evaluation on intended-use data before launch". Procedures say how, and guidelines are advice. Only the policy needs board approval.
- An AI system inventory. One record per system: owner, purpose, model and vendor, data categories, affected people, status, tier and evidence links. Without it you cannot answer a regulator, an auditor or your own incident commander.
- Risk tiering rules. A deterministic function from inventory fields to a tier. The tier decides which gates apply, who approves, and how often the system is re-reviewed.
- A control library. Each control has an ID, an objective, the tiers it applies to, the evidence it produces and the external requirements it satisfies. Examples: pre-launch evaluation, red-team exercise, human oversight for consequential actions, model card published, logging retained for a set period.
- An evidence store. Wherever artefacts live (a GRC tool, a repository, a wiki), each must be linked from the inventory record and stamped with date, system version and approver. Evidence without a version is unverifiable.
Risks that survive the controls go in a risk register with an owner and a review date, so accepted risk stays visible instead of being accepted once and forgotten.
Inventory and tiering as code
Treat the inventory as data and tiering as code. A YAML record per system, kept in a repository, gives you review history, ownership through the repository's code-owner rules, and CI that refuses incomplete records. The record below describes an internal assistant that drafts responses to credit-card disputes.
id: ai-0042
name: dispute-summary-assistant
owner: priya.n@example.com
status: pilot # proposed | pilot | production | retired
source: built # built | bought | used
model: {provider: vendor-x, family: general-llm, finetuned: false}
purpose: Summarise a customer's dispute file and draft a response for an agent
affected_people: customers
decision_role: advisory # none | advisory | automated
data: [pii, financial]
actions: [] # tools the system can invoke
jurisdictions: [EU, US]
evidence:
eval_report: evidence/ai-0042/eval-2026-09.pdf
dpia: evidence/ai-0042/dpia.pdfREQUIRED = {"id", "name", "owner", "status", "source", "purpose",
"affected_people", "decision_role", "data", "actions"}
SENSITIVE = {"pii", "financial", "health", "biometric", "children"}
CONSEQUENTIAL_DOMAINS = {"credit", "employment", "insurance", "education",
"essential_services", "law_enforcement"}
def tier(rec):
missing = REQUIRED - rec.keys()
if missing:
raise ValueError(f"{rec.get('id')}: missing {sorted(missing)}")
score = 0
if rec["decision_role"] == "automated":
score += 3
elif rec["decision_role"] == "advisory":
score += 1
if set(rec["data"]) & SENSITIVE:
score += 1
if rec["affected_people"] in ("customers", "public", "applicants"):
score += 1
if rec["actions"]: # can change the world, not only text
score += 2
domains = set(rec.get("domains", []))
if domains & CONSEQUENTIAL_DOMAINS:
return 3 # floor: always fully reviewed
return 1 if score <= 1 else 2 if score <= 3 else 3
GATES = {1: ["inventory", "champion_review"],
2: ["inventory", "champion_review", "eval_report", "privacy_review"],
3: ["inventory", "council_approval", "eval_report", "privacy_review",
"security_review", "human_oversight_design", "monitoring_plan"]}Two design choices matter more than the weights. The domain floor sends credit, employment and similar systems to tier 3 whatever their score, so editing a minor field cannot lower the tier. And the output is a list of gates that CI checks against the record's evidence links. A pull request moving status to production fails until every required gate has an artefact. Our example scores advisory (1) + sensitive data (1) + customers (1) = 3, which is tier 2. If the team later tags the system with the credit domain, it jumps to tier 3 and picks up council approval.
Lifecycle gates
Gates belong where work already happens, not in a separate governance calendar. A typical sequence:
| Gate | Trigger | Output |
|---|---|---|
| Intake | New idea, vendor purchase request, new use of existing system | Draft inventory record, provisional tier |
| Design review | Before significant build spend | Approved intended use, data plan, oversight design |
| Pre-deploy | Release candidate | Eval report, red-team findings closed or accepted, sign-offs |
| Monitor | Continuous after launch | Drift and incident metrics, scheduled re-review |
| Change | New model version, new data, new action, new user group | Re-tier; repeat gates the change affects |
| Retire | Decommission | Data disposal record, inventory status retired |
The change gate is the one most programs miss. One new tool definition can turn an advisory summariser into an agent that files refunds. Wire re-tiering into model-registry promotion, agent tool manifests and procurement. If no purchase order for AI-featured software clears without an inventory ID, you catch the shadow AI that never passes through engineering.
One control library, many frameworks
Frameworks describe what good looks like, but none gives you an operating model. Running a separate program for each framework is the commonest waste. Build one control library and tag each control with the requirements it meets.
The NIST AI Risk Management Framework is voluntary, and NIST AI 600-1 adds a generative AI profile to it. ISO/IEC 42001:2023 is a certifiable AI management-system standard that uses the same management-system structure as ISO/IEC 27001, giving the program a clause structure. The EU AI Act is law. It puts obligations such as risk management, logging and human oversight on high-risk providers and deployers, and it applies in stages. A 2026 amendment package (the AI "Digital Omnibus") pushed the main stand-alone high-risk obligations back to December 2027, while prohibitions and general-purpose model obligations kept their earlier dates. Check the current consolidated text before you plan against a date.
The cells below are coarse groupings, not clause or subcategory numbers.
| Control | NIST AI RMF | ISO/IEC 42001 | EU AI Act (high-risk) |
|---|---|---|---|
| AI system inventory with owners | Govern | AI system inventory and roles | Supports provider and deployer duties |
| Pre-deploy evaluation on intended-use data | Measure | Operational controls, verification | Accuracy and robustness requirements |
| Human oversight design | Govern, Map | Operational controls | Human oversight requirement (Art. 14) |
| Event logging retained | Manage | Operation and monitoring | Record-keeping requirement (Art. 12) |
| Post-market monitoring and incidents | Manage | Performance evaluation, improvement | Post-market monitoring, serious-incident reporting |
Worked example: the first two quarters
Take a payments company with 900 staff, around 30 teams shipping software, and no program. Here is a realistic first two quarters.
Weeks 1-4. The COO, as sponsor, approves a two-page policy and staffs the hub with a lead and an analyst. Discovery uses three sources: a survey of engineering leads, vendor spend on AI features, and egress logs to model APIs. That finds 61 candidates, 14 of them duplicates or dead experiments.
Weeks 5-8. The other 47 go into the inventory. Tiering gives 29 tier-1, 13 tier-2 and 5 tier-3 systems, including a fraud model that places automated holds. Champions are named in six product areas.
Weeks 9-16. The council reviews only the five tier-3 systems. Two get conditions: human review of automated holds above a threshold, and a quarterly bias evaluation. One vendor tool is paused until the vendor supplies evaluation evidence. The CI gate and the procurement check go live.
Weeks 17-26. Audit samples ten records. It finds three stale evidence links and one production system whose owner has left. Findings like these mean the program is working, because the inventory made the gaps visible.
Metrics that show the program works
Measure coverage, flow and outcome, and report all three.
- Coverage: share of discovered systems with a complete inventory record; share with a named current owner; share of tier-3 systems with every gate evidenced. Estimate the denominator from egress and spend data, not from self-reporting.
- Flow: median and 90th-percentile days from intake to decision per tier. If tier-1 takes more than a few days, teams will route around you.
- Outcome: AI-related incidents by tier and whether the system was inventoried, overdue re-reviews, open exceptions past expiry, and audit findings per sampled record.
Failure modes
- Principles without gates. Fairness and transparency statements that are never turned into a testable standard. Fix: every principle maps to at least one control with evidence.
- Inventory rot. Records created at launch and never touched again. Fix: re-review dates driven by tier, owner validation against the HR directory, and the change gate wired into the model registry.
- Hub as bottleneck. All reviews centralised, launches wait weeks, and teams stop declaring systems. Fix: champions own tier-1 and tier-2, and you publish your own cycle-time numbers.
- Exceptions that never expire. Conditions nobody follows up. Fix: every exception gets an owner and an expiry date, and an expired exception counts as an open finding.
- Vendor blind spot. Only built systems are governed. Fix: the procurement gate, plus contract clauses requiring notice of material model changes.
Trade-offs
Centralised versus federated. A central hub is consistent but becomes slow as the portfolio grows, while federated champions scale but drift. Combine central rules with federated execution, checked by audit.
Formula versus judgement. Scored tiering is consistent but can be gamed. Keep the formula and the domain floors. Any reviewer may raise a tier, but lowering one needs council sign-off.
Certify or align. ISO/IEC 42001 certification reassures customers but costs audit effort. Certify when contracts ask for it.
What to do next
- Write the scope sentence (built, bought, used) and get the sponsor to sign a two-page policy that mandates the inventory and tiering.
- Discover systems from three sources at once: team survey, AI-related vendor spend, and egress logs to model APIs. Reconcile the three lists.
- Create the inventory as reviewed data in a repository, with required fields enforced in CI.
- Implement tiering as code with domain floors, and map each tier to a list of required gates.
- Build a control library of 15-25 controls, each with evidence and a framework crosswalk.
- Wire the change gate into model-registry promotion, agent tool manifests and procurement.
- Name champions per product area, and stand up the council with a charter for tier-3 decisions only.
- Publish coverage, flow and outcome metrics quarterly, and invite internal audit to sample within six months.