Most organisations that deploy AI end up with an AI governance council, and many of those councils fail in one of two ways. Either they review every chatbot experiment and become a bottleneck that teams route around, or they meet monthly, look at slides, approve everything, and govern nothing. Neither changes what ships.
This article treats the council as a decision system to be engineered. It covers what the council should decide and what it should delegate, a charter with seats, quorum and conflicts, intake and tiering as code, decision records that tooling can enforce, a worked example, the metrics that show whether it works, and the failure modes. It assumes you have, or are building, an inventory of AI systems and a security owner for them, as covered in the AI CISO role.
What the council decides
A council exists to make the decisions that no single function can make alone, because they trade risk against value across legal, security, product and the people affected. Write its decision rights down. A typical set:
- Approve, approve with conditions, or reject the highest-risk AI uses, and set the conditions.
- Set risk appetite for AI: which uses are prohibited, which need full review, which are routine.
- Accept residual risk above a threshold, with a named owner and an expiry date.
- Own the AI policy set, including acceptable use, vendor models and data use, and approve exceptions to it.
- Order a pause or rollback of a deployed system after an incident or a failed condition.
Just as important is what it does not decide. It does not pick models, write prompts, or review every feature. Delegate those to the teams and to lightweight review paths, and keep the council for decisions that need judgement across functions. A useful rule of thumb is that if the council sees more than a handful of full reviews a month, the tiering is too coarse.
The charter
The charter is a short document, one or two pages, that the executive sponsor signs. It must answer who sits on the council, how decisions are made, and what happens when the council disagrees with a business owner.
| Seat | Why it is there | Typical decision input |
|---|---|---|
| Executive chair | Breaks ties and owns escalation to the board | Business value, strategic fit |
| Security | Abuse, injection, data exposure, incident response | Threat model, control baseline |
| Privacy and data protection | Lawful basis, retention, data subject rights | Data flow, impact assessment |
| Legal and compliance | Regulation, contracts, liability | Applicable obligations |
| Engineering or ML lead | Feasibility of conditions; evaluation quality | Evals, monitoring plan |
| Product or business owner (rotating) | Owns the use case being reviewed; no vote on it | Value, user impact |
| Risk or internal audit (observer) | Independence; checks the process works | Process evidence |
Set a quorum, for example the chair plus security, privacy and legal, and decide whether any seat holds a veto. Many councils give security and privacy a veto that only the executive sponsor can override in writing. Members must declare conflicts, and a sponsor of a proposal does not vote on it. Name a secretariat, usually one person in the risk or security team. Without someone who owns the queue, the records and the follow-ups, a council degrades into a meeting.
Intake and tiering as code
Tiering decides which path a request takes, and it should be code, not a judgement made in a meeting. Code makes it consistent, auditable and quick to change. Ask factual questions that a requester can answer honestly: what data, who is affected, what the system can do, and who sees the output. Avoid questions like is this high risk?
from dataclasses import dataclass
@dataclass
class Intake:
system_id: str
data: set # {"public", "internal", "personal", "special_category"}
audience: str # "internal" or "external"
actions: set # tools with side effects: {"send_email", "refund", ...}
decides_about_people: bool # hiring, credit, access to services, discipline
human_review_of_output: bool
third_party_model: bool
PROHIBITED_ACTIONS = {"autonomous_termination", "covert_monitoring"}
def tier(r: Intake) -> tuple[int, list[str]]:
reasons = []
if r.actions & PROHIBITED_ACTIONS:
return 4, ["prohibited use"] # 4 = reject without review
if r.decides_about_people:
reasons.append("consequential decisions about people")
if "special_category" in r.data:
reasons.append("special-category personal data")
if r.actions and r.audience == "external":
reasons.append("side-effecting tools reachable by outsiders")
if reasons:
return 3, reasons # full council
if "personal" in r.data:
reasons.append("personal data")
if r.actions and not r.human_review_of_output:
reasons.append("unreviewed side effects")
if r.audience == "external":
reasons.append("external users")
if r.third_party_model:
reasons.append("third-party model")
return (2, reasons) if reasons else (1, ["internal, low impact"])Every tier records its reasons, so a requester can see why they landed where they did, and the council can revisit a rule that sends too much to tier 3. Tier 1 is a checklist the team completes and the secretariat spot-checks. Tier 2 goes to two named reviewers, for example security and privacy, with a five-working-day target. Tier 3 goes to the council. Keep the rules in version control and treat changes to them like policy changes, because they are.
Decision records tooling can enforce
A decision that lives only in meeting minutes cannot be enforced. Record each one as structured data that deployment tooling can read.
{
"decision_id": "AIGC-2026-041",
"system_id": "support-agent",
"scope": {"model": "vendor-model-x", "tools": ["lookup_order", "issue_refund"],
"data": ["internal", "personal"], "audience": "external"},
"outcome": "approved_with_conditions",
"conditions": [
{"id": "C1", "text": "issue_refund capped at 100 EUR per call in tool code", "evidence": "test link"},
{"id": "C2", "text": "refunds above 50 EUR require agent-side human approval", "evidence": "config link"},
{"id": "C3", "text": "injection eval suite passes before each model change", "evidence": "CI job"}
],
"residual_risk_owner": "head-of-support",
"expires": "2027-04-01",
"rereview_triggers": ["model change", "new tool", "new data category", "sev1 incident"]
}Then make the deploy pipeline refuse to ship a system whose scope is not covered by an unexpired approval. Compare the deployed model, tool list and data categories with the record's scope. Any difference is a re-review trigger, not a judgement call for the deploying engineer. Store records next to the risk register entries they mitigate, so an auditor can walk from a risk to the decision, the conditions and the evidence.
Exceptions to policy use the same record with a different outcome. An exception names the policy clause it waives, the compensating control, the owner and an end date no more than a few months out. When the end date passes, the exception lapses on its own and the deploy gate starts refusing, rather than waiting for someone to remember. Track how many exceptions are renewed. A clause that is waived again and again is either a bad rule or a missing platform capability, and both are worth a council agenda item.
Record dissent as well. If a member voted against or a veto was overridden, the record should say who, why and on what authority. That protects the dissenting member, and it gives the next review a starting point.
Cadence, papers and the emergency path
Run three tempos. Triage is continuous: the secretariat tiers new requests within two working days. The full council meets on a fixed cadence, every two to four weeks, with papers circulated in advance and a decision expected in the meeting. An emergency path, the chair plus two quorum members within 24 hours, handles incidents and urgent pauses. Pair the emergency path with a tested technical ability to stop a system, such as an agent kill switch. A council that can order a pause but cannot execute one has no real authority.
Each paper should fit on two pages. It should state the use and its value, the tiering reasons, the threat model, the evaluation results, the proposed conditions, and the residual risk with its owner. Reviewers comment before the meeting, and meeting time goes to disagreements. Publish decisions, with sensitive details removed, so teams can learn what gets approved and design for it from the start.
Worked example: a refund-capable support agent
A support team wants an external-facing agent that answers order questions and can issue refunds, built on a vendor model. The intake says: personal data, external audience, actions {lookup_order, issue_refund}, no decisions about people in the regulated sense, no human review of each reply, third-party model. Tiering returns tier 3 for one reason: side-effecting tools reachable by outsiders.
The paper's threat model highlights prompt injection that pushes refunds through, and data leakage between customers. Evaluation shows the agent resists a standard injection suite in most cases but not all. The council does not ask for a perfect model. It approves with conditions that do not depend on model behaviour: a hard refund cap in tool code, human approval above a lower threshold, an order lookup scoped to the authenticated customer, and an injection suite in CI. It sets a six-month expiry and names the head of support as residual risk owner.
Three months later the team swaps to a newer model version. The deploy gate sees a scope mismatch and blocks the release. Because the change is a trigger, the request goes to tier 2 instead of a full sitting: security re-runs the injection suite, privacy confirms no new data use, and the record is amended in three days.
Where frameworks fit
Frameworks expect a body like this, but none of them dictates its exact shape. The NIST AI Risk Management Framework puts accountability structures, roles and policies under its GOVERN function, which underpins MAP, MEASURE and MANAGE. See the NIST AI RMF in depth. NIST AI 600-1, the generative AI profile, applies the same structure to generative systems. ISO/IEC 42001:2023, the AI management system standard, requires top management to assign roles, responsibilities and authorities. A council with a charter and decision records is a natural way to show that. If you are subject to the EU AI Act, map its obligations for your role, whether provider or deployer, into the tiering rules. Check the current application dates, because parts of the timetable have been subject to revision.
Failure modes
| Failure mode | Symptom | Fix |
|---|---|---|
| Bottleneck | Weeks-long queue; teams launch without asking | Tier in code; delegate tiers 1 and 2; publish SLAs |
| Rubber stamp | Near-100% approval with no conditions | Require threat model and evals in papers; track conditions set |
| No teeth | Conditions never verified | Deploy gate on decision records; evidence links per condition |
| Invisible inventory | Council reviews only what teams bring | Detect AI usage at the gateway and in spend; reconcile with intake |
| Policy drift | Approved scope differs from what runs | Scope diff in CI; re-review triggers |
| Conflicted votes | Sponsors approve their own projects | Declared conflicts; sponsor seat without vote |
| Slide theatre | Meetings present, nobody decides | Decision expected per paper; minutes record the outcome and the dissent |
Metrics that show it works
Measure the council like any other service. Useful indicators are median time to decision per tier, share of requests per tier, share of approvals that carry conditions, share of conditions with verified evidence, expired approvals still running, exceptions past their end date, and incidents in approved systems traced back to a missing or failed condition. Report them quarterly to the executive sponsor. A rising tier-3 share means the rules need tuning, and AI spend outside the inventory means people are routing around the process. Pair the council with an employee AI usage policy, so that everyday use has clear rules and the council can focus on systems.
What to do next
- Write a one-page charter listing decision rights, seats, quorum, vetoes, conflicts and the secretariat, and get the executive sponsor to sign it.
- Turn your tiering questions into code with recorded reasons, and run last quarter's AI projects through it to check the distribution.
- Define a decision record schema with scope, conditions, evidence, owner, expiry and re-review triggers.
- Add a deploy gate that compares the running scope with an unexpired decision record.
- Set SLAs per tier and an emergency path, and test the technical ability to pause a system.
- Publish anonymised decisions so teams can design for approval.
- Report the council's metrics quarterly, and retune the tiering rules when the tier-3 share climbs.