Most AI features do not need a committee. A summariser for internal wiki pages, an autocomplete in a code editor, a classifier that tags support tickets for routing: these can ship through ordinary engineering review. A small number of features are different. They take actions on their own, operate in domains where a wrong answer hurts someone, mix data across users, or publish generated content to the world. For those, a product team should not be the only party deciding whether the risk is acceptable, and an AI feature review board is the mechanism that brings a second, informed opinion in before launch.
This article is about the per-feature review itself: which features it catches, what evidence it demands, what decisions it can make, and, most importantly, how its decisions become binding on the deploy pipeline instead of living in meeting notes. The body that sets policy for the whole portfolio, with its charter and decision rights, is covered in AI Governance Council, in depth. The board described here is the operational arm that applies that policy one feature at a time.
What the board decides
A review board answers one question per feature: given what we know about this feature today, should it reach users, at what scale, and under which conditions? Everything else is somebody else's job. The board does not write the threat model, it reads it. It does not run the evaluations, it decides whether they are adequate. It does not own the feature after launch, but it does own the list of promises the team made to get it approved.
Membership follows the risks being judged. A typical full board has a security engineer, a privacy or data protection lead, someone from legal or compliance, a domain expert for the area the feature touches (a clinician for health, a credit risk analyst for lending), and an ML engineer who can read an evaluation report critically. The team presenting the feature is never part of the quorum for its own decision.
Triggers as code over a feature manifest
The stub version of this page listed triggers in prose: autonomous actions, high-risk domains, cross-user data, publicly shared generated content. Those are the right instincts, but prose triggers are applied inconsistently, and teams learn to describe their feature in whatever words avoid them. The fix is to make every AI feature declare a small manifest in its repository and to evaluate triggers as code against it.
# ai-feature.yaml, checked in next to the service
feature: portal-reply-drafter
owner: patient-experience
model: hosted-llm-2026-08 # pinned model version
autonomy: suggest # suggest | act_with_confirm | act
domain: healthcare
data_classes: [phi]
cross_user_data: false
output_visibility: clinician_reviewed # internal | clinician_reviewed | user | public
tools: []HIGH_RISK_DOMAINS = {"healthcare", "finance", "legal", "employment", "housing", "education"}
SENSITIVE = {"phi", "biometric", "minors", "financial_account"}
TRIGGERS = [
("autonomous_action", lambda m: m["autonomy"] == "act"),
("high_risk_domain", lambda m: m["domain"] in HIGH_RISK_DOMAINS),
("cross_user_data", lambda m: m["cross_user_data"]),
("public_generation", lambda m: m["output_visibility"] == "public"),
("sensitive_data", lambda m: bool(set(m["data_classes"]) & SENSITIVE)),
("tool_side_effects", lambda m: any(t.get("writes") for t in m["tools"])),
]
def review_route(manifest):
hits = [name for name, pred in TRIGGERS if pred(manifest)]
if not hits:
return "self_attest", hits
if len(hits) >= 2 or "autonomous_action" in hits:
return "full_board", hits
return "light_review", hitsThree routes come out of this. Features that fire no trigger self-attest: the owner confirms the manifest is accurate and the record is filed. One trigger sends the feature to a single reviewer with a short checklist. Two or more, or any autonomous action, goes to the full board. A CI check runs review_route on every change to the manifest, so the route is recomputed whenever the feature changes, not only at first launch.
The manifest is self-declared, so it can lie. Two cheap cross-checks catch most inaccuracies: a scan of the service's code for calls to model endpoints and tool registries that do not match the declared tools and model, and the data catalogue's classification of the tables the service reads. A mismatch fails the build with a message pointing at the manifest line.
The review packet
A full review is only as good as the packet the team brings. Boards that accept slide decks get marketing; boards that specify the packet get evidence. Each item below has a named template, and the board secretary rejects incomplete packets before the meeting rather than spending board time discovering gaps.
| Packet item | What it must contain | Common gap |
|---|---|---|
| Feature manifest | The file above, plus a one-paragraph plain description | Autonomy understated |
| Threat model | Assets, entry points, abuse cases, prompt injection paths; see threat modelling LLM systems | Only external attackers considered |
| Impact assessment | Who is affected, how, and the worst credible harm; see AI impact assessment | Harm to non-users ignored |
| Evaluation report | Datasets, sample sizes, metrics with intervals, slices, known failures | Aggregate accuracy only |
| Red team summary | Scope, techniques tried, findings and fixes | Run by the building team alone |
| Mitigations | Each risk mapped to a control and its test | Controls without tests |
| Rollout plan | Stages, cohort sizes, success and stop thresholds | No stop threshold |
| Rollback plan | Kill switch, who can pull it, last time it was exercised | Never exercised |
Two rules keep the packet honest: every metric comes with its sample size and an interval, and every mitigation names the test that proves it works.
Four outcomes and conditions that bind
The board has four possible outcomes, and naming them precisely prevents the most common failure, a vague yes. Approve means launch per the plan. Approve with conditions means launch is permitted, but specific conditions must be met before named rollout stages, or by named dates. Revise and resubmit means the evidence is insufficient to decide; the board states exactly what is missing. Reject means the risk cannot be brought within appetite by any change the team proposed, and the record says why, so the decision can be appealed to the governance council with reasons rather than politics.
Approve with conditions is the outcome that matters most, because it is where most high-risk features land, and it is where boards lose their teeth. A condition such as "improve escalation recall" is unenforceable. A usable condition has an identifier, a testable statement, an owner, the rollout stage it blocks, and a due date:
- id: C-214-1
text: "Urgent-symptom detector recall >= 0.97 on held-out set n >= 600"
evidence: eval/urgent_recall_report.json
owner: ml-platform
blocks: pilot
due: 2026-11-15
- id: C-214-2
text: "Kill switch exercised in staging within the last 14 days"
owner: patient-experience
blocks: pilot
- id: C-214-3
text: "Clinician edit-rate dashboard live with alert at 40 percent"
owner: patient-experience
blocks: ga
The launch gate
The decision record lands in a registry the deploy pipeline can read. Before any change to a feature flag's rollout stage, the pipeline calls a gate. The gate is deliberately boring: it refuses to advance while a condition blocking that stage is open, and it refuses everything while a condition is overdue.
from datetime import date
STAGES = {"internal": 0, "pilot": 1, "ga": 2}
def launch_gate(feature, target_stage, registry, flags, today=None):
today = today or date.today()
record = registry.decision(feature)
if record is None or record.outcome not in ("approve", "approve_with_conditions"):
return False, [f"no approving decision for {feature}"]
if record.manifest_hash != registry.current_manifest_hash(feature):
return False, ["manifest changed since decision; re-review required"]
blocking = []
for c in record.conditions:
if c.status == "closed":
continue
if STAGES[c.blocks] <= STAGES[target_stage]:
blocking.append(f"{c.id} open: {c.text} (owner {c.owner})")
elif c.due and c.due < today:
blocking.append(f"{c.id} overdue since {c.due}")
if not flags.kill_switch_exercised(feature, within_days=14):
blocking.append("kill switch not exercised in the last 14 days")
return not blocking, blockingTwo details carry most of the value. The manifest hash check means a team cannot get approval for a suggest-only feature and then quietly flip autonomy to act: the hash changes, the gate closes, and the trigger rules recompute the route. And closing a condition requires attaching the evidence artefact named in it, which a reviewer signs off. Owners cannot close their own conditions.
An emergency override needs two named board approvers, expires after 72 hours and opens an incident ticket automatically.
Worked example: a patient portal reply drafter
Consider a hospital group that wants to help clinicians answer patient portal messages. The feature drafts a reply from the patient's message and recent visit notes; a clinician edits and sends it. The figures that follow are illustrative, chosen to show how the board reasons rather than to describe a real deployment.
Routing: the manifest fires high_risk_domain and sensitive_data, so the feature goes to the full board. Autonomy is suggest, which keeps it short of the strictest treatment.
The threat model identifies three serious risks: a reassuring draft that nudges a busy clinician past a message needing urgent care; retrieval pulling another patient's notes; and prompt injection through the patient's own message.
The evaluation report shows clinicians rated 88 percent of 400 historical drafts as usable with minor edits. The board is not very interested in that number. It asks instead about the 52 messages in the set that a triage nurse had marked urgent: the drafting model flagged 47 of them, a recall of about 0.90. With 52 positives the 95 percent interval on that recall is wide, roughly 0.79 to 0.96, so the board cannot conclude the true recall is acceptable.
The decision is approve with conditions. A separate urgent-symptom detector must show recall of at least 0.97 on 600 or more held-out urgent messages before the pilot; when it fires, the draft is suppressed and the message goes to the nurse queue. Retrieval must be scoped by patient identifier at the query layer, with a test that seeds a decoy record and asserts it never appears. Drafts are labelled as AI-generated inside the clinician interface. The edit-rate dashboard must be live before general availability, because a falling edit rate is the earliest signal of automation complacency.
Re-review and post-launch review
Approval is for a specific feature, not a team or a model family. The manifest hash ties the decision to the declared shape, and a short list of changes reopens it: a new model or major version, a new tool with side effects, a new data class, an expansion of the user population or domain, a move from suggest to act, and any severity-one incident involving the feature. Minor prompt edits do not reopen review, but they do re-run the evaluation suite, and a regression beyond a threshold set in the decision record does.
Every full-board feature also gets a scheduled post-launch review, usually 90 days after general availability, comparing live metrics with the packet's predictions. If evaluation reports keep overstating production quality, the board tightens its evidence requirements.
Running the board as a queue
A board is a queue. For a stable queue, Little's law says the average number of features waiting equals the arrival rate times the average time in review. But if six full-board features arrive per week and the board can decide four, there is no steady state at all: the backlog grows without limit, and teams start describing features in ways that avoid the triggers.
Three levers keep the queue healthy. Accurate routing keeps light reviews off the board's agenda. Packet linting before the meeting means board time is spent deciding, not discovering missing documents. And an explicit service level, for example a decision within ten working days of a complete packet, makes delay visible. Track time to decision, the share of conditions closed on time, the share of packets rejected as incomplete, and the incident rate of reviewed versus unreviewed features. The last metric is the only one that tells you whether the board reduces harm.
Failure modes
- Conditions with no enforcement. Conditions live in minutes; launches proceed regardless. The launch gate is the fix.
- Trigger gaming. Teams declare
act_with_confirmwhen the confirmation is a pre-ticked box. Review the user interface, not just the manifest. - Evaluation by aggregate. Overall accuracy hides the slice that matters. Require slice metrics for the harms in the impact assessment.
- Stale approvals. A feature approved on one model runs on another a year later. The manifest hash and model pinning prevent it.
- Board as designer. Reviewers rewrite the feature in the meeting and become accountable for it. Ask for evidence; let the team design.
- Untested rollback. The kill switch exists on paper but fails under load. Require it to be exercised regularly, and record when.
Trade-offs
Every review adds latency, and latency has a cost: competitors ship, and teams route around slow processes. The trade-off is managed by tiering, not by lowering the bar for high-risk features. A second trade-off is between standardised criteria and judgement. Fully mechanical approval rules are predictable but miss novel risks; pure judgement is inconsistent between meetings. The pattern that works is mechanical routing and evidence requirements, with judgement reserved for weighing that evidence and written down as reasons. A third is centralisation: one board sees patterns across the company but becomes a bottleneck, while per-division boards scale but drift. Many organisations run divisional boards with a shared rulebook and periodic cross-review of a sample of decisions. For pre-launch adversarial testing practice, see the red team playbook, and for estimating the size of a risk rather than labelling it, see AI risk assessment.
What to do next
- Write the manifest schema and require it in every repository that calls a model.
- Encode your trigger rules as code and run them in CI on every manifest change.
- Publish packet templates and lint packets for completeness before any meeting.
- Adopt the four outcomes and a condition format with id, test, owner, blocking stage and due date.
- Add a launch gate to the deploy pipeline that reads the decision registry and the manifest hash.
- Define re-review triggers and schedule 90-day post-launch reviews for full-board features.
- Measure time to decision, on-time condition closure and incident rates, and review them quarterly.