Risk appetite is the amount and kind of risk an organization is willing to take on to pursue its goals. For AI systems it answers questions that otherwise get settled case by case, in meetings, by whoever argues hardest. How often may a customer-facing assistant be wrong before it is pulled? May an agent spend money without a human? Will we ship a feature that still fails a fifth of indirect prompt-injection tests if the blast radius is small? Without a written appetite every launch review re-litigates these questions. With a vague one, the answer is always "low appetite" and nothing changes.
This article shows how to turn appetite into something a pipeline can enforce. It separates appetite, tolerance and limit. It then writes appetite per AI risk category, converts statements into measurable tolerances with the sample sizes needed to prove them, and sets limits by autonomy and blast radius. Finally it encodes the whole cascade as a file a launch gate reads, and applies it to two systems. Appetite feeds the residual-score bands of an AI risk register and the alert thresholds of AI risk monitoring; this page is where those numbers should come from.
Appetite, tolerance and limit
Three words do different jobs, and mixing them is the commonest reason appetite statements fail. The appetite is a board-level statement of intent per risk category, such as "we accept moderate risk of minor inaccuracy in internal productivity tools to move quickly, and minimal risk of customer harm from automated decisions". A tolerance is a measurable bound on an outcome derived from that intent, for example "harmful-answer rate in customer channels at most 0.5%, demonstrated at 95% confidence". A limit is an operational constraint on a specific system that keeps it within tolerance, such as "refunds above 200 require a human approver".
Appetite changes rarely and is owned by leadership. Tolerances change when evidence or strategy changes and are owned by risk owners. Limits change often and are owned by system teams within the tolerances. Frameworks back this split. The NIST AI Risk Management Framework expects each organization to define its own risk tolerance rather than prescribing one, and enterprise risk practice generally treats appetite as a strategic statement that has to be cascaded before it can be applied.
Appetite per risk category
A single AI appetite is useless, because an organization can be eager about productivity gains and averse to discrimination at the same time. Write appetite per risk category on a short ordinal scale. Five levels from averse through minimal, cautious and open to eager are common, but the labels matter less than defining what each level permits.
| Risk category | Example appetite | What the level permits |
|---|---|---|
| Harmful or unsafe output to customers | Minimal | Ship only with demonstrated low rates and human escalation |
| Inaccuracy in internal assistive tools | Open | Ship with disclosure and feedback; fix forward |
| Personal data exposure | Averse | No launch with any known leakage path |
| Prompt injection leading to actions | Cautious | Allowed where actions are reversible and capped |
| Unfair outcomes in eligibility decisions | Averse | Decisions require human review; disparity bounds proven |
| IP and licensing of generated content | Cautious | Allowed with provenance filters and indemnified vendors |
| Vendor model dependency | Open | Allowed with an exit plan and a tested fallback |
Each row needs an owner and a date. A row with no owner will never be revisited, and an appetite nobody has revisited in two years is describing a different company. Setting the levels is a governance act; the AI governance council is the usual forum, with board sign-off for the averse and minimal rows.
Tolerances you can prove
A tolerance must name a metric, a population, a threshold and the evidence standard. The evidence standard is what most teams forget. If the tolerance is "harmful-answer rate at most 0.5%" and the team tests 100 prompts with zero failures, it has shown very little: the true rate could easily be 3%. The honest test is that the upper confidence bound of the failure rate is below the threshold. With zero observed failures, the number of independent samples needed for a one-sided 95% bound is ln(0.05) / ln(1 - t). That is 299 samples for a 1% tolerance, 598 for 0.5% and 2,995 for 0.1%. Each failure you observe raises the requirement.
import math
def binom_cdf(k, n, p):
lp, lq = math.log(p), math.log1p(-p)
return sum(math.exp(math.lgamma(n + 1) - math.lgamma(i + 1) - math.lgamma(n - i + 1)
+ i * lp + (n - i) * lq) for i in range(k + 1))
def upper_bound(failures, n, confidence=0.95):
"""Exact one-sided Clopper-Pearson upper bound on the failure rate."""
lo, hi = failures / n, 1.0
for _ in range(60):
mid = (lo + hi) / 2
if binom_cdf(failures, n, mid) > 1 - confidence:
lo = mid
else:
hi = mid
return hi
def samples_needed(tolerance, confidence=0.95):
"""Samples needed to prove the tolerance if zero failures are observed."""
return math.ceil(math.log(1 - confidence) / math.log(1 - tolerance))Two cautions apply. The bound is only as good as the sample: prompts must be drawn from the population the tolerance names, so a tolerance for customer channels needs customer-like prompts, including adversarial ones in realistic proportion. And when an LLM judge labels the failures, the judge's own miss rate must be measured on a human-labelled set and folded in, or the bound is optimistic.
Limits: autonomy and blast radius
Tolerances bound outcomes; limits bound what a system is allowed to do. The two levers are autonomy, meaning what the system can do without a human, and blast radius, meaning how much damage one failure can cause. A cautious appetite for injection-driven actions does not mean "no agents". It means limits that make each failure small and recoverable.
- Per-action caps: a maximum refund, a maximum number of records changed, no deletion without approval.
- Aggregate caps: a daily spend or action budget per agent and per user, so a compromised session cannot repeat a small action ten thousand times.
- Reversibility: irreversible actions such as payments out, emails sent and data deleted need a human or a delay window.
- Scope: least-privilege credentials per tool, and no access to data classes the appetite marks averse.
- Kill switches: a tested way to drop a system to read-only within minutes when a tolerance is breached.
The cascade as code
Put the cascade in version control so that changing a tolerance is a reviewed change with an owner, and so that the launch gate reads the same numbers the board approved.
# appetite.yaml
categories:
harmful_output_customer:
appetite: minimal
owner: head-of-trust
reviewed: 2026-09-01
applies_to: [customer_facing]
tolerance: {metric: harmful_rate, max: 0.005, confidence: 0.95}
injection_actions:
appetite: cautious
owner: ciso
reviewed: 2026-09-01
applies_to: [refund_tool]
tolerance: {metric: injection_success_rate, max: 0.02, confidence: 0.95}
limits: {refund_max: 200, daily_refund_budget: 5000, irreversible_needs_human: true}def launch_gate(appetite, system):
"""system: {"capabilities": set, "evidence": {metric: {...}}, "limits": {...}}"""
findings = []
for name, cat in appetite["categories"].items():
if not set(cat["applies_to"]) & system["capabilities"]:
continue # category not in scope
tol = cat["tolerance"]
ev = system["evidence"].get(tol["metric"])
if ev is None:
findings.append(f"BLOCK {name}: no evidence for {tol['metric']}")
continue
ub = upper_bound(ev["failures"], ev["n"], tol["confidence"])
if ub > tol["max"]:
findings.append(f"BLOCK {name}: upper bound {ub:.4f} > {tol['max']}")
for key, cap in cat.get("limits", {}).items():
actual = system["limits"].get(key)
if actual is None or (isinstance(cap, bool) and actual != cap) \
or (not isinstance(cap, bool) and actual > cap):
findings.append(f"BLOCK {name}: limit {key}={actual} exceeds {cap}")
return (not findings), findingsEach system declares its capabilities, and only categories in scope are checked, so a chat assistant is never blocked by refund limits. Missing evidence blocks rather than passes. That one rule changes behaviour more than any other line, because it forces teams to measure before they argue. Exceptions are possible but go through a separate, recorded path with an expiry, as in the risk register.
Worked example: two systems at the gate
Apply this to two systems. The first is a customer-support assistant under the minimal-appetite harmful-output tolerance of 0.5% at 95% confidence. The team runs 1,200 sampled prompts and the judge, after calibration against human labels, flags 6 harmful answers. The observed rate is 0.5%, but the exact upper bound is about 0.98%, so the gate blocks. The tempting response is to argue that 0.5% equals the tolerance. The correct response is that the evidence cannot yet show compliance. To pass, the team must either reduce failures or collect enough samples to tighten the bound. With 3 failures in 1,200 the bound would be about 0.65%, still blocked; with 4 failures in 2,000 it is about 0.46%, which passes.
The second system is a refund agent under the cautious injection appetite. Its tolerance is an injection success rate of at most 2%, and red-team evidence shows 2 successes in 600 attempts, for an upper bound of about 1.05%, which passes. But its configuration sets refund_max to 500, above the limit of 200, and declares no daily refund budget, so the gate blocks twice. The fix is not more testing. It is lowering the cap so that an injection that does succeed costs at most 200, and adding the daily budget so it cannot be repeated. (The agent is internal-only; if it also answered customers, the harmful-output tolerance would apply too.) The structured approach to finding those attack paths is covered in AI risk assessment.
Keeping appetite current
Appetite is not set once. Review it on a calendar, typically annually for the statement and quarterly for tolerances, and pull reviews forward on triggers. Triggers include a new class of capability such as agents gaining write access, a serious incident in your systems or a peer's, a regulatory change, repeated exceptions against the same tolerance, and tolerances that every system passes trivially. Repeated exceptions mean the appetite and the business strategy disagree, and one of them has to change openly. Tolerances that never bind are either well set or meaningless, and the review should say which. The governance program owns this cadence.
Keep the numbers consistent everywhere they appear. The residual-score bands in the risk register should be derived from the appetite level of each category: an averse category might allow only the low band without board sign-off, while an open category lets owners accept medium residual risk themselves. Monitoring thresholds should be set at or inside the tolerances, with a warning level that fires early enough to act before the tolerance is breached. When one of these documents changes, a check in CI should compare it against the appetite file and fail if they disagree, so that the board's statement, the register, the gate and the dashboards cannot drift apart one edit at a time. A breach in production then has a defined meaning: the system is outside what leadership agreed to carry, and the response, whether restricting limits, falling back to a safer model or switching to read-only, is decided in advance rather than during the incident.
Failure modes
- Platitudes. "We have a low appetite for AI risk" cannot be applied to any decision. Every statement must cascade into at least one measurable tolerance.
- Zero tolerance. A zero failure rate cannot be demonstrated by sampling. Write a small, provable number and pair it with limits that bound the damage of the failures that remain.
- Point estimates. Comparing an observed rate with the threshold lets small samples pass. Gate on upper confidence bounds.
- Averaging across systems. A portfolio rate within tolerance can hide one system far outside it. Apply tolerances per system and per material slice.
- Vendor defaults as appetite. A provider's safety settings encode the provider's appetite, not yours. Measure your own system with them in place.
- Exception creep. Exceptions without expiry become the real appetite. Report exception counts per category to the same forum that set the appetite.
Trade-offs
| Choice | Benefit | Cost |
|---|---|---|
| Strict tolerances | Strong assurance | Large eval sets, slower launches |
| Loose tolerances plus tight limits | Fast launches, damage still bounded | More engineering of caps and approvals |
| Per-category appetite | Precise, decisions are explainable | More statements to own and review |
| Gate blocks on missing evidence | Forces measurement | Friction for low-risk prototypes |
| Calendar plus trigger review | Stays current | Needs owners and trigger watchers |
What to do next
- List your AI risk categories and write an appetite level and an owner for each.
- For every category, write at least one tolerance with a metric, population, threshold and confidence.
- Compute the sample sizes those tolerances imply, and budget eval and red-team work accordingly.
- Define limits per system tier for autonomy, caps, reversibility, scope and kill switches.
- Commit the cascade to a version-controlled file and make the launch gate read it, blocking on missing evidence.
- Route exceptions through a recorded path with an expiry, and report them to the appetite owners.
- Feed tolerances into register bands and monitoring thresholds so the numbers match everywhere.
- Schedule annual and quarterly reviews and write down the triggers that pull them forward.