Every risk register ends up with the same four treatments: avoid the risk, reduce it, accept it, or transfer it. Transfer is the one that most often gets ticked without analysis. "We have cyber insurance" and "the vendor indemnifies us" are statements about documents, not about how much of a bad year you will actually pay. For LLM systems the gap is wide. Policies were written before generative AI, model vendors cap and condition their indemnities, and some of the largest AI losses, regulatory fines above all, are often legally impossible to insure.
This article treats risk transfer as an engineering problem. You stack the instruments in the order they respond to a loss, model a year of losses through that stack in about fifty lines of Python, and read off what you retain on average and in a bad year. The scenario and every number in the worked example are hypothetical. Replace them with your own register before deciding anything.
What transfer moves, and what it does not
Risk transfer moves the financial consequence of a loss to another party, in exchange for a price: a premium, a higher vendor fee, or a concession in a customer contract. It does not move the event. When your support agent leaks a customer's records, you still run the incident, notify regulators and customers, and live with the headlines. The insurer reimburses some covered costs, later, subject to the policy wording, notice conditions and its own investigation.
So transfer is the last treatment you apply, not the first. Reduce what you can with controls, decide what you are willing to accept against your appetite, and transfer the part of the remaining tail that you cannot afford to carry. The inputs come from a quantified register. If you do not have one, start with AI risk assessment and the AI risk register.
The transfer stack
A loss passes through the layers in a fixed order, and each layer has its own trigger, scope and cap:
- Customer contracts. Limitation-of-liability clauses and disclaimers cap what your own customers can recover from you. This is transfer to the customer. It is strong in B2B and weak against consumers, whose statutory rights you usually cannot contract away.
- Vendor indemnity. Model and platform vendors offer indemnities for named claim types, typically third-party IP claims over outputs. They are capped, often at fees paid over a period, and conditioned on using the vendor's filters and not modifying outputs in excluded ways. AI contracts in depth walks through those conditions.
- Retention. The deductible or self-insured retention you pay on each covered event before the policy responds. High-frequency small claims live almost entirely here.
- Insurance. Pays covered kinds of loss above the retention, up to sublimits and an annual aggregate limit. Once the aggregate is exhausted, the rest of the year is yours.
- Residual. Whatever no layer names, plus everything above every cap.
What cannot be transferred
Some exposure cannot be transferred at any price:
- Fines and penalties. Whether regulatory fines are insurable depends on the jurisdiction and the conduct. Many jurisdictions prohibit it on public-policy grounds. Under the EU AI Act the top tier of fines reaches 7% of worldwide annual turnover, and you should model that as retained.
- Regulatory duties. Obligations attached to your role as provider or deployer, such as risk management, logging, human oversight and incident reporting, stay with you whatever the contract says about who pays.
- Reputation and churn. Some policies cover narrowly defined reputational-harm costs, but lost customers and a damaged brand are mostly retained.
- Basis risk. This is the gap between the loss you suffer and the loss the wording describes. A hallucinated refund policy that costs you goodwill credits may be neither a liability claim nor a covered first-party loss.
Instruments compared
| Instrument | What it actually pays | Watch for |
|---|---|---|
| Cyber / tech E&O policy | Privacy breach costs, third-party claims from errors in your service | Silent-AI ambiguity and new generative-AI exclusions; see insurance for AI systems |
| Affirmative AI liability cover | Named AI failure modes such as hallucination, errors and drift | New, small market; Armilla, a Lloyd's coverholder with Chaucer, advertises limits up to $25M |
| Performance guarantee insurance | Lets a vendor warrant model metrics and pay if they miss | Examples: Armilla Guaranteed (backed by Chaucer, Greenlight Re and Swiss Re) and Munich Re's aiSure contractual-liability variant; aiSure also has own-damage variants for AI you build |
| Vendor indemnity | Named third-party claims, usually IP | Caps, conditions, and the vendor's own solvency |
| Customer liability cap | Limits what customers can recover | Unenforceable against some statutory claims |
Modelling the stack in code
The stack is easy to simulate. Each scenario gets an annual frequency and a lognormal severity, taken from your register. Each simulated year sends every event through the vendor layer (if the kind matches, until the cap is used), then the policy (if the kind is covered, above the retention, until the aggregate is used), and the remainder lands in retained:
import random, math
# Annual loss model for one LLM feature. Each scenario: (frequency per year, median $, sigma, kind)
SCENARIOS = {
"wrong_advice_claim": (2.0, 60_000, 1.2, "liability"),
"data_leak": (0.15, 900_000, 1.0, "privacy"),
"ip_claim": (0.05, 1_500_000, 0.8, "ip"),
"regulatory_fine": (0.03, 2_000_000, 0.9, "fine"),
}
RETENTION = 100_000 # per-event deductible on the insurance policy
POLICY_LIMIT = 5_000_000 # annual aggregate limit
COVERED = {"liability", "privacy"} # what the policy wording actually covers
VENDOR_CAP = 400_000 # indemnity cap: 12 months of fees
VENDOR_KINDS = {"ip"} # vendor indemnifies output IP claims only
def poisson(lam, rng):
k, p, L = 0, 1.0, math.exp(-lam)
while True:
p *= rng.random()
if p <= L:
return k
k += 1
def simulate_year(rng):
retained = insured = vendor = 0.0; by = {}
vendor_left, limit_left = VENDOR_CAP, POLICY_LIMIT
for name, (lam, median, sigma, kind) in SCENARIOS.items():
for _ in range(poisson(lam, rng)):
loss = median * math.exp(sigma * rng.gauss(0, 1))
if kind in VENDOR_KINDS:
paid = min(loss, vendor_left); vendor_left -= paid
vendor += paid; loss -= paid
if kind in COVERED and loss > RETENTION:
paid = min(loss - RETENTION, limit_left); limit_left -= paid
insured += paid; loss -= paid
retained += loss; by[kind] = by.get(kind, 0) + loss
return retained, insured, vendor, by
def pct(xs, q):
xs = sorted(xs); return xs[int(q * (len(xs) - 1))]
rng = random.Random(7)
years = [simulate_year(rng) for _ in range(200_000)]
for i, label in enumerate(["retained", "insured", "vendor"]):
col = [y[i] for y in years]
print(f"{label:9s} mean={sum(col)/len(col):>10,.0f} p99={pct(col, .99):>12,.0f}")
gross=[y[0]+y[1]+y[2] for y in years]
print(f"gross mean={sum(gross)/len(gross):>10,.0f} p99={pct(gross,.99):>12,.0f}")
tail = sorted(years, key=lambda y: y[0])[-2000:] # worst 1% of retained years
for k in ["liability","privacy","ip","fine"]:
print(k, f"{sum(y[3].get(k,0) for y in years)/len(years):,.0f}", f"tail share {sum(y[3].get(k,0) for y in tail)/sum(y[0] for y in tail):.2f}")
Worked example: an account-guidance agent
The hypothetical feature is a customer-facing LLM agent that gives account guidance. Its program is a policy with a $100,000 per-event retention and a $5M aggregate covering liability and privacy, plus a vendor IP indemnity capped at $400,000. 200,000 simulated years give:
| Layer | Mean per year | 99th percentile year |
|---|---|---|
| Gross loss | $666,313 | $6,687,670 |
| Paid by insurance | $312,073 | (not additive) |
| Paid by vendor | $19,393 | $400,000 (cap exhausted) |
| Retained | $334,847 | $4,937,459 |
Two readings matter. On average the program transfers about half the loss, which sounds healthy. But the 99th-percentile retained year is about $4.9M, while the 99th-percentile gross year is about $6.7M, so the tail is mostly still yours. The breakdown of the worst 1% of retained years shows why. 50% of the retained tail is regulatory fines, which no layer touches. 29% is IP claims above the vendor's cap. 20% is privacy, where retentions and aggregate erosion bite. Only 2% is the frequent wrong-advice claims. Those are the largest single share of the mean retained loss through deductibles ($121k of it) and barely register in the tail.
So the program is tuned for the wrong risk. Buying a higher insurance limit changes almost nothing, because the tail is driven by uninsurable fines and an uncovered category. The useful moves are a control that cuts fine frequency (reduce, not transfer), negotiating the IP indemnity cap up or buying cover that names IP, and checking whether the real wording gives privacy losses the full aggregate or a smaller sublimit, which this model does not include. Raising the retention to $250,000 would move the mean retained loss from $334,847 to $413,798 and the mean insured loss from $312,073 to $233,122. That is worth it only if the premium falls by more than about $79,000.
Deciding each layer
Decide each layer with two numbers. The first is price versus expected recovery: a premium divided by the mean insured loss is the loading you pay for volatility. Your broker can tell you what loading is typical for the line you are buying. Paying 5x for a layer that rarely triggers is buying comfort, not capacity. The second is the change in the retained tail: does the layer bring your 99th-percentile retained year inside the tolerance set in your AI risk appetite? A layer that improves the mean but leaves the tail unchanged usually is not worth buying. A layer that cuts a tail you cannot survive is worth a high loading.
Operating the programme
- Map every register scenario to the layer that responds and its exact clause. Any scenario without a clause is retained.
- Keep the evidence an insurer will ask for: prompt and output logs, model versions, evaluation results, and the control state on the loss date.
- Track notice deadlines. Read the policy's notice clause: some require prompt notice of a circumstance, not just a claim, and late notice can cost you cover.
- Monitor aggregate erosion during the year. After a large claim, the rest of the year has less cover than the dashboard implies.
- Re-run the model at renewal, after a model or vendor change, and after any new exclusion appears in the wording.
Failure modes
| Failure mode | How it shows up |
|---|---|
| Treating transfer as mitigation | Controls deferred because "insurance covers it"; frequency rises, renewals fail |
| Unread exclusions | A generative-AI exclusion turns a covered breach into a retained one |
| Indemnity conditions breached | Disabling the vendor's content filter voids the IP indemnity you modelled |
| Counterparty failure | A small vendor's indemnity is worth its balance sheet, not its cap |
| Flow-down mismatch | Your customer cap is higher than the vendor cap behind it; you carry the difference |
Trade-offs
Transfer smooths results and protects against ruin, at a price above expected loss, plus disclosure work, plus constraints such as conditions you must keep meeting. Retention is cheaper on average and keeps the incentive to invest in controls, but it exposes you to the tail. For high-frequency, low-severity AI errors, retain and reduce. For low-frequency, high-severity events that you can legally insure, transfer. For uninsurable events, reduce or avoid.
Two structural options sit between those poles. A captive is an insurer you own. It formalises retention, builds reserves from premiums you pay yourself, and gives access to reinsurance markets, but it only makes sense at a scale where the setup and regulatory costs are small against the premiums. Vendor concentration is the opposite trap. If three AI features rely on one model provider, a single indemnity cap and a single counterparty stand behind all of them, and the simulation should treat their IP losses as correlated, not independent. Ignoring correlation is a common modelling error in transfer programmes: independent scenarios make the tail look thinner than it is, and the layers bought against that tail look more adequate than they are. When unsure, rerun with correlated draws and keep the worse answer.
What to do next
- List the top AI loss scenarios from your register with frequency and severity ranges.
- For each, write the responding layer and clause: customer cap, vendor indemnity, retention, policy, or none.
- Run the simulation with your numbers; record mean and 99th-percentile retained loss, and the tail share per scenario.
- Compare the retained tail with your appetite; target the scenarios that drive it.
- For uninsurable drivers, fund controls; for insurable ones, price a layer and compute its loading.
- Ask your broker in writing how the current wording treats generative-AI losses, and re-run at every renewal.