Every company that ships an AI feature has taken on a new kind of risk, and very few have checked whether their insurance actually responds to it. A support agent that invents a refund rule, a retrieval tool that leaks another customer's records after a prompt injection, a marketing image that copies a photographer's work, a scoring model that disadvantages a protected group: each of these can produce a claim, and each falls awkwardly between the policies a technology company normally buys. This article is about insurance for the operators of AI systems. AI used by insurers for claims and underwriting is a different subject, covered in AI in Insurance.
The goal here is engineering, not legal advice. You will learn how AI losses map onto existing policy lines, what changed in the market in 2025 and 2026, how to build a loss-scenario register that exposes coverage gaps in code, what evidence underwriters want and how your logs and evaluations produce it, and how to handle a claim without voiding cover. Read your actual policy wording with a broker and counsel before relying on any of it; wording varies by carrier, state and renewal date.
Why AI losses fall between policies
Insurance is sold by line, and each line was written with a particular cause of loss in mind. AI failures do not respect those boundaries. Consider who is harmed and how.
| Loss | Who pays first | Usual line | Why it may not respond |
|---|---|---|---|
| Customer relies on a wrong answer | You, to the customer | Tech errors and omissions (E and O) | Wording may require a 'professional service' by a person, or exclude contractual liability |
| Data leaked through an agent tool | You, to data subjects and regulators | Cyber | Usually responds, but fines and some regulatory costs are often excluded or uninsurable |
| Model provider outage | You, in lost revenue | Cyber business interruption | Dependent BI may only cover named or IT providers; waiting periods apply |
| Generated content infringes IP | You, to the rights holder | Media liability | Many forms cover your own published content, not machine-generated output |
| Biased automated decision | You, to applicants or employees | Employment practices or general liability | Generative AI exclusions or 'professional services' carve-outs |
| Physical harm from an AI-driven device | You, to the injured party | General and product liability | The 2026 ISO generative AI endorsements can remove it |
The market calls the resulting uncertainty silent AI: a policy that neither grants nor excludes AI-caused loss, so nobody knows whether it pays until a claim is litigated. The term deliberately echoes silent cyber, the same ambiguity a decade ago when property and liability policies were silent about cyber attacks. Insurers resolved silent cyber by writing explicit exclusions into old lines and selling affirmative cyber policies. The same two moves are now under way for AI.
What changed in the market
Exclusions. On 1 January 2026, Verisk's ISO program made three optional generative AI exclusion endorsements available for commercial general liability: CG 40 47 excludes bodily injury, property damage and personal and advertising injury arising out of generative AI under Coverages A and B; CG 40 48 excludes only personal and advertising injury under Coverage B; and CG 35 08 applies to the products and completed operations coverage part. They are optional, so whether your general liability policy carries one depends on your carrier, your state and your renewal date. The practical consequence is that you must read the endorsement schedule at every renewal, not only the base form.
Affirmative cover. On 30 April 2025, Armilla announced an AI liability policy underwritten by Lloyd's underwriters including Chaucer, a standalone third-party liability product that names AI failures such as hallucinations, model drift and mechanical underperformance as covered events. Earlier products took a warranty approach instead: Munich Re's aiSure, for example, insures a model's performance against an agreed metric, so the payout is triggered by measured underperformance rather than by a lawsuit. Treat both approaches as specialist products: limits, retentions and exclusions differ considerably, and their wording deserves the same review as any other contract.
Vendor indemnities. In 2023 several large model providers, including Microsoft, Google and OpenAI, announced copyright indemnities for some paid offerings, typically conditioned on using the provider's safety filters and not deliberately prompting for infringing output. These are contract terms, not insurance, and they sit inside the provider's limitation-of-liability clause unless stated otherwise. They move part of the IP scenario off your balance sheet, but only for the products and conditions named.
A loss-scenario register in code
The most useful artifact you can build is a loss-scenario register: a short list of concrete failures, each with an estimated frequency and severity, mapped to the policy or contract that is supposed to pay. Keep it in version control next to your risk register so it changes when the system changes. The script below flags three kinds of gap: no responding line, an exclusion on the responding line, and a severity that exceeds the limit available after the retention.
from dataclasses import dataclass, field
@dataclass
class Policy:
name: str
limit: float # aggregate limit available to this scenario
retention: float # deductible or self-insured retention
exclusions: set = field(default_factory=set)
@dataclass
class Scenario:
sid: str
description: str
per_year: float # expected events per year
severity: float # plausible bad-case cost of one event
lines: list # policy names expected to respond
tags: set # e.g. {"genai", "contractual", "fines"}
def assess(scenarios, policies):
rows = []
for s in scenarios:
responding = [policies[n] for n in s.lines if n in policies]
live = [p for p in responding if not (p.exclusions & s.tags)]
if not responding:
gap = "NO LINE"
elif not live:
gap = "EXCLUDED"
elif s.severity > max(p.limit for p in live) + min(p.retention for p in live):
gap = "OVER LIMIT"
else:
gap = "ok"
retained = min((p.retention for p in live), default=s.severity)
rows.append((s.sid, gap, round(s.per_year * min(retained, s.severity))))
return rowsThe third column is the expected cost you carry yourself each year: frequency times the part of each event that falls inside the retention, or the whole event if nothing responds. Summed across scenarios, it is the number to put beside the premium when you decide whether to raise or lower retentions.
Worked example: an LLM support agent
Take a mid-sized software company with an LLM support agent that answers billing questions, can issue refunds up to a cap through a tool, and drafts help-centre images. It holds cyber with a 5 million limit and a 250,000 retention, tech E and O with a 3 million limit and a 100,000 retention whose wording excludes liability assumed under contract, and general liability that renewed with the CG 40 47 endorsement. It has no media liability policy. The register, with deliberately rough numbers, looks like this.
| ID | Scenario | Freq / yr | Severity | Expected line | Result |
|---|---|---|---|---|---|
| S1 | Agent states a refund rule that does not exist; customers demand it | 2 | 60,000 | Tech E and O | ok, fully retained: 120,000 per year |
| S2 | Injected document makes the agent reveal other customers' invoices | 0.1 | 1,500,000 | Cyber | ok, 25,000 per year retained |
| S3 | Generated help image copies a stock photo | 0.3 | 80,000 | Media | NO LINE: 24,000 per year |
| S4 | Refund tool bug pays out above the cap | 0.2 | 400,000 | Tech E and O | check: first-party loss, likely not covered |
| S5 | Agent advice leads to a customer's property damage claim | 0.02 | 800,000 | General liability | EXCLUDED by CG 40 47 |
Three findings fall out immediately. S1 is a frequency problem, not an insurance problem: every event sits under the E and O retention, so the fix is an output guard that checks refund statements against the policy table before they are sent, the lesson of the 2024 Moffatt v. Air Canada decision, where a tribunal held the airline liable for a bereavement-fare rule its chatbot had misstated. S3 needs either a vendor indemnity, a media policy that affirmatively names generated content, or a rule that generated images pass a similarity check. S4 is your own money leaving through your own tool, which liability policies do not cover at all; crime or a first-party AI performance product might, so ask. S5 is the case the new exclusions were written for, and it is exactly the scenario an affirmative AI liability policy exists to pick up.
The evidence underwriters want
Underwriters price what they can see. A submission that says 'we use GPT-class models responsibly' gets a conservative quote or a declination; a submission with evidence gets a real one. The good news is that the evidence is the same artifacts a sound security and governance program already produces.
| Underwriter question | Evidence that answers it |
|---|---|
| What decisions does the system make, and how autonomously? | Use-case inventory with autonomy level and human-review points |
| How often is it wrong, and how do you know? | Versioned evaluation results on a fixed test set, with error rates per task |
| Can you reconstruct what happened in a claim? | Tamper-evident logs of prompts, retrieved context, tool calls and outputs |
| How fast do you detect and contain failures? | Incident runbooks, detection alerts, mean time to containment from drills |
| Who else is liable? | Vendor contracts: indemnities, liability caps, data-use terms |
| Is the risk changing? | Change log of model and prompt versions with re-evaluation before release |
The logging row deserves emphasis because it is the one most teams cannot produce after the fact. A claim may arrive months after the conversation; if you cannot show the exact model version, system prompt and retrieved documents, you cannot defend it and the insurer cannot assess it. The design for that trail is in LLM audit logging architecture, and mapping your controls to a recognised framework such as the NIST AI RMF gives the underwriter a vocabulary they already use.
Claims mechanics that decide coverage
Most coverage disputes are lost on procedure, not on substance. Four mechanics matter for AI incidents.
- Claims-made triggers. Cyber and E and O are usually claims-made: the policy in force when the claim is first made responds, provided the incident happened after the retroactive date. Switching carriers without preserving the retroactive date can leave every earlier conversation uninsured.
- Notice of circumstances. If you discover the agent has been quoting a wrong policy for six weeks, you may be required to notify the insurer before any customer complains. Late notice is a standard defence.
- Consent to settle. Paying customers goodwill credits to make a problem go away can be treated as a voluntary payment the insurer never approved. Agree a protocol with your broker in advance.
- Evidence preservation. Freeze logs, prompts, model versions and retrieval indexes for the affected window before anyone 'fixes' the prompt. The same step belongs in your incident playbook; see LLM incident response.
Failure modes
- Reading the base form only. The generative AI exclusion arrives as an endorsement at renewal, and the schedule of forms is where it hides.
- Assuming the vendor's indemnity is unlimited. Indemnities typically sit under the provider's overall liability cap and carry conditions such as keeping safety filters enabled.
- Insuring frequency losses. Small, frequent errors under the retention are an engineering budget item; buying lower retentions to cover them is expensive.
- Forgetting first-party loss. Liability policies pay third parties. Money your own tool pays out by mistake needs a different product or a hard control.
- Misstating the application. Describing an autonomous agent as 'human reviewed' when review is sampled can give the insurer grounds to rescind.
- No change trigger. Adding a new tool such as payments to an agent changes the risk materially; some policies require disclosure of material changes.
Trade-offs
There is no single right structure, only a set of deliberate choices.
| Choice | Favour it when | Cost |
|---|---|---|
| Higher retention | Losses are frequent and small, and you have strong controls | More volatility in your own budget |
| Affirmative AI policy | AI output can cause third-party harm the old lines may not cover | Extra premium, new wording to negotiate, fewer carriers |
| Performance warranty | You can define a measurable metric and a contract trigger | Needs stable evaluations; payout tied to the metric, not your loss |
| Contractual risk transfer | Vendors offer meaningful indemnities | Capped, conditional, and only as good as the vendor's solvency |
| Engineering control | The loss is driven by a specific, checkable failure | Build and maintenance effort; residual risk remains |
In practice the order of operations is: remove what you can with controls, transfer what you can by contract, insure the severe and unpredictable remainder, and retain the rest on purpose. The legal theories behind the claims themselves are covered in AI Liability.
What to do next
- List every AI feature in production with what it can say, decide or do, and who is harmed if it is wrong.
- Write five to ten concrete loss scenarios with rough frequency and severity, and run them through a register like the one above.
- Pull every current policy and its full schedule of endorsements; search for generative AI exclusions, including the ISO CG 40 47, CG 40 48 and CG 35 08 forms.
- Ask your broker, in writing, which policy responds to each scenario, and record the answer next to the scenario.
- Collect vendor contracts and note each indemnity, its conditions and the liability cap it sits under.
- Close the cheapest gaps with controls first, such as output checks against policy tables and caps on tool actions.
- Assemble an underwriting evidence pack: inventory, evaluation results, logging design, incident drills and change log.
- Add 'notify the insurer?' and 'preserve evidence' as explicit steps in your AI incident runbook.
- Re-run the register whenever an agent gains a new tool, a new model, or a new class of user.