When an AI system's output hurts someone, somebody pays. Liability is the body of law that decides who, and for engineers it is less abstract than it sounds: the outcome of most disputes turns on what the system was designed to do, what its builders knew, what they tested, and what they can prove from their records. Those are engineering artefacts. This article explains the legal theories that apply to AI systems, the cases that show how courts and tribunals have treated them so far, the EU and US frameworks that are changing in 2026 and 2027, and the design and contract choices that reduce exposure.

It is written for teams that build or deploy AI products, not for lawyers, and it is not legal advice. Rules differ by country and sector, and several described here are recent or not yet in force; involve counsel before relying on any of it.

The legal theories

No single "AI liability law" exists in most places. Claims arrive through existing doctrines, and the same incident can support several at once:

TheoryWhat the claimant must usually showTypical AI example
Contract and warrantyA promise in the contract was brokenA vendor's accuracy or uptime commitment is not met
NegligenceA duty of care, a breach of the expected standard, causation, damageDeploying a model in a high-stakes flow without testing or oversight
Negligent misrepresentationA false statement relied on reasonably, causing lossA chatbot states a wrong policy and a customer acts on it
Product liabilityA defective product caused damage; fault not required in strict regimesSoftware in a device or service behaves unsafely
Anti-discrimination lawDisparate treatment or impact on a protected groupA screening model rejects older applicants at higher rates
Professional liabilityA professional fell below the standard of their professionA lawyer files a brief with fabricated citations
Data protectionUnlawful processing or a breach of data rightsPersonal data in prompts retained or exposed

Intellectual-property claims over training data and outputs form a separate, fast-moving area and are out of scope here. Note that the AI Act itself is not a liability regime: it sets obligations and fines, but compliance with it, or its absence, becomes strong evidence of the standard of care in negligence and product claims.

What the cases show

Four decisions show the pattern so far. In Moffatt v. Air Canada (British Columbia Civil Resolution Tribunal, 2024), the airline's website chatbot told a customer they could claim a bereavement fare retroactively, contradicting the airline's policy page. The airline argued that the chatbot was responsible for its own statements. The tribunal rejected that, found negligent misrepresentation, and ordered the airline to pay the difference, roughly C$650 plus interest and fees. The amount was small; the principle, that a business answers for what its chatbot says, was not.

In Mata v. Avianca (S.D.N.Y., 2023), lawyers filed a brief citing cases invented by a chatbot and were sanctioned $5,000. Liability sat with the professionals who signed the filing, not the tool. In Mobley v. Workday (N.D. Cal.), the court in 2024 allowed discrimination claims against the vendor of a screening tool to proceed on the theory that the vendor acted as the employers' agent, and in 2025 conditionally certified a nationwide age-discrimination collective. And in Garcia v. Character Technologies (M.D. Fla.), a May 2025 ruling let most claims over a teenager's death proceed, including product liability claims, and declined at that early stage to treat the chatbot's output as protected speech. That was a ruling on a motion to dismiss, not a finding of liability.

Read together: deployers cannot disclaim their own system's statements, professionals keep their duties when they use tools, vendors can be pulled in when their tool performs a function the law regulates, and courts are willing to analyse conversational AI as a product.

The EU: software becomes a product

The EU's revised Product Liability Directive, (EU) 2024/2853, is the most consequential change. Member states must transpose it by 9 December 2026, and it will apply to products placed on the market after that date. Its key moves for AI teams:

  • Software, including AI systems, is a product, whether embedded or supplied on its own. Free and open-source software supplied outside a commercial activity is excluded.
  • Liability is strict: a claimant shows defect, damage and causation, not fault. Defectiveness takes account of a product's ability to keep learning after deployment.
  • Manufacturers stay liable for defects introduced by updates they control, and for failing to supply security updates needed to keep the product safe.
  • Courts can order defendants to disclose relevant evidence, and defect or causation can be presumed when a defendant does not disclose, when a product breaks mandatory safety rules, or when technical complexity makes proof excessively difficult for the claimant.
  • Compensable damage includes medically recognised psychological harm and the destruction or corruption of data not used for professional purposes.

The separate AI Liability Directive, proposed in 2022 to ease fault-based claims, is gone: the Commission announced its withdrawal in its February 2025 work programme, and the withdrawal notice was published in the Official Journal in October 2025. Fault-based claims therefore stay under national law. For the AI Act's obligations and dates, see EU AI Act, in depth.

The United States

The United States has no federal AI liability statute. Claims run through state tort and contract law, consumer-protection enforcement and anti-discrimination statutes. Whether Section 230's protection for hosting third-party content covers text a model generates itself is unsettled. At state level, Colorado replaced its 2024 AI Act with a narrower law, SB 26-189, signed in May 2026 and effective 1 January 2027, which regulates automated decision-making technology that materially influences consequential decisions. Other states regulate specific uses, such as hiring tools and chatbots that talk to minors. Track the jurisdictions where your users are, not only where you are; AI Regulation Deep Dive shows one way to keep a dated obligation register.

Allocating liability along the supply chain

Where liability lands in an AI supply chain, and the levers each party holdsModel providerweights, API, usage policyIntegratorprompts, RAG, tools, UIDeployerputs it in front of peopleAffected personcustomer, applicantoutputClaims flow back from the harmed personcontract, tort, product, discriminationContracts reallocate between businesseswarranties, indemnities, caps, flow-down termsEvidence decides most disputesversioned decision records, logs, test results, documentation, update historyThe deployer usually faces the claim first; contracts decide whether the cost moves upstream.
Liability in a typical AI supply chain: claims start at the deployer; contracts and evidence decide where the cost ends up.

In practice the harmed person sues whoever they dealt with, usually the deployer. The deployer then looks upstream. Whether the cost moves depends on contracts and on facts: did the deployer use the model within its documented limits, did the integrator's prompt or retrieval cause the failure, did the provider's update change behaviour without notice. Large model providers commonly cap their liability and disclaim fitness for particular purposes, and several offer copyright indemnities with conditions. Read what your provider actually offers rather than assuming; most of the risk of wrong outputs is usually left with the customer.

Contract levers worth negotiating, in rough order of value: notice before model changes and a pinned version you can stay on; incident notification within a fixed time; access to the documentation you need for your own obligations; an indemnity for the provider's own defects; and caps that scale with the harm your use case can cause rather than with fees paid.

Engineering controls that reduce exposure

Most liability for AI outputs is created by a small number of design choices. The Air Canada pattern, a system making commitments the business never authorised, is the most common and the easiest to engineer against. Route anything that sounds like a commitment through an authoritative source:

import re

COMMITMENT = re.compile(
    r"\b(refund|reimburse|compensat\w*|guarantee\w*|waive\w*|discount|"
    r"you (are|will be) (eligible|entitled)|we will (pay|cover|honou?r))\b", re.I)

def guard_commitments(draft, retrieved_policy_ids, policy_store):
    """Block or rewrite drafts that commit the business without a cited policy."""
    if not COMMITMENT.search(draft):
        return draft, "pass"
    if not retrieved_policy_ids:
        return ("I can't confirm that here. Here is our policy page, or I can connect "
                "you with an agent who can.", "blocked_no_policy")
    cited = [policy_store[p]["url"] for p in retrieved_policy_ids]
    return draft + "\n\nPolicy: " + ", ".join(cited), "pass_with_citation"

A regular expression is a crude first filter, and that is fine: it errs toward blocking, every block is logged, and the logs tell you where to add retrieval or a better classifier. Other controls with high value per hour of work:

  • Scope limits: refuse topics outside the documented purpose, and say so plainly.
  • Human approval for consequential decisions: credit, hiring, medical or account termination decisions get a reviewer with authority and time; see human-in-the-loop for high-risk actions.
  • Accurate disclosure: tell users they are talking to AI and what it cannot do. Disclaimers do not cancel duties, as Air Canada learned, but clear instructions matter in product cases.
  • Change control: treat a prompt, model or retrieval change as a release with tests, because under the new product rules a harmful update is your defect.

Evidence: the decision record

When a dispute arrives, the question is what happened in one specific interaction months ago and whether your process was reasonable. Under the revised directive, a failure to disclose can itself create a presumption against you. Keep a decision record per consequential output:

{
  "record_id": "dec-2026-10-05-7f3a",
  "timestamp": "2026-10-05T09:14:22Z",
  "system": "returns-assistant",
  "versions": {"model": "provider-model@2026-08-14", "prompt_bundle": "sha256:9c1e...",
               "retrieval_index": "policies@2026-10-01"},
  "inputs_ref": "logstore://conv/88421#turn-12",
  "retrieved": ["policy-returns-v7"],
  "output_ref": "logstore://conv/88421#turn-13",
  "guards": {"commitment_guard": "pass_with_citation"},
  "human_review": null,
  "retention_until": "2033-10-05"
}

Store references to content rather than copies, so data-protection retention and legal retention can be managed separately; the design of tamper-evident trails is covered in LLM audit logging architecture. Keep test results, model documentation and the update history with dates: they are what shows the standard of care was met. Under the revised directive, claims generally expire ten years after the product was placed on the market, 25 for latent personal injury, so set retention by product lifetime, not by log volume.

Worked example: pricing a wrong-commitment risk

Take a retailer's returns assistant handling 200,000 conversations a month. Sampling shows 0.3% of answers state a refund term that contradicts policy: about 600 wrong commitments a month. Suppose 5% of those customers act on the answer and pursue it, and each claim costs the retailer about $120 once goods, handling and support time are counted. That is 30 claims and roughly $3,600 a month, before any regulator or class claim, which would multiply the figure.

After adding retrieval of the policy for every returns question and the commitment guard, sampling shows the wrong-commitment rate at 0.02%: 40 a month, 2 claims, about $240. The guard blocks about 1% of conversations, which go to human agents at, say, $4 each, or $8,000 a month. On these numbers the guard costs more than the claims it prevents, and the right decision is to improve retrieval until the block rate falls, not to remove the guard. The exposure that matters is not the average claim but the tail: a single wrong commitment to thousands of customers through a viral screenshot, or a regulator treating systematic misstatements as an unfair practice. Price that tail explicitly; it is why the guard stays.

Failure modes

FailureHow it creates liabilityMitigation
"The AI said it, not us"Tribunals hold the deployer to its system's statementsOwn the output; guard commitments
Silent provider updatesBehaviour changes; you cannot show what ran whenPin versions; contract for change notice
Logs too short or too thinCannot rebut a claim; disclosure gaps raise presumptionsDecision records with versions and retention
Rubber-stamp human reviewOversight exists on paper onlyMeasure override rates and review time
Use outside documented purposeBreaks the provider's terms and your own risk caseScope limits and usage monitoring
No security updatesUnpatched product becomes defective under the new directivePatch policy for the product's lifetime

What to do next

  1. List every place your AI output can commit the business or decide something about a person, and rank them by worst-case harm.
  2. Add a commitment guard or scope limit to the top three and log every block.
  3. Introduce decision records with version pins and set retention by product lifetime.
  4. Read your model provider's terms for caps, indemnities and change notice, and list the gaps for your next negotiation.
  5. If you sell into the EU, map which of your products will be placed on the market after 9 December 2026 and review update and patch policies for them with counsel.
  6. Document model limits for your users and staff, starting from Model Cards, in depth.
Key takeaway: AI liability mostly flows through existing law: misrepresentation, negligence, discrimination and, in the EU from December 2026, strict product liability for software. Deployers answer for what their systems say, so guard commitments, keep humans on consequential decisions, pin versions, and keep decision records long enough to prove what happened. Then use contracts to move the risk that belongs upstream back to the parties who control it.