Most writing about AI liability asks whether someone can be sued when a model causes harm. This article asks an engineering question instead: which facts about your build and release process decide whether you are the party a liability regime aims at, when that exposure starts and restarts, and what evidence the rules expect you to be able to produce. Those facts are made by engineers, usually without anyone noticing. A team that fine-tunes an open-weights model and ships it under its own brand has changed its legal role. A team that adds one tool to an assistant may have made a change that counts as a new product.
The legal theories, the case law so far and the overall shape of the revised EU Product Liability Directive are covered in AI Liability, in depth. Here they are a short recap. The core of this page is three mechanisms you can build: a role classifier, a release gate that labels changes by their liability effect, and an evidence pipeline that makes disclosure a query rather than an archaeology project. Nothing here is legal advice; the point is to make sure counsel has facts to work with.
Liability rules and conduct rules
It helps to separate two kinds of law. Conduct rules tell you what to do and are enforced by regulators with fines: the EU AI Act, data protection authorities enforcing the GDPR, a state attorney general. Liability rules decide who pays an injured person, and are enforced by that person in court. The two connect. Breaking a conduct rule is often the easiest way for a claimant to prove the liability case. Under the revised Product Liability Directive, a product is presumed defective when the claimant shows it does not comply with mandatory safety requirements in Union or national law that are meant to protect against the damage suffered. So an AI Act obligation you skipped is not only a fine risk. It can become the claimant's evidence.
So map conduct obligations to the harms they protect against, keep a record showing you met each rule at the time, and design evidence for liability periods that outlast normal log retention by an order of magnitude.
The instruments as engineering inputs
The instruments below are the ones most likely to matter to a team shipping language-model products in 2026. Dates were checked in October 2026. EU AI Act application dates for high-risk systems have been, or are being, changed by the Commission's Digital Omnibus package, so check the current dates before relying on them.
| Instrument | Kind | Who acts | Engineering trigger |
|---|---|---|---|
| EU Product Liability Directive (EU) 2024/2853 | Strict liability for defective products, software included | Injured persons, in national courts | Placing on the market, substantial modification, updates under your control. Transposition deadline 9 December 2026 |
| EU AI Act | Conduct rules and fines | Market surveillance authorities | Role (provider, deployer, importer) and risk class. A breach can feed the PLD defect presumption |
| GDPR Article 82 | Compensation for damage from infringement | Data subjects | Processing of personal data in prompts, logs and training sets |
| California SB 243 | Companion chatbot duties | Users, via a private right of action: actual damages or $1,000 per violation | Operating a companion chatbot. In force since 1 January 2026 |
| Colorado SB 26-189 | Duties for automated decision-making technology | Read the statute's enforcement provisions | Materially influencing consequential decisions. Effective 1 January 2027 |
| National negligence law | Fault-based liability | Anyone harmed | Everything: what a reasonable provider would have done |
The AI Liability Directive is absent from the table on purpose. The Commission withdrew it, so fault-based AI claims in the EU fall back to national tort law, with the Product Liability Directive as the harmonised strict-liability route. For the obligation-register approach to tracking all of this, see the AI regulation deep dive.
Who becomes the manufacturer
Liability regimes attach to roles, and engineering choices pick your role. The revised directive lists the economic operators who can be liable. The manufacturer comes first. So does the manufacturer of a defective component integrated into the product, and anyone who presents themselves as the manufacturer by putting their name or trademark on the product. Then come the importer, the authorised representative and, where none of those is established in the Union, the fulfilment service provider. A distributor becomes liable when it cannot name an EU operator or its own supplier within one month of a request. The rule that matters most for AI teams is this one: any person who substantially modifies a product outside the original manufacturer's control, and then makes it available or puts it into service, is treated as a manufacturer of that product.
The AI Act has a parallel rule. A distributor, importer or deployer that puts its name or trademark on a high-risk system, substantially modifies it, or changes its intended purpose so that it becomes high-risk takes on the provider's obligations. The two regimes use different words, but in both, fine-tuning, re-branding and re-purposing are the acts that move you up the chain. The role classifier below is a starting point to encode with counsel. Run it whenever a product's supply chain changes.
from dataclasses import dataclass
@dataclass
class Facts:
sells_under_own_brand: bool # name or trademark on the product
integrates_third_party_model: bool # model is a component you did not build
modified_weights: bool # fine-tune, adapter merge, distillation
changed_intended_purpose: bool # new use beyond the supplier's stated purpose
upstream_authorised_change: bool # supplier performed or consented to the change
established_in_eu: bool
places_on_eu_market: bool
def liability_roles(f: Facts) -> set[str]:
roles = set()
if f.sells_under_own_brand:
roles.add("manufacturer (own-brand)")
if f.integrates_third_party_model:
roles.add("integrator: supplier is a component manufacturer")
substantial = f.modified_weights or f.changed_intended_purpose
if substantial and not f.upstream_authorised_change:
roles.add("manufacturer (substantial modification) -- confirm with counsel")
if f.places_on_eu_market and not f.established_in_eu:
roles.add("importer or authorised representative needed")
return rolesTwo fields cause most mistakes. The first is upstream_authorised_change. A change the original manufacturer performs, authorises or consents to stays within that manufacturer's control, so contracts and model licences that grant or refuse consent to modification change the answer. That is one reason AI contracts matter to engineers. The second is changed_intended_purpose. Turning a general chat model into a medication-dosing assistant is a purpose change even if no weight moved.
Releases that restart exposure
Exposure under the revised directive is anchored to events: the product is placed on the market, it is substantially modified, it receives updates. Three rules make releases matter. Products placed on the market before 9 December 2026 stay under the old regime unless they are substantially modified afterwards. A manufacturer cannot use the 'defect did not exist when it left my hands' defence for defects caused by software updates or upgrades within its control, or by failing to supply updates needed to keep the product safe. And for a substantially modified product, the long-stop period runs from when the modified product was made available, not from the original launch. Every release is therefore a liability event, and some releases restart the clock.
The directive defines substantial modification mainly by reference to product safety rules. Where those rules are silent, it means a change that alters the product's original performance, purpose or type, was not foreseen in the initial risk assessment, and creates a new hazard or raises the level of risk. That maps well onto a three-class release gate:
| Class | Examples | Gate action |
|---|---|---|
| A: no new risk | Prompt wording fix, latency work, a bug fix that leaves behaviour the same | Standard tests. Manifest recorded |
| B: risk change within assessed scope | New model version from the same supplier, retrieval corpus refresh, a tighter guardrail | Re-run the safety eval suite against thresholds. Product owner signs off |
| C: candidate substantial modification | New tool with side effects, new user group (for example minors), new intended purpose, a fine-tune on new data | Fresh risk assessment, legal review, treat as a new placing on the market |
SIDE_EFFECT_TOOLS = {"delete_file", "send_payment", "send_email", "change_setting"}
def classify_release(diff: dict) -> str:
if diff.get("new_intended_purpose") or diff.get("new_user_population"):
return "C"
if set(diff.get("tools_added", [])) & SIDE_EFFECT_TOOLS:
return "C"
if diff.get("weights_changed") and diff.get("training_data_new_domain"):
return "C"
if diff.get("weights_changed") or diff.get("corpus_changed") or diff.get("guardrail_changed"):
return "B"
return "A"Err towards C: a false C costs a meeting, a false A can cost a defence and a clock.
Architecture: the evidence pipeline
The diagram shows the evidence pipeline that the release gate feeds. Every release produces an immutable manifest. It holds content hashes of weights, system prompt, tool schemas and guardrail configuration, the eval results with their thresholds, the release class, and who signed off. Every production call logs the manifest identifier it ran under, so any individual output can be tied to the exact configuration and its test record. Manifests and decision records go to a write-once vault. Retention follows the product lifetime: ten years after the product was placed on the market, and up to 25 years where latent personal injury is plausible.
Storing full prompts and outputs for 25 years conflicts with data minimisation, so split the record. Keep the manifest, the decision metadata (timestamp, manifest id, tool calls, guardrail verdicts, a hash of the output) and the evaluation evidence for the long period. Keep raw personal content only as long as the privacy basis allows. A hash proves an output existed without keeping its content.
Disclosure readiness
Courts applying the revised directive can order a defendant to disclose relevant evidence once the claimant has shown a plausible claim. The order must be necessary and proportionate, and trade secrets are protected. Failing to disclose lets the court presume defectiveness. Where technical or scientific complexity makes proof excessively difficult, defect or causation can be presumed on a showing of likelihood. For opaque models that is the realistic route for most claimants. The engineering response is to make a scoped, reviewable disclosure cheap to produce, so that the choice is never between dumping everything and producing nothing.
def build_disclosure(incident_id, vault, window, scope):
"""Assemble a scoped evidence package for counsel to review. Never auto-send."""
calls = vault.decisions(incident_id=incident_id, between=window)
manifests = {m: vault.manifest(m) for m in {c.manifest_id for c in calls}}
package = {
"decisions": [c.metadata() for c in calls], # no raw content by default
"manifests": {k: v.public_view() for k, v in manifests.items()},
"evals": {k: vault.eval_report(k) for k in manifests},
"release_history": vault.releases(between=window),
"known_issues": vault.issues(touching=list(manifests), between=window),
}
if scope.get("include_content"):
package["content"] = [vault.content(c.id) for c in calls if vault.content_retained(c.id)]
package["redactions"] = redact_secrets_and_third_party_pii(package)
return seal(package) # hash + timestamp so later edits are detectableThe known-issues register cuts both ways: an issue triaged and mitigated on a recorded date helps you, a report left open for months does not.
Worked example: the cleanup tool
This example is hypothetical. A company sells a home-assistant app in the EU. It fine-tunes an open-weights model on its own support transcripts and ships under its own brand, so the role classifier returns 'manufacturer (own-brand)' and, unless the model licence counts as the supplier's consent to modification, 'manufacturer (substantial modification)' too. Either way, the weight supplier cannot carry the product's liability for it. If the weights were released free and open source outside a commercial activity, the directive excludes that release altogether.
Version 3.2 adds a cleanup_storage tool that deletes files on the user's connected drive. The release gate marks it class C because the tool has side effects. A user asks the assistant to 'tidy up old stuff'. It deletes a personal photo library. The revised directive covers destruction or corruption of data not used for professional purposes, so this is compensable damage, not just bad service.
- The legal hold freezes deletion for the user's decision records and the 3.2 manifest.
- The disclosure builder pulls the call: manifest 3.2-r4, the tool call with its argument glob, and the guardrail verdict 'confirmation not required'. That last field is the problem.
- The 3.2 eval report shows destructive-tool tests ran with explicit file lists, never with vague instructions. The gap is documented, and it was within the manufacturer's control.
- The decision record is decisive whichever way it points. Here it shows the defect.
The fix is engineering, not legal. Every destructive tool now requires a confirmation that lists the exact items. Deletions go to a 30-day recycle bin, which turns data destruction into recoverable inconvenience. The eval suite gains ambiguous-instruction cases, all recorded in the 3.3 manifest.
Failure modes
- Role drift. A partnership re-brands a supplier's assistant, and nobody re-runs the role classifier.
- Silent class A. A tool or corpus change ships as a config push outside the gate, so it has no manifest.
- Log retention set by cost. 30-day logs against a ten-year claim window. Nothing remains to disclose, and silence feeds the presumption.
- Over-retention. Raw prompts kept for 25 years 'for liability' break data minimisation and create breach exposure.
- Unsealed evidence. Records that could have been edited after the incident carry little weight.
- Known-issue rot. A register full of untriaged reports is evidence against you.
- Update neglect. A deprecated model left serving without security fixes, when updates were within your control.
Trade-offs
| Choice | Cheaper option | Safer option | Guidance |
|---|---|---|---|
| Release gate | Two classes (ship or review) | Three classes with automatic C rules | Use three. Class C is where clocks restart |
| Evidence content | Full prompts and outputs | Metadata plus hashes | Keep metadata long, keep content only as long as privacy allows |
| Fine-tune or prompt | Fine-tune for quality | Prompt and retrieve on the supplier's model | Fine-tuning moves you up the chain, so price that in |
| Disclosure | Ad hoc exports | Builder with counsel review | Build it before the first claim |
What to do next
- Run the role classifier with counsel for every AI product and record the result with its date.
- Add a release class (A, B or C) to every deployment, and block deployments that lack one.
- Write side-effect tools, new user groups and purpose changes into automatic class C rules.
- Emit a signed manifest per release and log its id on every production call.
- Set evidence retention by product lifetime, split from raw-content retention.
- Build and dry-run the disclosure builder against a mock incident this quarter.
- Map each AI Act and sector obligation to the harm it protects against, and close the gaps that carry liability weight first. Audit the evidence once a year.