When a diagnostic AI is involved in a missed cancer, a late sepsis call or a misread scan, the first legal question is not whether the model was accurate on average. It is who owed the patient a duty, what a reasonable practitioner would have done with the information available, and whether anyone can reconstruct what that information was. Most organisations deploying diagnostic AI can answer the first question and fail the third.

This article explains how liability for AI-assisted diagnosis is allocated today among the clinician, the hospital and the vendor; why the standard of care creates a lopsided incentive to use AI only when it agrees with you; where the United States device boundary sits after the FDA's January 2026 clinical decision support guidance; and how to engineer a decision record and a review workflow that make good care provable. It closes with a worked example of a chest X-ray triage miss and a checklist. It is engineering guidance, not legal advice; the doctrine varies by jurisdiction, and your counsel should map it to your facts.

Who carries the exposure

Three parties carry most of the exposure, under different theories.

PartyMain theoryWhat the claimant must showWhat helps the defence
ClinicianMedical negligence (malpractice)Duty, breach of the standard of care, causation, damageEvidence of an independent assessment consistent with accepted practice
Hospital or health systemVicarious liability for staff; corporate negligence in selecting, configuring and supervising toolsThe institution chose, configured or monitored the tool unreasonablyDocumented validation, local testing, training and monitoring
VendorProduct liability (defect in design, manufacture or warnings), negligence, contractA defect, or warnings that failed to describe limits; causationClear intended use, labelled limitations, the learned intermediary doctrine

The learned intermediary doctrine matters in the United States: where a product reaches the patient through a professional, the manufacturer's duty to warn is generally owed to the professional. A vendor whose labelling clearly says 'triage aid, not for ruling out' has a stronger position than one whose sales deck promised the opposite. US courts have also been reluctant to treat software as a 'product' for strict liability, which historically pushed claims towards negligence. The European Union is moving the other way: the revised Product Liability Directive treats software, including AI, as a product, with member state transposition due by 9 December 2026. The general picture is covered in AI liability in depth; the rest of this article stays with diagnosis.

The standard of care and the follow-or-reject matrix

Malpractice asks whether the clinician met the standard of care: what a reasonable practitioner in the same specialty would have done. That standard is set by current practice, and current practice in most specialties does not yet include following an algorithm. Price, Gerke and Cohen (JAMA, 2019) worked through the consequences. Consider the four combinations of the AI being right or wrong and the clinician following or rejecting it:

Clinician follows AIClinician rejects AI
AI agrees with standard careLow exposure either wayExposure if the rejection departs from standard care and harms
AI departs from standard careExposure if the AI was wrong: the clinician followed a non-standard recommendationGenerally protected: the clinician followed standard care, even if the AI was right

The asymmetry is the point. Under a custom-based standard, the safe legal move is to treat AI as a second reader that can only confirm, which throws away exactly the cases where AI adds value: the subtle finding a human would miss. As tools become widely adopted, the standard of care may shift so that not consulting a validated tool becomes the breach, but no one can tell you the date. Engineering cannot change the doctrine; it can make sure that whatever the clinician did is documented as an independent, reasoned judgment, which is what the standard rewards in every cell of that table.

Device or decision support

Whether a diagnostic function is a regulated medical device changes who must prove what before launch, and it changes the liability story afterwards. In the US, section 520(o)(1)(E) of the Federal Food, Drug, and Cosmetic Act excludes clinical decision support from the device definition only if all four criteria hold: it does not acquire, process or analyse a medical image, an in vitro diagnostic signal or a pattern from a signal acquisition system; it displays or analyses medical information; it supports a health care professional's recommendation; and it lets that professional independently review the basis, so they do not rely primarily on it.

The FDA replaced its 2022 guidance on these criteria with a revised final guidance on 6 January 2026. As summarised by several law firms, it extends enforcement discretion to tools that give a single recommendation when only one is clinically appropriate, and clarifies what counts as a signal or pattern. Read the guidance itself before relying on any specific example. One thing did not change: imaging AI fails the first criterion, so a chest X-ray classifier is a device, typically cleared through 510(k) with a stated intended use. See FDA AI/ML regulation for pathways and change control plans.

Criterion four is also a liability hinge. A tool designed so the clinician can review its basis (the guideline, the lab values, the evidence it weighed) supports the argument that the clinician exercised independent judgment. A black-box score shifts the reliance story, and with it some exposure, towards the vendor and the institution that deployed it. In the EU, AI that is a medical device or its safety component, and that needs a notified body's conformity assessment under the Medical Device Regulation (in practice class IIa and above), is high-risk under Article 6(1) and Annex I; after the 2026 Omnibus amendment those duties apply from 2 August 2028.

Architecture: the diagnostic decision record

The diagnostic decision record: capture what was shown, not just what was computedinput snapshotstudy ID + content hashmodel registryversion, intended useinference servicescore, calibrated p, abstainpresentation layerwhat the clinician actually sawoutputclinician actionindependent read, accept, overrideshown_atappend-only decision recordhashes, versions, timestamps, reason codesacted_atoutcome linkagefinal diagnosis, 30-day follow-upmonitoringoverride rate, misses by subgroupDefensible care is provable care:the record answers who knew what, when.
Every inference writes one record that links input, model version, what was displayed, what the clinician did and when, and later the outcome. Monitoring reads the same records.

Litigation over a diagnostic miss often happens years after the event, when the model has been retrained twice and the UI redesigned. If the record only says 'AI score 0.12', nobody can show whether the clinician saw a red flag, a green tick or nothing at all. The decision record must capture four things the model log does not: the exact presentation, the timing of display relative to the clinician's own read, the action taken, and the reason for an override.

@dataclass(frozen=True)
class DiagnosticDecision:
    record_id: str
    study_id: str
    input_sha256: str            # hash of pixels or structured inputs actually scored
    model_id: str                # registry key, e.g. "cxr-triage"
    model_version: str           # immutable build, never "latest"
    intended_use: str            # copied from labelling at inference time
    in_envelope: bool            # device, protocol, population inside validated range
    output: dict                 # {"label": "no_ptx", "p": 0.07, "abstained": False}
    presentation: str            # "worklist_priority=routine; badge=none"
    independent_read_at: str | None  # clinician's own impression, before reveal
    shown_at: str
    acted_at: str | None
    action: str                  # "accepted" | "overridden" | "not_viewed"
    override_reason: str | None  # coded, plus free text
    clinician_id: str

def record(decision: DiagnosticDecision, ledger) -> None:
    if decision.action == "overridden" and not decision.override_reason:
        raise ValueError("override requires a reason code")
    ledger.append(decision)      # write-once store, retention set by counsel

Two fields deserve emphasis. in_envelope records whether the input was inside the validated range (scanner model, patient age, protocol); outputs outside it should be visually different or suppressed, and the record proves which. independent_read_at exists because sequential reading, where the clinician commits an impression before the AI is revealed, is the strongest evidence of independent judgment and the best defence against automation bias.

Designing against automation bias

Automation bias is the tendency to accept a machine's output and to stop searching once it has spoken. In diagnosis it shows up as two errors: commission (following a wrong flag) and omission (missing what the tool did not flag). Both are design problems as much as training problems, and the workflow controls below are what a court, a regulator or a quality committee will look for.

  • Reveal after the read for high-stakes tasks: show the AI result after the clinician records an impression, or at minimum log whether it was viewed first.
  • Abstain visibly. A calibrated model should say 'not assessed' for out-of-envelope or low-quality inputs instead of producing a confident negative.
  • Never let a negative deprioritise silently. Triage tools cleared to raise priority were not validated to lower it; configuring a worklist to bury negatives is an institutional decision that creates institutional exposure.
  • Show the basis. Heatmaps, cited findings or the guideline step that fired let the clinician check the reasoning, which also supports the non-device CDS argument where relevant.
  • Train on failure cases. Onboarding should include examples where the tool is wrong, drawn from local data. The ambient documentation version of this is in AI and clinician workflow.

Worked example: a triage miss

An emergency department deploys a cleared chest X-ray triage tool whose labelling says it flags suspected pneumothorax to prioritise reading. IT configures the radiology worklist so that studies the tool marks negative drop to routine priority. A 34-year-old with a small apical pneumothorax is scored 0.07 (negative). The study waits four hours; the patient deteriorates.

Walk the record. The input was a portable film from a scanner model outside the vendor's validation set, so in_envelope was false, but the UI showed no difference. The presentation was 'routine, no badge'. The radiologist never opened the study before the deterioration, so action was 'not_viewed'.

Who is exposed? The vendor's labelling said 'prioritise', not 'rule out', which helps it under the learned intermediary doctrine, though a claimant may argue the warnings about portable films were inadequate. The hospital used the tool beyond its intended use by letting a negative lower priority, and suppressed the envelope signal: a corporate negligence theory writes itself. The radiologist's exposure turns on whether a four-hour wait for a routine study met local practice. The fixes are configuration and display, not model accuracy: negatives never demote, out-of-envelope studies show 'not assessed' and keep their original priority, and the record proves both.

Fairness duties, contracts and insurance

Diagnostic models often perform unevenly across subgroups because training data under-represents some populations or uses proxies such as cost or access. In the US, the HHS Section 1557 rule published in 2024 (45 CFR 92.210) requires covered entities to make reasonable efforts to identify patient care decision support tools that use race, colour, national origin, sex, age or disability as inputs, and to mitigate the risk of discrimination, with compliance from 1 May 2025. Check current enforcement posture with counsel, but the engineering response is the same: an inventory of tools and their input variables, subgroup miss rates from the decision record, and a documented mitigation for each gap.

Contracts and insurance close the loop. Ask vendors for subgroup performance on populations like yours, notice before model changes, version pinning, access to logs, and indemnities that match who controls what. Tell your malpractice and cyber carriers which AI tools are in use; some policies now ask. AI and insurance covers the underwriting questions.

Failure modes

  • 'Latest' in production: the record names a mutable tag, so nobody can reproduce the output. Pin immutable versions.
  • Presentation not logged: the score is stored but not what the clinician saw. Log the rendered state.
  • Override without reason: overrides are your best evidence of independent judgment and your best signal of model drift; require a coded reason.
  • Scope creep: a triage aid becomes a de facto rule-out because a downstream configuration changed. Treat worklist and alert configuration as part of the regulated deployment.
  • Silent envelope breach: a new scanner or population arrives and nothing changes on screen. See AI medical devices for input integrity checks.
  • Retention too short: claims can be brought years later; set record retention with counsel against limitation periods, not storage budgets.

What to do next

  1. Inventory every diagnostic or triage AI in use, with its labelled intended use and who configured it.
  2. Compare each configuration with the intended use; remove any path where a negative lowers priority or suppresses review.
  3. Implement the decision record, including presentation, timing, envelope flag and override reasons, in a write-once store.
  4. Switch high-stakes reads to sequential reveal, or at least log viewing order.
  5. Report monthly override rates, abstention rates and misses by subgroup to the quality committee.
  6. Review vendor contracts for version pinning, change notice, log access and indemnity, and tell your insurers what you run.
Key takeaway: Liability for AI-assisted diagnosis still runs through malpractice, institutional negligence and product liability, and today's standard of care rewards documented independent judgment. So configure tools strictly within their intended use, never let a negative quietly lower priority, reveal AI output after the clinician's own read where stakes are high, and keep a decision record that proves what was shown, when and what the clinician did about it.