An AI medical device is software whose output can change what happens to a patient: a triage score that moves a scan up the reading list, a detector that marks a suspected lesion, an algorithm that flags an arrhythmia, a dose recommendation. When such a device is attacked or degrades silently, the harm is clinical: a missed finding, a delayed treatment, a wrong dose. That changes the security engineering. Confidentiality still matters, but integrity and availability of the model's decisions matter more, and a safe failure has to be designed in, because "the model is unsure" must turn into a visible, conservative behaviour rather than a quiet guess.
This article treats the device as a system to defend. It walks through the attack surface specific to machine learning in clinical settings, the engineering controls that address each part, a worked example of an integrity-checked inference service, how to tell drift from attack in the field, and where these controls meet the regulatory expectations. The regulatory pathway itself, from device status to change control plans, is covered in FDA AI/ML regulation, in depth, and this page links there rather than repeating it. Nothing here is legal or clinical advice; it is an engineer's threat model.
The attack surface
Start by listing what an attacker, or an accident, can touch. An imaging AI device typically receives files from scanners and archive systems over hospital networks, runs a model, and writes results back into a worklist or report. Each hop is an attack path, and the machine learning layer adds paths that ordinary software does not have.
| Surface | What can go wrong | Primary control |
|---|---|---|
| Input data in transit and at rest | Images or signals altered, swapped between patients, or replayed | Provenance checks, integrity metadata, patient and study binding |
| Model inputs at inference time | Evasion: inputs crafted or corrupted so the model misclassifies | Input validation, out-of-distribution detection, abstention |
| Training and tuning data | Poisoning: mislabeled or planted samples change behaviour | Data lineage, curation review, held-out clinical test sets |
| Model artifact and update channel | Substituted or tampered weights, rollback to a vulnerable version | Signed artifacts, version pinning, verified updates |
| Outputs and integrations | Results written to the wrong patient, or hidden by an integration fault | Output binding, acknowledgement tracking, audit log |
| Free-text features (LLM components) | Instructions embedded in notes steer a summariser or extractor | Treat text as data, constrained output schemas, human review |
FDA's January 2025 draft guidance on AI-enabled device software functions names a similar list of AI-specific threats: data poisoning, model inversion and theft, evasion, data leakage and manipulation that drives performance drift. It was issued as a draft, so check whether a final version has since replaced it before citing it as current expectations.
Input integrity and the intended-use envelope
Research has demonstrated two kinds of manipulation that make input integrity concrete. Finlayson and colleagues (Science, 2019), drawing on their own experiments across three clinical imaging tasks, warned that small, carefully chosen perturbations can flip a medical classifier's output while the image looks unchanged to a human. The CT-GAN work by Mirsky and colleagues (USENIX Security, 2019) showed that a generative model, given access to scans in transit on a hospital network, could add or remove convincing evidence of disease in CT volumes, fooling radiologists and an AI model alike. You do not need the details of either attack to draw the lesson: a pixel array arriving at your model is not evidence of what the scanner produced.
The controls are layered. First, authenticate the path: mutual TLS between the device and the archive, network segmentation so only expected systems can send studies, and logging of the sending system for every study. Second, bind identity: check that patient, study and series identifiers are consistent across the header and the order that triggered the analysis, and refuse studies whose identifiers do not match. Third, validate content against the model's intended use: modality, body part, acquisition parameters, pixel spacing, bit depth and value ranges that were in the validation data. Inputs outside that envelope should produce an explicit "not analysed, outside intended use" result, never a confident score. Fourth, add a cheap out-of-distribution signal, such as a density estimate on embeddings, and route high-novelty inputs to abstention.
Model artifacts, training data and the update channel
The model file is code in everything but name: whoever controls the weights controls the output. Treat it the way you treat firmware. Sign the model artifact and its manifest at release, and have the device verify the signature before loading. Pin the expected version and refuse silent rollbacks to an older, possibly vulnerable model. Avoid serialisation formats that execute code on load, such as Python pickle files, in favour of formats that store tensors only. Keep the preprocessing code, label maps and thresholds inside the same signed bundle, because changing a threshold changes clinical behaviour as surely as changing a weight.
Training data deserves the same provenance discipline. Record where every dataset came from, who labeled it, and which version trained which model, so that a poisoning report can be traced to the affected releases. The ingestion controls in data poisoning, in depth apply directly. Clinical test sets should be held out, versioned and access-controlled, because a test set that leaks into training gives inflated results that hide both overfitting and tampering. For connected devices, section 524B of the FD&C Act requires premarket submissions for cyber devices to include a plan for monitoring and addressing vulnerabilities and a software bill of materials. FDA's final premarket cybersecurity guidance of June 27, 2025, which superseded the 2023 version, explains what reviewers expect. Including model artifacts and ML runtime libraries in that SBOM is a sensible extension.
Worked example: an integrity-checked inference service
Here is a minimal inference wrapper that enforces the controls above in code: it verifies the model bundle before loading, rejects inputs outside the validated envelope, abstains on novelty, and writes an audit record for every decision. The signature check uses Ed25519 from the Python cryptography package; the model loader, embedding function and novelty scorer are placeholders for your own components.
import hashlib, json, time
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PublicKey
from cryptography.exceptions import InvalidSignature
ENVELOPE = {"modality": {"CT"}, "body_part": {"CHEST"},
"slice_mm": (0.5, 2.5), "kvp": (80, 140)}
NOVELTY_LIMIT = 0.99 # chosen on validation data, documented
def load_verified(bundle_path, sig_path, pubkey: Ed25519PublicKey, pinned_version):
blob = open(bundle_path, "rb").read()
try:
pubkey.verify(open(sig_path, "rb").read(), blob)
except InvalidSignature:
raise SystemExit("model bundle signature invalid: refusing to start")
manifest = json.loads(read_manifest(blob))
if manifest["version"] != pinned_version:
raise SystemExit("unexpected model version: refusing to start")
return load_model(blob), manifest, hashlib.sha256(blob).hexdigest()
def in_envelope(meta):
return (meta["modality"] in ENVELOPE["modality"]
and meta["body_part"] in ENVELOPE["body_part"]
and ENVELOPE["slice_mm"][0] <= meta["slice_mm"] <= ENVELOPE["slice_mm"][1]
and ENVELOPE["kvp"][0] <= meta["kvp"] <= ENVELOPE["kvp"][1])
def analyse(study, model, manifest, model_hash, audit):
record = {"study": study.uid, "model": manifest["version"],
"model_sha256": model_hash, "ts": time.time()}
if not study.ids_consistent():
record["result"] = "rejected: identifier mismatch"
elif not in_envelope(study.meta):
record["result"] = "not analysed: outside intended use"
else:
emb = embed(study.pixels)
if novelty(emb) > NOVELTY_LIMIT:
record["result"] = "abstain: unusual input, read normally"
else:
score = float(model(study.pixels))
record["result"] = {"score": round(score, 4),
"threshold": manifest["threshold"]}
audit.append(record) # append-only store, separate credentials
return record["result"]Walk one study through it. A chest CT at 1.25 mm slices and 120 kVp, with consistent identifiers, passes the envelope; its embedding scores 0.42 on novelty, under the limit, so the model runs and the score is reported together with the signed manifest's threshold. A knee MRI sent to the same endpoint by a misrouted rule is answered "outside intended use". A chest CT with an unusual reconstruction kernel that pushes novelty to 0.995 gets an explicit abstention, which the worklist shows as "read normally", the safe default. Every path writes an audit record naming the exact model hash, so an incident review can reconstruct what the device saw and decided.
Drift versus attack in the field
Once deployed, model performance changes for mundane reasons: a new scanner model, a protocol change, a different patient population. An attacker who manipulates inputs at scale produces some of the same symptoms. Monitoring has to catch both and help you tell them apart. Track input statistics per site and per device (acquisition parameters, intensity histograms, embedding novelty), output statistics (score distributions, positive rates, abstention rates) and, where available, agreement with confirmed outcomes. Alert on shifts against a baseline from the validation period.
Distinguishing causes is an investigation, not a formula, but a few signals help. Natural drift tends to arrive with a known change and to affect every study from the new source. Manipulation tends to be targeted: a cluster of studies from one sending system, inputs whose headers and pixel statistics disagree, or score changes concentrated near the decision threshold. Keep raw inputs for a retention window agreed with your privacy team, so suspicious studies can be examined later. The patient data protections in HIPAA for healthcare LLM applications apply to that archive as much as to the production path.
Generative and LLM components
Generative components are entering devices and the software around them: summarising prior reports, extracting findings into structured fields, drafting patient letters. These inherit the prompt-injection problem. A referral note or a scanned document can contain text that reads as an instruction to the model. Where a generative feature influences a regulated function, keep it on the data side of the boundary. Have it produce output only in a constrained schema that is validated before use, never let it call tools that change orders or results, and show its output to a clinician as a draft with sources. Patient-scoped retrieval and sensitivity labels, as described in AI and medical records, in depth, keep one patient's notes from leaking into another's summary. Whether a given generative feature is itself a device is a regulatory question for the FDA page and your counsel.
Failure modes
- Confident output on out-of-scope input: a model scoring images it was never validated on because nobody checked the envelope.
- Results written to the wrong patient after an integration change, with no identifier binding to catch it.
- A model or threshold changed by a configuration push outside the signed bundle, bypassing change control.
- Rollback to an older model with a known weakness because version pinning was not enforced.
- Audit logs that record the score but not the model hash, so an incident cannot be reconstructed.
- Monitoring that alerts only on aggregate accuracy, missing a targeted attack on one site.
- Abstention that silently drops the study from the worklist instead of returning it to normal reading.
- A generative summariser that follows instructions found inside a clinical document.
Trade-offs
Every control costs something. Tight input envelopes and abstention reduce coverage: some legitimate studies will go unanalysed, and clinicians will notice. Tune limits on validation data and report abstention rates as a product metric, not a hidden one. Adversarial training and input purification can improve robustness to perturbations, but they rarely give guarantees and can reduce clean accuracy. Treat them as defence in depth behind provenance controls, not as a replacement. Signed, pinned models slow emergency fixes unless you rehearse the release path. Logging raw inputs aids forensics but enlarges the sensitive data you must protect. The right balance depends on the device's risk: a worklist prioritiser that only reorders cases can tolerate more automation than a tool that recommends a dose.
What to do next
- Draw your device's data flow with explicit trust boundaries and list, for each crossing, how the data is authenticated and validated.
- Write the intended-use envelope as code (modality, anatomy, acquisition ranges) and return an explicit result for inputs outside it.
- Package model weights, preprocessing, label maps and thresholds in one signed, versioned bundle, verified at load, with rollback refused.
- Add a novelty score with a documented abstention threshold, and make abstention route the case back to normal reading.
- Log every decision with study id, model hash, inputs summary and result to an append-only store with separate credentials.
- Set up per-site monitoring of input statistics, score distributions and abstention rates, with a written triage playbook for drift versus attack.
- Extend your SBOM to model artifacts and ML runtimes, and check current FDA guidance status (including whether the January 2025 AI draft has been finalised) with regulatory counsel.