If you build machine learning for healthcare in the United States, the Food and Drug Administration decides whether your model is a regulated medical device, what evidence you must show before selling it, and what you may change afterwards without asking again. That last question is where most of FDA's AI-specific policy has landed.
This article walks the framework the way an engineering team meets it: first deciding whether software is a device at all (including the January 2026 changes that matter for LLM features), then the pathways, what an AI submission has to contain, how a predetermined change control plan lets you retrain on schedule, and what continues after clearance. It ends with code that enforces a change protocol, failure modes and a checklist. It is an engineer's map, not legal advice; regulatory counsel signs off on the real decisions.
The regulatory path at a glance
Is your software a medical device?
The Federal Food, Drug, and Cosmetic Act defines a device in section 201(h) by intended use: something intended for diagnosis, cure, mitigation, treatment or prevention of disease. Software qualifies on its own, which the international regulators' forum (IMDRF) calls Software as a Medical Device (SaMD). Intended use is read from your labeling, marketing, UI text and even sales decks, so a feature can become a device because of how it is described.
The 21st Century Cures Act added section 520(o), which removes several software functions from the definition: administrative support, general wellness, electronic health records, transferring and displaying data, and certain clinical decision support (CDS). The CDS carve-out applies only if all four criteria hold: the software does not acquire, process or analyse a medical image or a signal from a diagnostic device; it displays or analyses medical information; it supports a healthcare professional's recommendation rather than a patient's; and the professional can independently review the basis of the recommendation instead of relying on it.
On 6 January 2026 FDA published revised final versions of its CDS and general wellness guidances without a draft round. The notable CDS change: FDA now says it will exercise enforcement discretion for software that produces a single clinically appropriate recommendation, provided the other criteria are met and the function is not meant for time-critical decisions. The earlier version treated a single directive output as a sign of device status.
| Feature | Likely status | Why |
|---|---|---|
| LLM drafts discharge summaries a clinician edits | Often outside device scope | Administrative or documentation support; no diagnostic claim |
| LLM suggests one guideline-based next step, with cited sources, for a clinician | CDS enforcement discretion is plausible after Jan 2026 | Single recommendation, reviewable basis, not time-critical |
| Model flags suspected pneumothorax on a chest X-ray | Device | Analyses a medical image |
| Sepsis alert that pages the team within minutes | Device | Time-critical, and the basis is not practically reviewable |
| Chatbot tells a patient whether to go to the ER | Device | Patient-facing, diagnostic intent |
Risk classes and pathways
Devices fall into Class I, II or III by risk, and the class largely picks the pathway. Most cleared AI devices reach market through 510(k), which shows substantial equivalence to a legally marketed predicate. De Novo creates a new classification for a novel device of low to moderate risk, and the device it authorises can then be a predicate for others. Premarket Approval (PMA) is for Class III and demands the strongest clinical evidence. FDA keeps a public list of AI-enabled devices it has authorised, which now runs to more than a thousand entries, dominated by radiology.
Before you file, use the Q-Submission (Pre-Sub) programme to get written feedback on intended use, predicate, validation design and change plan; it is the cheapest moment to learn what FDA will accept.
What an AI submission must show
Three documents shape what reviewers expect from AI. The Good Machine Learning Practice guiding principles (FDA, Health Canada and the UK MHRA, 2021) are ten short principles; the ones engineers feel most are representative datasets, training sets independent of test sets, testing in clinically relevant conditions, a focus on the human-AI team, and monitoring of deployed models. Transparency principles followed in 2024.
The January 2025 draft guidance on lifecycle management and marketing submissions for AI-enabled device software functions turns those into submission content. It was still a draft at the time of writing, so treat it as FDA's current thinking. It asks for a clear device and model description; data management (sources, sites, collection dates, labelling process, how test data was kept independent); performance validation with subgroup results by sex, age, race and ethnicity, and acquisition device; analysis of bias; human factors for how users interpret outputs; cybersecurity threats specific to AI such as data poisoning; a postmarket performance monitoring plan; and labeling that reads like a model card. The engineering consequence: your training pipeline must record lineage well enough to regenerate every one of those tables on demand.
Predetermined change control plans
The Food and Drug Omnibus Reform Act of 2022 added section 515C, which lets FDA authorise a predetermined change control plan (PCCP) as part of the original submission. Changes made according to an authorised PCCP do not need a new 510(k), De Novo or PMA supplement. FDA finalised its PCCP guidance for AI-enabled device software functions in December 2024, widening the scope from machine-learning devices to all AI-enabled ones. A PCCP has three parts:
- Description of Modifications: the specific, bounded changes you plan, such as retraining on new data from the same modality, adding a supported scanner model, or adjusting an operating threshold within a range.
- Modification Protocol: how each change is developed, validated and rolled out: data management, retraining practices, performance evaluation with pre-specified acceptance criteria, and update procedures, including how users are told.
- Impact Assessment: the benefits and risks of each change, how the protocol mitigates them, and how changes interact with each other.
A PCCP cannot change intended use or indications, and it cannot cover a change whose validation you cannot specify in advance; a version outside the bounds needs a new submission. The acceptance criteria are therefore the key engineering artefact: they decide whether a candidate may ship.
Enforcing the modification protocol in code
Encode the Modification Protocol as a release gate in CI so a person cannot quietly ship a model that falls outside it. The script below compares a candidate against the authorised criteria, judges the lower confidence bound rather than the point estimate, treats a missing subgroup as a failure, and refuses to evaluate on anything but the sequestered test set whose hash the plan names.
import json, sys
from dataclasses import dataclass
@dataclass
class Criterion:
metric: str # e.g. "sensitivity"
subgroup: str # "overall", "sex=F", "scanner=VendorB" ...
floor: float # absolute minimum from the authorised PCCP
max_drop: float # allowed drop versus the currently released model
def gate(candidate: dict, released: dict, criteria: list[Criterion],
test_set_hash: str, locked_hash: str) -> list[str]:
"""Return a list of violations; empty means the protocol is satisfied."""
bad = []
if test_set_hash != locked_hash:
bad.append("test set differs from the sequestered set named in the PCCP")
for cr in criteria:
key = f"{cr.metric}/{cr.subgroup}"
new = candidate.get(key)
old = released.get(key)
if new is None:
bad.append(f"{key}: not reported") # missing subgroup = failure, not a pass
continue
if new["ci_low"] < cr.floor: # judge the lower 95% bound, not the point
bad.append(f"{key}: CI low {new['ci_low']:.3f} below floor {cr.floor}")
if old and old["point"] - new["point"] > cr.max_drop:
bad.append(f"{key}: dropped {old['point'] - new['point']:.3f}")
return bad
if __name__ == "__main__":
cfg = json.load(open("pccp_criteria.json"))
crit = [Criterion(**c) for c in cfg["criteria"]]
cand, rel = json.load(open("candidate_metrics.json")), json.load(open("released_metrics.json"))
problems = gate(cand, rel, crit, cand["_test_set_sha256"], cfg["locked_test_set_sha256"])
print("\n".join(problems) or "PCCP acceptance criteria met")
sys.exit(1 if problems else 0) # CI blocks the release on non-zeroKeep the criteria file under design control: editing a floor changes the authorised plan.
Worked example: a triage model and an LLM drafter
A company sells a Class II chest X-ray triage device cleared through 510(k) that flags suspected pneumothorax for radiologist prioritisation. Its PCCP allows quarterly retraining on new data from existing sites, adding two named scanner vendors, and moving the operating threshold within a stated range. Acceptance criteria: sensitivity lower bound at least 0.90 and specificity lower bound at least 0.85 overall, no subgroup (sex, age band, vendor, site) below 0.85 sensitivity, and no drop larger than 0.02 against the released model.
In the third quarter, the retrained model improves overall sensitivity from 0.93 to 0.94 but the gate fails: sensitivity for the newly added vendor's portable units has a lower bound of 0.83 because only 140 positive cases were available. The team cannot lower the floor, so they ship the retrained weights only for the original vendors, which the protocol allows, and keep collecting portable-unit cases. Two quarters later the new vendor passes and is enabled, with a labeling update that names it. Every step leaves a record that maps to the plan.
The same company adds an LLM that drafts the report impression from the radiologist's own findings. It does not analyse the image and a professional edits and signs the text, so they keep it outside the device, with labeling that makes no diagnostic claim. Letting it read the image would make it part of the device.
Cybersecurity and the quality system
Section 524B, also added by FDORA, applies to cyber devices: devices that include software, can connect to the internet, and could be vulnerable. Submissions must include a plan to monitor, identify and address postmarket vulnerabilities, processes giving reasonable assurance the device and related systems are secure, and a software bill of materials. For an AI device the SBOM includes the inference runtime, model-serving libraries and their transitive dependencies; threat models should cover model files as artefacts that can be tampered with and training data as an attack surface.
The quality system changed too. The Quality Management System Regulation (QMSR), effective 2 February 2026, replaced the old Quality System Regulation by incorporating ISO 13485:2016 by reference, plus FDA-specific requirements. Cloud providers and model vendors are suppliers you must control.
Postmarket monitoring and drift
Clearance starts obligations. Medical Device Reporting requires reports of deaths, serious injuries and certain malfunctions. Complaints must be investigated. For AI the extra risk is silent degradation: a hospital replaces a scanner, changes a protocol, or its patient mix shifts, and accuracy falls with no error. Monitor input distributions and, where labels arrive later, outcome-linked performance.
import numpy as np
def psi(expected, actual, bins=10):
"""Population stability index between validation-time and live inputs."""
edges = np.quantile(expected, np.linspace(0, 1, bins + 1))
e = np.histogram(expected, edges)[0] / len(expected) + 1e-6
a = np.histogram(np.clip(actual, edges[0], edges[-1]), edges)[0] / len(actual) + 1e-6
return float(np.sum((a - e) * np.log(a / e)))
# per site, per week: score distribution and an image-statistics proxy (mean pixel intensity)
if psi(val_scores, site_scores_this_week) > 0.2:
open_capa(site, "score distribution shift") # investigate before anyone retrainsA drift alert opens an investigation under your CAPA process; it does not trigger automatic retraining. Retraining that is not described in the PCCP is a new submission.
Where LLMs strain the framework
Generative models strain assumptions built for locked classifiers. Outputs are open-ended, so a fixed test set does not bound behaviour; the same input can produce different outputs; and the foundation model often comes from a vendor that updates it on its own schedule; FDA's Digital Health Advisory Committee devoted its November 2024 meeting to generative AI-enabled devices. Practical consequences: pin the model version and treat a vendor update as a change you must evaluate; build task-specific evaluation suites with clinician-graded rubrics and hallucination checks; log inputs and outputs for postmarket review; and design the UI so the clinician sees the sources behind any recommendation, which is also what the CDS criteria reward.
Failure modes
- Marketing turns a tool into a device. A sales deck claims the documentation assistant catches missed diagnoses. Intended use now includes diagnosis. Review all external claims against the regulatory strategy.
- Test set leakage. One patient's images land in train and test. Split by patient and site.
- Subgroups too small to pass. Confidence intervals on 50 cases are wide, so criteria fail on noise. Size the test set for the subgroup bounds you promised.
- A vague PCCP. It will not be authorised or auditable. Name data sources, metrics and thresholds.
Trade-offs
A broad PCCP buys release speed but needs more validation and scrutiny; a narrow one is easier to authorise and sooner outgrown. Staying outside device scope avoids premarket review, but it limits what you can claim and leaves you exposed if the feature drifts towards diagnosis. De Novo costs more than 510(k) but gives you a classification others must follow. The usual answer is to clear a narrow first version with a PCCP aimed at the changes your roadmap actually needs, then widen.
Related reading on this site: the EU AI Act for LLM systems, the NIST AI Risk Management Framework, HIPAA for LLM applications, writing model cards and AI over medical records.
What to do next
- Write the intended use statement for each AI feature in one sentence and check it against the four CDS criteria and the January 2026 guidance.
- Audit marketing, UI copy and sales material for claims that widen intended use.
- Identify a predicate or decide on De Novo, then book a Pre-Sub that includes your draft PCCP.
- Split data by patient and site, sequester the test set, record its hash, and size it for subgroup confidence bounds.
- Turn the Modification Protocol into a CI gate like the one above and put its criteria file under design control.
- Produce an SBOM for the inference stack and a threat model covering model files and training data.
- Map your QMS to ISO 13485 under the QMSR, including cloud and model vendors as suppliers.
- Deploy per-site input drift and outcome monitoring, with alerts routed into CAPA, before the first customer goes live.