Clinical trials are one of the most regulated data environments there is, and language models are arriving in almost every part of them: drafting protocols, matching patients to eligibility criteria, coding adverse events, reconciling data queries, summarising safety narratives. Each of these touches protected health information, and several touch the trial's integrity itself: the blinding that stops anyone knowing who received the drug, the audit trail inspectors rely on, and the safety reporting clocks set in regulation.
This article treats AI in trials as a security and data-integrity problem. It maps use cases by risk, summarises the rules engineers actually meet, lays out a reference architecture, walks through the threats that are specific to trials, and builds a worked eligibility pre-screening pipeline with a blinding firewall and a hash-chained audit log. Device regulation for AI is covered separately in FDA Regulation of AI, in depth. This is an engineer's map, not legal advice; your quality, privacy and regulatory functions own the final decisions.
Where AI touches a trial, ranked by risk
The useful question for each use is the one the FDA's draft credibility framework asks: how much does the model's output influence a decision, and how bad is a wrong decision? A protocol drafting assistant whose output is rewritten by medical writers has low influence. An adverse-event triage model that decides which reports a human reads first has high influence on a high-consequence decision, because a mis-triaged serious event can miss a reporting deadline.
| Use | Data touched | Influence | Consequence of error | Minimum controls |
|---|---|---|---|---|
| Protocol and document drafting | Usually none | Low | Low to moderate | Human authorship, version control |
| Eligibility pre-screening | PHI from records | Moderate | Moderate: wrong referral or missed patient | De-identification or covered deployment, evidence spans, coordinator decides |
| Adverse event coding (MedDRA) | Trial data | Moderate | Moderate: wrong term skews safety tables | Suggest only, coder confirms, agreement metrics |
| Safety case triage | Trial data, narratives | High | High: missed expedited report | Never deprioritise without a human; clock starts on receipt |
| Data query generation in EDC | Trial data | Moderate | Moderate | Audit trail entries, site confirms |
| Analysis or endpoint derivation | Trial data | High | High: affects efficacy conclusions | Full credibility assessment, pre-specified in the analysis plan |
Notice that the last row is a different kind of system. Once a model produces data that supports a regulatory decision on safety or effectiveness, it is squarely in scope of the credibility framework and needs the most evidence. Most teams start with the middle rows, where the model assists a qualified person.
The rules engineers actually meet
- ICH E6(R3) Good Clinical Practice was adopted by ICH in January 2025, took legal effect in the EU on 23 July 2025, and was published by the FDA as guidance in September 2025. It makes data governance and computerised system validation explicit sponsor and investigator duties: systems must be fit for purpose, validated in proportion to risk, and their audit trails must let an inspector reconstruct who did what to which record and when.
- 21 CFR Part 11 applies to electronic records in FDA-regulated trials. Section 11.10(e) requires secure, computer-generated, time-stamped audit trails of operator entries and actions that create, modify or delete records, and record changes must not obscure what was there before. A model writing into a record is an operator action that needs the same trail.
- HIPAA, for US covered entities, gives research uses a few paths: participant authorisation, an IRB or privacy board waiver, a limited data set under a data use agreement, or de-identified data under Safe Harbor or expert determination. Reviews preparatory to research let a researcher look at records to design a study or find candidates, but the PHI may not leave the covered entity, so that provision does not cover sending charts to an outside model service.
- The FDA draft guidance on using AI to support regulatory decision-making for drugs and biologics was issued in January 2025 and, at the time of writing, remains a draft; check the FDA guidance page for a final version. It sets out seven steps: define the question of interest, define the context of use, assess model risk from influence and consequence, write a credibility assessment plan, execute it, document the results, and decide adequacy for the context of use.
- The joint FDA and EMA guiding principles of good AI practice in drug development, published on 14 January 2026, list ten principles including a risk-based approach, a clear context of use, data governance and documentation, risk-based performance assessment and life cycle management.
EU trials add the Clinical Trials Regulation and GDPR on top; pseudonymised trial data is still personal data under GDPR, which matters when choosing where a model runs. For how HIPAA applies to LLM systems generally, see HIPAA for LLM Systems, in depth.
Reference architecture
The architecture has one rule that everything else follows from: the model proposes, a qualified person decides, and the system can prove both. Data reaches the model only through a minimisation gateway that de-identifies or allowlists fields for the specific use. Anything from the randomisation system is stopped at a blinding firewall. The model runs on a pinned version inside an environment you have validated, whether that is your own deployment or a provider under a business associate agreement. Its output is checked by a verifier before a human sees it, and every step writes to an audit trail.
Pinning the model version is not optional in a trial. A model that changes behind an alias in the middle of enrolment means patients in month two were screened by a different system from those in month one, and you cannot show that the process was consistent. Treat a model or prompt change as a change to a validated system: impact assessment, regression run, documented approval.
Threats specific to trials
General LLM risks apply, but several are sharper in trials.
- Unblinding. If a model can see treatment assignment, its summaries can leak it to blinded staff, even indirectly through phrasing. Field-level blocking stops direct leaks; it does not stop inference from clinical signals, such as a lab pattern typical of the active drug. Keep unblinded data in a separate tenant and do not let the same model context serve blinded and unblinded users.
- Re-identification. Trial cohorts are small, and rare diseases make them smaller. Free-text notes carry dates, places and family details that Safe Harbor removal can miss. Prefer structured extraction to sending whole notes, and test de-identification with an adversarial sample.
- Prompt injection through source documents. Clinical notes, referral letters and patient-reported outcomes are untrusted text. A note that contains instructions should not change what the model does; constrain outputs to a schema and verify claims against spans.
- Hallucinated eligibility. A fluent statement that a patient meets an inclusion criterion with no supporting text is worse than no answer. See LLM hallucination risk for general controls; here the control is that every claim must cite a span that the verifier checks.
- Safety clock delays. Under 21 CFR 312.32, serious and unexpected suspected adverse reactions must be reported in an IND safety report within 15 calendar days of the sponsor determining that the information qualifies, and unexpected fatal or life-threatening ones within 7 calendar days of initial receipt. A triage model must never move a case out of a human queue, and time spent waiting on it counts against both clocks.
- Recruitment bias. A pre-screener trained or prompted on records that under-document some groups will refer fewer of them. Measure referral rates by group against the eligible population.
Worked example: evidence-checked eligibility pre-screening
Consider a site that screens its records for a heart failure trial. Some criteria are structured: age between 40 and 80, an ejection fraction below a threshold, no eGFR below a cut-off. Others live in notes: no planned cardiac surgery, able to complete a six-minute walk. The pipeline evaluates structured criteria with deterministic rules first, and only asks the model about the rest. The model never outputs "eligible"; the best it can do is refer a patient to a coordinator, with evidence.
from dataclasses import dataclass
@dataclass
class CriterionResult:
criterion_id: str
verdict: str # "met" | "not_met" | "unknown"
evidence: list # [(doc_id, start, end)] into the de-identified text
source: str # "rule" | "model"
def prescreen(patient, criteria, llm, model_meta, audit):
results = []
for cr in criteria:
if cr.structured: # age, labs, coded diagnoses
results.append(cr.evaluate(patient.structured))
continue
out = llm.assess(criterion=cr.text, docs=patient.deidentified_docs,
schema=CriterionResult) # structured output only
if not out.evidence or not all(span_supports(patient, e, cr) for e in out.evidence):
out = CriterionResult(cr.id, "unknown", [], "model") # unsupported claim dropped
results.append(out)
if any(r.verdict == "not_met" and r.source == "rule" for r in results):
decision = "exclude_by_rule"
else:
decision = "refer_to_coordinator"
append(audit, actor="prescreen-service", action=decision, record=patient.pseudo_id,
after=[r.__dict__ for r in results], model=model_meta)
return decision, resultsThree design choices carry the safety argument. Rules exclude, the model never does, so a model error can only cost a coordinator a few minutes of review, not a patient their chance at the trial. Claims without verified evidence become unknown rather than met. And the model sees de-identified documents under a pseudonymous id; re-identification for contact happens inside the site, by staff entitled to do it. For reading records safely in general, see AI over medical records, in depth.
The blinding firewall
The blinding firewall is a few lines of code with a large effect. It sits in front of every model call and fails closed: if a context contains any allocation field and the caller's role is blinded, the call does not happen.
ALLOCATION_FIELDS = {"arm", "treatment_assignment", "randomization_code", "dose_level", "ip_allocation"}
class BlindingViolation(Exception):
pass
def flatten_keys(obj, prefix=""):
if isinstance(obj, dict):
for k, v in obj.items():
yield k.lower()
yield from flatten_keys(v, prefix + k + ".")
elif isinstance(obj, list):
for v in obj:
yield from flatten_keys(v, prefix)
def to_model_context(record, role):
if role != "unblinded_statistician":
leaked = ALLOCATION_FIELDS & set(flatten_keys(record))
if leaked:
raise BlindingViolation(sorted(leaked))
return recordAdd a test that seeds a record with each allocation field under nested keys and asserts the violation, and alert on any violation in production, because one means a data feed changed shape upstream.
A tamper-evident audit trail
Part 11 asks for an audit trail that is computer-generated, time-stamped, and does not obscure earlier values. For model-assisted actions, record the model name, version and prompt version alongside the human actor, and make the log tamper-evident by chaining hashes, so any edit to an old entry breaks every later hash. Store values in the validated data store; the chain proves they were not altered. For logging design at scale, see Audit logging for LLM systems.
import hashlib, json
from datetime import datetime, timezone
def append(log, actor, action, record, after, before=None, model=None):
prev = log[-1]["hash"] if log else "0" * 64
entry = {"ts": datetime.now(timezone.utc).isoformat(), "actor": actor,
"action": action, "record": record, "before": before, "after": after,
"model": model, "prev": prev}
entry["hash"] = hashlib.sha256(json.dumps(entry, sort_keys=True, default=str).encode()).hexdigest()
log.append(entry)
return entry
Validation and change control
Validation follows the credibility steps. Write down the question and context of use, for example: identify candidates for coordinator review, with no automated exclusion. Rate the risk. Build a reference set of charts labelled by two coordinators, and measure the pipeline against it: sensitivity for eligible patients matters most, because the cost of a false referral is a short review while the cost of a miss is a lost participant. Record agreement between the labellers too, since the model cannot be held to a standard the humans do not meet. Lock the model and prompt versions, re-run the reference set on any change, and keep the results with the trial master file documentation your quality team requires.
Failure modes
| Failure | What it looks like | Control |
|---|---|---|
| Allocation data in context | Blinded staff learn arms from summaries | Firewall on every call; separate unblinded tenant |
| PHI sent to an external service | Charts leave the covered entity without a legal basis | De-identify, or deploy under a BAA inside the entity |
| Unsupported eligibility claim | Referral cites text that does not say that | Span verifier; unsupported becomes unknown |
| Model updated mid-trial | Screening behaviour shifts with no record | Pinned versions, change control, regression set |
| Triage hides a serious case | Expedited report filed late | Model can only raise priority; humans clear queues |
| Injected instruction in a note | Output deviates from schema or task | Schema-constrained output, untrusted-text handling |
What to do next
- List every AI use in your trials with its data, influence and consequence, and decide which need a full credibility assessment.
- Put a minimisation gateway and a fail-closed blinding firewall in front of every model call, with tests.
- Pin model and prompt versions, and route any change through your computerised-system change control.
- Require evidence spans for every model claim and verify them before a human sees the output.
- Write model-assisted actions to a hash-chained audit trail with model and prompt versions.
- Build a labelled reference set, measure sensitivity and subgroup referral rates, and file the results.