Pharmaceutical companies use AI across the whole lifecycle: target discovery, molecule design, trial recruitment, manufacturing, regulatory writing and drug safety. Security in pharma means more than keeping attackers out. It also means data integrity: being able to prove that every record supporting a safety, quality or efficacy decision is accurate, attributable and unchanged. Regulators inspect for this, and a model that writes or changes such records is in scope.
This article is a practical guide for engineers and quality leads. It explains which AI uses fall under GxP rules and which do not, what current and draft guidance says about AI, a reference design for pharmacovigilance case intake with an LLM, where every extracted field is checked against its source, the threats that are specific to pharma, how to build validation evidence, and how to keep the system validated as models change. Medical device rules are out of scope here; see FDA AI/ML regulation for those.
Map uses to rules
The first question for any pharma AI system is whether it is GxP-relevant: whether its output feeds a record or decision governed by good practice rules (GMP for manufacturing, GCP for clinical trials, GVP for pharmacovigilance). That decides how much validation and data-integrity control it needs.
| Use | Usually GxP? | Main security concern |
|---|---|---|
| Target and molecule discovery | No | Intellectual property: structures and assay data leaking to vendors |
| Trial recruitment and feasibility | Partly (GCP) | Patient privacy, unblinding, consent scope |
| Clinical data review | Yes (GCP) | Data integrity, unblinding through access to allocation data |
| Manufacturing quality, batch review | Yes (GMP) | Data integrity, validated state, change control |
| Pharmacovigilance case intake | Yes (GVP) | Missed adverse events, patient data, late reports |
| Regulatory and medical writing | Partly | Hallucinated claims entering submissions or labels |
Non-GxP does not mean uncontrolled. A discovery team pasting unpublished structures into a public chatbot can lose patent novelty or trade secrets. It simply means the controls are about confidentiality, not validation.
What the rules and drafts say
Several stable requirements already apply to any computerised system in GxP work. US 21 CFR Part 11 covers electronic records and signatures, including audit trails. EU GMP Annex 11 covers computerised systems. ALCOA+ is the data integrity standard regulators use: data must be attributable, legible, contemporaneous, original and accurate, plus complete, consistent, enduring and available. ISPE's GAMP 5 (second edition, 2022) gives the industry's risk-based approach to validating computerised systems and includes guidance on machine learning.
AI-specific guidance is newer and partly still in draft. In January 2025 the FDA published a draft guidance, Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products. It proposes a seven-step, risk-based credibility framework built around the model's context of use, where model risk combines model influence (how much the output drives the decision) and decision consequence (how bad a wrong decision would be). In July 2025, the EU and PIC/S opened consultation on a new GMP Annex 22 on AI, alongside a revised Annex 11. The consultation draft allows static, deterministic models in GMP-critical uses with validation, and excludes generative AI and LLMs from critical uses. It is a draft, not in force, and may change. Check the current status of both documents before you cite them in a validation plan.
The practical reading for LLMs is consistent across these documents. Keep them in roles where a qualified person reviews every output that matters, pin the version, and build evidence that the combined human and model process performs at least as well as the process it replaces.
Reference design: safety case intake
Pharmacovigilance is a good first GxP use for LLMs because the work is heavy and well defined. Safety teams receive reports of possible adverse events by email, call-centre notes, literature and partner feeds. Each one must be assessed for validity. A valid individual case needs four elements: an identifiable patient, an identifiable reporter, a suspect product and an adverse event. Each one must then be coded (MedDRA for events), assessed for seriousness and expectedness, entered into the safety database and, if serious and unexpected, reported to regulators within short deadlines, typically 15 calendar days. Reports are exchanged in the ICH E2B(R3) format.
The LLM drafts the structured case from unstructured text. It does not decide validity, seriousness or reportability on its own. Those remain human decisions, supported by checks that are deliberately biased toward flagging.
Span-grounded extraction
The key control is span-grounded extraction. Every field the model returns must quote the exact source text it came from, and code verifies that the quote exists. A field without a verifiable source is not accepted. This turns hallucination from a silent error into a visible gap.
import hashlib, json, re
from datetime import datetime, timezone
REQUIRED = ["patient", "reporter", "suspect_product", "adverse_event"]
SERIOUS_HINTS = re.compile(r"\b(died|death|fatal|hospitali[sz]ed|admitted|life[- ]threatening|"
r"disabilit|birth defect|congenital|intensive care|ICU)\b", re.I)
def norm(t):
return " ".join(t.split()).lower()
def verify_extraction(source_text, extraction, model_version, prompt_version):
src = norm(source_text)
fields, gaps = {}, []
for name, item in extraction.items(): # item = {"value": ..., "evidence": "..."}
ev = item.get("evidence", "")
if ev and norm(ev) in src:
fields[name] = item
else:
gaps.append(name) # unverifiable: shown to the human as missing
missing = [r for r in REQUIRED if r not in fields]
serious_flag = bool(SERIOUS_HINTS.search(source_text)) or extraction.get("serious", {}).get("value") is True
return {
"source_sha256": hashlib.sha256(source_text.encode()).hexdigest(),
"model_version": model_version,
"prompt_version": prompt_version,
"fields": fields,
"unverified": gaps,
"missing_minimum_criteria": missing,
"possible_serious": serious_flag, # OR of model and keyword screen: biased to recall
"created_utc": datetime.now(timezone.utc).isoformat(),
"status": "awaiting_human_review", # never auto-submitted
}Three choices are deliberate. The seriousness flag is the logical OR of the model and a keyword screen, because a missed serious case is far worse than an extra review, so the combination favours recall. A case missing one of the four minimum criteria is not discarded; it is routed for follow-up, because the missing element may simply not have been extracted. The status is always awaiting review. The record captures source hash and model and prompt versions, which is what makes an output attributable and reproducible under ALCOA+. Patterns for the trail itself are in LLM audit logging.
Threats specific to pharma
- IP leakage. Unpublished structures, assay results and trial designs are among a pharma company's most valuable assets. Use enterprise agreements that exclude training on your data, keep the most sensitive work on models you host, and classify data before it reaches any prompt. Policy patterns are covered in LLM data governance.
- Injection through documents. Literature PDFs, partner emails and patient messages are untrusted input. A crafted document can try to tell the extractor to drop the adverse event. Span verification helps, because a dropped event shows up as a missing field, and so does the keyword screen running outside the model.
- Unblinding. A trial assistant with retrieval over all trial data may reveal treatment allocation to blinded staff. Enforce blinding in retrieval permissions, not in the prompt.
- Patient re-identification. Case narratives contain health data. Minimise what reaches the model and control where logs live; see PII leakage.
- Silent model change. A hosted model updated by the vendor is a changed system under GxP. Pin versions and treat any change as a change-control event that needs re-testing.
- Hallucinated claims in submissions. Writing assistants can invent references or numbers. Require every factual statement in a regulatory draft to trace to a source document.
Building validation evidence
Validation shows that the system is fit for its intended use. For the intake design above, the evidence package typically contains the following.
- Intended use and context of use. "Drafts structured case fields from English email reports for review by a trained case processor." Narrow wording keeps validation tractable.
- Risk assessment. Model influence is moderate, because a person reviews every case. Decision consequence is high, because a missed serious event delays a required report. That combination drives the depth of testing.
- Reference test set. Historical reports with gold-standard fields agreed by experienced processors, including hard cases: several products, events buried in long threads, other languages, and deliberately adversarial text.
- Acceptance criteria set in advance. For example, recall of possible serious cases at least as high as the current manual process, and zero accepted fields without verified evidence. Set the numbers from your own baseline, not from a vendor brochure.
- Human-plus-model comparison. Measure the combined process against the old one, including whether reviewers miss errors in pre-filled fields more often than in fields they type themselves.
- Change control. Define what counts as a change (model version, prompt, schema, keyword list) and which tests rerun for each.
Evaluation methods for LLM components are covered in LLM security evaluations.
Worked example
An email arrives from a pharmacist: "Patient (F, 67) started product X 50 mg last Tuesday. Three days later she was admitted with severe rash and fever. Discontinued. Recovering. J. Rao, community pharmacist." The extractor returns patient (female, 67, evidence "Patient (F, 67)"), reporter (pharmacist, evidence "J. Rao, community pharmacist"), suspect product (product X 50 mg), and event (severe rash and fever). The verifier finds every evidence string in the source. The keyword screen matches "admitted", so the case is flagged as possibly serious, because hospitalisation is a seriousness criterion. The processor confirms hospitalisation, codes the event in MedDRA, checks expectedness against the product label and starts the expedited reporting clock.
Now suppose the model had dropped the admission from its output. The keyword screen still flags it, and the processor sees a mismatch between the flag and the extracted fields. That disagreement is what you want the design to surface.
Operating a validated system
In operation, re-run the reference set on a schedule and on every change. Sample a fixed share of reviewed cases for second review, and track the rate at which reviewers change pre-filled fields; a falling edit rate can mean the model improved or that reviewers stopped checking, so compare against the second-review results. Watch the share of cases with unverified fields by source type, because a new partner feed or email format often shows up there first. Keep the old manual path working, so a model outage or a failed re-validation does not stop safety reporting.
Failure modes
- Model treated as decision maker. Seriousness or validity accepted without review. Keep status awaiting review and audit it.
- Unverified fields accepted. Plausible but invented values enter the safety database. Reject fields without matching evidence.
- Recall traded for speed. Thresholds tuned to cut review load miss serious cases. Measure recall first and keep the keyword screen.
- Unpinned vendor model. Validation silently goes stale. Pin and re-test on change.
- Logs outside the quality system. Audit trails in a vendor console that cannot be retained or inspected. Export them to a controlled store.
- Over-broad intended use. Validated for English email, used on call-centre transcripts. Enforce source types at intake.
Trade-offs
Span grounding rejects correct answers that paraphrase or infer, for example an age computed from a birth date, so reviewers fill more gaps by hand; in a safety system that is the right trade. Hosting models yourself protects IP and makes version pinning easy, but costs operations effort and limits model choice. A narrow intended use makes validation affordable and limits scope, and every extension needs new evidence. Keyword screens add false positives, which cost minutes, against false negatives that can cost a late safety report.
What to do next
- Inventory AI uses and mark each as GxP-relevant or not, with the governing rule set.
- Block unapproved tools for confidential research data and offer an approved alternative.
- For the first GxP use, write a narrow intended use and a risk assessment using model influence and decision consequence.
- Build span-grounded extraction with code verification and a recall-biased screen outside the model.
- Assemble a reference test set with hard and adversarial cases, and set acceptance criteria before testing.
- Pin model and prompt versions, and define what changes trigger re-validation.
- Export audit trails into a controlled, retained store.
- Track the FDA draft guidance and EU Annex 22 to final form and update your validation approach when they change.