Connecting an LLM to an electronic medical record (EMR, also called an electronic health record or EHR) is mostly an integration and identity problem, not a modelling one. The model needs patient context to be useful. Its output must land where clinicians work, and every read and write has to happen with the right authority for the right patient. Get the plumbing wrong and a helpful summariser becomes a way to read any chart in the hospital, or to write unreviewed text into a legal record.
This article covers the integration surfaces the HL7 standards define: SMART on FHIR app launch, SMART backend services, CDS Hooks, and FHIR write-back. For each one it explains the identity involved, what to verify, and where an LLM adds new risk. It ends with a worked example, failure modes and a checklist. The read-side design of a chart assistant (patient-scoped retrieval, sensitivity labels, citations) is covered in AI + Medical Records, and the regulatory mapping is in HIPAA for healthcare LLM applications. Vendor EHRs implement these standards with their own registration processes and extensions, so treat the specifics below as the standard baseline and check your vendor's documentation.
The integration surfaces
FHIR (Fast Healthcare Interoperability Resources) exposes clinical data as typed resources such as Patient, Observation, MedicationRequest and DocumentReference, over a REST API. Around it, four paths matter for AI features:
| Path | Who starts it | Identity | Typical AI use |
|---|---|---|---|
| SMART app launch | a clinician opens your app inside the EHR | the user, through OAuth 2.0 scopes | chart summary panel, drafting assistant |
| CDS Hooks | the EHR, at a workflow event | the EHR authenticates to you with a signed JWT | suggestions while ordering or opening a chart |
| SMART backend services | your service, on a schedule or queue | your system, with no user present | batch coding review, quality measures |
| Write-back | your app or service | whichever identity holds write scopes | draft notes, proposed orders |
The rule that runs through the rest of the article: the identity on the arrow sets the ceiling for what the AI feature can touch, and the model itself never holds a credential.
SMART app launch and scopes
In an EHR launch, the EHR opens your app's launch URL with two parameters: iss, the EHR's FHIR base URL, and launch, an opaque handle for this launch's context. Your app fetches .well-known/smart-configuration from the FHIR base to find the authorisation and token endpoints, then runs an OAuth 2.0 authorisation-code flow.
The SMART App Launch guide sets several requirements that matter for security. Apps must use PKCE, and servers must support the S256 method and must not support plain. The state value must be unpredictable, with at least 122 bits of entropy, and validated on return. The authorise request carries an aud parameter naming the FHIR server, which stops a real token being sent to a fake resource server. One more check is not spelled out as a requirement but follows from how the launch works: iss arrives as a query parameter. An app that follows any iss can be pointed at an attacker's server. Keep an allowlist of FHIR base URLs you were registered with and refuse the rest.
Scopes decide what the token can read. SMART v2 scopes have the form context/ResourceType.permissions, with context patient, user or system and permission letters c r u d s (create, read, update, delete, search) in that order. Ask for the minimum your feature needs:
launch openid fhirUser
patient/Patient.r
patient/Condition.rs
patient/MedicationRequest.rs
patient/Observation.rs?category=http://terminology.hl7.org/CodeSystem/observation-category|laboratoryThe last line uses a query-parameter suffix, which the v2 syntax allows, to limit Observation access to laboratory results. Not every server enforces these suffixes, so confirm that yours does before relying on one. Avoid wildcard scopes such as user/*.cruds for an AI feature. They make every prompt-injection or bug a chart-wide incident. Request offline_access only if the feature must work while the user is away, because refresh tokens extend the blast radius of a leak.
Backend services: the system identity
Batch jobs, such as nightly review of discharge summaries, run with no user. SMART backend services use the OAuth 2.0 client-credentials grant, and the client authenticates with a signed JWT rather than a shared secret. Per the SMART asymmetric client authentication profile, the assertion's iss and sub are the client ID, aud is the token URL, exp is no more than five minutes ahead, and jti is a unique nonce the server checks for replay. Clients must support RS384 and ES384.
import time, uuid, jwt, requests # PyJWT with the cryptography package
def backend_token(client_id, token_url, private_key_pem, kid, scopes):
now = int(time.time())
assertion = jwt.encode(
{"iss": client_id, "sub": client_id, "aud": token_url,
"exp": now + 300, "jti": str(uuid.uuid4())},
private_key_pem, algorithm="ES384", headers={"kid": kid})
r = requests.post(token_url, data={
"grant_type": "client_credentials",
"client_assertion_type": "urn:ietf:params:oauth:client-assertion-type:jwt-bearer",
"client_assertion": assertion,
"scope": " ".join(scopes), # e.g. ["system/DocumentReference.rs"]
}, timeout=10)
r.raise_for_status()
return r.json()["access_token"]A system token is the most powerful credential in the design, because system/ scopes are not limited to one patient. Keep its private key in a key management service, scope each job's client to the resources that job needs, and never let a user-facing request path reach code that holds a system token. That last rule is the confused deputy problem: a clinician asking a chat feature about one patient must not be served by a credential that can read every patient.
CDS Hooks: the EHR calls you
CDS Hooks lets the EHR call your service at defined points in the clinician's workflow. The standard hooks include patient-view, order-select, order-sign, order-dispatch, encounter-start, encounter-discharge and appointment-book. The EHR POSTs a request with hook, a unique hookInstance, the hook's context (for example the user and patient IDs and the draft orders), optional prefetch data, and optionally fhirServer with a fhirAuthorization access token. The specification limits that token's scope to the service being invoked and the current user.
Verify every call before doing any work. The EHR sends a JWT bearer token. Check its signature against the EHR's published keys, check that iss is on your allowlist and that aud is your service URL, and reject expired tokens and reused jti values. The specification forbids the none algorithm and symmetric algorithms. Then check that the patient in context matches the patient in any prefetched resources. Pin that patient ID for the rest of the request and check every FHIR read against it.
Your service answers with cards. Each card has a summary under 140 characters, an indicator of info, warning or critical, and a source. It can also carry suggestions, whose actions are create, update or delete on FHIR resources, and links of type absolute or smart, the latter launching a SMART app.
{
"cards": [{
"summary": "Possible duplicate anticoagulant: apixaban already active",
"indicator": "warning",
"source": {"label": "Medication safety assistant (AI-generated, verify)"},
"detail": "Active MedicationRequest for apixaban 5 mg twice daily, started 2026-09-12.",
"links": [{"label": "Review medication list", "type": "smart",
"url": "https://ai.example.org/launch"}]
}]
}Two cautions apply to LLM-backed cards. Hooks sit in the clinician's critical path and the EHR will not wait long, so precompute or fall back to returning no cards rather than blocking an order. And a suggestion the clinician can accept with one click is effectively a write. Keep LLM output in the card's text, and only attach suggestions built by deterministic code from structured data.
Write-back: drafts, provenance and idempotency
Writing to the chart is where an AI feature carries the most risk. Three rules keep it manageable.
- Drafts, not records. Write generated notes as a DocumentReference with
docStatusset topreliminary, or into the EHR's own draft or in-basket mechanism. A clinician edits and signs, and only signing makes it final. - Provenance. Record a Provenance resource that targets the draft and names your software and model version as an agent. When a summary is later found to be wrong, you can find every document that model version touched.
- Idempotency. Retries happen. Give each draft a business identifier derived from the hook instance or job ID and the patient, and search for it before creating, or use a conditional create if your server supports it. Otherwise a timeout produces two drafts in the chart.
Never let the model choose the target patient or encounter of a write. Those IDs come from the launch or hook context that your code pinned, not from model output.
Worked example: discharge instructions
Take a discharge-instructions assistant for an emergency department.
- The clinician opens the patient's chart. The EHR fires
patient-view. Your service verifies the JWT, pins the patient ID, and returns oneinfocard with asmartlink: "Draft discharge instructions". - The clinician clicks it. The EHR launches your app with
issandlaunch. The app checksissagainst its allowlist and completes the authorisation-code flow with PKCE, asking for the scopes listed earlier pluspatient/DocumentReference.cs(create, and search to find an earlier draft) andpatient/Provenance.c. - The backend reads active conditions, medications and the latest laboratory results for the pinned patient only. It builds a prompt from structured fields and short excerpts, and leaves out identifiers the task does not need, such as address and insurance numbers.
- The model call goes to an endpoint covered by a business associate agreement. No OAuth token, FHIR URL or internal ID is in the prompt.
- The response is checked: every medication it names must exist in the patient's active list, and any dose must match the structured record. A failure blocks the draft and shows the clinician why.
- The app shows the draft. The clinician edits it and presses save. The app writes a preliminary DocumentReference and a Provenance resource, keyed by an identifier built from the launch and patient, so a retry finds the existing draft.
- Audit records hold the user, the patient ID, the resource types read, the model version and a hash of the prompt, not the prompt text.
Each step is boring by design. The model writes prose; everything that decides who, which patient and what gets stored is ordinary code.
Failure modes
- Tokens in prompts or logs. A debugging change logs the full request, including the bearer token, to a vendor's observability tool. Redact authorisation headers at the HTTP client and treat any token found in a log as compromised.
- Patient drift. The user switches charts while a long request runs, and the result lands in the wrong patient's view. Tie every response to the pinned patient ID and drop mismatches.
- Injected instructions in notes. Free-text notes, referral letters and patient messages can carry text that steers the model. See indirect prompt injection. Because writes are drafts and targets are pinned, injection degrades the draft rather than taking an action.
- Over-broad scopes. A feature registered with wide read scopes during a pilot keeps them in production. Review granted scopes per release.
- Unverified callers. A CDS service that accepts unsigned requests lets anyone on the network query patient data through your
fhirAuthorizationhandling. - Silent latency failures. A slow model makes the hook time out. The EHR shows nothing, and nobody notices the feature is dead. Track hook latency and empty-response rates as service-level indicators.
Trade-offs
User-delegated SMART tokens give the strongest guarantee, because the AI feature can never see more than the clinician could, but they need a user present. Backend services allow batch work at the cost of a powerful credential you must contain. CDS Hooks puts help inside the workflow but caps your latency and makes every suggestion a near-write. Drafts with human sign-off slow adoption slightly and remove most of the liability. Sending less context to the model reduces exposure and sometimes quality; decide that per task with the clinical owner, and record the decision. Personal data minimisation is covered in more depth in PII leakage in LLM systems.
What to do next
- Draw your four arrows: list every path your AI feature uses into the EHR and the identity on each.
- Replace any wildcard scopes with per-resource v2 scopes, and test whether your server enforces query-parameter suffixes before relying on them.
- Add an allowlist for
isson SMART launch and for the JWT issuer on CDS Hooks, and verifyaud,expandjtion every hook call. - Move system credentials into a key management service, and prove in a test that no user-facing endpoint can reach them.
- Make every write a preliminary draft with Provenance and an idempotency identifier.
- Add a check that every patient ID in a response matches the pinned context, and alert on mismatches.
- Track hook latency, timeouts and empty-card rates on a dashboard with an owner.