Clinicians in many health systems spend a large share of their working day on documentation, much of it after clinic hours. Ambient documentation tools, often called AI scribes, attack that burden directly: a device listens to the consultation, speech is transcribed, a language model drafts the clinical note, and the clinician reviews and signs it. Microsoft's Dragon Copilot, announced in March 2025 and combining the earlier Dragon Medical One dictation and DAX ambient products, Abridge and Suki are well-known examples, and many electronic health record vendors now ship their own.

This article is about engineering that workflow safely: where meaning can change between what was said and what is signed, how protected health information flows and where it must stop, why clinician review is a control that has to be designed rather than assumed, and how to monitor quality once thousands of notes a day go through the pipeline. Patient-facing chat is a different problem, covered in AI Health and LLMs.

The workflow you are changing

Before any AI, a typical outpatient note is written from memory and scribbles after the patient leaves, or dictated, or typed during the visit while the clinician half-faces a screen. The ambient workflow moves the clinician's job from author to editor. That change is the whole point and also the main risk: an editor reading a fluent draft is more likely to miss an error than an author is to make one, a pattern known as automation bias.

The other shift is in data. A dictated note contains what the clinician chose to say. An ambient recording contains everything: the patient's partner discussing their own health, a phone call taken mid-visit, a sensitive disclosure the patient asked to keep out of the record. The pipeline has to treat that raw material as more sensitive than the note it produces.

Architecture of an ambient pipeline

A defensible architecture has seven stages. The figure shows them; the rest of the article walks through the controls at each.

Ambient documentation pipeline: every arrow is a place where data can leak or meaning can changeConsent checkrecorded per encounterAudio capturedevice, encryptedSpeech to textdiarised transcriptContext fetchproblems, meds, FHIRNote drafter (LLM)sentence-level evidence links to transcriptVerifierunsupported claims, meds, laterality, negationsClinician reviewflagged sentences must be toucheddraft + flagsSign and write backDocumentReference, finalTelemetryedit rate, omission audits, deletesAudio and transcript are deleted on a fixed schedule; only the signed note becomes part of the legal record.
Consent gates capture; transcript and chart context feed the drafter; a verifier flags risky sentences; the clinician must address flags before signing; telemetry watches edits and omissions.
StageMain riskControl
ConsentRecording without permission; recording third partiesPer-encounter consent captured in the record; pause and redact controls
CaptureAudio on unmanaged devices; loss of a phoneManaged app, encrypted at rest, no local copy after upload
Speech to textMisheard drug names, wrong speaker attributionMedical vocabulary models; speaker labels; confidence scores carried forward
Context fetchPulling the wrong patient's chartContext bound to the encounter ID from the EHR launch, never to a name
DraftingHallucinated findings, dropped negations, invented plansSentence-level evidence links; template constraints
VerificationUnsupported claims reaching the clinician unmarkedIndependent checks that flag, never silently fix
Write-backUnsigned drafts treated as finalDraft status until signature; provenance records the AI's role

Consent and capture

Consent is the first control because nothing downstream can undo an unlawful recording. Rules on recording conversations vary by jurisdiction, and health privacy law adds its own requirements, so the legal standard must come from counsel. Engineering's job is to make the chosen standard impossible to skip: the capture app starts only after a consent flag for that encounter is set, the flag records who obtained consent and when, and a visible pause control lets the patient or clinician stop recording for part of the visit.

On the device, audio should stream to the service or sit encrypted in app storage until upload, then be deleted. Personal phones running a managed app are common; the management profile must allow remote wipe and block audio export. Under HIPAA, the vendor processing audio is a business associate, so a business associate agreement must be in place before the first recording, and the PHI handling rules described in HIPAA and LLM Applications apply to every hop.

From speech to draft: where errors enter

Errors enter at two points. Speech recognition mishears: drug names that sound alike, doses, left and right, and who said what. The drafter then adds its own class of error: stating a finding that was never examined, turning 'no chest pain' into 'chest pain', merging two complaints, or writing a plan the clinician only mentioned as a possibility.

The most effective structural control is to make the drafter cite. Each sentence in the draft carries the transcript span or chart element that supports it. A sentence with no support is either a template phrase, which is allowed and labelled, or a candidate hallucination. The verifier then checks the highest-risk categories independently of the drafter.

HIGH_RISK = ("medication", "dose", "allergy", "laterality", "negation", "plan_order")

def verify(draft, transcript, chart):
    flags = []
    for s in draft.sentences:
        if not s.evidence and not s.is_template:
            flags.append((s.id, "unsupported"))
            continue
        for kind in s.categories & set(HIGH_RISK):
            if kind == "medication" and not chart.has_med(s.drug) and not transcript.mentions(s.drug):
                flags.append((s.id, "med not in chart or transcript"))
            if kind == "dose" and s.dose != transcript.dose_near(s.drug, s.evidence):
                flags.append((s.id, "dose differs from what was said"))
            if kind == "negation" and transcript.polarity(s.evidence) != s.polarity:
                flags.append((s.id, "negation flipped"))
            if kind == "laterality" and transcript.side(s.evidence) not in (None, s.side):
                flags.append((s.id, "left/right mismatch"))
        if any(span.asr_confidence < 0.80 for span in s.evidence):
            flags.append((s.id, "low-confidence audio"))
    return flags

The verifier flags and never silently rewrites. A silent correction hides the fact that the pipeline was uncertain, and the clinician is the only party who heard the conversation. The 0.80 confidence threshold is illustrative; tune it on your own audited notes.

Designing review that resists automation bias

Clinician review is the control everyone relies on, and it fails quietly. A busy clinician who has approved fifty good drafts will approve the fifty-first without reading it. Treat review as a user interface problem with measurable behaviour.

  • Flagged sentences must be touched. The sign button stays disabled until each flagged sentence is accepted, edited or deleted. Highlighting alone is ignored.
  • Evidence on demand. Clicking a sentence plays or shows the supporting transcript span, so checking costs seconds rather than minutes.
  • Orders are never auto-placed. Draft orders and prescriptions are suggestions that go through the normal ordering workflow, with its existing safety checks.
  • Attestation is honest. The signature text states that the note was AI-drafted and reviewed, and the clinician remains its author of record.

Sign-off can be enforced in code with a simple gate:

def can_sign(note, flags, actions):
    unresolved = [f for f in flags if f.sentence_id not in actions]
    if unresolved:
        return False, f"{len(unresolved)} flagged sentences need review"
    if note.has_pending_orders and not note.orders_routed_to_cpoe:
        return False, "orders must be placed through the ordering screen"
    return True, "ok"

Writing back to the record

When the note leaves the pipeline it becomes part of the legal medical record, so write-back needs care. With a FHIR R4 interface, a note is commonly stored as a DocumentReference. Its docStatus field distinguishes preliminary from final, and a Provenance resource can record that an AI system drafted the text and which clinician signed it. Write drafts as preliminary or keep them out of the EHR entirely until signature.

{
  "resourceType": "DocumentReference",
  "status": "current",
  "docStatus": "final",
  "type": {"text": "Outpatient progress note"},
  "subject": {"reference": "Patient/123"},
  "context": {"encounter": [{"reference": "Encounter/abc"}]},
  "author": [{"reference": "Practitioner/dr-lee"}],
  "content": [{"attachment": {"contentType": "text/plain", "data": "<base64 note>"}}]
}

Bind the write to the encounter identifier received at launch, and reject the write if the patient in the draft does not match the patient on the encounter. Wrong-patient documentation is one of the oldest errors in health IT, and an asynchronous pipeline that finishes after the clinician has moved on to the next patient makes it easier. Chart access controls are covered in AI and Medical Records.

Security issues specific to ambient pipelines

Several security issues are specific to ambient pipelines.

  • Retention of raw material. Audio and full transcripts are more sensitive than the note. Keep them only as long as needed for review and quality audit, delete on a fixed schedule, and log the deletion.
  • Spoken prompt injection. The transcript is untrusted input. A patient or anyone in the room can say text that reads like an instruction. The drafter must not have tools that act on the record, and its output is only ever a draft.
  • Third-party speech. Relatives and interpreters are recorded. Diarisation and a redact control let clinicians remove content that should not enter the patient's record.
  • Vendor model use. Contracts should state whether audio or notes are used to train models, and if so under what de-identification standard.
  • Access logging. Every playback of audio and every view of a transcript is a PHI access and belongs in the audit trail.

Worked example: a primary care visit

A primary care visit: a patient with a cough asks about a rash on the left forearm, mentions their sister's diabetes, and says they stopped one of their blood pressure tablets. The draft arrives with four flags. 'Lisinopril 20 mg daily' is flagged because the transcript says the patient stopped it; the clinician edits it to 'lisinopril, stopped by patient two weeks ago'. 'Rash on right forearm' is flagged for laterality, since the transcript says left. 'History of diabetes' is unsupported for the patient, because the transcript attributes it to the sister, so it is moved to family history. A low-confidence span around a cough medicine name is checked by playing four seconds of audio. Review takes ninety seconds instead of the several minutes needed to write the note, and every risky sentence was looked at.

Monitoring in production

Monitoring has to run continuously because drafting models, templates and speech engines all change. Track per clinician and per specialty:

  • Edit rate: the share of characters changed between draft and signed note. A sudden drop can mean better drafts or rubber-stamping; sample to find out which.
  • Flag resolution: how often flagged sentences are edited versus accepted unchanged.
  • Omission audits: a reviewer compares a sample of recordings with signed notes for missing findings, because edit rate cannot see what the draft left out.
  • Wrong-patient and wrong-side incidents reported through the normal safety system.

Roll out in stages. Start with a pilot group of clinicians who agree to have a sample of their notes compared against recordings, publish the audit results to them, and expand only when omission and flip rates are stable. Vendor marketing figures for minutes saved per encounter are self-reported; measure your own time-in-notes before and after, using the EHR's activity logs rather than surveys.

Whether a scribe is regulated as a medical device depends on its intended use and claims; see FDA AI/ML Regulation before adding features such as suggested diagnoses.

Failure modes and trade-offs

Failure modeWhy it happensMitigation
Rubber-stamp signingFluent drafts and time pressureForced flag resolution; sample audits
Omitted findingsDrafter summarises away detailTemplate sections that must be addressed; omission audits
Flipped negationSpeech recognition or model errorNegation-specific verifier check
Wrong patientAsynchronous completionEncounter-bound writes with patient match
Over-retentionAudio kept for model improvement by defaultContracted deletion schedule

The core trade-off is time saved against review depth. More flags mean safer notes and slower review; tune verifier sensitivity per specialty, using audits rather than complaints as the signal.

What to do next

  1. Map your current documentation workflow and decide which note types are in scope first; start with low-acuity outpatient visits.
  2. Agree the consent standard with counsel and build it as a hard gate on capture.
  3. Require sentence-level evidence links from any vendor or internal drafter you adopt.
  4. Implement a verifier for medications, doses, negation and laterality that flags without rewriting.
  5. Disable signing until flagged sentences are addressed, and route orders through the normal ordering screen.
  6. Bind write-back to the launch encounter and record AI provenance.
  7. Set and enforce deletion schedules for audio and transcripts, and log every playback.
  8. Run monthly omission audits on a sample of recordings and review edit-rate trends per clinician.
Key takeaway: Ambient documentation turns clinicians from authors into editors, which saves time and invites automation bias. Gate capture on consent, make the drafter cite the transcript for every sentence, verify high-risk content independently, block signing until flags are addressed, bind write-back to the encounter, delete raw audio on schedule, and audit for omissions that edit rates cannot see.