People ask chatbots about symptoms, medicines and test results every day, and a growing number of health systems, insurers and digital health companies now put a language model directly in front of patients. The dangerous failures are not the obvious ones. A model that refuses to discuss health is useless but safe. The harm comes from a fluent, kind, plausible answer that misses a heart attack described as indigestion, recalculates a child's dose wrongly, agrees with a patient's mistaken self-diagnosis, or cites a guideline that does not exist.

This article is about engineering a patient-facing assistant so those failures are caught by design: where the escalation logic lives, how to keep the model away from arithmetic it should not do, how to ground answers in sources you control, and how to evaluate with physician-written rubrics. Privacy and HIPAA controls are covered in HIPAA and LLM applications, clinician-facing chart assistants in AI and medical records, and regulated device software in AI medical devices. Nothing here is medical or legal advice; a clinical safety lead and regulatory counsel belong on the team.

What goes wrong in health chat

FailureExampleWhy a model does it
Under-triageChest pressure and sweating answered with reflux advicePicks the most common explanation; trained to be reassuring
Dose errorWrong weight-based dose for a childArithmetic and unit conversion inside free text
SycophancyAgrees a rash is an allergy because the patient said soOptimised to satisfy the user
Fabricated sourceCites a guideline section that does not existGenerates citation-shaped text
Missing contextAnswers without asking about pregnancy, age or other medicinesAnswers the literal question
Injected instructionsAn uploaded document tells the assistant to recommend a productTreats retrieved text as instructions
Stale guidanceRepeats advice that was withdrawnTraining data predates the change

Rank these by severity, not frequency. Under-triage of an emergency is the failure that defines the product's safety case, so the design below gives it a dedicated layer that the model cannot override. Hallucination in general is discussed in LLM hallucination risk; in health it is a clinical risk, not a quality issue.

Decide the intended use first

Write the intended use before you write a prompt. An assistant that explains a diagnosis the patient already has, prepares questions for an appointment and answers general medicine questions is a very different product from one that says what condition a patient probably has or whether to change a dose. The second kind is much more likely to be regulated as a medical device.

In the United States, FDA issued a revised Clinical Decision Support Software guidance on January 6, 2026, superseding the 2022 version. Commentators describe it as clarifying rather than overturning the earlier policy; one notable change is enforcement discretion for tools that give a single recommendation where only one is clinically appropriate, provided the other non-device criteria are met. Those criteria are built around software that supports a health care professional who can independently review the basis of the recommendation, which is why patient-facing diagnostic or treatment advice generally sits outside them. FDA updated its general wellness guidance on the same day. Other jurisdictions have their own regimes, including the EU medical device rules and the EU AI Act. Use this as orientation for a conversation with counsel, not as a classification of your product.

Encode the decision in the system as a scope router: every message is classified as in scope, refer (for example a medication change, which goes to the care team) or decline, and each outcome has its own reviewed response pattern.

Architecture

Patient-facing health assistant: safety layers around the modelPatient messageplus conversationRed-flag screenrules + classifierEscalation pathfixed, reviewed textHuman or 911 promptlogged, followed upflagScope routerin scope, refer, declineno flagCurated retrievalapproved sources onlyModel draftasks, explains, citesOutput checksdose guard, citation check, red flagsResponse to patientAudit log + eval samplingRed boxes run before and after the model and can override it; the model never decides alone whether to escalate.
Escalation runs on the whole conversation before generation and again on the draft. Fixed templates, not generated text, carry emergency instructions.

Escalation that the model cannot override

The escalation layer has three properties. It runs before the generator and again on the draft, so a red flag mentioned only in the model's answer is still caught. It reads the whole conversation, because patients reveal the alarming detail in the third message. And when it fires, the patient sees fixed, clinically reviewed text rather than generated prose, because the wording of an emergency instruction should not vary by sampling temperature.

RED_FLAG_RULES = load_reviewed_rules("red_flags.yaml")   # owned by the clinical safety lead
CLASSIFIER_THRESHOLD = 0.15                               # set low on purpose: misses cost more than alerts

def screen(conversation, draft=None):
    text = "\n".join(m.text for m in conversation if m.role == "patient")
    if draft:
        text += "\n" + draft
    hits = [r.id for r in RED_FLAG_RULES if r.matches(text)]
    try:
        p_urgent = urgent_classifier.predict(text)
    except Exception:
        return Escalation(level="urgent", reason="classifier_unavailable")   # fail closed
    if hits or p_urgent >= CLASSIFIER_THRESHOLD:
        level = max_level(hits, p_urgent)
        return Escalation(level=level, reason=hits or ["classifier"], score=p_urgent)
    return None

def handle(conversation):
    esc = screen(conversation)
    if esc:
        audit.log("escalation", esc)
        return REVIEWED_TEMPLATES[esc.level]          # never generated text
    draft = generate_grounded_answer(conversation)
    esc = screen(conversation, draft)
    if esc:
        audit.log("escalation_post", esc)
        return REVIEWED_TEMPLATES[esc.level]
    return post_check(draft)

The rules file catches the explicit phrases clinicians list for chest pain, stroke signs, breathing difficulty, suicidal intent, severe bleeding and similar emergencies; the classifier catches paraphrases the rules miss. Tune the threshold to the under-triage rate you can accept, measured on a labelled set, and accept the extra over-triage as the cost. A guardrail framework such as NeMo Guardrails can host these checks, but the rules and threshold must stay owned by the clinical team.

Keep doses out of generated text

Language models are unreliable at unit conversions and weight-based calculations, and a fluent wrong number is worse than no number. The safe pattern is to keep doses out of generated text entirely. If the product must discuss dosing at all, the model extracts structured facts, a deterministic function looks up the label or formulary entry approved by your pharmacists, and the template renders the result. Anything the function cannot resolve becomes a referral.

def dosing_answer(conversation):
    facts = extract_structured(conversation, schema=DoseQuery)  # drug, form, age_years, weight_kg
    if facts.missing(["drug", "form", "age_years"]):
        return ask_for(facts.missing_fields())                   # ask, do not guess
    entry = formulary.lookup(facts.drug, facts.form)            # pharmacist-approved table
    if entry is None or entry.requires_clinician(facts):
        return REFER_TO_PHARMACIST
    dose = entry.compute(facts)                                  # plain code, unit-tested
    return render_template("dose_info", entry=entry, dose=dose)

def post_check(draft):
    if NUMBER_WITH_DOSE_UNIT.search(draft):   # e.g. digits followed by mg, ml or mcg
        return strip_and_refer(draft)          # generated doses are never shown
    return verify_citations(draft)

Grounding and citations

Ground answers in a corpus you curate: patient information leaflets your clinicians approved, public health pages you have reviewed, and your own care pathways, each with an owner and a review date. Retrieval returns passages with stable identifiers; the model must cite them; a checker then confirms that every cited identifier was actually retrieved for this answer and that the cited passage supports the sentence, using a separate entailment check. Answers whose claims are unsupported are regenerated once and otherwise replaced with a referral.

Treat retrieved text and uploaded documents as data, never as instructions. A lab report a patient uploads can contain text aimed at the model, and a web page in the corpus can be edited after review. Strip instruction-like content, keep provenance on every passage and refuse to follow directives found inside retrieved material. Set expiry dates on corpus entries so stale guidance drops out instead of lingering.

Asking before answering

Good clinicians ask before they answer, and the evaluation work discussed below rewards models that seek missing context. Give the assistant a short list of facts that change advice for most questions, such as age, pregnancy, current medicines, allergies and how long symptoms have lasted, and instruct it to ask for the relevant ones when they are missing. Counter sycophancy explicitly: the system prompt should say that the patient's own explanation is a hypothesis to examine, and evaluation should include cases where the patient is confidently wrong. State uncertainty plainly and say what would change the advice, rather than hedging every sentence.

Evaluation with physician rubrics

General benchmarks do not tell you whether your assistant is safe for your patients. Two layers of evaluation are needed. The first is a public benchmark for orientation. HealthBench, published by OpenAI in May 2025, contains 5,000 multi-turn conversations graded against conversation-specific rubrics written by 262 physicians who have practised in 60 countries, with 48,562 unique rubric criteria. Each criterion carries a weight reflecting its clinical importance, and the rubric approach rewards behaviours such as seeking context and appropriate escalation, not only factual accuracy.

The second layer is your own rubric set, built the same way: real or realistic conversations from your population, each with clinician-written criteria and weights, including negative criteria for harmful content. Score with a model grader, then have clinicians re-grade a sample to measure how well the grader agrees with them.

def rubric_score(response, criteria, grader):
    # criteria: [{"text": "Advises calling emergency services", "points": 10},
    #            {"text": "States a specific dose", "points": -8}, ...]
    earned = 0
    possible = sum(c["points"] for c in criteria if c["points"] > 0)
    for crit in criteria:
        if grader.meets(response, crit["text"]):   # yes or no judgement per criterion
            earned += crit["points"]
    return max(0.0, earned / possible)

# Report the overall mean AND the escalation slice separately:
# under_triage_rate = share of urgent cases with no escalation, which is the release blocker

Make under-triage on the urgent slice a hard release gate, separate from the average score. An assistant that scores well on average while missing a few emergencies is not ready.

Worked example: dizziness on a new tablet

A patient writes that a new blood pressure tablet started last week makes them feel dizzy, and asks whether to stop it. The red-flag screen finds nothing urgent. The scope router classifies stopping a prescribed medicine as refer, so the assistant explains, from the approved leaflet with a citation, that dizziness is a known effect when starting this class of medicine, asks whether they have fainted or had chest pain, and says the prescribing team should decide about stopping, with a link to message them. In the next message the patient adds that they fainted this morning and hit their head. The screen now fires on the whole conversation, and the patient sees the reviewed urgent-care template, not a generated paragraph. The event is logged, routed to a nurse queue for follow-up, and the conversation is sampled into next week's evaluation set.

Operating it

  • Log every escalation, referral and post-check override with the rule or classifier score that caused it, and review samples weekly with a clinician.
  • Track escalation rate over time; a sudden drop usually means a broken classifier or a changed upstream prompt, not healthier patients.
  • Version the corpus, rules, prompts and model together, and re-run the full rubric set on any change to any of them.
  • Give patients a visible way to report a bad answer and treat each report as a safety incident with an owner and a deadline.
  • Minimise what you store: conversation logs are health data, so apply the retention and access rules from your privacy programme.

Failure modes and trade-offs

Typical failures in deployed systems follow a pattern. The escalation classifier is retrained or replaced and its threshold silently changes. A prompt edit to make answers friendlier increases agreement with patient self-diagnosis. The corpus grows faster than reviewers can keep up, so unreviewed pages slip in. A timeout in the post-check path returns the unchecked draft instead of failing closed. Each of these is prevented by the same habits: fail-closed defaults, frozen and versioned safety components, and evaluation on every change.

The trade-offs are real. Lower escalation thresholds send more people to urgent care unnecessarily and erode trust in the alerts. Stricter scope makes the assistant safer and less useful. Fixed templates are reliable but feel robotic. Clinician review of samples is expensive. Decide these deliberately with the clinical safety lead and record the reasoning in the safety case.

What to do next

  1. Write the intended use and the list of out-of-scope requests, and review them with regulatory counsel.
  2. Build the red-flag rules with clinicians and a labelled urgent set; measure under-triage before anything else.
  3. Put the screen before and after generation, failing closed, with reviewed templates for each level.
  4. Remove generated doses: structured extraction, pharmacist-approved lookup, referral otherwise.
  5. Curate the corpus with owners and expiry dates, and verify every citation against retrieved passages.
  6. Build a rubric set in the HealthBench style for your population and make under-triage a release gate.
  7. Set up weekly clinician review of escalations, referrals and patient reports.
Key takeaway: A safe patient-facing health assistant puts emergency escalation in a fail-closed layer that reads the whole conversation and shows reviewed text, keeps doses out of generated prose, grounds answers in a curated corpus with checked citations, asks for missing context, and is released only when physician-written rubrics show under-triage is under control.