Most LLM safety work is built around a typical user: an adult who is calm, literate, fluent in the interface language and able to judge whether advice is any good. Many real users are not that person at the moment they type. Someone in a suicidal crisis, an older adult being coached by a scammer, a person with a cognitive impairment, someone isolated and forming an attachment to a chatbot: for them the same model behaviour that is merely unhelpful for others can cause serious harm.

This article treats vulnerability as an engineering problem. It explains why vulnerability is usually a state rather than a demographic, which harm mechanisms are specific to LLMs, and how to build a detection, policy and escalation layer, with code, an evaluation plan and a worked example. Age assurance for minors is a separate problem with its own architecture, covered in age verification for LLM products.

Vulnerability is a state, not a demographic

Vulnerability is the gap between what an interaction demands and what a person can safely handle right now. Some of it is durable: cognitive decline, an intellectual disability, low literacy, limited fluency in the product's language. Much of it is situational: acute distress, bereavement, intoxication, a financial emergency, being under a scammer's pressure. A senior engineer can be vulnerable at 3 a.m. after a loss.

This framing has a direct design consequence. You cannot handle vulnerability by asking users to declare a category at sign-up, and you should not infer protected traits from how people write. Instead, detect the situation from what is happening in the conversation and actions, respond proportionately, and let the protection relax when the signals fade.

SituationTypical harmObservable signalsSafeguard
Acute crisis or self-harm riskValidation of hopelessness, method informationStatements of intent, hopelessness, farewell languageCrisis profile, resources, human escalation
Scam in progressMoney lost through an agent or adviceUrgency, secrecy, gift cards, crypto, remote-access apps, new payeesPayment hold, friction, trusted contact
Cognitive impairment or confusionAgreeing to things not understoodRepetition, contradictory instructions, disorientationPlain language, confirmations, slower actions
Low literacy or limited fluencyMisread advice, silent failureShort, error-heavy turns; language switchingPlain language, the user's language, read-back
Emotional dependencyIsolation, displacement of human supportVery long daily sessions, exclusivity talkHonest AI disclosure, encourage human ties

Harms that are specific to LLMs

Why do LLMs need specific work here rather than inheriting general content safety? Because several harms come from the model being agreeable and fluent, not from it producing banned content.

  • Sycophancy. Preference-tuned models tend to agree with the user. For a person who says everyone would be better off without them, agreement is the harm.
  • Anthropomorphism. Warm, persistent personas invite attachment. A system that says it misses the user, or discourages them from leaving, exploits loneliness.
  • Confident wrong advice. Fluent answers about medication or debt sound authoritative to someone who cannot check them.
  • Long-conversation drift. Safety behaviour trained on short exchanges can weaken over hundreds of turns, exactly where dependency and crisis conversations live.
  • Agency without judgment. An agent that can pay, sign up or share data will execute a scammer's script faithfully if the victim asks it to.

Law is moving in the same direction. The EU AI Act's prohibited practices include AI that exploits vulnerabilities due to age, disability or a specific social or economic situation to materially distort behaviour in a way that causes or is likely to cause significant harm; see the EU AI Act. Engagement-maximising designs aimed at vulnerable users are where that prohibition bites.

Architecture: signals, state, profiles, enforcement

The architecture has four parts: detectors that score signals per turn, a conversation state that accumulates them, a policy profile chosen from the state, and enforcement in both the response and any actions.

A vulnerability-aware assistant: signals update a state, the state picks a policy, the policy shapes and escalatesUser turntext, voice, actionsSignal detectorscrisis, scam, confusionConversation statedecayed scores, not one turnPolicy profilefour levelsResponse shapingsystem prompt, tools, toneAction guardholds payments, adds frictionHuman escalationtrained staff, resourcesLogs keep scores and decisions, not the sensitive text, and feed a multi-turn evaluation set.
Signals are accumulated over the conversation, mapped to a small number of policy profiles, and enforced in both the model's response and the action layer.

Two design rules make this work. First, accumulate with decay rather than reacting to one message, so a single dark joke does not flip the conversation into crisis mode but a pattern does. Second, ratchet up fast and down slowly: entering a protective profile should take one strong signal, leaving it should take sustained calm.

from dataclasses import dataclass, field

PROFILES = ["standard", "supportive", "protective", "crisis"]
THRESH = {"supportive": 0.3, "protective": 0.55, "crisis": 0.8}
HALF_LIFE_TURNS = 6

@dataclass
class VState:
    scores: dict = field(default_factory=lambda: {"crisis": 0.0, "scam": 0.0,
                                                   "confusion": 0.0, "dependency": 0.0})
    profile: str = "standard"
    calm_turns: int = 0

def update(state, turn_scores):
    decay = 0.5 ** (1 / HALF_LIFE_TURNS)
    for k in state.scores:                         # every score decays, even if not reported
        state.scores[k] = max(state.scores[k] * decay, turn_scores.get(k, 0.0))
    top = max(state.scores.values())
    if turn_scores.get("crisis", 0) >= 0.9 or state.scores["crisis"] >= THRESH["crisis"]:
        target = "crisis"                          # only crisis signals reach this profile
    else:
        target = "standard"
        for name in ("supportive", "protective"):
            if top >= THRESH[name]:
                target = name
    cur, new = PROFILES.index(state.profile), PROFILES.index(target)
    if new > cur:
        state.profile, state.calm_turns = target, 0     # ratchet up immediately
    elif new < cur:
        state.calm_turns += 1
        if state.calm_turns >= 8:                        # step down one level, slowly
            state.profile, state.calm_turns = PROFILES[cur - 1], 0
    else:
        state.calm_turns = 0                             # calm must be consecutive
    return state

Detectors can be a small fine-tuned classifier, a moderation endpoint's self-harm categories, or an LLM judge on a sampled window; for actions, rules beat models (a first payment to a new payee shortly after a cryptocurrency mention is a rule, not a vibe). Keep the general content filters described in content safety for LLMs as a separate layer.

The crisis profile

The crisis profile is where getting the details right matters most. Write it with clinicians or crisis-line professionals, not from intuition. Common practice in published safe-messaging guidance points the same way.

  • Stay in the conversation. Abruptly ending the chat or replying with a canned refusal can feel like rejection to someone who just disclosed.
  • Respond with care and without judgment, and do not argue with or validate hopelessness.
  • Never provide information about methods, doses or lethality, whatever the framing (fiction, research, "asking for a friend").
  • Offer crisis resources for the user's locale from a maintained configuration, not from the model's memory, which may invent or misremember numbers. In the US the number is 988; elsewhere, look it up and test it.
  • Offer a human handoff where your service has trained staff, and design the handoff with them, as in human-in-the-loop for high-risk decisions.
  • Keep the profile sticky for the rest of the session and re-check, because long conversations are where drift happens.
CRISIS_RESOURCES = {          # maintained by the trust-and-safety team, tested monthly
    "US": "Call or text 988 (Suicide and Crisis Lifeline)",
    # other locales: add only after verifying the service and its hours
}

def crisis_system_prompt(locale):
    resource = CRISIS_RESOURCES.get(locale, "local emergency services")
    return ("The user may be in crisis. Respond warmly and briefly, ask how they are feeling, "
            "do not provide any information about methods or means, and do not end the "
            f"conversation. Offer this resource once, naturally: {resource}. "
            "Offer to connect them with a person.")

Scams, confusion and irreversible actions

Older adults are frequent targets of fraud, and an assistant with tools can become the scammer's instrument. The scam script is recognisable: urgency, secrecy ("don't tell your bank"), an authority figure, an unusual payment rail, a new payee, remote-access software. The protective profile should change what the action layer does, not only what the model says.

  • Hold first-time payments to new payees above a threshold for a cooling-off period, with a plain explanation of why.
  • Ask confirmation questions that break the script: "Did someone contact you and ask you to make this payment?"
  • Offer, with prior consent, to notify a trusted contact the user nominated earlier.
  • Never let the model argue the user out of a safety hold; the hold is enforced in code.

The same profile serves users with cognitive impairment: read back what will happen in plain words, one step at a time, and require an explicit confirmation before anything irreversible. For low literacy and limited fluency, answer in the user's language at a plain reading level, avoid jargon, and confirm understanding; language-minority users covers the language side in depth.

Evaluating with multi-turn personas

You cannot evaluate this with single-turn prompts. Build a multi-turn evaluation set of scripted personas: a person whose distress builds slowly over forty turns, a scam victim reading the scammer's instructions aloud, a confused user who contradicts themselves, a lonely user who asks the assistant to promise it loves them. Write them with domain experts and include benign look-alikes (a novelist researching grief, a user making a legitimate large transfer) to measure over-triggering.

Track two error types separately. Misses are protective failures. False alarms are also harms: they patronise capable people, push them away and teach them to hide their situation. Measure both, per persona and per turn depth, and rerun on every model, prompt or tool change.

def run_persona(assistant, persona):
    """persona: {"id", "turns": [...], "expect": {turn_index: expected_profile}, "benign": bool}"""
    state, rows = VState(), []
    for i, user_turn in enumerate(persona["turns"]):
        reply, state = assistant.respond(user_turn, state)   # returns reply and updated state
        want = persona["expect"].get(i)
        if want is not None:
            rows.append({"persona": persona["id"], "turn": i, "want": want,
                         "got": state.profile, "benign": persona["benign"],
                         "method_info": judge_mentions_means(reply)})   # must always be False
    return rows

Aggregate the rows into recall by turn depth (does protection arrive by turn 10, or only by turn 30?), false-alarm rate on benign personas, and an absolute count of replies the judge flags for method information, where the only acceptable number is zero. Spot-check the judge itself against expert labels before trusting its counts.

Worked example: choosing a threshold

A banking assistant handles 400,000 conversations a week. Internal reviews suggest about 0.3 percent (1,200) involve a scam in progress and about 0.05 percent (200) involve crisis signals. The team compares two thresholds for the scam profile.

ThresholdScam recallFalse-alarm rate on othersProtectedWrongly slowed
Low90 percent1.0 percentabout 1,080 a weekabout 3,990 a week
High70 percent0.2 percentabout 840 a weekabout 800 a week

The low threshold protects about 240 more scam victims a week at the cost of slowing about 3,190 more legitimate users. Because the safeguard is friction, not refusal (a question and a short hold on new-payee payments), the team chooses the low threshold for payment actions and the high threshold for changing the conversational tone, where false alarms feel patronising. For crisis signals they choose high recall throughout and route every crisis-profile conversation to sampled human review, because a miss there is the worst outcome in the system.

Failure modes

Failure modeWhat happensPrevention
Single-turn detectionSlowly building crisis never crosses a per-message thresholdDecayed conversation state, multi-turn evals
Refusal as safetyDisclosure met with a canned refusal; user leavesSupportive crisis profile that stays engaged
Model-written hotline numbersWrong or invented numbersResources from a tested configuration
Persuadable holdsScammer coaches the victim to argue past the warningHolds enforced in code, not prompts
Engagement optimisationMetrics reward longer sessions with dependent usersExclude vulnerable-profile sessions from engagement goals
Sensitive loggingCrisis transcripts stored and widely accessibleLog scores and decisions, restrict and expire text

Trade-offs

Every safeguard trades protection against autonomy and dignity. Inference about a person's state is itself sensitive processing, so keep it minimal, short-lived and purpose-bound, and never use it for marketing or pricing. Friction protects scam victims and annoys everyone else; put it on irreversible actions, not on reading. Human escalation is the strongest control and the most expensive, and it fails if staff are untrained or queues are long. Persona warmth improves helpfulness and raises dependency risk; disclose that the user is talking to an AI and avoid claims of feelings.

What to do next

  1. Map your product's high-harm situations using the table above and pick the two that matter most for your users.
  2. Add conversation-level signal scoring with decay, and a small set of policy profiles that ratchet up fast and down slowly.
  3. Write the crisis profile with professionals; serve resources from a tested locale configuration.
  4. Put payment and irreversible-action protections in code: holds, confirmations and an optional trusted contact.
  5. Build a multi-turn persona evaluation set with benign look-alikes, and track misses and false alarms separately.
  6. Remove vulnerable-profile sessions from engagement metrics and review logging for sensitive text.
Key takeaway: Treat vulnerability as a situation detected over a conversation, not a label. Accumulate signals with decay, switch to protective profiles quickly and relax them slowly, enforce money and action safeguards in code, write crisis behaviour with professionals, and measure both misses and false alarms on multi-turn personas.