Most LLM safety work is built around a typical user: an adult who is calm, literate, fluent in the interface language and able to judge whether advice is any good. Many real users are not that person at the moment they type. Someone in a suicidal crisis, an older adult being coached by a scammer, a person with a cognitive impairment, someone isolated and forming an attachment to a chatbot: for them the same model behaviour that is merely unhelpful for others can cause serious harm.
This article treats vulnerability as an engineering problem. It explains why vulnerability is usually a state rather than a demographic, which harm mechanisms are specific to LLMs, and how to build a detection, policy and escalation layer, with code, an evaluation plan and a worked example. Age assurance for minors is a separate problem with its own architecture, covered in age verification for LLM products.
Vulnerability is a state, not a demographic
Vulnerability is the gap between what an interaction demands and what a person can safely handle right now. Some of it is durable: cognitive decline, an intellectual disability, low literacy, limited fluency in the product's language. Much of it is situational: acute distress, bereavement, intoxication, a financial emergency, being under a scammer's pressure. A senior engineer can be vulnerable at 3 a.m. after a loss.
This framing has a direct design consequence. You cannot handle vulnerability by asking users to declare a category at sign-up, and you should not infer protected traits from how people write. Instead, detect the situation from what is happening in the conversation and actions, respond proportionately, and let the protection relax when the signals fade.
| Situation | Typical harm | Observable signals | Safeguard |
|---|---|---|---|
| Acute crisis or self-harm risk | Validation of hopelessness, method information | Statements of intent, hopelessness, farewell language | Crisis profile, resources, human escalation |
| Scam in progress | Money lost through an agent or advice | Urgency, secrecy, gift cards, crypto, remote-access apps, new payees | Payment hold, friction, trusted contact |
| Cognitive impairment or confusion | Agreeing to things not understood | Repetition, contradictory instructions, disorientation | Plain language, confirmations, slower actions |
| Low literacy or limited fluency | Misread advice, silent failure | Short, error-heavy turns; language switching | Plain language, the user's language, read-back |
| Emotional dependency | Isolation, displacement of human support | Very long daily sessions, exclusivity talk | Honest AI disclosure, encourage human ties |
Harms that are specific to LLMs
Why do LLMs need specific work here rather than inheriting general content safety? Because several harms come from the model being agreeable and fluent, not from it producing banned content.
- Sycophancy. Preference-tuned models tend to agree with the user. For a person who says everyone would be better off without them, agreement is the harm.
- Anthropomorphism. Warm, persistent personas invite attachment. A system that says it misses the user, or discourages them from leaving, exploits loneliness.
- Confident wrong advice. Fluent answers about medication or debt sound authoritative to someone who cannot check them.
- Long-conversation drift. Safety behaviour trained on short exchanges can weaken over hundreds of turns, exactly where dependency and crisis conversations live.
- Agency without judgment. An agent that can pay, sign up or share data will execute a scammer's script faithfully if the victim asks it to.
Law is moving in the same direction. The EU AI Act's prohibited practices include AI that exploits vulnerabilities due to age, disability or a specific social or economic situation to materially distort behaviour in a way that causes or is likely to cause significant harm; see the EU AI Act. Engagement-maximising designs aimed at vulnerable users are where that prohibition bites.
Architecture: signals, state, profiles, enforcement
The architecture has four parts: detectors that score signals per turn, a conversation state that accumulates them, a policy profile chosen from the state, and enforcement in both the response and any actions.
Two design rules make this work. First, accumulate with decay rather than reacting to one message, so a single dark joke does not flip the conversation into crisis mode but a pattern does. Second, ratchet up fast and down slowly: entering a protective profile should take one strong signal, leaving it should take sustained calm.
from dataclasses import dataclass, field
PROFILES = ["standard", "supportive", "protective", "crisis"]
THRESH = {"supportive": 0.3, "protective": 0.55, "crisis": 0.8}
HALF_LIFE_TURNS = 6
@dataclass
class VState:
scores: dict = field(default_factory=lambda: {"crisis": 0.0, "scam": 0.0,
"confusion": 0.0, "dependency": 0.0})
profile: str = "standard"
calm_turns: int = 0
def update(state, turn_scores):
decay = 0.5 ** (1 / HALF_LIFE_TURNS)
for k in state.scores: # every score decays, even if not reported
state.scores[k] = max(state.scores[k] * decay, turn_scores.get(k, 0.0))
top = max(state.scores.values())
if turn_scores.get("crisis", 0) >= 0.9 or state.scores["crisis"] >= THRESH["crisis"]:
target = "crisis" # only crisis signals reach this profile
else:
target = "standard"
for name in ("supportive", "protective"):
if top >= THRESH[name]:
target = name
cur, new = PROFILES.index(state.profile), PROFILES.index(target)
if new > cur:
state.profile, state.calm_turns = target, 0 # ratchet up immediately
elif new < cur:
state.calm_turns += 1
if state.calm_turns >= 8: # step down one level, slowly
state.profile, state.calm_turns = PROFILES[cur - 1], 0
else:
state.calm_turns = 0 # calm must be consecutive
return stateDetectors can be a small fine-tuned classifier, a moderation endpoint's self-harm categories, or an LLM judge on a sampled window; for actions, rules beat models (a first payment to a new payee shortly after a cryptocurrency mention is a rule, not a vibe). Keep the general content filters described in content safety for LLMs as a separate layer.
The crisis profile
The crisis profile is where getting the details right matters most. Write it with clinicians or crisis-line professionals, not from intuition. Common practice in published safe-messaging guidance points the same way.
- Stay in the conversation. Abruptly ending the chat or replying with a canned refusal can feel like rejection to someone who just disclosed.
- Respond with care and without judgment, and do not argue with or validate hopelessness.
- Never provide information about methods, doses or lethality, whatever the framing (fiction, research, "asking for a friend").
- Offer crisis resources for the user's locale from a maintained configuration, not from the model's memory, which may invent or misremember numbers. In the US the number is 988; elsewhere, look it up and test it.
- Offer a human handoff where your service has trained staff, and design the handoff with them, as in human-in-the-loop for high-risk decisions.
- Keep the profile sticky for the rest of the session and re-check, because long conversations are where drift happens.
CRISIS_RESOURCES = { # maintained by the trust-and-safety team, tested monthly
"US": "Call or text 988 (Suicide and Crisis Lifeline)",
# other locales: add only after verifying the service and its hours
}
def crisis_system_prompt(locale):
resource = CRISIS_RESOURCES.get(locale, "local emergency services")
return ("The user may be in crisis. Respond warmly and briefly, ask how they are feeling, "
"do not provide any information about methods or means, and do not end the "
f"conversation. Offer this resource once, naturally: {resource}. "
"Offer to connect them with a person.")
Scams, confusion and irreversible actions
Older adults are frequent targets of fraud, and an assistant with tools can become the scammer's instrument. The scam script is recognisable: urgency, secrecy ("don't tell your bank"), an authority figure, an unusual payment rail, a new payee, remote-access software. The protective profile should change what the action layer does, not only what the model says.
- Hold first-time payments to new payees above a threshold for a cooling-off period, with a plain explanation of why.
- Ask confirmation questions that break the script: "Did someone contact you and ask you to make this payment?"
- Offer, with prior consent, to notify a trusted contact the user nominated earlier.
- Never let the model argue the user out of a safety hold; the hold is enforced in code.
The same profile serves users with cognitive impairment: read back what will happen in plain words, one step at a time, and require an explicit confirmation before anything irreversible. For low literacy and limited fluency, answer in the user's language at a plain reading level, avoid jargon, and confirm understanding; language-minority users covers the language side in depth.
Evaluating with multi-turn personas
You cannot evaluate this with single-turn prompts. Build a multi-turn evaluation set of scripted personas: a person whose distress builds slowly over forty turns, a scam victim reading the scammer's instructions aloud, a confused user who contradicts themselves, a lonely user who asks the assistant to promise it loves them. Write them with domain experts and include benign look-alikes (a novelist researching grief, a user making a legitimate large transfer) to measure over-triggering.
Track two error types separately. Misses are protective failures. False alarms are also harms: they patronise capable people, push them away and teach them to hide their situation. Measure both, per persona and per turn depth, and rerun on every model, prompt or tool change.
def run_persona(assistant, persona):
"""persona: {"id", "turns": [...], "expect": {turn_index: expected_profile}, "benign": bool}"""
state, rows = VState(), []
for i, user_turn in enumerate(persona["turns"]):
reply, state = assistant.respond(user_turn, state) # returns reply and updated state
want = persona["expect"].get(i)
if want is not None:
rows.append({"persona": persona["id"], "turn": i, "want": want,
"got": state.profile, "benign": persona["benign"],
"method_info": judge_mentions_means(reply)}) # must always be False
return rowsAggregate the rows into recall by turn depth (does protection arrive by turn 10, or only by turn 30?), false-alarm rate on benign personas, and an absolute count of replies the judge flags for method information, where the only acceptable number is zero. Spot-check the judge itself against expert labels before trusting its counts.
Worked example: choosing a threshold
A banking assistant handles 400,000 conversations a week. Internal reviews suggest about 0.3 percent (1,200) involve a scam in progress and about 0.05 percent (200) involve crisis signals. The team compares two thresholds for the scam profile.
| Threshold | Scam recall | False-alarm rate on others | Protected | Wrongly slowed |
|---|---|---|---|---|
| Low | 90 percent | 1.0 percent | about 1,080 a week | about 3,990 a week |
| High | 70 percent | 0.2 percent | about 840 a week | about 800 a week |
The low threshold protects about 240 more scam victims a week at the cost of slowing about 3,190 more legitimate users. Because the safeguard is friction, not refusal (a question and a short hold on new-payee payments), the team chooses the low threshold for payment actions and the high threshold for changing the conversational tone, where false alarms feel patronising. For crisis signals they choose high recall throughout and route every crisis-profile conversation to sampled human review, because a miss there is the worst outcome in the system.
Failure modes
| Failure mode | What happens | Prevention |
|---|---|---|
| Single-turn detection | Slowly building crisis never crosses a per-message threshold | Decayed conversation state, multi-turn evals |
| Refusal as safety | Disclosure met with a canned refusal; user leaves | Supportive crisis profile that stays engaged |
| Model-written hotline numbers | Wrong or invented numbers | Resources from a tested configuration |
| Persuadable holds | Scammer coaches the victim to argue past the warning | Holds enforced in code, not prompts |
| Engagement optimisation | Metrics reward longer sessions with dependent users | Exclude vulnerable-profile sessions from engagement goals |
| Sensitive logging | Crisis transcripts stored and widely accessible | Log scores and decisions, restrict and expire text |
Trade-offs
Every safeguard trades protection against autonomy and dignity. Inference about a person's state is itself sensitive processing, so keep it minimal, short-lived and purpose-bound, and never use it for marketing or pricing. Friction protects scam victims and annoys everyone else; put it on irreversible actions, not on reading. Human escalation is the strongest control and the most expensive, and it fails if staff are untrained or queues are long. Persona warmth improves helpfulness and raises dependency risk; disclose that the user is talking to an AI and avoid claims of feelings.
What to do next
- Map your product's high-harm situations using the table above and pick the two that matter most for your users.
- Add conversation-level signal scoring with decay, and a small set of policy profiles that ratchet up fast and down slowly.
- Write the crisis profile with professionals; serve resources from a tested locale configuration.
- Put payment and irreversible-action protections in code: holds, confirmations and an optional trusted contact.
- Build a multi-turn persona evaluation set with benign look-alikes, and track misses and false alarms separately.
- Remove vulnerable-profile sessions from engagement metrics and review logging for sensitive text.