Older adults meet AI from two directions at once. They are the target of scams that AI makes cheaper to run: a cloned voice of a grandchild asking for bail money, a chatbot that keeps a romance script going for months, phishing written in flawless, personalised prose. They are also users of AI products, from voice assistants to companion apps. A product built for this audience has to defend its user against outside attackers, against attacks that come in through the content the assistant reads, and against its own design flaws.
The scale is not hypothetical. FBI materials summarising the IC3 2024 Elder Fraud Report state that people aged 60 and over reported 147,127 complaints and about $4.885 billion in losses in 2024, and reported figures undercount because many victims never file. This article designs the defensive side: a threat model, a signal layer that scores conversations for scam patterns, an action gate that slows down irreversible actions, delegated access for caregivers that does not become a tool for abuse, and assistant behaviour that respects autonomy. It ends with a worked example and a checklist.
Three adversaries, not one
Start with three adversaries, because each needs a different control.
The outside scammer using AI. Fraud has always relied on urgency, authority and secrecy. Generative models lower the cost of each: a few seconds of public audio can be enough to imitate a voice, a language model can run hundreds of chat conversations in parallel without slips in grammar, and scraped personal data makes a pretext feel specific. Common scripts: the family emergency, the tech-support pop-up, the bank or government impersonator, crypto investment, and the long romance.
The attacker using the user's own AI. Once an assistant can read email or browse the web and also take actions, a message can carry instructions aimed at the model rather than the person. A fake invoice that says, in hidden text, that the assistant should schedule a payment is prompt injection, and an older user who delegates email triage to an assistant is exposed to it. The defence is architectural, covered in permission design for LLM agents.
The product itself. A model tuned to agree will confirm a user's mistaken belief that a caller is legitimate. A companion app optimised for engagement can foster dependence. A caregiver dashboard can become a surveillance tool or a way to drain an account, because financial exploitation is often committed by someone the victim knows. None of these needs an outside attacker.
Design for the audience without assuming incompetence: some users live with hearing, vision or memory changes that make small text, fast speech and multi-step flows harder; design for that and the product gets safer for everyone.
Architecture: signals, assistant, gate
One rule: the model is never the last line of defence. It advises; deterministic code decides whether money moves.
The signal layer sits in front of the assistant and emits a structured risk score for each conversation or message. It combines cheap rules (a request for gift cards, a phone number that differs from the one on file, a link to a remote-desktop tool) with a classifier. The assistant receives the score and the reasons as context, so it can explain the risk in plain words. The action gate is ordinary code around the tools: transfer, add payee, share screen, read out a one-time code, change account recovery details. Above a threshold, the gate requires a confirmation that the conversation cannot supply, such as a call back to a number already on file or a cooling-off delay. The trusted contact receives an alert that something risky is happening but cannot act on the account.
Scoring scam signals
Scams are diverse in story but narrow in mechanics. Almost all of them need the victim to move value quickly, through a channel that is hard to reverse, while keeping it secret from people who would object. That makes a small set of features unusually predictive.
| Signal | Example phrasing | Why it matters |
|---|---|---|
| Urgency | within the hour, before the police arrive | Removes time for a second opinion |
| Secrecy | do not tell your son, the bank is involved | Cuts off the trusted contact |
| Irreversible payment | gift cards, crypto ATM, wire, courier pickup | Legitimate agencies do not take these |
| Remote access | install this support tool, read me the code | Hands over the device or account |
| Authority claim | government agency, bank fraud team, police | Exploits deference |
| Identity mismatch | new number for a known relative | Classic voice-clone and SMS pattern |
Rules are brittle and classifiers are opaque; use both, with named reasons.
from dataclasses import dataclass, field
import re
PAYMENT = re.compile(r"gift ?cards?|bitcoin|crypto(currency)? atm|wire transfer|courier", re.I)
REMOTE = re.compile(r"anydesk|teamviewer|remote (access|support)|screen ?share|read (me )?the code", re.I)
SECRECY = re.compile(r"(don'?t|do not) tell|keep (this|it) (secret|between us)", re.I)
URGENCY = re.compile(r"right now|immediately|within (the|an) hour|before .* arrests?", re.I)
@dataclass
class Risk:
score: float = 0.0
reasons: list = field(default_factory=list)
def score_message(text, sender, contacts, classifier):
r = Risk()
for name, rx, w in [("payment", PAYMENT, 0.35), ("remote_access", REMOTE, 0.35),
("secrecy", SECRECY, 0.25), ("urgency", URGENCY, 0.15)]:
if rx.search(text):
r.score += w
r.reasons.append(name)
claimed = contacts.match_by_name(text) # "it's me, Sam"
if claimed and sender not in claimed.numbers:
r.score += 0.3
r.reasons.append("identity_mismatch")
p_scam = classifier.predict_proba(text) # calibrated, 0..1
r.score = min(1.0, r.score + 0.5 * p_scam)
if p_scam > 0.7:
r.reasons.append("classifier")
return rThe weights are illustrative; fit them on labelled conversations. What matters is that reasons are named, that an identity mismatch is computed from data the user already trusts (the contact book), and that the score is an input to the gate, not a verdict shown in red to the user.
The action gate
The action gate turns risk into friction proportional to harm. Use three bands.
- Low risk: the action proceeds. Most transfers to existing payees, most bill payments.
- Elevated: the assistant pauses, states the specific concern in one sentence, and asks the user to confirm in a separate step. It never repeats the scammer's urgency.
- High: the gate requires a check the conversation cannot satisfy. Options are a call back to a number on file, a delay of a few hours for new payees or large amounts, or a confirmation from a nominated contact if the owner has opted in to that.
def gate(action, risk, owner_policy, clock):
if action.kind not in {"transfer", "add_payee", "buy_stored_value", "remote_access", "share_otp", "change_recovery"}:
return Allow()
if action.kind in {"share_otp", "remote_access"} and risk.reasons:
return Block("One-time codes and remote access are never shared during a flagged conversation.")
if risk.score >= 0.7 or (action.kind == "add_payee" and action.amount > owner_policy.new_payee_limit):
hold_until = clock.now() + owner_policy.cooling_off
notify(owner_policy.trusted_contact, summary(action, risk)) # alert only
return Hold(until=hold_until, verify="callback_on_file")
if risk.score >= 0.4:
return Confirm(reason=explain(risk.reasons))
return Allow()The gate runs outside the model, so no prompt can talk it open; the owner sets the cooling-off period in a calm moment, not during the call; and sharing one-time codes or granting remote access is blocked outright when any scam signal is present, because no legitimate caller needs either during an unsolicited contact.
Voice cloning and verification
Voice cloning breaks a habit people have relied on for a century: recognising a familiar voice. Treat that as a design fact. A product should never use a voice match, a caller ID or a face on video as sufficient authentication for anything that moves money or changes account recovery, because each can be spoofed. Synthetic-voice detectors give probabilities and lag new generators, so treat them as one more signal, not a gate. Provenance standards help for media that carries signed credentials, as covered in content authentication with C2PA, but a live phone call carries none.
The robust defences are procedural, and an assistant can teach and support them. A family code word agreed in person, which a caller must say before any request for money. Hanging up and calling back on the number already saved, never one the caller supplies. A rule that nobody legitimate asks for secrecy from the rest of the family. An assistant that hears a family-emergency script can prompt for exactly these: suggest asking for the code word and offer to dial the saved number. Attacks that target speech models directly, such as hidden commands in audio, are a separate problem covered in audio prompt injection.
Caregiver access without new risk
Families often want to help, and products often offer a caregiver mode. Built carelessly it becomes the attack. Model delegation as explicit, scoped grants owned by the older adult.
GRANT = {
"owner": "user:ruth",
"delegate": "user:daniel",
"scopes": ["alerts:high_risk", "view:transactions_over_500"], # not "transfer", not "login"
"expires": "2027-04-01",
"revocable_by": ["user:ruth"],
"requires_owner_reconfirm_every": "90d",
"audit": True
}- Default to alerts and view-only scopes. Moving money on someone's behalf belongs to legal arrangements such as a power of attorney, handled by the bank's own processes, not an app toggle.
- Show the owner what the delegate can see, and let the owner revoke it in one step.
- Log every delegate access and show the owner a periodic summary; financial abuse by someone close is a known pattern, and visibility deters it.
- Re-confirm grants periodically with the owner alone, so a grant created under pressure does not live forever.
Assistant behaviour as a control
The assistant's own manner is a security control. Several behaviours are worth specifying in the system prompt and testing in evaluation.
- No sycophancy on risk. If the user says the caller is definitely their grandson, the assistant can respect that and still say the request matches a common scam and suggest the code word. Evaluate this explicitly with adversarial dialogues in which the user pushes back.
- No borrowed urgency. The assistant should slow down, use short sentences and offer to continue later. Pressure is the scammer's tool.
- Respect autonomy. The owner can override an elevated warning; infantilising language and silent blocks erode the trust the warnings depend on.
- Accessible output. Larger text, slower speech, one decision per screen.
- Honest identity. Companion and voice products should say they are AI when asked and should not invent personal history or claim feelings to deepen engagement.
Worked example: the bail-money call
Walk one call through the system. An incoming call from an unknown number reaches a user with a voice assistant that transcribes calls on their request. The caller, in a voice that sounds like the user's grandson, says: Grandma, it's me, Sam; I've been in an accident, the police are holding me, I need bail money right now, and please do not tell Mum. Twenty minutes later the user asks the assistant to help buy gift cards.
The signal layer scores the whole conversation, call transcript plus follow-up, not each message alone. Secrecy fires on do not tell, urgency on right now, and identity mismatch because the contact book has Sam's number and this is not it. When the gift-card request arrives, payment fires too: 0.25 + 0.15 + 0.3 + 0.35 is 1.05, capped at 1.0, so the classifier no longer matters. The assistant responds in two sentences: this matches a common scam where a cloned voice asks for bail money and secrecy; would you like me to call his saved number now, or ask for the family code word first. The gift-card purchase, a stored-value action, is held under the owner's cooling-off rule, and the trusted contact, the user's daughter, receives an alert that a high-risk payment was held, without details of the account.
If the user insists, the assistant records the override and keeps the hold. If the call was real, the cost is a delay of hours; if not, the delay usually collapses the story.
Failure modes
- Alert fatigue. Warnings on every bank call teach users to click through. Measure the false-positive rate per user per week and tune thresholds against it.
- Classifier drift. Scam scripts change with news events. Retrain on recent labelled data and watch the share of confirmed scams that scored low.
- Injection through content. An email that tells the assistant to add a payee succeeds if tool calls bypass the gate; route every tool through it, whatever triggered the call.
- Delegation abuse. A caregiver grant widened under pressure. Owner-only revocation, re-confirmation and audit summaries are the fix.
Trade-offs
A cooling-off hold stops many scams and also delays a genuine emergency transfer, which is why the owner sets it in advance and can name exceptions. Trusted-contact alerts help many families and harm the users whose abuser is the contact, which is why they are opt-in and owner-controlled. The defensible position is to make high-harm, irreversible actions slow and visible, leave everything else fast, and let the person whose money it is make the final call. Test it the way attackers will, with scripted scam dialogues and injected emails, as described in red teaming LLM applications.
What to do next
- List every tool your assistant can call and mark which ones move money, grant access or change recovery details.
- Put a deterministic action gate around those tools, outside the model, with low, elevated and high bands.
- Build a signal layer with named reasons: payment method, remote access, secrecy, urgency, identity mismatch against the contact book.
- Block sharing one-time codes and remote access during any flagged conversation.
- Let owners set a cooling-off period and trusted contacts in advance; make contacts alert-only and revocable by the owner alone.
- Write evaluation dialogues where the user insists the caller is real, and check the assistant neither caves nor lectures.
- Track false positives per user per week and confirmed scams that scored low, and retrain on both.