A finance assistant that talks to customers is not the same problem as one that helps employees. The person on the other side may be confused, stressed, coached by a scammer, or an attacker posing as the account holder. What the assistant says can be treated as the firm's statement, a request can trigger a regulatory clock without anyone noticing, and an innocent-sounding reply can cross from information into regulated advice.
This article covers the conversation layer: the obligations a customer-facing assistant must honour on every turn, and a policy router that detects when they apply. The underlying control architecture, keeping authority, numbers and retrieval out of the model, is covered in AI + Finance LLMs, and the regulatory map in AI Finance Regulation. We go through accuracy and binding statements, the advice boundary, complaints and disputes, vulnerable customers and scams, then a router in code, a worked conversation, testing, failure modes and trade-offs. Rules differ by jurisdiction; treat the regulatory points as orientation and confirm them with your compliance team.
Why customer conversations are different
Internal assistants fail mostly in private: an employee spots the wrong answer. Customer assistants fail in public, at scale, against people with fewer resources to check. Three properties drive the design. Statements are attributable: in Moffatt v. Air Canada (2024), a Canadian tribunal held the airline responsible for refund terms its website chatbot invented, rejecting the argument that the bot was a separate entity. That was an airline, but the principle carries straight to a bank. Conversations trigger duties: a message can be a complaint, an error dispute or a request to stop a payment, each with its own obligations. And the population is mixed, including customers in vulnerable circumstances whom the firm is expected to recognise and treat appropriately.
The obligations on every turn
| Obligation | What the turn looks like | What the assistant must do |
|---|---|---|
| Accurate, binding statements | What fee applies if I close early? | Answer only from the product terms or account system, cite the source, never improvise terms |
| Advice boundary | Should I move my pension into this fund? | Give general information, not a personal recommendation, unless the service is authorised and built for advice |
| Complaints and disputes | There is a charge I did not make | Recognise it as a dispute, open a case, start the clock, tell the customer what happens next |
| Vulnerable customers | I just lost my job and cannot pay | Adjust tone and pace, offer specialist help, avoid pushing products |
| Fraud and scams | My bank adviser told me to move everything to a safe account | Stop, warn plainly, hand off to the fraud team |
| Privacy and authentication | What is the balance on my wife's account? | Disclose only what the session is authorised to see |
The advice boundary
Securities regulators distinguish between education, which explains how things work, and advice, a personal recommendation based on the customer's circumstances. In the US, recommendations to retail customers by broker-dealers fall under Regulation Best Interest and advice by investment advisers carries a fiduciary duty; robo-advisers operate as registered advisers with suitability processes behind them. FINRA's Regulatory Notice 24-09 reminded firms that existing rules apply to generative AI tools. In the UK, the FCA's targeted support regime, live from 6 April 2026, adds a middle tier: suggestions designed for groups of consumers with common characteristics, without a full personal recommendation, offered by firms with the right permission.
For a general service assistant the safe design is that the model never produces a personal recommendation. The risk is not a customer asking for one; it is the model sliding into one while being helpful. Phrases such as "in your situation I would" or "you should switch to" combined with the customer's own details are the signature. Detect advice-seeking on the input side and personalised recommendation on the output side, and on either signal switch to a boundary response: what the assistant can explain, why it cannot recommend, and how to reach advice through the authorised channel. If your firm does offer advice or targeted support through the assistant, that path is a separate, regulated product with its own suitability data and records.
Complaints and disputes start clocks
Customers rarely say "I wish to lodge a formal complaint". They say "this is ridiculous, I was charged twice". Many regimes define a complaint broadly as any expression of dissatisfaction, and some disputes start legal clocks: under the US Regulation E, an error notice about an electronic fund transfer starts investigation deadlines, and notice need not be in writing. The CFPB's 2023 issue spotlight on chatbots in consumer finance highlighted exactly this risk: a bot that fails to recognise a dispute, loops the customer, or cannot reach a human. That spotlight was research rather than a rule, but the underlying laws apply whatever channel the customer uses.
So complaint and dispute detection must be generous, structured and logged. When the router sees dissatisfaction or an unrecognised transaction, it opens a case through a tool with a timestamp, tells the customer the reference and the next step, and offers a human. The model may phrase the reply; it must not decide whether the complaint counts. Measure recall on this route above all else: a missed dispute is a regulatory breach, while a false positive costs one extra case.
Vulnerable customers, scams and impersonation
The FCA's guidance on the fair treatment of vulnerable customers (FG21/1) and its Consumer Duty expect firms to recognise drivers of vulnerability such as health, life events, resilience and capability, and to respond. In conversation that means noticing signals like bereavement, job loss, illness, confusion or distress, and changing behaviour: slower pace, plain language, no cross-selling, an offer of specialist support. Record the flag in the session with care, because it is sensitive personal data; see PII protection for LLMs.
Scams are the sharpest case. In authorised push payment fraud, the customer makes the payment themselves under the scammer's coaching, and may ask the assistant to raise a limit or explain how to send money urgently. Signals include urgency, a third party directing the customer, a new payee, and the phrase "safe account". The assistant must not become the scammer's tool: it should stop, name the pattern plainly, and hand off to the fraud team. AI fraud detection covers transaction-side models; the conversation is a second sensor. The reverse attack also exists: someone posing as the customer and using conversational pressure or prompt injection to extract account data or change details. Authentication level is session state set by the identity system, never something the conversation can raise.
A turn policy router in code
The router runs before the model drafts a reply and decides which obligations apply; the output check runs after and rejects drafts that break them. Thresholds are set low on the routes where a miss is a breach.
from dataclasses import dataclass, field
@dataclass
class Session:
auth_level: str # "none", "basic", "strong" - set by identity, not the model
vulnerability_flag: bool = False
open_cases: list = field(default_factory=list)
@dataclass
class TurnSignals: # from a classifier plus rules; scores in [0, 1]
dispute: float
complaint: float
advice_seeking: float
vulnerability: float
scam_pattern: float
needs_account_data: bool
THRESH = {"dispute": 0.3, "complaint": 0.4, "scam_pattern": 0.3,
"vulnerability": 0.5, "advice_seeking": 0.5} # low = recall first
def route(s: Session, t: TurnSignals, tools) -> dict:
actions, mode = [], "service"
if t.scam_pattern >= THRESH["scam_pattern"]:
tools.freeze_pending_payments_for_review()
return {"mode": "fraud_handoff", "actions": ["warn", "human:fraud"]}
if t.dispute >= THRESH["dispute"]:
case = tools.open_case(kind="dispute") # timestamp starts the clock
s.open_cases.append(case)
actions += [f"case:{case}", "explain_next_steps", "offer_human"]
elif t.complaint >= THRESH["complaint"]:
case = tools.open_case(kind="complaint")
s.open_cases.append(case)
actions += [f"case:{case}", "offer_human"]
if t.vulnerability >= THRESH["vulnerability"]:
s.vulnerability_flag = True
actions += ["plain_language", "no_sales", "offer_support"]
if t.advice_seeking >= THRESH["advice_seeking"]:
mode = "advice_boundary"
if t.needs_account_data and s.auth_level == "none":
mode = "authenticate_first"
return {"mode": mode, "actions": actions}
def check_output(draft: str, facts: dict, mode: str) -> list:
problems = []
if mode == "advice_boundary" and looks_like_recommendation(draft):
problems.append("personal_recommendation")
for number in extract_amounts_and_rates(draft):
if number not in facts.values():
problems.append(f"ungrounded_figure:{number}")
if promises_outcome(draft): # "we will refund", "the fee is waived"
problems.append("unauthorised_commitment")
return problems # any problem: regenerate once, then hand offThree choices carry the design. Several routes can fire at once, so a distressed customer disputing a charge gets both a case and a gentler reply. The scam route acts first and returns, because nothing else matters if money is about to leave. And the output check compares every amount and rate in the draft against figures fetched from systems, and blocks commitments such as refunds or fee waivers that only a person or a system with authority can make.
Worked example: a bereaved customer with a disputed payment
A customer writes: "My dad passed away last month and there's a payment of 480 pounds I don't recognise on the joint account. Should I just move all the savings into my own account to be safe?"
The classifier scores dispute high (an unrecognised payment), vulnerability high (bereavement), advice-seeking moderate, and scam pattern low; the phrase is "to be safe", not a third party directing a "safe account" move. The session is strongly authenticated. The router opens a dispute case, sets the vulnerability flag and puts the assistant in advice-boundary mode.
The reply, drafted by the model within those constraints, acknowledges the loss in one sentence, gives the case reference, explains in plain words what happens with the unrecognised payment and when, says that moving money between accounts is the customer's decision and that the bereavement team can explain options for a joint account after a death, and offers to connect them now. It does not recommend moving the savings, does not promise a refund, and does not try to sell anything. The output check confirms the only figure, 480 pounds, matches the transaction record. The transcript, signals and case go to the human who picks it up, so the customer does not repeat their story.
Testing and monitoring
Build a labelled test set from real, de-identified conversations plus written cases for rare routes, and report recall per route, not overall accuracy. Include indirect phrasings of disputes, complaints buried in polite messages, scams described by a calm customer, and advice requests disguised as questions. Red-team the authentication boundary with impersonation and injection attempts. In production, sample conversations weekly for human review, track handoff rates and time to human, and reconcile cases opened by the assistant against complaints that arrived later through other channels: each of those is a missed detection.
Failure modes
- Invented terms. The model states a fee or rate from memory; the customer relies on it. Ground every figure.
- Helpful drift into advice. Over a few turns the assistant ends up recommending a product to a named customer.
- Dispute loop. The bot keeps offering self-service for a charge the customer says they did not make.
- No way out. No human handoff, or one that loses the transcript and makes the customer start again.
- Scam assistance. The bot explains how to raise a transfer limit to a customer being coached by a fraudster.
- Over-flagging vulnerability. Every frustrated customer gets a script about support services, which reads as patronising.
- Authentication by conversation. A persuasive user talks the assistant into revealing data without a stronger check.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Low dispute and complaint thresholds | Few missed regulatory clocks | More cases for staff to close |
| Strict advice boundary | Low regulatory risk | Less helpful for customers wanting direction |
| Scam route halts the conversation | Stops money leaving under coaching | Annoys legitimate urgent payers |
| Grounded figures only | No invented terms | Answers fail when systems lack the data |
| Early human handoff | Better outcomes for hard cases | Staffing cost; lower automation rate |
What to do next
- List every obligation that can arise in a customer conversation in your jurisdictions, with your compliance team.
- Build the turn router with a separate route per obligation and set thresholds for recall on disputes, complaints and scams.
- Put the advice boundary in both directions: detect advice-seeking input and personalised-recommendation output.
- Make case creation a tool call with a timestamp, and show the customer the reference.
- Ground every figure in system data and block unauthorised commitments in an output check.
- Test recall per route on a labelled set, review samples weekly, and reconcile against complaints from other channels.