Governments are putting language-model assistants in front of the public: help with city business permits, tax helpline chat, benefits and housing questions, 311 service requests. These systems differ from a commercial chatbot in ways that change the security design. Residents cannot pick another provider. An answer carries the authority of the state. A wrong answer about a legal obligation can cost someone a fine, a tenancy or a benefit.
This page is about securing that citizen-facing service layer: the channels, sessions, retrieval, answer checks and handoffs between a resident and an agency's guidance and case systems. Decision systems that approve or deny claims raise different questions, covered in AI in government. Here you get a reference architecture, a threat model, an answer gate you can run, session-scoped tools, records handling, a worked example and an operating checklist.
What went wrong in New York
The failure that defines this area is confident, wrong guidance. In October 2023 New York City launched the MyCity chatbot, built on Microsoft's Azure AI services, to help business owners navigate city rules. On 29 March 2024 The Markup reported that it gave answers contradicting the law. It said employers could take a cut of workers' tips. It said there was no rule requiring notice of schedule changes. It told a landlord they did not have to accept tenants with rental assistance, although the city's own website says source-of-income discrimination has been illegal since 2008, with some exceptions. The mayor acknowledged it was wrong in some areas and called it a pilot.
Nobody had to attack the system for that to happen. The guidance existed, but the model was free to answer from its own general knowledge, and nothing checked the answer against the source before a resident saw it. So the first control is architectural: the assistant may only state what an in-force official passage says, and a deterministic check enforces that. Prompt instructions alone will not.
Reference architecture
The architecture separates three things a generic chatbot mixes together: what the rules are, who the caller is, and what the system may do for them.
- Guidance corpus. Agency-owned passages, each with a stable id, jurisdiction, effective date and superseded date. Content teams publish to it through review, and the model reads it but never writes it.
- Session tier. Anonymous callers get general information only. Signed-in callers, authenticated through the agency's existing login service, can also reach case tools. The model never sees passwords, codes or identity documents.
- Answer gate. Code that runs after generation and before display. It checks that every citation exists, belongs to the right jurisdiction and is in force today, and that every number in the answer appears in a cited passage.
- Handoff. A failed gate, a high-stakes topic (eviction, benefit loss, immigration status, safety) or a request for a human routes to staff, with the transcript attached so the resident does not have to repeat themselves. Design the queue with human-in-the-loop patterns.
- Records store. Transcripts, gate decisions and corpus versions, kept on a retention schedule with personal data redacted where the law allows.
Threat model
| Threat | How it happens here | Primary control |
|---|---|---|
| Wrong rule stated as fact | model answers from training data or a superseded passage | answer gate with effective dates |
| Prompt injection | instructions hidden in an uploaded letter, a pasted web page or a form field | treat uploads as data, scan them, no tools on anonymous tier |
| Account takeover via chat | attacker persuades the assistant to change an address or bank details | no identity changes in chat; agency login plus step-up |
| Cross-resident leakage | a case tool queried with an id the model chose | tools bound to the session subject, never a model argument |
| Corpus tampering | an editor account or a scraped page alters guidance | two-person review, signed corpus versions |
| Look-alike scam bots | fake 'official assistants' collect fees or personal data | one published channel list; the bot never asks for payment in chat |
| Unequal quality | answers degrade in some languages or for some groups | per-language golden sets; group-level metrics |
Injection matters most where the assistant reads text residents supply, such as photographed letters and pasted notices. Run inputs through a detector (see prompt injection scanners), but rely on the structural rule: on the anonymous tier there are no tools to hijack, and on the signed-in tier every tool is limited to the caller's own case.
The answer gate
The gate is ordinary code with no model in it, so it can be tested, reviewed and reasoned about. The model returns its answer plus the passage ids it relied on. The gate rejects the answer if a citation is unknown, belongs to another jurisdiction or is not in force today, or if the answer contains a number that does not appear in the cited text.
import re
from dataclasses import dataclass
from datetime import date
@dataclass(frozen=True)
class Passage:
pid: str
text: str
jurisdiction: str
effective: date
superseded: date | None = None
def live(p, today):
return p.effective <= today and (p.superseded is None or today < p.superseded)
NUM = re.compile(r"[$]?[0-9][0-9,.]*%?")
def gate(answer, cited_ids, corpus, jurisdiction, today):
reasons = []
cited = [corpus.get(i) for i in cited_ids]
if not cited_ids:
reasons.append("no citation")
for i, p in zip(cited_ids, cited):
if p is None:
reasons.append(f"unknown passage {i}")
elif p.jurisdiction != jurisdiction:
reasons.append(f"{i} is {p.jurisdiction} guidance")
elif not live(p, today):
reasons.append(f"{i} not in force on {today}")
support = {n.rstrip(".,") for p in cited if p for n in NUM.findall(p.text)}
for n in NUM.findall(answer):
if n.rstrip(".,") not in support:
reasons.append(f"number {n.rstrip('.,')} not in cited text")
return (not reasons, reasons)The number check is deliberately crude. It will not catch a wrong verb ("may" for "must"), so add an entailment check or a second model as a further layer, and still route high-stakes topics to people. It does catch the commonest damaging error, a wrong fee, deadline or threshold, and it never fails silently: every rejection has a reason that goes into the records store.
Session-scoped tools for signed-in residents
Signed-in residents want actions: check a case status, book an appointment, upload a document. The rule is that the subject of every tool call comes from the authenticated session, not from the model. The model can choose which tool to call. It cannot choose whose record the tool touches.
READ_ONLY = {"case_status", "appointment_slots"}
WRITE = {"book_appointment", "attach_document"}
NEVER_IN_CHAT = {"change_address", "change_bank_details", "close_case"}
def call_tool(session, name, args):
if name in NEVER_IN_CHAT:
return handoff(session, reason=f"{name} requires the account portal")
if session.tier != "signed_in":
raise PermissionError("anonymous sessions have no case tools")
if name in WRITE and not session.recent_step_up():
return ask_step_up(session) # re-authenticate before writes
args = {k: v for k, v in args.items() if k not in ("person_id", "case_id")}
return TOOLS[name](subject=session.subject_id, **args) # subject from session only
# handoff, ask_step_up and TOOLS are your platform's own functions.Changes to identity or payment details stay out of chat entirely, because a convincing conversation is exactly what social engineering produces. Send residents to the account portal, which has its own verification flow.
Records, privacy and accessibility
Assistant transcripts are government records in many jurisdictions. Depending on local law they may be open to freedom-of-information requests, subject to retention schedules, or both. Confirm the rules with records and legal staff before launch, not after the first request arrives. Practical steps that hold up in most places:
- Collect the minimum. The anonymous tier needs no name, address or case number. Tell residents not to paste identifiers, and redact common patterns before storage.
- Store the gate decision, the cited passage ids and the corpus version with each turn. When a resident later says the assistant told them something, you can reconstruct exactly what the assistant saw and said.
- Keep model-provider processing inside the agency's contract terms: data location, no training on resident data, and deletion on request.
- Meet accessibility obligations. Many jurisdictions apply WCAG-based requirements to public digital services, and a chat widget is part of the service. Keep a phone and in-person route for residents who cannot or will not use it.
Worked example: a fee that changed in July
A city's business-permit assistant answers questions about mobile food vending. On 1 July 2026 the annual permit fee rose from $200 to $260. The content team published passage fees/mobile-food#4 with an effective date of 1 July 2026 and set a superseded date on #3. Retrieval still returns the old passage for some phrasings, because it is older and more heavily linked. On 5 October 2026 a vendor asks what the permit costs.
corpus = {
"fees/mobile-food#3": Passage("fees/mobile-food#3",
"The annual mobile food vending permit fee is $200.", "city",
date(2025, 7, 1), superseded=date(2026, 7, 1)),
"fees/mobile-food#4": Passage("fees/mobile-food#4",
"From 1 July 2026 the annual mobile food vending permit fee is $260.", "city",
date(2026, 7, 1)),
}
today = date(2026, 10, 5)
gate("The annual permit fee is $200.", ["fees/mobile-food#3"], corpus, "city", today)
# (False, ['fees/mobile-food#3 not in force on 2026-10-05'])
gate("The annual permit fee is $260.", ["fees/mobile-food#4"], corpus, "city", today)
# (True, [])
gate("The annual permit fee is $250.", ["fees/mobile-food#4"], corpus, "city", today)
# (False, ['number $250 not in cited text'])The first draft cites the superseded passage and is blocked. The system retries with superseded passages filtered out of retrieval, gets the $260 answer, and replies with the source and its effective date. The third call shows a generation error, a fee that appears in no passage, being caught. Over a month the gate log shows 3% of permit answers rejected for stale citations, which tells the content team to fix retrieval filtering rather than rely on the gate forever. The fees, dates and rejection rate are invented for this example. The gate outputs are from running the code.
Evaluation and operations
Evaluate before launch and continuously after. Build a golden set from the agency's top call-centre intents, with answers written by policy staff, in every supported language. Run it on every model, prompt or corpus change, and report results per language. A multilingual assistant that is only tested in English is not tested.
Measure the right outcomes. Deflection rate (questions answered without staff) is the metric vendors sell. Pair it with escalation rate, gate-rejection rate by topic, complaint rate and a sampled accuracy audit. Map the programme to a framework such as the NIST AI RMF so oversight bodies can follow it.
Have a kill switch. When guidance changes suddenly, as with an emergency order or a court ruling, you need to switch a topic to handoff-only within minutes, without a deploy.
Failure modes
- Disclaimers instead of controls. "Answers may be inaccurate" does not help a resident who acted on one. Gate the answer.
- Effective dates missing from the corpus. Without them the gate cannot tell old rules from current ones. Backfill dates before launch.
- Tools that accept ids from the model. One injected instruction becomes a data breach across residents.
- Handoff that drops context. Residents repeat everything to a person, then stop using the assistant.
- Measuring deflection alone. A bot that confidently answers everything wrongly has excellent deflection.
Trade-offs
A strict gate refuses more often, so residents see more "let me connect you" replies and staff take more contacts. A loose gate saves staff time and spends resident trust. Strict is the right default for fees, deadlines, eligibility and legal duties. Looser settings suit opening hours and directions. Anonymous-only assistants are far easier to secure than signed-in ones, so add case tools only when the evidence shows they save residents real effort.
What to do next
- Inventory the guidance the assistant will draw on and give every passage an id, jurisdiction, effective date and owner.
- Implement the answer gate and log every rejection with its reason.
- Define high-stakes topics with policy staff and route them to people by default.
- Bind every case tool to the session subject and keep identity and payment changes out of chat.
- Agree retention, redaction and records-request handling with records and legal staff.
- Build per-language golden sets from real call-centre intents and run them in CI.
- Publish the official channel list so residents can tell the real assistant from imitations.
- Add a per-topic kill switch and test it.