LLM assistants are being deployed in agricultural extension, community health work, government services and mobile money support across low- and middle-income countries. The security engineering for those deployments differs from that of a web app in a well-connected market, and the differences are concrete. Users reach the system over SMS, USSD or a messaging platform instead of an authenticated app. Devices are shared. Connectivity is intermittent, so models may run offline on phones. Languages are often under-served by safety tooling. Data protection law may restrict sending data to a cloud region abroad. Fraud ecosystems are built around phone numbers and mobile wallets.
This page is a threat model and control set for that environment. It is not a policy essay. It covers channel security, signed offline model updates, language-aware moderation routing, data residency, and impersonation and fraud. A worked example follows a health-worker assistant through a realistic incident. The page ends with failure modes and a checklist. The specific gap in safety behavior across languages has its own page, linked below, so here it gets only the architectural treatment.
Architecture and trust boundaries
Start from the assets. They are the users' personal data, often health or financial data; the assistant's reputation as a trusted source; any actions it can trigger, such as a referral, an appointment or a balance query; and the model and content bundles shipped to devices. Then list who sits between the user and the model. The mobile carrier sees SMS and USSD in clear text. A messaging platform hosts the conversation and its metadata. The device may belong to someone else, a family member or a community health worker shared across a village. The host country's regulator decides where data may be processed.
Threat model
| Threat | Where it enters | Primary control |
|---|---|---|
| Impersonation of the official assistant | Look-alike numbers, cloned bot profiles | One published verified channel; the assistant never asks for PINs or one-time codes |
| Account takeover via SIM swap | Phone number used as identity | No sensitive actions on number alone; step-up check through a second factor |
| Disclosure on shared devices | Chat history on a borrowed phone | Short on-device retention, app PIN, no sensitive data in notifications |
| Tampered offline model or content | Side-loaded APKs, mirrored weights | Signed bundles, pinned hashes, safetensors or GGUF rather than pickle |
| Safety bypass in under-served languages | Prompts in languages the filters handle poorly | Language ID routing, per-language evaluation, conservative defaults |
| Unlawful cross-border transfer | Cloud region outside the country | In-country or approved-region hosting; minimized logs |
| Harmful advice at scale | Hallucinated dosages or agronomy | Retrieval over vetted documents, refusal outside scope, human escalation |
Channels: SMS, USSD and messaging bots
SMS and USSD have no end-to-end confidentiality and no strong sender authentication. Treat everything on them as visible to the carrier and possibly spoofable, and keep sensitive content off them. A phone number is an address, not an identity. SIM-swap fraud lets an attacker receive a victim's messages, so a number alone must never authorize anything with consequences. Messaging platforms are better, but they still see metadata, and bot profiles can be cloned.
The gateway enforces three rules. Identifiers are hashed with a keyed hash before they reach logs or the model. Sessions are bound to the channel and expire quickly. Anything sensitive needs a step-up: a code sent through a different channel, or confirmation by a registered health worker or agent.
import hmac, hashlib, time
PEPPER = load_secret("gateway-pepper") # from a secret store, never from code
SESSION_TTL = 15 * 60
SENSITIVE = {"referral", "balance", "record_lookup"}
def subject_id(msisdn: str) -> str:
# keyed hash: logs and the model never see the raw number
return hmac.new(PEPPER, msisdn.encode(), hashlib.sha256).hexdigest()[:24]
def handle(channel: str, msisdn: str, text: str, store):
sid = subject_id(msisdn)
if not store.rate_ok(sid, limit=30, window=3600):
return "Too many messages. Please try again later."
sess = store.session(sid, channel)
if sess is None or time.time() - sess.started > SESSION_TTL:
sess = store.new_session(sid, channel)
intent = classify_intent(text)
if intent in SENSITIVE and not sess.stepped_up:
store.send_step_up(sid, via="registered_agent") # different channel, never the same SMS thread
return "For your safety we need to confirm this request. A health worker will contact you."
return answer(sess, redact_pii(text))Two user-facing rules matter more than any of this code. Publish them everywhere the service is advertised: the assistant has exactly one number or verified profile, and it never asks for a PIN, password or one-time code. Scammers will clone a popular free assistant. A rule users can remember is the control that defeats the clone.
Offline models and signed bundles
Where connectivity is poor, small quantized models run on phones or a clinic laptop, together with a retrieval bundle of approved documents. That moves the model and its content into the supply chain. Weights get copied over Bluetooth, re-shared on memory cards and downloaded from mirrors. Assume someone will try to ship a modified bundle that, for example, recommends a particular product or gives harmful advice.
Sign bundles, and verify them on the device before loading. Ship weights in a format that cannot execute code. Python pickle-based checkpoints can, so prefer safetensors or GGUF, and never load pickle files from an untrusted source. GGUF is not inert either: its metadata can carry a Jinja chat template, and an unsandboxed template renderer was exploitable in llama-cpp-python (CVE-2024-34359). Keep loaders patched, render templates in a sandbox, and cover the template with the signed manifest like any other file. Pin the public key in the app at build time, so a re-signed bundle from someone else fails verification.
import hashlib, json
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PublicKey
from cryptography.exceptions import InvalidSignature
PINNED_KEY = Ed25519PublicKey.from_public_bytes(bytes.fromhex(BUILD_TIME_PUBKEY_HEX))
def verify_bundle(manifest_bytes: bytes, signature: bytes, files: dict, installed_version: int):
try:
PINNED_KEY.verify(signature, manifest_bytes)
except InvalidSignature:
raise RuntimeError("bundle signature invalid: refusing to install")
manifest = json.loads(manifest_bytes)
if manifest["version"] <= installed_version:
raise RuntimeError("rollback or replay of an older bundle")
for name, want in manifest["sha256"].items():
if hashlib.sha256(files[name]).hexdigest() != want:
raise RuntimeError(f"{name} does not match the signed manifest")
return manifestThe version check blocks rollback. Without it, an attacker can replay an older, correctly signed bundle containing guidance that has since been withdrawn. On the device, encrypt conversation history at rest, keep it for days rather than months, and expose a visible clear-history action, because the next person to hold the phone may not be the user. Plan for the offline model being weaker than the hosted one. Give it a narrower scope and more conservative refusal rules.
Language coverage as a security boundary
Safety training and moderation classifiers are strongest in high-resource languages. Research such as Yong, Menghini and Bach's 2023 study on low-resource languages showed that translating harmful requests into under-served languages could bypass a frontier model's safeguards far more often than the English originals. Architecturally, that means three things. Run language identification on every message and log the result. Evaluate refusal and harmful-output rates per language, not in aggregate. And where a classifier is weak for a language, apply it to a machine translation as well as the original, and take the stricter verdict. Code-switching between two languages in one sentence is the norm in many regions, so test it explicitly. The companion page on language minorities covers measurement in depth.
Data protection and residency
Many countries now have comprehensive data protection laws, including Kenya's Data Protection Act (2019), Nigeria's Data Protection Act (2023), India's Digital Personal Data Protection Act (2023), South Africa's POPIA and Brazil's LGPD. They differ in cross-border transfer rules, registration duties and how they treat health data. Confirm the current obligations and any implementing regulations with local counsel, because these regimes are still being operationalized. Several engineering defaults hold across all of them.
- Host inference and logs in-country or in an explicitly approved region. Know which subprocessors, including the model API provider and the messaging platform, see which fields.
- Minimize at the gateway. Redact names, numbers and locations before text reaches the model or the logs. Log hashed subject IDs, not phone numbers.
- Set retention per data class, and enforce it with deletion jobs that are themselves monitored.
- Design consent for low literacy and shared devices. Short voice or local-language explanations work better than linked policy pages, and consent should be re-confirmed when the purpose changes.
Impersonation and mobile-money fraud
Where mobile money is the main financial rail, fraud runs on social engineering by phone. Generative AI makes that cheaper: cloned voices of officials or relatives, fluent scam scripts in local languages, fake support bots. A legitimate assistant connected to payments or benefits becomes both a target and a template for imitation. Controls follow from that. The assistant never initiates a payment conversation. Any money movement is confirmed out of band through the provider's own channel. Agents and health workers get a way to verify that a message really came from the service. Reports of impersonation feed a takedown process with the messaging platform and carriers.
Worked example: a health-worker assistant under attack
Consider a maternal-health assistant for community health workers in a rural district. It answers questions from approved national guidelines over a messaging bot, with an offline Android app for areas without coverage. Here is one plausible incident, traced through the controls.
- A look-alike bot profile starts messaging mothers, offering free supplements in exchange for a mobile money registration fee. Because the program's posters say there is one official profile and it never asks for payment, several health workers report the clone within a day.
- The operations team files the takedown and sends a broadcast through the verified channel. The gateway's logs, keyed by hashed IDs, show no data leaked from the real service, which keeps the incident report factual.
- In the same week, a modified offline bundle circulates on memory cards. On install, its manifest fails signature verification. The app refuses it and reports the failure the next time the device is online, so the team learns how far the bundle has spread.
- Per-language evaluation then shows that refusal rates for dosage questions are lower in one local language. Until the classifier is fixed, the team routes that language's dosage intents to a fixed guideline excerpt plus human escalation.
None of these controls is novel. What the example shows is that each one maps to a property of the environment: clonable channels, physical sharing of models, and uneven language coverage.
Failure modes
- Copying the high-connectivity design. App logins and email resets assume things users may not have. Teams then fall back to phone-number identity without stepping up sensitive actions.
- Logging everything for debugging. Raw transcripts with phone numbers in a foreign cloud region turn a quality tool into a cross-border transfer problem and a breach waiting to happen.
- Aggregate safety metrics. A 99% refusal rate on harmful prompts can hide a much lower rate in the language most users speak.
- Unsigned offline updates. Once a bundle spreads hand to hand, there is no recall without signature and version checks.
- No impersonation plan. A successful free assistant will be cloned. Without a published rule and a takedown path, the clone borrows the real service's trust.
Trade-offs
| Decision | Option A | Option B | Guidance |
|---|---|---|---|
| Hosting | Global cloud region | In-country host | Prefer in-country for health and financial data; verify transfer rules |
| Offline capability | Online only | On-device model | Offline only with signed bundles and a narrower scope |
| Channel | SMS / USSD | Messaging platform or app | Keep SMS for notifications, not sensitive content |
| Moderation | One multilingual classifier | Per-language routing | Route by language where per-language evals show gaps |
What to do next
- Write the threat table for your deployment: channels, devices, intermediaries and the actions the assistant can trigger.
- Put a gateway in front of the model that hashes identifiers, rate-limits, expires sessions and requires a step-up for sensitive intents.
- Publish the one-channel, never-asks-for-PINs rule, and set up a takedown path with your messaging platform and carriers before launch.
- If you ship offline models, sign manifests, pin the key, reject older versions and use non-executable weight formats.
- Add language ID to every log line, and report safety metrics per language and for code-switched input.
- Map data flows against the applicable law with counsel, then minimize, localize and set retention to match.
- Keep learning: measuring safety gaps per language, voice and video impersonation fraud, jailbreak defense architecture and an AI regulation applicability engine.