In the 2023 round of the OECD Survey of Adult Skills (PIAAC), 28 percent of US adults scored at or below Level 1 in literacy, up from 19 percent in 2017. Adults at that level can usually read short, simple texts and find a single piece of information, but struggle with long, dense or multi-step documents. Add people reading in a second language, people under stress, people with cognitive or visual impairments, and older adults new to digital services, and a large share of any consumer AI product's users will not read its answers, warnings or consent screens the way the designers assumed.
That is a security problem, not only an accessibility one. Controls that depend on the user reading carefully, such as disclaimers, confirmation dialogs, cited sources and privacy notices, quietly fail for these users, and attackers know it. This article builds a threat model for LLM products serving low-literacy users, then turns each threat into an engineering control with code: a secret guard, a readability gate on outputs, read-back confirmation for risky actions, plain-language consent with a comprehension check, and an evaluation plan that uses real users rather than formulas alone.
A threat model for low-literacy users
| Threat | Why low literacy amplifies it | Control |
|---|---|---|
| Fluent but wrong answers | The user cannot check the cited source or spot a hedge buried in paragraph three | Short answers, uncertainty stated first, human handoff |
| Unread consent and disclosures | A wall of legal text is accepted without understanding | Layered plain-language consent and a teach-back question |
| Scams using AI | Urgent, official-sounding messages and cloned voices are hard to evaluate | Scam check on forwarded content; stated invariants |
| Risky agent actions | An ambiguous request is executed and the confirmation screen is not read | Action tiers with spoken read-back |
| Injected instructions in forwarded content | Users paste letters and screenshots they could not read themselves | Treat forwarded content as data, never as instructions |
| Oversharing secrets | Users paste one-time codes, PINs or ID numbers to ask what they mean | Detect and redact before the model and logs see them |
The fifth row is the one most teams miss. A user who struggles with a letter photographs it and asks the assistant to explain it. Whatever text is in that letter is now in the prompt, so this is exactly the indirect prompt injection problem, aimed at the users least able to notice the assistant behaving oddly. Low-literacy users are not less intelligent; they are more dependent on the system telling the truth in a form they can use, which raises the cost of every failure.
Architecture: controls around the model
Everything here sits around the model, not inside it. That is deliberate: you can test a readability gate or a confirmation rule deterministically, while a system prompt that says use simple words is only a request.
Keep secrets out of the conversation
Start with what must never reach the model or the logs. Users who cannot read a bank message will paste it whole, one-time code included. A guard that recognises those patterns, redacts them and explains why protects the user even if the conversation is later compromised.
import re
PATTERNS = {
"one_time_code": re.compile(r"\b(code|otp|passcode|verification)\D{0,20}(\d{4,8})\b", re.I),
"card_number": re.compile(r"\b(?:\d[ -]?){13,19}\b"),
"us_ssn": re.compile(r"\b\d{3}-\d{2}-\d{4}\b"),
}
def guard_secrets(text: str):
found = []
for name, rx in PATTERNS.items():
if rx.search(text):
found.append(name)
text = rx.sub("[removed]", text)
return text, found
NOTICE = ("I removed a secret number from your message. "
"Never share this number with anyone, even me. "
"Your bank will never ask you for it.")The patterns are a starting point, not a complete detector; card numbers in particular should be confirmed with a Luhn check to cut false positives. The notice matters as much as the redaction: it states an invariant, never share the code, in words a Level 1 reader can act on.
A readability gate on every answer
The Flesch-Kincaid grade level estimates the US school grade needed to read a text: 0.39 x (words / sentences) + 11.8 x (syllables / words) - 15.59. It only measures surface features, sentence length and word length, but that makes it cheap, deterministic and good at catching the most common failure, which is the model drifting back into long, abstract sentences.
import re
VOWEL_GROUPS = re.compile(r"[aeiouy]+")
def syllables(word: str) -> int:
w = word.lower().strip(".,;:!?\"'()")
n = len(VOWEL_GROUPS.findall(w))
if w.endswith("e") and not w.endswith(("le", "ee")) and n > 1:
n -= 1 # silent final e
return max(n, 1)
def fk_grade(text: str) -> float:
sentences = max(len(re.findall(r"[.!?]+", text)), 1)
words = re.findall(r"[A-Za-z][A-Za-z'-]*", text)
if not words:
return 0.0
syl = sum(syllables(w) for w in words)
return 0.39 * len(words) / sentences + 11.8 * syl / len(words) - 15.59
def gated_answer(llm, question, must_keep, target=6.0, attempts=2):
original = answer = llm(question)
for _ in range(attempts):
if fk_grade(answer) <= target:
break
answer = llm(f"Rewrite for a reader at grade {target:.0f}. Short sentences. "
f"One idea per sentence. Keep these facts exactly: {must_keep}\n\n{answer}")
missing = [f for f in must_keep if f not in answer]
if missing: # simpler but wrong is worse than hard
answer = original
return answer, fk_grade(answer)Worked example. Compare two messages that say the same thing. First: Your transaction could not be authorized because of a verification failure at the issuing institution. That is 15 words, one sentence and 30 syllables, so the grade is 0.39 x 15 + 11.8 x 2.0 - 15.59 = 13.86, college level. Second: Your payment did not go through. Call the number on the back of your card. That is 15 words, two sentences and 17 syllables: 0.39 x 7.5 + 11.8 x 1.133 - 15.59 = 0.71. The code above produces both figures. Its heuristic counts two words wrongly in the first message, giving four syllables to authorized and two to issuing, and the errors happen to cancel; expect that kind of noise and set the target with margin.
The second message is not just easier, it is safer: it names one action, and the action is one the user can verify independently. Notice also the must-keep check. A rewrite that drops a dosage, a deadline or an amount is worse than a hard answer, so the gate refuses it. Do not apply this English formula to other languages; use a language-specific measure or human review there.
Read-back instead of confirmation dialogs
Confirmation dialogs assume reading. Replace them, for anything that moves money or data, with a read-back: the assistant restates the action in one short sentence, reads amounts and names in a form that is hard to mishear, and requires an explicit yes.
import time
TIERS = {"check_balance": "read", "explain_letter": "read",
"pay_bill": "money", "send_money": "money",
"share_document": "data", "change_phone_number": "account"}
def say_amount(cents: int) -> str:
dollars, c = divmod(cents, 100)
return f"{dollars} dollars and {c} cents"
def confirm(action: str, params: dict, ask) -> bool:
tier = TIERS.get(action, "account") # unknown actions get the strictest tier
if tier == "read":
return True
if action in ("send_money", "pay_bill"):
line = f"You want to send {say_amount(params['cents'])} to {params['payee_name']}."
if params.get("first_time_payee"):
line += " You have never sent money to this person before."
else:
line = f"You want to {action.replace('_', ' ')}."
def said_yes(reply: str) -> bool:
return reply.strip().lower() in ("yes", "yes please", "correct")
if not said_yes(ask(line + " Is that right? Say yes or no.")):
return False
if tier == "account" or params.get("first_time_payee"):
started = time.monotonic() # enforced cooling-off, not a request
reply = ask("For your safety, please wait one minute, then say yes again.")
return time.monotonic() - started >= 60 and said_yes(reply)
return TrueTwo rules make this robust. Actions are only ever proposed from the user's own turns: if the request came from a forwarded letter or message, the assistant may explain it but may not act on it. And the read-back describes the action in the user's terms, a name and an amount, never an account number or an internal identifier.
Consent people can actually give
A consent screen nobody can read is not consent. Rewrite it in layers: a first layer of three or four short sentences covering what is collected, who sees it, and how to stop, with the full policy one tap away. Read it aloud on request. Then ask a single teach-back question, for example who can read your messages, and record the answer with the consent version. A wrong answer triggers a simpler explanation, not a block.
Disclosure that the user is talking to an AI belongs in the same plain register; the AI disclosure norms and consent flows for AI data articles cover the legal and product side. In the US, the Plain Writing Act of 2010 already requires federal agencies to write public documents clearly, which makes the federal plain language guidelines a useful style source.
Scams, voices and stated invariants
Low-literacy users often rely on voice and on the apparent authority of a message. Cloned voices, covered in deepfakes and synthetic media, and chatbots posing as banks or agencies exploit that. Three controls help. State invariants early and often, in one line: we will never ask for your code, your PIN or gift cards. Run a scam check on forwarded content that looks for urgency, payment by gift card, cryptocurrency or wire transfer, threats of arrest or account closure, and requests for codes, and explain a positive result in plain words: this message wants you to pay with gift cards; real companies do not ask for that. And offer a human handoff that is easy to reach by voice, because the user who is most at risk is also the one least able to navigate a menu.
Evaluate with people, not only formulas
Formulas tell you whether text is short; only people tell you whether it works. Recruit testers through adult education programmes and community organisations, pay them, and test tasks, not pages: can the user correctly say what a letter asks them to do, refuse a staged scam message, and cancel a payment they did not mean to make? Track task success, comprehension answers, read-back abandonment and handoff rate per user group, and compare them with your general population. Re-run the panel when you change model or prompt, because a model upgrade can quietly raise the reading level of every answer.
Failure modes
- Simplification that changes meaning. The rewrite drops a deadline or reverses a condition. Use must-keep facts and spot-check with humans.
- Formula gaming. Short sentences full of jargon still score well. Pair the grade with a list of banned terms for your domain.
- Read-back fatigue. Confirming every action trains users to say yes. Keep read-backs for money, data and account changes.
- Acting on forwarded content. The assistant pays the bill described in a fraudulent letter. Only the user's own request can trigger an action.
- Secrets in logs. Redaction runs after logging. Put the guard before every sink.
- Condescension. Users notice being talked down to and stop using the help. Plain is not childish; test tone with the panel.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Grade target of 6 | Most adults can read the answer | Some nuance moves to a follow-up |
| Rewrite pass | Consistent register | Extra model call and latency on long answers |
| Spoken read-back | Errors caught before money moves | Slower flows; abandonment to monitor |
| Cooling-off for first-time payees | Interrupts scam pressure | Friction for legitimate urgent payments |
| Aggressive secret redaction | Codes never stored | Occasional false positives on order numbers |
What to do next
- Add the secret guard in front of the model and every log sink, and show the notice when it fires.
- Measure the grade level of a week of real answers, then set a target and a must-keep list for your highest-stakes intents.
- Classify every action into tiers and replace confirmation dialogs with spoken read-backs for money, data and account changes.
- Mark forwarded letters, screenshots and messages as untrusted and block actions that originate from them.
- Rewrite consent as a layered, plain first screen with one teach-back question, and version it.
- Recruit a paid low-literacy test panel and track task success and scam refusal before every model or prompt change.