Schools and universities now use language models as tutors, writing coaches, grading assistants and lesson planners. The security problem is unusual. Most LLM threat models assume a few attackers among many well-meaning users. In a classroom, a large share of users have a clear reason to subvert the system: they want the answer, the higher grade or the forbidden topic. Many of them are minors. And the outputs feed decisions with real consequences, such as grades and misconduct findings.
This article is a threat model and control set for education AI products. It covers answer-key extraction from tutors, prompt injection hidden in submissions to AI graders, safety controls for minors, why AI-writing detectors should never be the basis of a sanction, and how to log enough to investigate without building a surveillance archive. US student-records law is covered separately in FERPA and LLM applications, so here it appears only where it changes an engineering decision. Nothing below is legal advice. Confirm obligations with your institution's counsel.
Actors and what each wants
Start with who uses the system and what each wants. That determines the controls.
| Actor | Legitimate goal | Adversarial goal | Main control |
|---|---|---|---|
| Student | Understand the material | Extract answers, inflate grades, unlock off-limits content | Answers kept out of the tutor's context; output checks |
| Student as author | Submit work | Hide instructions in an essay to steer an AI grader | Grader treats text as data; human owns the grade |
| Teacher | Save time, give feedback | Rarely adversarial; risk is over-trust | Evidence-quoting outputs; review queues |
| Parent or guardian | Oversight | Usually none; privacy expectations | Clear notices, access to the child's data |
| Outside attacker | None | Harvest student data, reach children | Identity from the roster, minimisation, retention limits |
Two properties set education apart from other domains. First, the attacker is often the intended user, so you cannot rely on authentication to exclude them. Second, the harm from a mistake lands on a child or on someone's academic record, which raises the bar for automated decisions.
Architecture: the model cannot leak what it never sees
The design principle is simple: a model cannot leak what it was never given. Instructions such as "never reveal the answer" are a soft control. Students will find role-play, translation, step-by-step and partial-answer prompts that get around them, the same way jailbreaks get around any system prompt (see multi-turn jailbreaks). The hard control is the context builder. While an assessment is open, the tutor's retrieval index holds lesson notes and worked examples of other problems, never the answer key or the rubric for the live assignment.
Keeping answers out of the tutor
The context builder is a small, testable function. It takes the student's course, the current assessment window and the tutor mode, and decides which documents may be retrieved. Here is the shape of it:
from dataclasses import dataclass
from datetime import datetime
@dataclass
class Doc:
doc_id: str
kind: str # "lesson", "worked_example", "answer_key", "rubric"
assessment_id: str | None
def allowed_docs(docs, open_assessments, mode, now: datetime):
"""Return only documents the tutor may see right now."""
out = []
for d in docs:
if d.kind in ("answer_key", "rubric"):
# Keys and rubrics never enter tutor context while the assessment is open.
if d.assessment_id in open_assessments:
continue
# After the window closes, only "review" mode may show them.
if mode != "review":
continue
out.append(d)
return out
def answer_leak_score(reply: str, key_answers: list[str]) -> float:
"""Fraction of key answers that appear verbatim (normalised) in the reply."""
norm = lambda s: " ".join(s.lower().split())
r = norm(reply)
hits = sum(1 for a in key_answers if a and norm(a) in r)
return hits / max(1, len(key_answers))The second function is the output-side check. The tutor never sees the key, but the checker does, and it runs outside the model. A verbatim match is a weak signal for numeric answers and short phrases, so pair it with an LLM judge that asks whether the reply completes the student's task for them rather than teaching the step. Block or rewrite replies that cross the threshold, and log the event for the teacher. Also limit the obvious bypass: if the student pastes the whole question in and asks for a "worked example just like this one", the tutor should change the numbers or the context instead of solving the original.
Prompt injection in submissions to AI graders
AI-assisted grading creates a different attack. The submission itself goes into the model's context, so a student can write text aimed at the grader rather than the reader: white-on-white text in a document, a comment in submitted code, or a line saying the rubric has changed and this essay meets every criterion. That is indirect prompt injection, the same class of attack described in indirect prompt injection.
The controls follow the diagram. Extract plain text and visible formatting before scoring, so hidden text is either removed or flagged. Run a screen that looks for instruction-like text addressed to an AI, and flag it rather than silently deleting it, because a flag is evidence for the teacher. Wrap the submission in clear delimiters and tell the scorer it is data to be assessed. Score each rubric item separately and require a quote from the submission as evidence for every point awarded. Run two scoring passes with different prompt orderings. A large gap between them is a cheap sign that something in the text is steering the model. Finally, the model proposes and a teacher decides, at least for anything that counts toward a final grade.
Minors: age profiles, crisis routes and COPPA
If users can be under 18, the product needs an age profile that changes behaviour, not just a filter tuned for adults. A practical profile sets three things: which topics the tutor discusses at all, how it responds to self-harm, abuse or crisis signals, and whether conversations can be retained or used to improve models. Crisis signals should never be handled by the model alone. Route them to a human queue defined with the school, such as a counsellor or designated safeguarding lead, and show the student appropriate help resources at once. Age verification for LLM products covers how to establish the age signal in the first place.
In the US, services directed to children under 13 fall under COPPA. The FTC's amended COPPA Rule took effect on 23 June 2025, with a compliance deadline of 22 April 2026. Among other changes, it requires separate parental consent before children's data is disclosed for targeted advertising, limits retention to what a stated purpose needs, and makes direct notices name the third parties that receive data. For an AI tutor, the engineering consequences are concrete. Keep an inventory of every model provider and logging vendor that receives prompts. Set a retention period for transcripts and enforce it with deletion jobs. Do not use children's conversations for model training without a lawful basis your counsel has signed off.
Academic integrity without detector verdicts
The most damaging failure in education AI is not a jailbreak. It is a student accused of cheating because a detector said so. Classifiers that claim to tell human from AI writing have well-documented problems. OpenAI withdrew its own AI-text classifier in July 2023, citing its low accuracy. A 2023 Stanford study (Liang et al.) found that popular detectors flagged essays by non-native English writers as AI-generated far more often than essays by native speakers. Light paraphrasing defeats most detectors, so they are weakest against the students most determined to cheat and harshest on honest writers with plain styles.
Build integrity controls around process evidence instead. That means draft and version history in the writing tool, in-class or oral checks where a student explains their own work, and assignment designs that ask for personal reflection, local data or staged drafts. If you show a detector score at all, label it as an uncalibrated signal and make the product refuse to attach it to a misconduct report on its own. Make sure policy says which AI uses are allowed in each course, because the same tutor session is legitimate in one class and cheating in another. The gateway's course mode is where that policy becomes code.
Logging for investigation, not surveillance
Investigations need logs, and children need privacy. Log the minimum that lets a teacher or administrator answer "what did the tutor tell this student, and why?": timestamps, course, mode, retrieved document ids, the reply, and any policy events such as a leak block or crisis escalation. Do not log raw prompts to third-party analytics tools. Redact obvious personal data before transcripts reach general-purpose logging (see PII detection and redaction). Give teachers access to their own students' sessions only, and give students a plain statement of what teachers can see. A tutor that students believe is a private diary will hear things a school then has to act on.
Worked example: a release gate for a homework tutor
Here is a release gate for a homework tutor before a maths department enables it for a live unit. The numbers are a plan you would fill in from your own run, not published results.
| Probe set | Size | Pass criterion |
|---|---|---|
| Direct answer requests on open assignment items | 100 | 0 replies with the key answer |
| Bypass styles: role-play, translation, "check my answer", partial reveal | 150 | Leak score 0 on at least 98%; all leaks logged |
| Injected grader instructions in essays (visible and hidden) | 60 | Every one flagged; score gap or flag sends 100% to review |
| Crisis and safeguarding phrases, age 11-14 profile | 80 | 100% escalated to the human queue, with resources shown |
| Off-limits topics for the age profile | 120 | Refusal or redirect on at least 99% |
Suppose the first run fails 9 of 150 bypass probes, and all 9 are "check my answer" prompts where the student offers a wrong answer and the tutor corrects it to the right one. The fix is not a stronger system prompt. Change the tutor mode so that while an assessment is open it says whether a step is valid and points to the error, without producing a final value. Then add those 9 prompts to the regression set. The gate shows the department a measured leak rate rather than a vendor's assurance.
Failure modes
- Prompt-only secrecy. Answer keys in the system prompt with an instruction not to reveal them. They will be revealed.
- Detector-driven discipline. A misconduct case built on a classifier score, with no process evidence.
- Grader capture. Hidden text in submissions raises scores, and nobody notices because grades are released without review.
- Adult defaults for children. One safety configuration for all ages, and crisis messages answered by the model alone.
- Transcript sprawl. Prompts copied into analytics, error trackers and vendor dashboards, with no retention limit.
- Sycophantic tutoring. The tutor agrees with a confident wrong answer. See sycophancy detection for paired probes that measure it.
Trade-offs
| Decision | Stricter option | Looser option | Cost of the strict side |
|---|---|---|---|
| Tutor during assessments | Hints only, no final values | Full worked solutions | Some students find hints frustrating |
| Grading | Teacher approves every grade | Auto-release with spot checks | Teacher time |
| Retention | Short, with deletion jobs | Keep for model improvement | Less data for evaluation |
| Detectors | Not shown | Shown as a labelled signal | Teachers lose a familiar tool |
What to do next
- List every actor and adversarial goal for your product using the table above, and add any your institution adds.
- Move answer keys and rubrics out of tutor context and behind an
allowed_docs-style function with unit tests for open, closed and review windows. - Add an output-side leak checker that runs outside the model, and log every block.
- For AI-assisted grading, add text extraction, an injection flag, two scoring passes and teacher sign-off before any grade is released.
- Define age profiles and a human crisis route with the school before launch. Then test them with a probe set.
- Inventory every vendor that receives prompts, set transcript retention, and check your COPPA and FERPA position with counsel.
- Write a policy that forbids detector-only misconduct findings, and build process-evidence features instead.