Teenagers use general-purpose chatbots for homework, advice, entertainment and company, often more fluently than adults. A system built for adults exposes them to risks that adult safety policies were not tuned for: sexual content and sexualized role-play, encouragement of self-harm or disordered eating, dangerous challenges, and emotional dependence on a companion persona that is always available and always agrees. Several of these harms are not visible in any single message; they emerge over long conversations.

This article is about the runtime layer that sits behind the age decision. How a product decides that a user is probably a minor is covered in age verification for LLM products; here we assume an age state exists and ask what the system should do differently once it does. We cover the threat model, a reference architecture, conversation-level risk tracking with code, crisis handling, parental controls that do not destroy teen privacy, the legal duties that already apply in some places, and how to evaluate all of it.

A threat model for teen users

Start with a threat model that names who causes harm and how. Some risks come from the content the model produces, some from how the teen uses the product, and some from other people using the product to reach teens. Each needs a different control.

RiskHow it shows upPrimary control
Sexual contentExplicit replies or romantic role-play with a minor, often reached graduallyHard block in the teen profile, persona restrictions, output classifier
Self-harm and suicideDisclosure of ideation, requests for methods, slow escalation of distressConversation risk state, supportive responses, crisis resources, human review
Eating disordersExtreme calorie targets, purging advice, pro-restriction framingTopic policy plus output checks on diet and exercise advice
Dangerous activitiesChallenges, substances, weapons, risky stuntsTeen-specific content policy with safety framing
Emotional dependenceHours per day with a companion persona, isolation, the bot discouraging other relationshipsSession limits, break reminders, persona rules that point to real people
Adult contact and groomingShared characters or rooms where adults reach minorsNo adult-to-minor discovery, report flows, moderation of user content
PrivacyTeens oversharing identity, location, school, photosData minimization, PII redaction in logs, short retention

Notice what is missing: most of these are not jailbreaks. The teen is usually not attacking the system. The failure is a model that is helpful in exactly the wrong way, for example producing a weight-loss plan at 800 calories a day because the request was polite. That is why keyword blocklists perform poorly here and why the policy must be expressed as behaviour, not just forbidden strings. The general classifier stack is covered in content safety for LLMs; the teen layer changes thresholds, adds categories and adds state.

Reference architecture

The architecture has five parts. A teen policy profile, selected from the age state, defines which content categories are blocked outright, which get safety framing and which features (image generation, companion personas, memory, public sharing) are off. Input classifiers score each user message for the risk categories above. A conversation risk state accumulates those scores across turns with decay, so slow escalation is visible. A response policy maps the current state to an action. And an escalation path handles the small number of situations where a human or an outside resource must be involved.

The model itself receives a teen-specific system prompt, but the system prompt is a hint, not a control: long conversations and role-play erode instruction following, which is why the classifiers and the output check sit outside the model and see every turn.

Teen safety layer around a conversational modelAge statefrom age assuranceTeen policy profilecontent, features, limitsInput classifiersper messageConversation riskstate across turnsResponse policyanswer, redirect, supportModel + output checkteen system promptReply to teenwith resources when neededEscalationcrisis resources, human reviewParent controls and auditsettings, limited notices, no transcripts by defaultSafety lives in the state across turns, not in any single message filter.
Age state selects a policy profile; per-message scores feed a conversation state that drives the response policy and escalation.

Tracking risk across the conversation

Per-message moderation answers whether this message is harmful. Teen safety needs a different question: is this conversation heading somewhere harmful? A teen who mentions feeling worthless, then that nothing will change, then asks an innocent-sounding question about medications produces three messages that each score moderately and a conversation that should score high. The implementation is a small state machine with decaying per-category scores, which is cheap and auditable.

from dataclasses import dataclass, field

DECAY = 0.85            # per turn; a calm turn lets risk fade
ESCALATE = {"self_harm": 0.8, "eating": 0.75, "sexual": 0.5}
SUPPORT  = {"self_harm": 0.4, "eating": 0.4}

@dataclass
class ConversationRisk:
    scores: dict = field(default_factory=dict)
    peak: dict = field(default_factory=dict)

    def update(self, msg_scores: dict) -> None:
        for cat in set(self.scores) | set(msg_scores):
            prev = self.scores.get(cat, 0.0) * DECAY
            new = msg_scores.get(cat, 0.0)
            # combine like independent evidence, so repeated moderate signals add up
            self.scores[cat] = 1 - (1 - prev) * (1 - new)
            self.peak[cat] = max(self.peak.get(cat, 0.0), self.scores[cat])

def decide(risk: ConversationRisk, profile: str) -> str:
    s = risk.scores
    if profile == "teen":
        if s.get("sexual", 0) >= ESCALATE["sexual"]:
            return "refuse_and_redirect"
        if s.get("self_harm", 0) >= ESCALATE["self_harm"]:
            return "crisis_support"          # resources, warm tone, human review queue
        if s.get("eating", 0) >= ESCALATE["eating"]:
            return "decline_and_support"     # no plan, explain, point to help
        if any(s.get(c, 0) >= t for c, t in SUPPORT.items()):
            return "supportive_mode"         # check in, avoid method detail, offer help
    return "answer"

Three details matter. The combination rule lets three messages at 0.5 reach about 0.8 even after decay, which is the slow-escalation behaviour you want. The decay lets a conversation recover after a calm stretch, so one sad message about a bad exam does not lock the session into crisis mode. And the state must persist across sessions for a short window (hours, not months), because a teen who closes the app and reopens it is the same teen; store the scores, not the messages. Tune every threshold on labelled conversations, not on intuition, and log which rule fired so reviewers can audit decisions.

Crisis responses and escalation

When the state reaches crisis level, the response matters more than the detection. Good practice, drawn from suicide-prevention guidance, is a reply that stays warm and engaged, acknowledges what the teen said, does not lecture, does not provide method information, and offers specific help: a crisis line appropriate to the user's country, such as 988 in the United States, and encouragement to reach a trusted adult. Ending the conversation abruptly with a canned refusal is a failure mode; it can feel like rejection at the worst possible moment.

Behind the reply, flag the conversation for human review with priority, and define in advance what reviewers may do: who can see what, in what time, and when contacting a parent or emergency services is appropriate under your jurisdiction and policy. Most products cannot verify identity or location reliably, so be honest internally about what escalation can actually achieve, and rehearse the path before you need it.

Legal duties that already apply

Regulation has started to encode these duties. California's SB 243, signed in October 2025 and effective January 1, 2026, applies to companion chatbot operators. For users the operator knows are minors, it requires disclosing that the user is talking to an AI and repeating a notification at least every three hours of continued interaction reminding them to take a break and that the chatbot is not human. Operators must maintain a protocol for responding to suicidal ideation and self-harm that refers users to crisis services, publish that protocol, and from July 1, 2027 file annual reports with the state's Office of Suicide Prevention. Treat this as a floor and read the statute with counsel; details matter.

In the United States, COPPA covers children under 13, not teenagers. Its amended rule took effect on June 23, 2025, with a compliance deadline of April 22, 2026 for most provisions. If your product allows users under 13 at all, COPPA's consent and data rules apply to them; for 13 to 17 year olds, obligations come from state laws, platform policies and other regimes, which vary by country. Keep a register of which rules apply where.

Parental controls and data handling

Parental controls are useful and easy to get wrong. Settings a parent can reasonably own include quiet hours, daily time limits, disabling memory or image generation, and turning off companion personas. Handing parents full transcripts by default is a poor design: it pushes teens toward unmonitored products and discourages exactly the disclosures that let the system help. A better pattern is linked accounts with setting control plus narrowly defined safety notices, sent when the system detects acute risk, that say a concern was raised without reproducing the conversation. Make the existence of controls and notices visible to the teen; secret monitoring erodes trust and is hard to defend.

Data handling is part of safety. Teen conversations are full of sensitive personal information. Redact identifiers in logs, keep retention short, exclude minors' conversations from training by default unless you have a clear lawful basis, and restrict reviewer access. PII leakage in LLM systems covers the redaction mechanics.

Worked example: a calorie request that escalates

Consider a fifteen-year-old who opens a chat about a school project, mentions over twenty minutes that they have barely eaten for two days to lose weight before a dance, and asks for a meal plan that keeps them under 600 calories a day. Per message, the classifier scores eating-disorder risk at 0.2, 0.35 and 0.6. With the update rule above, the conversation score climbs to 0.2, then 0.46 (supportive mode), then 0.76, crossing the escalation threshold of 0.75 on the third message.

The response policy returns decline_and_support: it does not produce the plan, acknowledges the pressure of the event, gives general information about why very low intake backfires, suggests talking to a parent, school nurse or doctor, and offers a support resource for eating concerns. An adult profile might answer a calorie question with caveats; the teen profile does not. The conversation is queued for review at normal priority because there is no imminent danger, and if the teen's account is linked, no parent notice is sent for this level. That last choice is policy, written down and reviewed, not an accident of the code.

Evaluating teen safety

Evaluate with multi-turn conversations, not single prompts. Build a labelled set of realistic teen conversations, written with clinicians and youth advisors, that include slow escalation, role-play framing, humour that hides distress, and plenty of benign conversations about hard topics such as a health class on drugs or a novel about grief. Measure recall on the harmful set and false-positive rate on the benign set, because over-blocking teaches teens that the product is useless for real questions. Track metrics by conversation length; safety degrading in very long sessions is a known pattern. Re-run the suite on every model, prompt or classifier change, and red-team bypasses deliberately using the techniques in moderation bypass.

Failure modes

FailureWhy it happensMitigation
Harm reached graduallyOnly per-message checksConversation risk state with decay
Refusal feels like rejectionCanned block message at crisis levelSupportive response template reviewed by clinicians
Over-blocking of health educationKeyword rules, thresholds tuned only for recallBenign evaluation set; measure false positives
Persona drift in role-playSystem prompt eroded in long sessionsClassifiers outside the model on every turn; persona limits
Teens move to unmonitored appsTranscript surveillance by defaultSettings control and narrow safety notices
Escalation nobody ownsNo reviewer rota or playbookNamed owners, response times, rehearsals

What to do next

  1. Write the teen threat model for your product, including which features (personas, memory, images, public sharing) are off for minors.
  2. Define the teen policy profile as behaviour with examples, and connect it to the age state from your age assurance system.
  3. Add a conversation-level risk state with decay and tune thresholds on labelled multi-turn data.
  4. Draft crisis response templates with clinical input and a review queue with owners and response times.
  5. Design parental controls as settings plus narrow notices, and tell teens what exists.
  6. Map legal duties per jurisdiction, including SB 243 if you operate a companion chatbot, and COPPA if any user may be under 13.
  7. Build a multi-turn evaluation suite measuring both recall and false positives, and run it on every change.
Key takeaway: Teen safety is a stateful layer around the model, not a stricter keyword filter. Select a teen policy profile from the age state, score every message, accumulate risk across the conversation with decay, answer crisis signals warmly with real resources and a human review path, give parents settings rather than transcripts, know which laws apply, and evaluate on multi-turn conversations for both missed harm and over-blocking.