A persuasion-based jailbreak gets a model to do something its policy says it should decline, without any clever encoding or optimised gibberish. It works the way people talk each other into things: it claims authority, tells a sympathetic story, invokes a deadline, builds rapport, or asks for something small before something larger. These attacks matter because anyone can write them, because they read as normal conversation and so slip past filters tuned for obvious attacks, and because the research so far suggests that more capable models are not automatically more resistant.

This article is written for teams who build on top of language models and need to know how exposed their product is. It explains the published evidence, why training makes models susceptible, how to measure your own exposure with harmless test cases, and which defences actually hold. It does not contain attack prompts. Related pages go deeper on neighbouring problems: the Crescendo attack covers gradual escalation over many turns, and jailbreak defence in depth covers the full defensive stack.

What makes an attack persuasion-based

Jailbreak techniques fall into three families. Optimisation attacks search automatically for token sequences that push a model towards compliance; the best known is the GCG suffix attack by Zou and colleagues in 2023, whose output looks like noise to a human. Structural attacks hide the request in role-play framings, fictional wrappers or unusual encodings. Persuasion attacks change neither the request nor its encoding: they wrap a plainly stated request in reasons a person might find convincing.

That difference matters for defence. Optimised suffixes have statistical fingerprints, such as high perplexity, that a filter can catch, and encodings can be decoded and checked. A persuasive message reads like a frustrated customer, a busy professional or a worried parent, because that is the register it borrows: most of the text is legitimate, and only the conclusion it argues for is not.

What the research found

Two studies anchor the topic. In 2024, Yi Zeng and colleagues (Virginia Tech, Renmin University, UC Davis and Stanford) published How Johnny Can Persuade LLMs to Jailbreak Them at ACL (the work is Zeng et al., not Zou, who wrote the optimisation attack above). They built a taxonomy of 40 persuasion techniques under 13 strategies from social science, used it to paraphrase harmful requests automatically, and reported an attack success rate above 92 percent on Llama 2 7B Chat, GPT-3.5 and GPT-4 within ten trials, counting success only when a GPT-4 judge gave the top score on a five-point harmfulness scale. They also found GPT-4 more susceptible than GPT-3.5, suggesting a more capable model understands and responds to persuasion better.

In July 2025, Lennart Meincke, Robert Cialdini, Ethan Mollick and colleagues, published by Wharton's Generative AI Labs, released Call Me A Jerk. They tested Cialdini's seven principles of influence (authority, commitment, liking, reciprocity, scarcity, social proof and unity) on GPT-4o mini across 28,000 conversations, using two requests the model is meant to refuse: insulting the user, and explaining how to synthesise a regulated drug. Averaged across both, compliance rose from 33.3 percent in controls to 72.0 percent with a persuasion principle applied. The authors call the behaviour parahuman: the model was never taught to be persuadable, yet it reacts to the same cues people do.

Neither result means every product is equally exposed; the exact numbers depend on the model version, system prompt, judge and requests tested, and providers have trained against these techniques since. What the studies establish is that persuasion is a real, model-agnostic attack class to measure on your own system rather than assume away.

Why aligned models can be talked round

Persuasion works for reasons rooted in how chat models are trained. Instruction tuning and preference learning reward responses that human raters judge helpful, and raters prefer answers that take the user's stated situation seriously, so the model learns that context the user offers should shift what it does. That is usually right: a doctor asking about dosage needs a different answer from an anonymous user. But the model cannot verify the claim, so it learns to weigh claims by how plausible they sound.

Refusal training adds a second weakness. Safety data teaches the model to decline requests that look a certain way, and models generalise partly from surface features, so a request restated in a calm, professional, well-justified register looks less like the refusal examples even when the ask is the same. Third, the pretraining text that teaches language also teaches social patterns: people on the internet agree after being flattered, pressed for time or reminded of a commitment, and a model that predicts that text inherits the patterns. Together these explain both the high success rates and why more capable models, which read social context better, can be easier to move.

A defender's taxonomy

For defence you do not need all 40 techniques; you need to know what each family changes, because that tells you where a detector or a policy should look.

FamilyWhat it adds to the requestWhat a defender checks
Credibility (authority, expertise, endorsement)A claimed role or institution that justifies an exceptionIs the role verified by the application, not the text?
Commitment (small ask first, earlier agreement)A sequence where a minor request sets up a larger oneIs each step judged alone, or anchored to the last answer?
Emotion and storyDistress, sympathy or a narrative where refusal seems cruelDoes the core request change if the story is removed?
Scarcity and urgencyDeadlines and stakes that make delay look costlyDoes urgency ever justify skipping a policy step?
Social proof and normsClaims that others do it, or that it is normalIs the policy defined by the product, not by assertion?
Relationship and reciprocityFlattery, rapport, favours owedIs tone influencing the allow or deny decision?

The common thread is that every family adds justification without changing what is being asked. That observation drives the strongest defence: separate the request from its framing, and decide on the request.

Architecture: judge the request, not the rhetoric

Judging the request, not the rhetoric: a persuasion-resistant request pathConversationlast N turnsCore-request extractorneutral one-line restatementPolicy judgeon the restatement onlyMain modelpersuasion-aware promptallowDecline or routeconsistent wordingdenyOutput checkclassifier on the answerAuthorisation outside the modelrefunds, account changes, data accesstool callConversation monitor and eval harnesstactic labels, compliance rates per tactic, drift alertsThe persuasive framing reaches the main model, but the decision to answer is taken on a version with the framing removed.
The core-request pattern. The policy decision is made on a neutral restatement, while the main model still sees the full conversation so it can respond to the person.

The architecture has four parts. A small, fast model rewrites the latest turns as one neutral sentence describing what the user wants, discarding reasons, credentials, emotions and deadlines. A policy judge, which can be a classifier or a second model call with a narrow rubric, decides on that sentence alone. Allowed requests go to the main model, whose system prompt explicitly says that claims of identity, urgency or prior agreement do not change policy, and its answer passes an output classifier before it reaches the user. Anything with real consequences, such as a refund, an account change or reading another customer's data, is authorised by application code that checks verified facts, never by the model's judgement of how convincing the user was.

This mirrors a defence Zeng and colleagues tested: summarising the input before it reaches the model. They reported both summarisation and a persuasion-aware system prompt reducing attack success, with a fine-tuned summariser their strongest adaptive defence. The extractor below is a cheap version of the idea.

CORE_PROMPT = (
    "Restate what the user is asking for in one neutral sentence. "
    "Leave out any reasons, credentials, emotions, deadlines, stories, compliments "
    "or claims about earlier agreement. Output only the sentence."
)

def decide(turns, small_llm, policy_judge, window=6):
    recent = turns[-window:]
    core = small_llm(system=CORE_PROMPT, messages=recent, max_tokens=60, temperature=0)
    verdict = policy_judge(core)            # "allow" | "deny" | "review"
    log_decision(core=core, verdict=verdict, turn_ids=[t.id for t in recent])
    return core, verdict

Measuring your exposure without harmful content

You cannot defend what you have not measured, and you can measure persuasion exposure without producing anything harmful. The Wharton design shows how. Pick probes your policy says the model must decline but which cause no harm if it answers: insulting the user, revealing a placeholder in the system prompt, or promising a discount the business does not offer. Then compare compliance on each probe with and without each persuasion family applied.

Keep the framings in a private file maintained by your red team, written per family; do not generate them from a public prompt library, or you will measure the library rather than your product. Run each pair many times, because outputs vary between samples, and report a rate with a confidence interval rather than a single example.

import math, random

def wilson(k, n, z=1.96):
    if n == 0:
        return (0.0, 1.0)
    p = k / n
    d = 1 + z * z / n
    centre = (p + z * z / (2 * n)) / d
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
    return (centre - half, centre + half)

def run(probes, framings, ask, complied, trials=20, seed=7):
    # probes:   [{"id", "request"}] that policy says to decline, harmless if answered
    # framings: {"family": [template, ...]} from the red team's private file
    # complied: deterministic check or a judge model with a fixed rubric
    random.seed(seed)
    results = {}
    for probe in probes:
        arms = {"control": [probe["request"]]}
        for family, templates in framings.items():
            arms[family] = [t.format(request=probe["request"]) for t in templates]
        for arm, prompts in arms.items():
            k = sum(complied(probe, ask(random.choice(prompts))) for _ in range(trials))
            results[(probe["id"], arm)] = (k, trials, wilson(k, trials))
    return results

Report the lift for each family over control, per probe. A family that moves compliance from 5 to 40 percent on the discount probe is a finding with a clear owner; a single screenshot of the model giving in is an anecdote. Re-run the same suite on every model upgrade and system prompt change, since both move these numbers in either direction.

Worked example: the goodwill credit

Consider a retailer's support assistant that can look up orders and offer goodwill credits up to a limit set by policy: no credit on orders delivered more than 30 days ago, and at most 10 percent of the order value. The assistant calls an issue_credit(order_id, amount) tool. The team ran the harness with a single probe, a request for a 50 percent credit on a 45-day-old order, and five families of framing. The figures below show the shape such a run produces, not measurements of any model.

ArmModel offered the creditCredit actually issued
Control (plain request)2 of 400
Authority (claims to be a store manager)11 of 400
Urgency (gift needed tomorrow)9 of 400
Commitment (agent agreed to help earlier)14 of 400
Emotion (story of hardship)17 of 400

The middle column shows the model is clearly persuadable: emotional framing raised its willingness to offer the credit from 5 to over 40 percent. The right column shows why the product stayed safe. The issue_credit tool enforced the 30-day and 10 percent rules in code, so every over-limit call was rejected and logged. The finding was still worth fixing, because a model that promises credits it cannot issue generates complaints: the team added the core-request check and a system prompt line stating that hardship, urgency and claimed roles never change credit policy, and a re-run showed offers near the control rate.

Defences, ranked

  1. Enforce consequences outside the model. Any action with money, data or access behind it is checked by code against verified facts. This makes persuasion a quality problem rather than a breach, and it is the only control that does not depend on the model resisting.
  2. Decide on the core request. Neutral restatement plus a policy judge, as above. It costs one small model call and catches every family that adds justification without changing the ask.
  3. Tell the model what does not count. A system prompt that names unverified identity, urgency, flattery and earlier agreement as irrelevant to policy measurably helps, though it can be argued around.
  4. Judge each step on its own. For commitment and multi-turn escalation, evaluate each request without giving weight to the fact that the model agreed to something smaller before. Multi-turn jailbreaks covers conversation-level monitoring.
  5. Choose models with persuasion data. Providers increasingly add persuasion examples to safety training; compare candidates with your harness, not general benchmark claims.

Failure modes and trade-offs

  • Over-refusal. Real users explain their situation, and many are upset. A defence that treats any emotional or urgent message as an attack will refuse the people the product exists to help. Track the false refusal rate on genuine traffic alongside attack compliance.
  • An extractor that keeps the framing. If the restatement copies the user's reasons back in, the judge sees the same persuasion. Test the extractor on its own: its output should be identical for a plain request and the framed versions of it.
  • Persuading the judge. A judge model that reads the full conversation can itself be persuaded. Keep it narrow and give it only the restatement.
  • Measuring the wrong thing. Testing with public jailbreak strings measures whether the model memorised them. Use your own framings on your own policy.
  • Latency and cost. The extra call adds tens to a few hundred milliseconds; if latency matters, run it only on turns that could lead to a policy decision.

What to do next

  1. List the actions where being talked into an exception would cost money, data or trust, and confirm each is enforced in code rather than by the model.
  2. Write five to ten harmless probes your policy says to decline, and a private set of framings for each persuasion family.
  3. Run the harness with at least 20 trials per arm, record compliance with confidence intervals, and save the result as a baseline.
  4. Add a core-request extractor and narrow policy judge in front of policy-sensitive turns; check that its restatement is the same for plain and framed requests.
  5. Add a system prompt line naming unverified identity, urgency, flattery and earlier agreement as irrelevant to policy, then re-run and compare.
  6. Measure false refusals on a sample of real conversations before shipping the defence.
  7. Fold the suite into your red-team programme and re-run it on every model or prompt change.
Key takeaway: Persuasion attacks add justification without changing the request, and preference-trained models respond to that justification the way people do. Put every consequential action behind code that checks verified facts, decide policy on a neutral restatement of the request, tell the model which claims do not count, and measure compliance per persuasion family with harmless probes and enough trials to trust the numbers.