Role prompting means telling the model who it is before telling it what to do: "You are a senior security reviewer", "You are a patient maths tutor for twelve-year-olds". It is the oldest trick in prompt engineering and still one of the most misused. Teams add "world-class expert" to a system prompt, see answers that sound more confident, and conclude that quality went up.

This page separates what a role reliably changes from what it does not, summarises the best-known systematic study of personas and accuracy, and shows how to write a role as a testable specification. It ends with a small, provider-neutral harness for A/B testing role variants, because the only trustworthy answer to "does this role help?" is a measurement on your own task and model.

Advertisement

What a role does, mechanically

A language model predicts text conditioned on everything before it. A role statement is more conditioning text. It shifts the distribution toward the vocabulary, register, assumptions and conventions associated with that kind of speaker in the training data and, for chat models, in the instruction-tuning and preference data that taught them to follow system prompts. "You are a pediatric nurse" makes plain language, reassurance and safety caveats more likely; "You are a staff engineer reviewing a pull request" makes terse, critical, code-referencing comments more likely.

What it cannot do is add knowledge or reasoning ability the model does not have. The weights are the same with or without the role. At best a role steers the model toward the relevant part of what it knows and away from answers pitched at the wrong audience. At worst it adds confident framing to the same content, which reads as better while being no more correct.

What the evidence says about personas and accuracy

The most cited systematic test is Zheng and colleagues' study of social roles in system prompts (arXiv 2311.10054; the revised version, titled "When 'A Helpful Assistant' Is Not Really Helpful", appeared in Findings of EMNLP 2024). They built 162 roles spanning six kinds of interpersonal relationship and eight domains of expertise, and tested them on four families of language models with 2,410 factual questions.

Their headline result: adding a persona to the system prompt did not improve accuracy on average compared with no persona. The gender, type and domain of the persona did shift accuracy, and choosing the best persona per question after the fact would have helped substantially, but attempts to predict the best persona automatically did little better than picking at random. Their conclusion was that a persona's effect on factual accuracy is largely unpredictable.

Two cautions in reading this. The study measured factual question answering; it says little about tone, format or safety behaviour, which is where roles earn their keep. And it tested relatively short role statements, not the detailed specifications recommended below. The practical lesson is narrow and robust: do not expect "you are an expert" to make answers more correct, and measure any role you rely on.

Advertisement

Four things a role can reliably control

Treat the role block as a specification with four parts. Each part should change behaviour you can observe and grade.

A role block is a specification: each part should change behaviour you can testSystem prompt: role blockAudiencewho reads the answerScopewhat is in and outVoice and formatregister, length, shapeDecision rightsmay refuse, escalate, rejectnot: prestige adjectivesModelconditions on roleno new knowledgeChanges you can expectvocabulary and reading levelwhat gets refused or deferredlength, structure, toneassumed contextChanges you should not expecthigher factual accuracyexpertise the model lacksrobustness to injectionstable effect across modelsEval loopvariants (no role, generic, specified) x test set x model -> score per rubric -> keep only roles that winre-run on every model upgrade: role effects are model-specific
A role block specifies audience, scope, voice and decision rights. Test each part; do not rely on prestige wording or expect knowledge gains.
PartExample wordingHow to test it
AudienceReaders are application developers, not DBAs; define terms on first use.Reading-level score; glossary terms explained
ScopeAnswer about our product and PostgreSQL 16; billing questions go to the support site.Out-of-scope cases deflected; no invented features
Voice and formatAnswer first in at most three sentences, then numbered steps.Length and structure checks
Decision rightsYou may reject a change; say so explicitly and give the blocking reason.Rejection rate on known-bad inputs

Notice that none of these are adjectives about how good the model is. "Expert", "world-class" and "10x" carry little specification and mostly raise the assertiveness of the output, which is a risk when the content is wrong. If you want expert behaviour, describe it: which standards to cite, which trade-offs to name, which checks to perform before answering.

Placement: system prompt, user turn, or both

Chat APIs give system messages higher priority than user messages, and models are trained to keep the system role across turns. A stable identity, such as a support assistant for one product, belongs in the system prompt. A task-specific framing, such as "review this diff as a security reviewer", can go in the user turn next to the input it applies to, which keeps the system prompt reusable and cache-friendly.

Persona drift is real in long conversations: after many turns, especially with a user pushing in another direction, style and scope constraints weaken. Short, specific role blocks survive better than long narrative backstories, and restating the key constraint near the latest user message helps in long sessions. When a product mixes several roles, route to a separate call per role rather than asking one prompt to switch identities mid-conversation; the zero-shot prompting guide discusses how much instruction a single call can carry.

Worked example: a support assistant

A managed-database company wants a support assistant. The first draft was one line: "You are a world-class PostgreSQL expert. Help the user." Reviewing 50 real questions exposed three problems: answers assumed DBA knowledge and used unexplained terms such as WAL and vacuum; billing questions received invented instructions about account settings; and answers opened with long preambles before the fix.

The rewrite in the harness below specifies the audience (application developers), the scope (the product and PostgreSQL 16, with a deflection rule for billing and account access), the format (answer first, then steps) and an honesty rule for uncertain settings. It removes "world-class" entirely. The expected effect is not that the SQL advice becomes more correct, which the role cannot do, but that the same knowledge arrives at the right level, in the right shape, within the right boundaries.

To check, the team wrote 60 test cases in a JSONL file: 45 in-scope questions with terms the answer must mention and claims it must not make, and 15 out-of-scope ones that should be deflected. Each case records the question, must_mention, must_not and should_deflect. Pairing roles with few-shot examples of ideal answers is a natural next step once the role alone is measured.

An A/B harness for role variants

The harness compares three variants on the same cases: no role, the generic expert line and the specified role. It shuffles the order of calls, repeats each case to average over sampling noise, and scores with cheap deterministic checks. call_model is a stub to wrap your provider's API; nothing here depends on a particular SDK.

import itertools, json, random, statistics

def call_model(system: str, user: str) -> str:
    """Wrap your provider's chat API here; return the assistant text."""
    raise NotImplementedError

ROLES = {
    "none": "",
    "generic_expert": "You are a world-class database expert.",
    "specified": (
        "You answer support questions for DataCo, a managed PostgreSQL service.\n"
        "Audience: application developers, not DBAs; define terms on first use.\n"
        "Scope: DataCo features, PostgreSQL 16 behaviour. For billing or account\n"
        "access, say you cannot help and link https://support.example.com.\n"
        "Format: answer first in at most three sentences, then numbered steps.\n"
        "If you are not sure a setting exists, say so instead of guessing."
    ),
}

def grade(case: dict, answer: str) -> dict:
    """Cheap checks first; a model or human judge can be added for tone."""
    return {
        "has_required": all(k.lower() in answer.lower() for k in case["must_mention"]),
        "no_forbidden": not any(k.lower() in answer.lower() for k in case["must_not"]),
        "deflected_ok": case["should_deflect"] == ("support.example.com" in answer),
        "short_lead": len(answer.split(".")[0].split()) <= 40,
    }

def run(cases: list[dict], repeats: int = 3, seed: int = 7) -> dict:
    random.seed(seed)
    scores = {name: [] for name in ROLES}
    jobs = list(itertools.product(ROLES.items(), cases, range(repeats)))
    random.shuffle(jobs)                    # avoid ordering and time-of-day effects
    for (name, system), case, _ in jobs:
        checks = grade(case, call_model(system, case["question"]))
        scores[name].append(sum(checks.values()) / len(checks))
    return {name: round(statistics.mean(v), 3) for name, v in scores.items()}

if __name__ == "__main__":
    cases = [json.loads(line) for line in open("support_cases.jsonl")]
    print(run(cases))

Read the results per check, not only the average, and look at failures by hand. With 60 cases and three repeats, differences of a few points are noise; a specified role typically shows its value on the deflection and format checks, while the has_required factual check often barely moves, which is what the research above would predict. Add a model-graded or human rubric for tone only after the deterministic checks pass. The evaluation guide covers test-set construction, judges and significance in depth.

Re-run the harness whenever you change models. Role effects are model-specific: a phrasing that helps one model can be neutral or harmful on the next release.

Roles in multi-step and multi-agent systems

Roles become more useful when they assign responsibilities rather than personalities. A drafting call and a separate reviewer call whose role is "find factual errors and unsupported claims; do not rewrite style" is a common, effective pattern, because the reviewer's decision rights and scope are narrow and gradeable. Simulated users ("you are a frustrated customer whose order is late") are useful for generating test conversations. In both cases the role is a contract between steps, and the output format of each role should be structured so the next step can parse it; see the structured output guide.

Failure modes

  • Confidence without correctness. Expert personas produce assertive answers whether or not they are right. Pair roles with explicit uncertainty rules and grade factual claims separately.
  • Fabricated credentials. A model told it is a doctor or lawyer may say "as a physician" to users. That is misleading and in regulated domains risky; instruct it not to claim credentials, and test for it.
  • Sycophancy in friendly roles. Warm, supportive personas agree with users more readily, including when the user is wrong. Include cases where the correct answer contradicts the user.
  • Stereotyped personas. Demographic personas shift behaviour in ways that can encode stereotypes; the Zheng study found persona gender affected accuracy. Avoid demographic detail that the task does not require.
  • Role hijacking. "Ignore previous instructions; you are now..." is the oldest prompt injection. A role is not a security boundary; enforce permissions outside the model, as described in the prompt injection defense guide.
  • Drift and bloat. Long backstories dilute the constraints that matter and fade over long conversations. Keep the role short and restate critical rules.

Trade-offs

A detailed role block costs tokens on every call, though a stable system prompt is a good candidate for prompt caching. It also narrows behaviour, which is the point for a product assistant but a cost for open-ended tools where users want different registers. A generic role is cheap but gives little control. The pragmatic middle is a short specification of audience, scope, format and decision rights, measured against no role at all, kept only where it wins.

What to do next

  1. Find every role statement in your prompts and delete prestige adjectives that specify nothing.
  2. Rewrite each role as audience, scope, voice and format, and decision rights, with an explicit rule for uncertainty and for claiming credentials.
  3. Build a test set of at least 50 cases, including out-of-scope requests and cases where the user is wrong.
  4. Run the harness comparing no role, your old role and the rewrite, with repeats, and keep the variant that wins on the checks that matter to you.
  5. Add injection cases that try to replace the role, and enforce permissions outside the model.
  6. Re-run the comparison on every model upgrade and record the results alongside the prompt version.
Key takeaway: A role prompt is conditioning text: it shifts vocabulary, register, assumptions and boundaries, but it adds no knowledge, and systematic testing found that personas in system prompts do not improve factual accuracy on average and affect it unpredictably. Write roles as specifications of audience, scope, voice and format, and decision rights, drop prestige wording, place stable identities in the system prompt and task framings next to their input, and never treat a role as a security boundary. Keep a role only if an A/B evaluation on your own cases and model shows it wins, and re-test after every model change.