DREAD is a mnemonic for rating how dangerous a security threat is. Each threat gets a score on five factors: Damage, Reproducibility, Exploitability, Affected users and Discoverability. The scores are combined into one number used to rank remediation work. It came out of Microsoft in the early 2000s and was popularised in Howard and LeBlanc's Writing Secure Code. Microsoft later moved away from it because different raters produced very different scores for the same bug. That history is the most useful thing to know about DREAD: it is fast and easy to explain, and it is unreliable unless you pin every factor to an observable anchor.
LLM systems make that criticism sharper in one place and easier to fix in another. Reproducibility stops being a yes-or-no question, because a prompt injection that works on one run may fail on the next. But that also means it can be measured, by running the attack many times and counting. This article rebuilds DREAD for LLM applications. It covers an anchored rubric for each factor, reproducibility scored from a measured attack success rate with a confidence bound, an aggregation rule that cannot let a low-damage nuisance outrank a data leak, scoring code you can run, a worked example on a retrieval-augmented support agent, and the failure modes that make DREAD numbers lie.
The five factors, anchored for LLM systems
Use a 0 to 10 scale and write down what 0, 5 and 10 mean for your system before anyone scores anything. Without anchors, one rater's 7 is another's 4, and the ranking becomes a record of who was in the room. The anchors below are a starting point for an LLM application with tools and retrieval. Each one names a property you can check, not a feeling.
| Factor | 0 | 5 | 10 | LLM-specific notes |
|---|---|---|---|---|
| Damage | no security impact | one user's data exposed, or a reversible wrong action | cross-tenant data, money moved, code executed, irreversible actions | bounded by what the model's tools and credentials can reach, not by what the text says |
| Reproducibility | never observed | works on a meaningful fraction of attempts | works essentially every time | measured as attack success rate over N trials; score the lower bound |
| Exploitability | needs insider access plus custom tooling | needs crafted prompts and some persistence | a pasted sentence, no account needed | natural-language attacks lower the skill bar; indirect injection needs only a web page or email |
| Affected users | none | a single tenant or a cohort | every user or every tenant | shared system prompts, shared caches and shared vector indexes widen the blast radius |
| Discoverability | requires internal knowledge | found by a motivated tester | public technique | most LLM attack classes are published; default to 10 |
Discoverability deserves a firm rule. Scoring it low rewards security through obscurity, and for LLM systems the attack classes are documented in public catalogues such as the OWASP Top 10 for LLM Applications. Many teams fix it at 10 for every threat. That turns DREAD into four real factors plus a constant, and that is fine. It also means two threats can only differ on the four factors that carry information.
Reproducibility is a probability
In classical software a bug either reproduces or it does not. An LLM pipeline has sampling temperature, retrieval that varies with index contents, and model updates that shift behaviour without any code change. So define reproducibility as a measured attack success rate. Run each attack N times under production settings and count successes. Then score the lower end of a confidence interval rather than the point estimate, because 3 successes in 10 trials and 30 in 100 are the same rate with very different certainty.
The Wilson score interval behaves well at small N and at rates near 0 or 1. The mapping from lower bound to score is a policy choice. Write it down once and reuse it, so that two teams scoring the same evidence get the same number. Re-measure after every model, prompt or retrieval change. A model upgrade can turn a threat that scored 2 into one that scores 8, with no code diff to review.
import math
from dataclasses import dataclass
def wilson_lower(successes: int, trials: int, z: float = 1.96) -> float:
"""95% lower confidence bound on an attack success rate."""
if trials == 0:
return 0.0
p = successes / trials
denom = 1 + z * z / trials
centre = p + z * z / (2 * trials)
margin = z * math.sqrt(p * (1 - p) / trials + z * z / (4 * trials * trials))
return (centre - margin) / denom
def reproducibility(successes: int, trials: int) -> int:
lb = wilson_lower(successes, trials)
for floor, score in ((0.5, 10), (0.2, 8), (0.05, 6), (0.01, 4)):
if lb >= floor:
return score
return 2 if successes else 0
@dataclass
class Threat:
name: str
damage: int
successes: int
trials: int
exploitability: int
affected: int
discoverability: int = 10 # assume the attacker knows
def scores(self):
r = reproducibility(self.successes, self.trials)
return (self.damage, r, self.exploitability, self.affected, self.discoverability)
def dread(self) -> float:
return sum(self.scores()) / 5
def priority(self) -> str:
s = self.dread()
band = "P1" if s >= 7.5 else "P2" if s >= 5.5 else "P3"
if self.damage >= 8 and band != "P1":
band = "P1 (damage floor)"
return band
Aggregation without hiding severity
The classic aggregate is the mean of the five scores. Its flaw is that the factors trade off against each other linearly. A threat that is trivially repeatable, easy and broad, but low in damage, can outrank a rarer threat that leaks customer data. Two adjustments fix most of this. First, a damage floor: any threat with damage at or above a threshold (8 here) is never ranked below the top band, whatever its average. Second, keep the vector next to the number in the register, written as D/R/E/A/D, so a reviewer can see why a threat ranks where it does. Some teams use damage times likelihood instead, treating R, E and A as likelihood inputs. That is a reasonable variant, but it needs its own anchors and shouldn't be mixed with averaged scores in one register.
Worked example: a RAG support agent
Consider a customer-support agent. It answers questions using retrieval over a shared vector index that holds every tenant's help-desk tickets, it can browse public web pages linked from tickets, and it has a send_email tool for follow-ups. A STRIDE pass over the data flow found five threats. The red team ran each attack 50 times against production settings. The cross-tenant retrieval threat comes from a missing tenant filter on one query path, so it reproduced deterministically. Feeding the evidence through the code above gives this ranking, taken directly from the script's output:
| Threat | Trials (succ/N) | D/R/E/A/D | Mean | Priority |
|---|---|---|---|---|
| Cross-tenant retrieval (missing filter) | 50/50 | 9/10/5/9/10 | 8.6 | P1 |
| Denial of wallet via max-length prompts | 50/50 | 5/10/9/7/10 | 8.2 | P1 |
Indirect injection to send_email exfiltration | 14/50 | 8/6/7/6/10 | 7.4 | P1 (damage floor) |
| System prompt extraction | 31/50 | 3/8/8/2/10 | 6.2 | P2 |
| Jailbreak to abusive reply | 6/50 | 4/6/6/3/10 | 5.8 | P2 |
Three things in this table are worth reading closely. First, the exfiltration attack succeeded on 14 of 50 runs, a point estimate of 28 percent. But the 95 percent lower bound is about 0.175, so it scores a 6 rather than an 8. That is the price of honesty at 50 trials, and running more trials is how you earn the higher score. Second, on the plain mean, denial of wallet (8.2) outranks exfiltration (7.4), even though one costs money and the other leaks customer data through an outbound email. The damage floor is what puts exfiltration back in P1. Without it the average would have pushed a data breach below a billing problem. Third, system prompt extraction reproduces often (lower bound about 0.48) but has low damage. It belongs in P2 as a reason to keep secrets out of prompts, not as an incident.
The mitigations follow from the vector, not the mean. The cross-tenant leak is fixed by enforcing tenant filters in the retrieval service, below the model, where no prompt can remove them. Exfiltration is cut by restricting send_email to the ticket's verified address and requiring confirmation for any new recipient. That lowers damage, not reproducibility, which is the right lever when the attack is a prompt you cannot fully block. Denial of wallet is handled with input token caps and per-tenant budgets. After each fix, re-run the same 50 trials and re-score. The register should show the before and after vectors.
Calibrating raters
Rater disagreement is DREAD's best-known weakness, and you can measure it directly. Have two people score each threat independently from the same evidence packet: the data-flow diagram, the attack transcript, the trial counts and the tool permissions. Any factor where they differ by more than 2 points goes to a short discussion. The outcome of that discussion is a sharper anchor in the rubric, not just a compromise number. Track the share of factors needing discussion over time. If it does not fall, your anchors are not specific enough.
Keep evidence attached to every score. A reproducibility score without trial counts, or a damage score without a named tool or data store, cannot be audited and will drift. For regulated environments the same packet doubles as the justification an assessor will ask for.
Failure modes
| Failure mode | What it looks like | Countermeasure |
|---|---|---|
| Unanchored scales | the same threat scores 4 and 8 from different raters | written 0/5/10 anchors per factor; dual scoring |
| Averaging hides severity | a nuisance outranks a data breach | damage floor; keep the vector in the register |
| Reproducibility guessed | a scary demo scored 10 after one success | measured success rate with a lower confidence bound |
| Stale scores | model or prompt changed, scores did not | re-score on every model, prompt, tool or retrieval change |
| Discoverability as a shield | low score because the bug is not public | fix discoverability at 10 |
| Damage judged from text | an output looks harmless but a tool call was made | score from the tools and data the model can reach |
| Scoring instead of enumerating | only known threats ever reach the register | pair DREAD with STRIDE, PASTA or attack trees |
DREAD against the alternatives
DREAD's advantage is speed and legibility. A product team can score twenty threats in an hour and explain the ranking to people outside security. CVSS is better for published vulnerabilities in shared software, because its vectors are standardised and comparable across organisations. But it was not designed for probabilistic model behaviour or for business-specific damage. The OWASP Risk Rating Methodology splits likelihood and impact into more factors and suits teams that need more granularity. Quantitative approaches that estimate loss in money, with ranges and simulation, are the right tool when you need to justify spending, but they cost far more per threat. A practical stack is to enumerate with STRIDE or attack trees, rank with anchored DREAD, and quantify only the few P1 items that need a budget decision.
Continue with STRIDE for LLM systems to enumerate threats, PASTA for LLM threat modeling for a risk-centric process, the AI risk register to track the scores, AI risk assessment for quantitative estimates, and threat modeling for LLM applications for attack trees on a similar RAG agent.
What to do next
- Write 0/5/10 anchors for each factor in terms of your system's tools, data stores and tenants, and get them signed off before scoring.
- Fix discoverability at 10 and record that decision in the rubric.
- For every threat, run the attack at least 50 times under production settings and score reproducibility from the Wilson lower bound.
- Adopt a damage floor so that no high-damage threat can be averaged out of the top band.
- Score independently with two raters, and discuss any factor where they differ by more than 2 points.
- Store the D/R/E/A/D vector, the trial counts and the evidence in the risk register next to the priority.
- Re-run trials and re-score after every model, prompt, tool or retrieval change, and track before and after vectors for each mitigation.