Most teams agree threat modeling is worth doing, then do it once, in a meeting that produces a diagram nobody opens again. The method is rarely the problem. The practice is: who comes, what comes out, how findings become work, and what brings the model back when the design changes.
This page is about running that practice. The concepts (data-flow diagrams, STRIDE, attack trees and residual risk) are covered in threat modeling architecture, and this page uses them without re-teaching them. Here you will run a complete session on one realistic system, a webhook delivery service, go deep on the threat that matters most for it, rank and decide, store the model as code next to the service with a CI check, and set the triggers that keep it current.
The four questions as a session agenda
The Threat Modeling Manifesto reduces the method to four questions: What are we working on? What can go wrong? What are we going to do about it? Did we do a good enough job? Each question maps onto part of a meeting, and each part has an input and an output. If you can't name the output of a step, that step is where sessions drift.
| Question | Session step | Input | Output |
|---|---|---|---|
| What are we working on? | Draw or correct the data-flow diagram (15 min) | Design doc, API list, deployment diagram | DFD with numbered flows and trust boundaries |
| What can go wrong? | Walk each boundary-crossing flow with STRIDE (30-40 min) | The DFD | Threat list, one row per threat, tied to a flow |
| What are we going to do about it? | Rank and decide (15 min) | Threat list | Each threat is mitigate, eliminate, transfer or accept, with an owner |
| Did we do a good enough job? | Review after the build, and on triggers | Merged code, tests, incidents | Evidence per mitigation; updated model |
Keep the room small: the design owner, a builder, a security person or trained champion, and someone who knows the deployment. Scope one change, not "the platform", and book 60 to 90 minutes. If the DFD alone takes more than 20, the scope is too wide; split it.
Worked example: a webhook delivery service
The system: customers register an HTTPS URL and receive a signing secret. When an event such as order.paid occurs, a delivery worker reads the subscription, signs the JSON body with the customer's secret, POSTs it to the URL, and retries with backoff if the call fails. Every attempt is logged. Flows are numbered because every threat will cite one, which is what lets a tool check coverage later.
F5 is what makes this design interesting: a server inside our network makes an outbound request to an address chosen by an untrusted party.
What can go wrong: eliciting threats per flow
Walk the boundary-crossing flows first, using each STRIDE letter as a prompt and recording only design-specific threats. "An attacker could exploit a vulnerability" is not a threat; "a customer registers http://169.254.169.254/ and the worker posts to the cloud metadata service" is. The first rows of a real list:
| ID | Flow | STRIDE | Threat | Candidate response |
|---|---|---|---|---|
| T1 | F5 | Elevation / Info disclosure | Customer URL resolves to an internal or metadata address (SSRF) | Resolve-and-check plus egress firewall |
| T2 | F5 | Spoofing | Attacker posts fake events to the customer's endpoint | HMAC signature over timestamp and body |
| T3 | F5 | Tampering / replay | Captured delivery replayed later | Timestamp in signed payload; receiver rejects old ones |
| T4 | F1 | Spoofing | Tenant A registers a URL on tenant B's subscription | Authorisation check on tenant id, not on the subscription id alone |
| T5 | F4 | Info disclosure | Signing secrets readable in a DB dump or backup | Envelope-encrypt secrets; workers decrypt via KMS |
| T6 | F5 | Denial of service | Slow customer endpoint ties up all workers | Per-request timeout, per-tenant concurrency cap |
| T7 | F5 | Denial of service (outbound) | Our service used to flood a victim URL | Verify the endpoint with a challenge before activating; rate limits |
| T8 | F6 | Info disclosure | Payloads with personal data copied into the delivery log | Log status and hashes, not bodies |
SQL injection, XSS and dependency CVEs are missing on purpose. Scanners and code review find those (see SAST and DAST); the session is for threats that come from design decisions.
Deep dive: SSRF in outbound delivery
T1 deserves its own section because the obvious fix is wrong. Checking the URL string for "localhost" or private IP literals at registration time fails three ways. A hostname can resolve to a private address. DNS can return a public address at registration and a private one at delivery (DNS rebinding). And IPv6-mapped forms such as ::ffff:10.0.0.1 get past naive string checks. The check has to happen at delivery time, on the resolved addresses, and the connection has to go to the address that was checked.
import ipaddress
import socket
from urllib.parse import urlsplit
def resolve_public(url: str) -> tuple[str, str, int]:
"""Return (hostname, checked_ip, port) or raise if any resolved address is non-public."""
parts = urlsplit(url)
if parts.scheme != "https" or not parts.hostname:
raise ValueError("only https URLs with a hostname are accepted")
port = parts.port or 443
infos = socket.getaddrinfo(parts.hostname, port, proto=socket.IPPROTO_TCP)
addrs = []
for info in infos:
addr = ipaddress.ip_address(info[4][0])
if addr.version == 6 and addr.ipv4_mapped:
addr = addr.ipv4_mapped # ::ffff:10.0.0.1 is 10.0.0.1
if not addr.is_global: # private, loopback, link-local, CGNAT, ...
raise ValueError(f"{parts.hostname} resolves to non-public {addr}")
addrs.append(addr)
# Reject if ANY answer is internal; then connect to one we checked, sending the
# original hostname as SNI and Host so TLS verification still uses the name.
return parts.hostname, str(addrs[0]), portThe function rejects a host if any answer is non-public, because an attacker controls the DNS records and can mix a public and a private answer. The caller must then open the TCP connection to the returned IP while presenting the hostname for TLS. Re-resolving inside the HTTP client gives rebinding its window back. Redirects need the same treatment: turn off automatic redirect following, or run every Location target through the same check.
Code is the second line of defence; the network is the first. Run delivery workers in a subnet whose egress allows only the public internet, with internal ranges and the metadata endpoint blocked, ideally through an egress proxy. Each layer covers the other's mistakes, the layered thinking zero trust applies to outbound traffic.
Spoofing and replay: signing deliveries
T2 and T3 share one mitigation. Sign a string that includes a timestamp, send the timestamp in a header, and document that receivers must reject deliveries older than a few minutes and compare signatures in constant time:
import hashlib, hmac, time
def sign(secret: bytes, body: bytes, ts: int | None = None) -> dict[str, str]:
ts = ts or int(time.time())
mac = hmac.new(secret, f"{ts}.".encode() + body, hashlib.sha256).hexdigest()
return {"X-Webhook-Timestamp": str(ts), "X-Webhook-Signature": f"v1={mac}"}
def verify(secret: bytes, body: bytes, headers: dict[str, str], max_skew: int = 300) -> bool:
ts = int(headers["X-Webhook-Timestamp"])
if abs(time.time() - ts) > max_skew:
return False
expected = sign(secret, body, ts)["X-Webhook-Signature"]
return hmac.compare_digest(expected, headers["X-Webhook-Signature"])The v1= prefix lets you rotate algorithms later. A mitigation that depends on the receiver verifying correctly is only partly yours, so record it as an assumption in the model.
Ranking and deciding
Ranking needs consistency, not precision. Score likelihood and impact from 1 to 3, multiply, and debate only the rows where people disagree. Many-factor scores invite arguments about arithmetic instead of risk.
| Threat | Likelihood | Impact | Score | Decision |
|---|---|---|---|---|
| T1 SSRF | 3 (well-known, scanned for) | 3 (cloud credentials) | 9 | Mitigate before launch: code check + egress rules |
| T4 cross-tenant | 2 | 3 | 6 | Mitigate before launch; add an authorisation test |
| T5 secrets at rest | 2 | 3 | 6 | Mitigate: envelope encryption |
| T6 slow endpoint | 3 | 2 | 6 | Mitigate: timeouts and per-tenant caps |
| T2/T3 spoofing, replay | 2 | 2 | 4 | Mitigate on our side; receiver guidance in docs |
| T7 flooding a victim | 2 | 2 | 4 | Accept for launch (rate limits exist); add activation challenge; review 2027-01-15 |
| T8 personal data in logs | 2 | 2 | 4 | Eliminate: stop logging bodies |
Every row ends in mitigate, eliminate (remove the feature or data), transfer (a contract or provider takes the risk) or accept. Eliminate is underused: T8 is fixed for free by not storing bodies. Acceptance is legitimate only when written down with an owner and a review date.
Keeping the model as code
A model in a slide deck goes stale the day the design changes. Keep it in the service's repository as data, so it is reviewed in pull requests and checked by CI. Tools exist for this (OWASP pytm generates DFDs and threats from Python, and OWASP Threat Dragon stores models as JSON), but a hundred lines of your own are enough to start, and they encode your team's rules:
# threatmodel/webhooks.py
from dataclasses import dataclass
from datetime import date
FLOWS = { # id: (source, destination, crosses_trust_boundary)
"F1": ("customer", "webhook-api", True),
"F2": ("webhook-api", "subscription-db", False),
"F3": ("event-bus", "delivery-worker", False),
"F4": ("subscription-db", "delivery-worker", False),
"F5": ("delivery-worker", "customer-url", True),
"F6": ("delivery-worker", "delivery-log", False),
}
@dataclass
class Threat:
id: str
flow: str
title: str
status: str # "open" | "mitigated" | "accepted" | "eliminated"
owner: str
evidence: str = "" # test name or PR proving the control exists
review_by: date | None = None
THREATS = [
Threat("T1", "F5", "SSRF to internal ranges", "mitigated", "payments-platform",
evidence="tests/test_delivery.py::test_rejects_private_resolution"),
Threat("T7", "F5", "Service used to flood a victim URL", "accepted", "payments-platform",
review_by=date(2027, 1, 15)),
# ...
]
def check(today: date) -> list[str]:
errors = []
covered = {t.flow for t in THREATS}
errors += [f"{f} crosses a boundary but has no threats" for f, (_, _, crosses)
in FLOWS.items() if crosses and f not in covered]
for t in THREATS:
if t.flow not in FLOWS:
errors.append(f"{t.id} refers to unknown flow {t.flow}")
if t.status == "mitigated" and not t.evidence:
errors.append(f"{t.id} is mitigated without evidence")
if t.status == "accepted" and (t.review_by is None or t.review_by < today):
errors.append(f"{t.id} acceptance has no review date or has expired")
return errors
if __name__ == "__main__":
problems = check(date.today())
print("\n".join(problems) or "threat model ok")
raise SystemExit(1 if problems else 0)The rules are deliberately few: a new boundary-crossing flow fails the build until someone considers it, a mitigation must cite evidence (ideally a test that fails if the control goes), and an accepted risk expires. People still judge whether the threats are right; the check only stops silent drift.
Did we do a good enough job? Triggers for re-review
"Did we do a good enough job?" is answered twice: once after the build, by checking the evidence column against merged code, and again whenever something happens that could invalidate the model. Write the triggers into the team's definition of done:
- A new flow that crosses a trust boundary: a new external API, a new third party, a new queue consumer in another account.
- A new kind of data in an existing flow, especially credentials or personal data.
- A change in authentication or tenancy model.
- An incident or a penetration test finding in this service. Ask why the model missed it, not only how to fix it.
- An accepted risk reaching its review date (the CI check enforces this one).
A re-review is fifteen minutes of STRIDE on the changed flows only, which is cheap because the model is a file, not a picture.
How threat-modeling practice fails
How the practice fails, and the countermeasure for each failure:
- Security team as author. If security writes the model alone, engineers don't own the findings. The design owner writes it; security facilitates and challenges.
- Checklist theatre. Filling every STRIDE cell for every element produces hundreds of rows, most of them noise. Start with boundary-crossing flows and stop when new rows stop being specific.
- Findings that never become work. A threat without a ticket and an owner is a note, not a decision. The session ends only when every row has a decision and an owner.
- Mitigations without proof. "We validate URLs" is a claim. A test named in the evidence column is proof. Without one, a later refactor can quietly remove the control.
Choosing a method
STRIDE per flow fits most product engineering because it is fast and engineers can learn it in an afternoon. Other methods suit other questions:
| Method | Asks | Use when | Cost |
|---|---|---|---|
| STRIDE per flow | How can each crossing be abused? | Default for services and features | Low; one session |
| LINDDUN | How can this system harm users' privacy? | Personal data, analytics, tracking | Medium; needs privacy expertise |
| Attack trees | How exactly can an attacker reach one goal? | Deep analysis of one high-value goal, such as T1 | Medium; one tree per goal |
| PASTA | Which threats matter to the business, via a seven-stage process? | Regulated, risk-led programmes | High; days, many stakeholders |
They combine: STRIDE for breadth, then an attack tree for the threats that score 9. A light model on every significant change beats a heavy model once a year.
What to do next
- Pick one change that ships in the next month and adds a trust-boundary crossing. Book 90 minutes with four people: the design owner, a builder, someone from security and someone from operations.
- Before the session, draw the DFD with numbered flows and dashed boundaries. Spend the meeting correcting it, not drawing it.
- Run STRIDE on boundary-crossing flows only. Write threats specific enough that someone could write a test for each.
- Score with 1-3 by 1-3, and give every row a decision (mitigate, eliminate, transfer or accept), an owner and a ticket.
- If any flow fetches an address supplied by a user, apply the resolve-and-check pattern and block internal ranges at the network layer. Do both.
- Commit the model as data next to the code, with a CI check for uncovered flows, missing evidence and expired acceptances.
- Write the re-review triggers into your definition of done, and after the next incident ask why the model missed it.