A data flow diagram (DFD) is the cheapest security artifact you can make for an LLM application, and often the most useful. It shows where data comes from, where it goes, what stores it, and which arrows cross a trust boundary. Most LLM incidents happen on one of those arrows. A customer's email address goes to a model vendor nobody approved. A retrieved web page carries instructions into the prompt. A trace log keeps every conversation for ever because nobody set a retention period.

This article explains the notation from first principles, then the parts LLM systems add: the context window, embeddings, prompt caches, tools and traces. It also covers what to write on every arrow. We then model a support assistant as code and run a short checker over it. The checker turns data class, boundary and retention into a list of missing controls. The aim is a diagram that answers real questions in a design review and stays true after the next release. Finding attack paths through the graph is a separate step, covered in LLM-specific threat modeling; the DFD is the map that step walks over.

The four elements and the boundary

DFDs come from structured analysis. Threat modelers use a small subset, the one behind STRIDE-per-element analysis and tools such as the Microsoft Threat Modeling Tool and OWASP Threat Dragon: four element types and one annotation.

  • External entity (rectangle): something you do not control and cannot change, such as a user, a partner system or a hosted model API. You can only constrain what you send it and how you treat what comes back.
  • Process (circle or rounded box): code you run that transforms data, such as the chat backend, a retriever, a tool executor or a guard model.
  • Data store (two parallel lines): data at rest, such as a database, vector index, object bucket, cache or log. Stores outlive requests, which is why they hold most of your privacy risk.
  • Data flow (arrow): data in motion between two elements. Flows have a direction. A request and its response are two flows, because they carry different data.
  • Trust boundary (dashed line): a place where the level of trust changes, for example a network edge, another tenant, another company or another privilege level. A flow that crosses a boundary is where you check authentication, authorisation, validation and encryption.

Two rules keep diagrams honest. Data never moves without a process, so a store writing straight to another store means a process is missing. And every store needs a flow in and a flow out; if nothing reads it, consider deleting it.

Three levels of zoom

Draw at three zoom levels and stop when the next level would not change a decision.

  • Context diagram (level 0). The whole system is one process, surrounded by the external entities it talks to. Use it for privacy and vendor reviews: it shows every organisation that touches the data.
  • Level 1. Open the system into its main processes and stores. For an LLM app this usually means the front end, an orchestrator or prompt builder, a retriever, a tool layer, the model call, and stores for documents, embeddings, conversations and traces. Most threat modeling happens here.
  • Level 2. Open a process further only where it hides a boundary. The usual candidate is the tool layer, where one box can hide a code interpreter, a mail sender and a database writer.

Number every flow (F1, F2, ...) and keep the numbers stable, because tickets and review notes refer to them.

What to write on every arrow

An arrow labeled just data tells a reviewer nothing. Give each flow a short record. In practice five fields carry most of the value:

FieldExampleQuestion it answers
Data classespii, untrusted_text, model_output, secretWhat could leak, and what could steer the model?
Protocol and authHTTPS, workload identity, API keyWho can send this, and can it be read in transit?
Trust of contentuser-supplied, retrieved web, internalCan this text contain instructions you did not write?
Retention at destination0 days, 30 days, 10 years, unknownHow long does this outlive the request?
Controls in placeredact, schema_check, tagged_as_dataWhat already stands between the data and harm?

Data classes must be a closed list with written definitions, otherwise teams will invent their own labels. For LLM work, add two classes that ordinary DFDs lack. untrusted_text is any text an outsider could have written, such as user messages, retrieved documents, web pages, emails and tool results. It matters because the model cannot reliably tell data from instructions, as prompt injection through retrieval shows. model_output is anything the model produced. It matters because it inherits the risk of every input that shaped it, so a downstream tool must treat it as untrusted too.

Write unknown when you do not know a value, and treat unknown as a finding. A blank field hides the gap.

What LLM systems add to the diagram

Several parts of an LLM stack are easy to leave off a diagram, and each one matters:

  • The context window is a process that mixes trust levels. Draw the prompt builder as its own process. System prompt, user message, retrieved passages and tool results all flow into it, and one flow leaves it. That single outbound flow carries the union of every inbound data class. If any input is pii, the prompt is pii. If any input is untrusted_text, the prompt is untrusted_text.
  • Embeddings are derived data, not anonymous data. Research on embedding inversion, such as Vec2Text by Morris and colleagues in 2023, recovered much of the source text from some embedding models. Label vector stores with the classes of the text they were built from.
  • The model vendor is an external entity with its own stores. Prompt caching, abuse monitoring and logging on the vendor side are stores you cannot see. Draw them inside the vendor boundary with the retention your contract states. If you cannot find it in the contract, write unknown.
  • Tools are processes with privileges. The tool call is a flow from your orchestrator to a tool, and its arguments are model_output. The privileges of the tool are what turn a bad completion into an incident.
  • Traces and evaluation sets are stores. Observability pipelines often copy full prompts and completions to a third-party tracing service. That is a new vendor boundary and a new retention clock; see audit logging for LLM systems for what to keep.

Worked example: a support assistant

Take a customer support assistant. A customer chats on the website. The app, a chat backend with a prompt builder inside it, asks a retriever for help-centre passages, then sends a prompt to a hosted model. The model can call one tool that opens a ticket. Every turn is written to a trace log. The figure leaves off the job that writes the knowledge base and the readers of the tickets and trace stores, to stay small; a full diagram includes them.

Level 1 DFD: customer support assistantInternetApplication zoneModel vendorData zoneCustomerexternalAppchat backend + prompt builderHosted LLMexternalRetrieverprocessTicket toolprocessKB indexstore, 10 yTickets DBstore, ? daysTrace logstore, ? daysF1F5 promptF6F2 queryF4 passagesF3F7 tool callF8F9 traceEvery arrow that crosses a dashed line is a boundary crossing and gets its own row in the review.
The support assistant at level 1. Dashed lines are trust boundaries. F5 and F9 cross into zones with different owners and retention, so they get the closest review.

Walk the flows in order. F1 carries pii and untrusted_text. F3 comes from an internal knowledge base, but many staff can edit it, so mark it untrusted_text too, and F4 carries those passages back to the app. F5 is the interesting one. It is the union of F1 and F4, it leaves your organisation, and it lands in vendor stores whose retention you must look up. F6 returns model_output. F7 is that output turned into tool arguments. F8 writes pii to the tickets database, and F9 copies whole conversations into a trace store that nobody owns.

The diagram as code, with a checker

A whiteboard diagram drifts from the system within a few releases. Keep a machine-readable model next to the code and run a checker in CI. Each rule below is a predicate over a flow and its endpoints, plus the controls the flow must declare. It checks each arrow against policy; it does not search paths.

from dataclasses import dataclass, field

@dataclass
class Node:
    name: str
    kind: str              # "external", "process", "store"
    zone: str              # trust zone; a zone change is a boundary crossing
    retention_days: int = 0  # 0 means "not declared"

@dataclass
class Flow:
    src: str
    dst: str
    label: str
    data: set = field(default_factory=set)      # data classes on the wire
    controls: set = field(default_factory=set)  # controls the owner declares

RULES = [
    (lambda f, s, d: s.zone != d.zone, {"authn", "tls"},
     "every boundary crossing needs authenticated, encrypted transport"),
    (lambda f, s, d: "untrusted_text" in f.data and d.name == "llm",
     {"tagged_as_data"}, "untrusted text entering the model must be marked as data, not instructions"),
    (lambda f, s, d: "pii" in f.data and d.zone == "vendor", {"redact", "dpa"},
     "personal data leaving to a vendor needs redaction and a processing agreement"),
    (lambda f, s, d: "pii" in f.data and d.kind == "store" and d.retention_days == 0,
     {"retention_declared"}, "a store receiving personal data must declare retention"),
    (lambda f, s, d: "model_output" in f.data and d.name.endswith("_tool"), {"schema_check"},
     "model output that drives a tool must be schema-validated first"),
]

def check(nodes, flows):
    by = {n.name: n for n in nodes}
    findings = []
    for f in flows:
        s, d = by[f.src], by[f.dst]
        for pred, need, why in RULES:
            if pred(f, s, d):
                missing = need - f.controls
                if missing:
                    findings.append((f.label, sorted(missing), why))
    return findings

The worked example, encoded:

nodes = [
    Node("customer", "external", "internet"),
    Node("app", "process", "app"),
    Node("retriever", "process", "app"),
    Node("kb", "store", "data", retention_days=3650),
    Node("tickets", "store", "data"),
    Node("llm", "external", "vendor"),
    Node("trace_log", "store", "data"),
    Node("ticket_tool", "process", "app"),
]
flows = [
    Flow("customer", "app", "F1 question", {"pii", "untrusted_text"}, {"authn", "tls"}),
    Flow("app", "retriever", "F2 query", {"pii"}),
    Flow("kb", "retriever", "F3 passages", {"untrusted_text"}, {"authn", "tls"}),
    Flow("retriever", "app", "F4 passages", {"untrusted_text"}),
    Flow("app", "llm", "F5 prompt", {"pii", "untrusted_text"}, {"authn", "tls", "dpa"}),
    Flow("llm", "app", "F6 completion", {"untrusted_text", "model_output"}, {"authn", "tls"}),
    Flow("app", "ticket_tool", "F7 tool call", {"pii", "model_output"}),
    Flow("ticket_tool", "tickets", "F8 write", {"pii"}, {"authn", "tls"}),
    Flow("app", "trace_log", "F9 trace", {"pii", "untrusted_text"}, {"authn", "tls"}),
]
for label, missing, why in check(nodes, flows):
    print(f"{label}: missing {', '.join(missing)} -- {why}")

The chat backend and prompt builder are one node, app, because they share a process and a zone. Running it prints five findings. F5 lacks tagged_as_data and redact, F7 lacks schema_check, and F8 and F9 lack retention_declared. No rule fires on the other flows: either their declared controls already satisfy policy, or no rule applies to them.

From findings to decisions

Each finding is a design decision, not a ticket to close with a comment. For F5, decide what personal data the model really needs. Order numbers usually help it answer. Full names and email addresses rarely do. Mask the rest before the call, as described in PII leakage in LLM systems. For F7, define a JSON schema for the ticket tool, reject calls that do not match it, and check the ticket owner against the authenticated session, not against anything the model says. For F8 and F9, an owner must write down a retention period and a deletion job must enforce it.

Declared controls are claims, so test them. Once a flow says redact, a CI test should send a canary email address through the prompt builder and assert it never appears in the outbound request. Egress logs are the other half. Compare actual outbound destinations with the external entities on the diagram, and treat any unknown destination as a missing flow. Egress control gives you the logs to do this.

Failure modes

The ways DFD work goes wrong are predictable:

  • Drawing the intended system. The diagram shows the architecture from the design document, not the one in production with its debug logger and the analytics SDK someone added. Build it from code, config and egress logs, then compare it with the design.
  • One arrow for request and response. The response often carries different classes, such as model_output, and goes to a different place, such as the trace log.
  • The model as a trusted process. Drawing the model inside your application zone hides that its output is shaped by untrusted input. Treat it as an external entity, even when you host it yourself.
  • Boundaries only at the network edge. Tenants sharing one vector index, or a tool running with a service account broader than the user, are boundaries with no firewall. Draw them.
  • Diagram rot. Run the checker on every change to prompts, tools, vendors or stores.

Trade-offs

ChoiceGainCost
Whiteboard DFDFast, good for a first conversationRots within weeks; nothing can test it
Diagram as code plus checkerReviewable diffs, CI enforcementSomeone must own the rule list
Generated from traces or egress logsShows what really happensMisses flows that rarely run; needs cleanup
Coarse data classes (3 to 5)Teams actually label flowsLess precise policy

What to do next

  1. Draw the level 0 diagram today, listing every organisation that receives prompt or completion data.
  2. Write a closed list of four or five data classes, including untrusted_text and model_output.
  3. Draw level 1, number every flow, and fill in the five fields. Write unknown wherever you are not sure.
  4. Encode it with the checker above and add rules for your own policies.
  5. Turn every unknown retention into a ticket with an owner and a due date.
  6. Add a CI test with a canary value for each flow that claims redaction.
  7. Compare egress logs with the diagram each month and add every flow you missed.
Key takeaway: A DFD for an LLM app is four element types, trust boundaries, and a short record on every arrow: data classes, auth, content trust, retention and controls. Treat the prompt builder as a process that merges all its inputs, the model as untrusted, and embeddings, caches and traces as stores. Keep the diagram as code, check each flow against policy in CI, and compare it with egress logs so it describes the system you run.