When an assistant reads your mail, browses the web and calls tools on your behalf, the person in front of the screen is often the last line of defence. They are the one who notices that a summary of an inbox somehow involved sending a message, or that an answer about a contract came from a page nobody uploaded. They can only notice what the interface shows them. Transparency UX is the design of that showing, and in a system exposed to prompt injection it is a security control, not a courtesy.
This article treats it that way: what to show, how to build it so that the model cannot forge it, how attackers abuse it, and how to measure whether it works. Two neighbouring topics have their own pages. The anatomy of an approval dialog is in Agent Permission Prompt Patterns, and proving that citations are real is in Citation Verification. Here the focus is the continuous surface around them: the activity, provenance and data-flow information a user sees while the agent works.
Why transparency is a security control
Prompt injection works by getting untrusted text into the model's context and letting it steer actions. Most defences try to stop that before it acts. Transparency handles what slips through: if the effect of an injected instruction is visible at the moment it happens, the user can stop it, report it, or at least not trust the result. That only works if three conditions hold.
- Completeness. Every consequential action appears, including the ones the model would rather not mention. An injected instruction will often ask the model to keep quiet; a summary the model writes will comply.
- Integrity. What appears was produced by code that observed the action, not by the model describing it. Model text can be shaped by the attacker; the runtime's record of a tool call cannot.
- Legibility. The user can tell at a glance what is ordinary and what is unusual. A log of four hundred lines is complete and useless.
There is also a regulatory floor. Article 50 of the EU AI Act requires, among other things, that people be told when they are interacting with an AI system unless it is obvious. Law-firm summaries checked on 2026-10-04 report those duties applying from 2 August 2026 and not deferred by the 2026 omnibus, apart from a later date for marking synthetic content from systems already on the market. That is disclosure, the thinnest layer; this is not legal advice, and the security value comes from the layers above it.
What to show: six questions
A useful way to organise the surface is by the question each element answers for the user.
| Question | Element | Source of truth |
|---|---|---|
| Am I talking to a machine? | persistent AI label on every reply and on outbound messages | product config |
| What did it read? | source list with a trust label per item: you, your org, the open web, an inbound message | retrieval and tool events |
| What did it do? | activity timeline: tool name, target, parameters summary, outcome | tool executor events |
| What left my boundary? | egress summary: recipients, domains contacted, files uploaded | network and tool events |
| What will it remember? | memory writes shown when they happen, with undo | memory store events |
| How sure is it? | explicit markers when an answer rests on a single untrusted source or none | retrieval metadata |
Every row's source of truth is the runtime, never the model. The model may add a natural-language explanation next to an element, but the element itself, the badge, the recipient list, the count of web pages contacted, is rendered from structured events. The "how sure" row is the softest: model self-reported confidence is poorly calibrated and easy to manipulate, so base the marker on facts the runtime knows, such as how many independent sources supported the answer and whether any were untrusted.
Architecture: two channels
The architecture follows from the integrity condition: two channels, and only one of them can produce interface chrome.
Concretely, the tool executor, the retriever and the outbound HTTP layer each emit typed events to an activity log. A trust labeller attaches the origin of every piece of content as it enters the context. The renderer has two inputs and different rules for each.
from dataclasses import dataclass
from enum import Enum
import html
class Trust(Enum):
USER = "you"; ORG = "your organisation"; EXTERNAL = "outside source"
@dataclass(frozen=True)
class Activity:
kind: str # "read" | "tool" | "egress" | "memory"
name: str # tool or source name, from the registry, never from model text
target: str # recipient, URL host, file path
trust: Trust
consequential: bool # sends, writes, pays, shares
triggered_by_external: bool # an EXTERNAL item was in context when the call was planned
def render_activity(a: Activity) -> str:
cls = "act-warn" if (a.consequential and a.triggered_by_external) else "act"
badge = f'<span class="trust trust-{a.trust.name.lower()}">{a.trust.value}</span>'
return (f'<li class="{cls}">{html.escape(a.name)} '
f'<span class="target">{html.escape(a.target)}</span> {badge}</li>')
def render_model_text(md: str) -> str:
# Markdown subset only: no raw HTML, no images, links shown with their real host.
return safe_markdown(md, allow_images=False, allow_html=False, show_link_host=True)The triggered_by_external flag is the most useful bit in the schema. It does not prove an injection, but a consequential action planned while untrusted text was in context is exactly the case the user should look at, and it can drive emphasis in the timeline and stricter confirmation. Computing it is cheap because the runtime already knows what it put in the context window.
Attacks on the transparency surface
Once transparency is a control, attackers target it. The mechanisms below are described so you can test for them; none needs a working payload to understand.
- Forged status in model text. Injected content asks the model to print something that looks like interface chrome: a verified badge, a fake "action cancelled" line, an imitation system notice. If the content pane renders rich formatting, the forgery can be pixel-close. Defence: the content pane cannot render the styles, icons or components the activity pane uses, and the two panes are visually distinct.
- Concealment. The injection asks the model not to mention an action. Any summary written by the model obeys. Defence: actions are listed from the activity log, so the model's silence changes nothing.
- Exfiltration through rendering. If model output may contain images or auto-loading links, the model can be steered to build a URL with conversation data in its query string, and the browser sends it the moment the reply renders, with no tool call at all. Data exfiltration via LLMs covers this channel; for transparency the lesson is that rendering is egress and belongs in the egress summary, or better, is disabled.
- Link text versus destination. A link reading as a familiar site points elsewhere. Defence: show the real host beside every link in model text.
- Flooding. Many harmless actions bury one harmful one. Defence: group routine reads, and never group consequential actions.
- Over-disclosure. Showing raw retrieved text or the agent's full reasoning can surface the system prompt, other users' data or the injected text itself in a trusted-looking pane. Show sources by title, origin and trust label; keep raw content behind an explicit expand, still escaped.
Worked example: an injected email
A user asks an email assistant to summarise unread mail. One message, from an outside sender, contains text telling the assistant to forward the latest invoice to an external address and not to mention it. Follow the two designs.
Model-authored transparency. The reply is a tidy summary of five messages. If the injection worked, an invoice was forwarded and the summary says nothing, because the model was told not to. The user has no reason to look.
Runtime-authored transparency. The reply is the same summary, but the activity pane under it shows five reads, four labelled your organisation and one labelled outside source, then one consequential line in warning style: forward_message to an external domain, triggered while outside content was in context. The egress summary reads one message sent to one external recipient. The user sees an action they did not ask for, in a place the model cannot edit. If forwarding also needs approval, the dialog from the permission prompt article appears first; the timeline is what makes the approval request make sense in context.
Two design choices made the difference: the action list came from the executor, and the trust label of the source was carried through to the action it influenced. Neither depended on detecting the injection.
Measuring whether it works
Transparency that users do not read is decoration, so measure it like any control. Run seeded drills: plant benign canary instructions in test content that cause a visible but harmless action, and measure how often participants notice and report it, and how long it takes. Compare designs with the same drill rather than asking people which they prefer. Track in production how often users open the activity pane, undo memory writes and stop runs mid-way; a near-zero rate may mean the surface is invisible, not that nothing went wrong. Audit completeness automatically: for a sample of sessions, every consequential executor event must have a rendered activity line, and a test should fail the build if a new tool lacks an activity renderer.
def test_every_tool_is_visible(registry, renderer):
missing = [t.name for t in registry.tools()
if t.consequential and not renderer.has_template(t.name)]
assert not missing, f"tools with no activity line: {missing}"
def audit_session(events, rendered_lines):
# events: executor log for one session; rendered_lines: what the UI actually showed
shown = {line.event_id for line in rendered_lines}
hidden = [e for e in events if e.consequential and e.id not in shown]
return len(hidden) == 0, hiddenThe first check runs in CI and stops a new tool shipping invisibly, which is the most common way completeness erodes: someone adds a capability, the executor runs it, and nobody wrote the line that shows it. The second runs on sampled production sessions and compares the executor's log with what the client reported rendering. A mismatch is a bug in the interface, or a client that has been tampered with, and either deserves a ticket.
Design the timeline for scanning rather than reading. Group consecutive routine reads into one collapsible line such as "read 12 messages, 1 from outside your organisation", keep every consequential action on its own line, sort by time, and put the trust badge in a fixed column so the eye can run down it. Use wording a non-specialist understands: the tool's registry name is for engineers, and the line a user reads should say "sent an email to" rather than naming a function. Test the wording with the same drills, because a line that people misread is as useless as a missing one.
Failure modes
- Activity pane collapsed by default and never opened. Keep consequential lines visible in the reply itself.
- Trust labels lost at chunking. A retrieval pipeline that merges sources forgets origin; carry it per chunk, as in prompt injection via RAG.
- Outbound messages without disclosure. Email or chat sent by the agent to third parties should say an AI drafted it.
- Warning fatigue. If every action is flagged, none is; reserve emphasis for consequential actions influenced by outside content.
- Logs that leak. The activity log holds recipients and file names; apply the same retention and access rules as the data itself.
Trade-offs
Transparency costs screen space and attention, and over-explaining reduces trust as fast as hiding does. It also does not replace authorisation: a user who sees an action after it happened can only clean up. Pair it with least-privilege tools and approvals for irreversible actions, as in human-in-the-loop approval gates, and with the confused-deputy analysis in the confused deputy problem. The strongest designs use transparency for the many reversible actions and approvals for the few that are not.
What to do next
- Inventory every tool and outbound channel; mark which are consequential.
- Emit a typed activity event from the executor for each call, with target and trust, and render the timeline only from those events.
- Strip images, raw HTML and auto-loading content from model output, and show the real host next to every link.
- Carry a trust label from ingestion to context to action, and compute the triggered-by-external flag.
- Make the content pane incapable of rendering the activity pane's components and styles.
- Add an egress summary and AI disclosure on outbound messages.
- Run a seeded canary drill, measure notice rate and time to notice, and add a completeness check to CI.