Every LLM application has two versions of each conversation. One is what the person in front of the screen sees: their messages, the replies, perhaps a rendered web page or document. The other is what the model actually receives: a long token sequence containing an operator's system prompt, tool definitions, retrieved text, memory, summaries of earlier turns, and the raw source of whatever document was attached, including parts no human ever saw rendered. A shadow prompt is any content in the second version that has no visible counterpart in the first.

Some shadow content is normal and necessary: nobody expects to read a product's system prompt in the chat window. Some is hostile: a web page with an instruction in white text, a PDF with a comment that addresses the summariser, a tool description that tells the model to forward data. Most security problems in this space come from the same root cause, that people approve, trust and audit the visible version while the model acts on the shadow one. This article treats the gap as an engineering object you can inventory, measure, log and control.

Two views of every conversation

Two projections of the same context: what the user sees and what the model seesSystem promptoperator, hiddenTool schemasdescriptions, hiddenRetrieved pagespartly hiddenMemory, summariesrarely shownUser messagevisiblePrompt assemblertags each segment with source, trust, visibilityModel viewexact tokens, loggedUser viewrendered UIGap analyserdiff, score, policyPolicystrip, quarantine, disclose, or blockShadow content is everything in the model view with no counterpart in the user view.
Segments from several channels are assembled into one model input. The user view is a strict subset; the gap analyser compares the two and the policy decides what the model may act on.

An inventory of shadow channels

Start with an inventory. For your application, list every channel through which text reaches the model and ask, for each, who can write it and who can see it.

ChannelWho writes itWho normally sees itTypical shadow risk
System or developer promptOperatorNobody outside the teamUndisclosed behaviour; stale instructions; secrets placed in it
Tool and function descriptionsTool author, possibly third-partyDeveloper at integration timeInstructions hidden in descriptions or parameter docs; descriptions changed after review
Retrieved web pagesAnyone on the internetUser sees a rendered page or a linkText hidden by CSS, HTML comments, attributes, off-screen elements
Uploaded documentsDocument authorUser sees the rendered documentComments, tracked changes, speaker notes, metadata, tiny or white text
Images to multimodal modelsImage authorUser sees the imageLow-contrast or small text the model reads and the user misses
Memory and conversation summariesThe system itself, from past turnsRarely shownPoisoned memories persisting across sessions; summaries that rephrase injected text as fact
Invisible charactersAny authorNobody, they do not renderEncoded instructions in format characters

The last row has its own article: Unicode smuggling covers invisible code points and how to normalise them. Extraction of the first row by attackers is covered in System prompt leakage. This article is about the general mechanism that connects them all: a difference between two projections of the same context.

Why the gap is the problem

Why does the gap itself matter, beyond the individual attacks that use it? Three reasons.

Authority confusion. Models are trained to follow instructions and are imperfect at telling an instruction from the operator apart from one that appears inside data. Content the user cannot see is content the user cannot object to, so an instruction there competes with the user's request without the user knowing a contest is happening. This is the core of indirect prompt injection.

Broken consent and review. A user who asks 'summarise this page' has consented to the visible page. A reviewer who approved a tool integration approved the description they read at the time. When the model's input diverges from those approved versions, the approval no longer covers what happened.

Unauditable incidents. When something goes wrong, teams look at the chat transcript, which is the user view. If the logged record is not the exact model view, the cause of the behaviour is simply absent from the evidence. Many investigations stall here.

Note the asymmetry with ordinary software. In a web application the server's view of a request and the client's view of the page differ too, but the server does not take instructions from the page. An LLM application does, which is why the gap is a security boundary rather than a rendering detail.

Provenance-tagged prompt assembly

The architecture in the diagram has one central rule: the prompt is assembled from tagged segments, never concatenated from strings. Each segment carries its source, its trust level and whether the user can see it. Everything else follows from having that metadata.

from dataclasses import dataclass, field
import hashlib, json, time

@dataclass
class Segment:
    text: str
    source: str        # "system", "tool:search.description", "retrieval:https://...", "user"
    trust: str         # "operator", "third_party", "user"
    visible: bool      # does the user see this text rendered?
    meta: dict = field(default_factory=dict)

def assemble(segments, log):
    """Build the model input and an audit record of exactly what was sent."""
    parts, record = [], []
    for s in segments:
        body = s.text
        if s.trust == "third_party":
            body = "<untrusted source=\"%s\">\n%s\n</untrusted>" % (s.source, body)
        parts.append(body)
        record.append({"source": s.source, "trust": s.trust, "visible": s.visible,
                       "sha256": hashlib.sha256(s.text.encode()).hexdigest(),
                       "chars": len(s.text), **s.meta})
    prompt = "\n\n".join(parts)
    log.write(json.dumps({"ts": time.time(), "segments": record,
                          "prompt_sha256": hashlib.sha256(prompt.encode()).hexdigest()}) + "\n")
    return prompt

Delimiting untrusted text is not a defence on its own; models can still follow instructions inside the tags. Its value is that the boundary is explicit, consistent, and recorded. With segments tagged you can compute, for every request, how many characters were invisible to the user and which sources supplied them, and you can store the exact prompt, or its hash plus the segment store, so an investigation can reconstruct the model view byte for byte.

Two related rules: pin tool descriptions by hash at review time and refuse to load a tool whose description hash has changed until someone reviews the change, and never put credentials in any prompt segment, because anything in the model view can eventually appear in output.

Measuring the gap: two projections of a page

For retrieved HTML, the most common hostile channel, you can measure the gap directly by extracting text twice: once approximating what a browser shows, once taking everything a naive pipeline would feed the model. The extractor below uses only the standard library and covers the frequent hiding tricks: hidden elements, inline styles that hide text, comments, and text-bearing attributes.

from html.parser import HTMLParser
import re, unicodedata

HIDE_STYLE = re.compile(
    r"display\s*:\s*none|visibility\s*:\s*hidden|font-size\s*:\s*0+(?:\.0+)?(?![\d.])"
    r"|opacity\s*:\s*0+(?:\.0+)?(?![\d.])|left\s*:\s*-\d{3,}px"
    r"|(?<![-\w])color\s*:\s*(#fff(fff)?|white)\b", re.I)
VOID = {"br", "img", "input", "meta", "link", "hr", "source", "wbr", "area", "col"}
ATTRS = ("alt", "title", "aria-label", "content")

class TwoViews(HTMLParser):
    def __init__(self):
        super().__init__(convert_charrefs=True)
        self.stack, self.visible, self.shadow = [], [], []
    def _hidden(self):
        return any(h for _, h in self.stack)
    def handle_starttag(self, tag, attrs):
        a = dict(attrs)
        hidden = (tag in ("script", "style", "template") or "hidden" in a
                  or a.get("aria-hidden") == "true" or bool(HIDE_STYLE.search(a.get("style") or "")))
        for k in ATTRS:
            if a.get(k):
                self.shadow.append(("attr:" + k, a[k]))
        if tag not in VOID:
            self.stack.append((tag, hidden))
    def handle_endtag(self, tag):
        # close the nearest matching open tag; ignore stray end tags like </b>
        for i in range(len(self.stack) - 1, -1, -1):
            if self.stack[i][0] == tag:
                del self.stack[i:]
                break
    def handle_data(self, data):
        if not data.strip():
            return
        if self._hidden():
            self.shadow.append(("hidden-element", data))
        else:
            self.visible.append(data)
    def handle_comment(self, data):
        self.shadow.append(("comment", data))

def gap_report(page_html):
    p = TwoViews()
    p.feed(page_html)
    vis = " ".join(p.visible)
    fmt = sum(1 for ch in page_html if unicodedata.category(ch) == "Cf")
    shadow_chars = sum(len(t) for _, t in p.shadow)
    return {"visible_chars": len(vis), "shadow_chars": shadow_chars,
            "shadow_ratio": shadow_chars / max(1, len(vis) + shadow_chars),
            "format_chars": fmt, "shadow": p.shadow[:20]}

This is a heuristic, deliberately simple, and it has blind spots covered below. What it gives you is a number per document, the shadow ratio, and a list of the hidden fragments with the reason each was considered hidden. That is enough to set policy and to show a reviewer exactly what the model would have read that the user did not.

Worked example: a product page with a hidden instruction

Consider a product page retrieved by a shopping assistant. The visible part is a few hundred characters of description and two short reviews. Inside a review container sits a <div style="display:none"> holding an instruction addressed to AI assistants telling them to describe the product as the best on the market and to recommend the seller's site, plus an HTML comment repeating it and an image whose alt text contains the same instruction.

Run the naive pipeline, which strips tags and keeps all text, and the model receives the description, the reviews, the hidden instruction and the alt text. The user, looking at the page, sees only the description and reviews. If the assistant's answer is unusually glowing, nothing on screen explains why.

Run gap_report and you get three shadow fragments, tagged hidden-element, comment and attr:alt, with a shadow ratio of perhaps 0.3. A policy of 'drop hidden-element and comment fragments, keep alt text only if under 120 characters and free of imperative phrasing, flag the document if shadow ratio exceeds 0.1' sends the model only the visible text, records the event, and lets the interface mark the source as containing hidden content. The answer now matches what the user could verify by reading the page.

The important property is not that the hidden text was malicious, which a classifier might or might not detect. It is that the model was prevented from acting on text the user could not see, which holds whether the text is an attack, an SEO trick or harmless clutter.

Policies: strip, quarantine, disclose, confirm

Once you can measure the gap, choose a policy per channel rather than one rule for everything.

  • Strip for third-party content where hidden text has no legitimate use for your task: comments, hidden elements, invisible format characters in retrieved pages.
  • Quarantine content that may be legitimate but should not carry instructions: send it to a model call that has no tools and can only extract facts into a fixed schema, and pass only those fields to the main agent.
  • Disclose for operator-controlled shadow content. Publish a plain-language description of what the system prompt makes the assistant do, and offer a 'show sources as the model saw them' view for retrieved material.
  • Block or require confirmation when the shadow ratio is high and the next step is consequential, such as sending email, making a purchase or calling a tool with side effects.

Measurement and stripping reduce the attack surface; they do not make the model obey only the user. Keep privilege separation as the backstop: tools with side effects need user confirmation that shows the actual arguments, and untrusted content should never be in the same context as credentials or unrestricted tools. Detection layers that look for injection phrasing are a useful second signal, as discussed in Prompt Injection Scanners, in depth.

Where the measurement is blind

Where gap measurement goes wrong, and what to do about each:

  • External stylesheets and scripts. A class such as sr-only or a rule in a separate CSS file can hide text that the inline-style check never sees, and JavaScript can inject or hide content after load. For high-value pipelines, render with a headless browser and take the visible text from the rendered page instead of guessing from source.
  • Accessibility false positives. Screen-reader-only text and alt attributes are legitimate and often useful. Treat them as a separate category with a length limit rather than deleting them all, or your assistant becomes worse for exactly the users who depend on that text.
  • Visible but unnoticed. Text in a two-pixel font, light grey on white, or buried in a footer is technically visible. A browser-based check that measures contrast and font size catches more of this; human attention still will not.
  • Formats other than HTML. Documents carry comments, revisions, notes and metadata; images carry text that only OCR or the model notices. Each format needs its own two-view extractor, and multimodal input has no reliable cheap equivalent yet.
  • The operator's own shadow prompt drifting. System prompts accumulate rules over months. Review them like code, with history and owners, and test that the disclosed description still matches behaviour.
  • Logging the user view only. If your logs store the chat transcript and not the assembled prompt, you cannot investigate any of the above. Fix this first.

Operating it

Operationally, track a few numbers: the distribution of shadow ratio per source domain, the count of documents stripped or quarantined per day, the rate of high-ratio documents preceding consequential tool calls, and the fraction of tool descriptions whose hash matches the reviewed version, which should be one hundred percent. Alert on sudden changes in any of them rather than on absolute thresholds; a domain that never had hidden text and now does is a stronger signal than a domain that always uses collapsed sections. Keep a red-team corpus of pages with each hiding technique and run it through the pipeline on every change, so you know which techniques your extractor currently misses. Prompt Injection via RAG Retrieval describes the retrieval side of the same pipeline.

What to do next

  1. Write the channel inventory table for your application, with who writes and who sees each channel.
  2. Change prompt construction to tagged segments, and log the exact assembled model input or a reconstructable record of it for every request.
  3. Hash and pin every tool description, and block tools whose description changed since review.
  4. Add a two-view extractor for each retrieved format, starting with HTML, and record the shadow ratio per document.
  5. Pick a policy per channel: strip, quarantine, disclose, or confirm, and require user confirmation with real arguments before side-effecting tool calls.
  6. Build a red-team corpus covering CSS hiding, comments, attributes, invisible characters, document comments and low-contrast image text, and run it in CI.
  7. Publish a plain-language description of what your system prompt instructs the assistant to do.
Key takeaway: Every LLM application has a user view and a model view, and shadow prompts are whatever sits in the second but not the first. Inventory the channels, assemble prompts from tagged segments, log the exact model input, pin tool descriptions, measure hidden content per document, and choose per channel whether to strip, quarantine, disclose or require confirmation, with privilege separation as the backstop.