A citation in an LLM answer looks like evidence, and readers treat it as evidence. That is exactly why it is a security problem. A marker such as [S3] or a footnoted URL tells the user "this sentence came from a source you can trust". If the marker points at a document that was never retrieved, at a file the user is not allowed to see, at a poisoned web page, or at an attacker's domain, the citation does real harm. It launders a false claim, leaks data, or sends the reader somewhere hostile.
Two other articles on this site cover the neighbouring problems. Output grounding verification asks whether each claim is supported by its evidence. Grounding and citations covers the prompt-side citation contract and verbatim quote checks. This article covers the reference itself. It explains how to verify that every citation an answer emits is real, permitted, intact and safe to render, and how to build that as a pipeline stage that fails closed.
Six ways a citation goes wrong
Start by listing what can go wrong with a reference, separately from what can go wrong with a claim. There are six distinct failures, and each needs a different check.
- Fabricated reference. The model invents a document ID, URL, DOI, case name or paper title. This is the classic hallucinated citation. Models produce plausible-looking identifiers because the format is easy to imitate and the content behind it is not.
- Misattributed reference. The ID is real and was retrieved, but it is attached to a sentence it does not support. This is the grounding problem, and the support stage handles it.
- Unauthorised reference. The cited source exists, but the user should not know it exists. One example is a board memo title surfaced to a contractor. The answer text can be harmless while the citation title, snippet or URL leaks the secret.
- Tampered or stale reference. The source changed between retrieval and display, or was retracted, or the cited version is not the version the model saw.
- Laundered reference. An attacker plants content in an indexed source. The model retrieves it, repeats it and cites it, and the citation lends it the authority of your product. See prompt injection via RAG for how poisoned chunks win retrieval.
- Weaponised reference. The citation becomes the delivery vehicle. Examples are a markdown link whose query string carries user data, an auto-loading image URL, a lookalike domain, or a URL that makes your own verifier fetch an internal address (SSRF).
The pipeline: the model proposes, the system grants
The design rule is simple. The model proposes citations; the system grants them. The model output is untrusted text. A citation reaches the user only after deterministic code has matched it to an object the system retrieved, checked that the caller may see it, and rendered it from trusted metadata rather than from model-written strings.
The stages, in order of cost:
- Parse. Extract every reference-shaped token from the answer: bracket markers, markdown links, bare URLs, DOIs and anything that looks like a citation in prose ("according to the 2023 Smith report"). Anything you do not parse, you cannot verify, so the renderer must neutralise unparsed links too.
- Membership. Each marker must map to a source in the evidence ledger. This is the immutable list of chunks the retriever actually returned for this turn. A marker outside the ledger is fabricated by definition.
- Authorisation. Re-check access for each cited source against the caller's identity at answer time. Do not rely only on the check the retriever made, because caches and shared sessions break that assumption.
- Integrity. Compare the content hash and version recorded in the ledger with the current source. Also apply source-trust rules: domain allowlists, retraction lists and minimum trust tiers.
- Span and support. If the answer quotes, the quote must appear verbatim in the cited chunk. If it paraphrases, an entailment check scores support. This is the expensive stage, so it runs only on citations that survived the cheap ones.
- Policy and render. Decide per citation whether to keep it, strip it, or block the answer. Then render links from ledger metadata, never from the model's own URL text.
The evidence ledger
Everything depends on the evidence ledger, so build it before generation. When the retriever returns chunks, assign each a short opaque handle (S1, S2 and so on). Store the handle with the source ID, version, content hash, ACL snapshot, trust tier and canonical URL. Give the model only the handles and the text. The model never sees a raw internal URL, so it cannot copy one, mutate one or invent a near miss.
Opaque handles also defeat a common evasion. If you show the model real document IDs, an injected instruction can tell it to cite doc-48213, a guessable ID the user was never given. With per-turn handles, an unknown handle is simply unknown.
from dataclasses import dataclass
import hashlib, secrets
@dataclass(frozen=True)
class Evidence:
handle: str # "S1", shown to the model
source_id: str # internal ID, never shown to the model
version: str
sha256: str
acl_groups: frozenset
trust_tier: int # 0 = untrusted web, 3 = curated internal
canonical_url: str # rendered by the system, not the model
text: str
def build_ledger(chunks):
ledger = {}
for i, ch in enumerate(chunks, start=1):
ledger[f"S{i}"] = Evidence(
handle=f"S{i}", source_id=ch.id, version=ch.version,
sha256=hashlib.sha256(ch.text.encode()).hexdigest(),
acl_groups=frozenset(ch.acl), trust_tier=ch.trust,
canonical_url=ch.url, text=ch.text)
return ledger
The verifier
The verifier itself is ordinary code. Notice that every branch that cannot confirm a citation returns a failing verdict. There is no "probably fine" path.
import re
MARKER = re.compile(r"\[(S\d{1,3})\]")
MD_LINK = re.compile(r"!?\[[^\]]*\]\(([^)\s]+)[^)]*\)")
BARE_URL = re.compile(r"https?://\S+")
def verify(answer, ledger, user, store, entail):
verdicts = []
for handle in set(MARKER.findall(answer)):
ev = ledger.get(handle)
if ev is None:
verdicts.append((handle, "fabricated")); continue
if not (ev.acl_groups & user.groups):
verdicts.append((handle, "unauthorised")); continue
cur = store.current(ev.source_id)
if cur is None or cur.retracted:
verdicts.append((handle, "retracted")); continue
if cur.sha256 != ev.sha256:
verdicts.append((handle, "changed_since_retrieval")); continue
if ev.trust_tier < 1:
verdicts.append((handle, "untrusted_source")); continue
for sent in sentences_citing(answer, handle):
if entail(premise=ev.text, hypothesis=sent) < 0.7:
verdicts.append((handle, "unsupported")); break
else:
verdicts.append((handle, "ok"))
# Model-written links are never rendered as links.
for url in MD_LINK.findall(answer) + BARE_URL.findall(answer):
verdicts.append((url, "model_url_stripped"))
return verdictsThree details matter. First, the membership check is a dictionary lookup, so it costs nothing and removes the whole class of invented references. Second, the ACL check uses the caller's current groups. If someone lost access between indexing and the question, the citation goes. Third, the entailment threshold (0.7 here) is illustrative. Calibrate it on your own labelled data, as the grounding verification article describes.
External references: DOIs, URLs and SSRF
Some products must cite the open web or scholarly literature, where the cited object is not in your ledger. Here the verifier has to resolve the reference, and resolving is where most teams introduce a new vulnerability.
DOIs and identifiers. Resolve them through the registry's resolver and compare the returned title and authors with what the answer claims. A DOI that resolves to a different paper is as bad as one that does not resolve. Cache results and rate-limit them, because a model asked about a busy topic can emit hundreds of identifiers.
URLs and SSRF. Server-side request forgery happens when attacker-influenced input makes your server issue a request to a destination the attacker chooses. A citation URL is exactly that kind of input. If a planted page convinces the model to cite http://169.254.169.254/... or an internal hostname, a naive verifier fetches cloud metadata or an internal admin page. The verifier then shows the response, or its title, back to the user. Defences:
- Run the resolver in an isolated egress zone with no route to internal networks or metadata endpoints.
- Allow only http and https. Resolve DNS once, reject private, loopback and link-local addresses, then connect to that resolved IP so a rebinding DNS answer cannot switch targets.
- Follow redirects manually, at most a few hops, and re-apply every check at every hop.
- Fetch with a byte cap and a short timeout, and never forward cookies or credentials.
- Treat fetched content as untrusted input. It must not flow back into the model as instructions.
Lookalike domains. Normalise the host by lowercasing it, converting internationalised names to their ASCII form and stripping trailing dots. Then check it against an allowlist, or at least against a list of your own brand's domains, so that a near-identical spelling cannot pose as an official source.
Rendering citations safely
The final stage is where citations most often turn into exfiltration channels. A markdown renderer that turns  into an image tag makes the user's browser send SECRET to the attacker. No click is needed. An injected instruction only has to persuade the model to put conversation data into that URL. Data exfiltration via LLMs walks through the attack, and insecure output handling covers the general rule that model output is untrusted HTML.
Defensive rendering for citations:
- Render each verified citation as a link built from the ledger's
canonical_urland title. Never use the model's text for either. - Do not render model-authored images at all, or proxy them through a fetcher that strips query strings and checks an allowlist.
- Show the destination host next to every external link, so a lookalike domain is visible.
- Escape citation titles and snippets as text. A document title is attacker-controlled if anyone can upload documents.
- Return snippets only after the ACL check, and trim them to the cited span rather than the whole chunk.
Worked example: a policy assistant
An internal assistant answers "What is our policy on contractor laptop encryption?" The retriever returns four chunks. The ledger maps S1 to the security standard v7 (tier 3), S2 to an IT FAQ (tier 2), S3 to a wiki page any employee can edit (tier 1), and S4 to a draft from a restricted policy folder. S4 entered the ledger only because a shared cache served a result built for an administrator.
The model answers: "Contractor laptops must use full-disk encryption [S1]. Exceptions require CISO approval [S4]. Use the self-service tool at the link in the FAQ [S2], or see [the policy portal](https://securlty-portal.example/login) [S5]."
The verifier produces:
| Reference | Verdict | Why |
|---|---|---|
| S1 | ok | In ledger, authorised, hash matches, sentence entailed |
| S4 | unauthorised | User is not in the restricted policy group; the cache leak is caught here |
| S2 | unsupported | The FAQ chunk does not mention a self-service tool; entailment 0.21 |
| S5 | fabricated | No S5 in the ledger |
| securlty-portal.example | model_url_stripped | Model-written URL; also a lookalike host |
Policy then decides the outcome. One reasonable rule: drop sentences whose only citation failed, keep the rest, and block the whole answer if any verdict is unauthorised. A leaked existence is already a disclosure, and partial redaction can still reveal it through context. Here the answer is blocked and regenerated without S4. The incident log records the cache key, which is the real bug. In this design the verifier is also a detector for retrieval-layer access bugs.
Failure modes
Failure modes that recur in production:
- Trusting the retriever's ACL check alone. Shared caches, background summarisation jobs and multi-user sessions all break it. Re-check at render time.
- Citing chunks, rendering documents. The chunk was permitted but the document link opens a larger file with restricted sections. Cite and link at the same granularity you authorised.
- Verifying markers but rendering raw markdown. The marker checks pass while a model-written image tag sends the data out anyway. Verification without safe rendering is incomplete.
- Soft-failing on timeouts. The entailment model is slow, so someone adds "if the check times out, keep the citation". Attackers can cause the timeouts. Fail closed, and fall back to stripping the citation rather than keeping it.
- Unbounded resolution. An external resolver without caps becomes a request amplifier, and one prompt fans out to hundreds of outbound fetches.
- Treating citation count as quality. Rewarding more citations teaches the model to cite generously. Track the strip rate instead; a rising one signals retrieval drift or an injection campaign.
Operations and trade-offs
Log one record per citation: handle, source ID, version, verdict, reason, latency and policy action. These records let you answer "which answers cited the poisoned page?" after an incident, and they give you the metrics that matter: fabricated-reference rate, unauthorised-citation rate, unsupported rate and strip rate per answer. Alert on any non-zero unauthorised rate, because each one is a data-exposure event.
Latency budget: membership, ACL, hash and trust checks are in-memory lookups plus one batched metadata read. The cost is in entailment and external resolution. Run entailment only on citations that pass the cheap checks, batch it across sentences, and cache verdicts by (source hash, sentence hash). For streaming answers, either hold citation rendering until the sentence completes or render markers as plain text first and upgrade them to links once verified.
Trade-offs. Strict fail-closed policy removes some correct citations, especially paraphrases that score just under the entailment threshold, and users notice missing sources. Blocking whole answers on one unauthorised citation is safe but costs a regeneration. Restricting external citations to an allowlist removes most SSRF and lookalike risk but narrows what the assistant can cite. Choose per product, but never relax the ACL check on sensitive corpora.
What to do next
- Build an evidence ledger per turn and give the model opaque handles, never internal IDs or URLs.
- Reject any citation handle not in the ledger, before running any model-based check.
- Re-check access for each cited source with the caller's current identity at answer time.
- Store content hashes and versions in the ledger, and compare them before rendering.
- Render links and titles from ledger metadata, strip model-written URLs, and never auto-load model-authored images.
- If you must resolve external URLs, do it from an isolated egress zone with IP checks, per-hop redirect checks and byte caps.
- Add entailment for surviving citations and calibrate its threshold on labelled examples.
- Log a verdict per citation and alert on any unauthorised citation.