A retrieval-augmented system answers questions by pasting retrieved text into a model's prompt. When one of those answers turns out to be wrong, leaked or manipulated, someone has to answer a forensic question quickly: which bytes, from which version of which source, processed by which pipeline, were in the prompt, and why were they chosen over everything else? A list of cited URLs does not answer it. URLs point at documents that have since changed, and citations name what the model chose to mention, not everything it read.
RAG provenance tracking is the engineering that makes that question answerable. This article designs the lineage record end to end: content-addressed chunk identifiers, ingestion runs, retrieval decision logs and generation events, chained so that tampering is evident. It then uses the record for the three jobs that justify its cost: replaying an answer, computing the blast radius of a poisoned document, and propagating deletions. Showing citations to users and verifying them is a separate topic, covered in the citation verification pipeline.
Why provenance is a security control
Provenance is a security control because the retrieval corpus is an input channel. Anyone who can write to a wiki, a ticket system or a shared drive can put text in front of the model, which is the basis of RAG poisoning and indirect prompt injection. Prevention controls such as trust labels and permission-aware retrieval (see RAG defense in depth) reduce the odds. Provenance is the detective and corrective layer: when prevention fails, it tells you what was affected and lets you prove it.
It also serves non-adversarial needs: debugging a bad answer ("was it retrieval or generation?"), honouring erasure requests, and demonstrating to an auditor which data a regulated decision relied on. The same record does all of these, which is the argument for building it once and properly.
Be clear about what provenance does not do. It does not decide whether a document is trustworthy, detect an injection, or check that an answer is supported by its sources. It records facts so that the controls which do those things can be audited and their misses traced afterwards. Output-side provenance, such as citations, watermarks and content credentials, is a related but different problem, surveyed in LLM output provenance architecture. The two connect at one point: every citation shown to a user should resolve to a chunk id in this log.
The lineage model
Borrow the vocabulary of the W3C PROV data model: entities are things (a document version, a chunk, a prompt, an answer), activities transform them (ingestion, retrieval, generation), and agents are responsible (a pipeline service, a user). The relations you need are "used", "was generated by" and "was derived from". You do not have to emit PROV documents; adopting the model keeps the schema honest.
| Record | Key | Must contain |
|---|---|---|
| source_version | sha256 of raw bytes | source URI, fetch time, owner, ACL snapshot, trust label |
| ingest_run | run id | parser, chunker and embedding model versions, config hash |
| chunk | hash(source_version, span, chunker) | byte span, text hash, embedding model, index id |
| retrieval_event | request id | query hash, index snapshot, candidates with scores, filter drops, rerank scores, selected ids |
| generation_event | request id | prompt hash, model id, parameters, output hash, cited chunk ids |
The key design decision is content addressing. A chunk's identifier is derived from the hash of the source version it came from, its byte span, and the chunker configuration, so the same text from a different document version gets a different identifier, and nothing can be silently overwritten in place. Mutable identifiers like doc-42#chunk-3 make every downstream record ambiguous the moment doc-42 is edited.
Recording ingestion and retrieval
The ingestion and retrieval sides need only a few lines each. The important property is that the retrieval log records decisions, including candidates that were filtered out and why, because "why did the model not see the right policy?" is as common a question as "why did it see the wrong one?".
import hashlib, json, time
def h(*parts: bytes) -> str:
d = hashlib.sha256()
for part in parts:
d.update(len(part).to_bytes(8, "big")); d.update(part) # length-prefixed
return d.hexdigest()
def ingest(uri, raw: bytes, run, chunker, log):
sv = h(raw)
log.append("source_version", sv, uri=uri, fetched=time.time(),
trust=run.trust_for(uri), acl=run.acl_for(uri))
for start, end in chunker.spans(raw):
cid = h(sv.encode(), f"{start}:{end}".encode(), chunker.version.encode())
log.append("chunk", cid, source_version=sv, span=[start, end],
text_sha=h(raw[start:end]), ingest_run=run.id,
embed_model=run.embed_model)
yield cid, raw[start:end]
def retrieve(req_id, query, index, filters, reranker, k, log):
cands = index.search(query, top=50) # [(cid, score)]
kept, dropped = [], []
for cid, s in cands:
reason = filters.reject_reason(cid, req_id) # ACL, trust, tombstone
(dropped if reason else kept).append((cid, s, reason))
ranked = reranker.score(query, [c for c, _, _ in kept])
chosen = [cid for cid, _ in ranked[:k]]
log.append("retrieval_event", req_id, query_sha=h(query.encode()),
index_snapshot=index.snapshot_id,
candidates=[[c, round(s, 4)] for c, s, _ in kept],
dropped=[[c, r] for c, _, r in dropped],
rerank=[[c, round(s, 4)] for c, s in ranked], selected=chosen)
return chosenNotice what is hashed rather than stored. The query text is often personal data; the log keeps its hash, and the raw query lives (if at all) in a separately governed store with shorter retention. The same applies to the assembled prompt. Hashes are enough to prove that a later replay used identical inputs.
Making the log tamper-evident
A lineage log that an attacker or a careless operator can edit is not evidence. Make it tamper-evident by chaining: each entry stores the hash of the previous entry, and a periodic checkpoint of the head hash is written somewhere the pipeline cannot modify (a write-once bucket, a separate account, or a transparency service).
class ChainedLog:
def __init__(self, sink, head="0" * 64):
self.sink, self.head = sink, head
def append(self, kind, key, **fields):
body = json.dumps({"kind": kind, "key": key, **fields},
sort_keys=True, separators=(",", ":")).encode()
entry_hash = h(self.head.encode(), body)
self.sink.write({"prev": self.head, "hash": entry_hash, "body": body.decode()})
self.head = entry_hash
def verify(entries, checkpoint_head):
head = "0" * 64
for e in entries:
if e["prev"] != head or h(head.encode(), e["body"].encode()) != e["hash"]:
return False, e
head = e["hash"]
return head == checkpoint_head, NoneChaining detects edits and deletions between checkpoints; it does not prevent them, and it cannot vouch for an entry that was false when written. Shard the chain per service or per hour so that writes do not serialise across the whole system, and keep the checkpoint interval short enough that the window of undetectable tampering is acceptable to you.
Replay, blast radius and erasure
Replay. Given a request id, load the retrieval event, open the index snapshot it names, and re-run retrieval with the recorded query. If the selected chunk ids match, retrieval is reproduced; then rebuild the prompt from the chunk texts, check its hash against the generation event, and regenerate with the recorded model and parameters. Generation may still differ at non-zero temperature or after a model update, but you have isolated the variable. Without index snapshots, replay degrades to "retrieval today", which answers a different question.
Blast radius. When a document is found to be poisoned, the question is which answers it reached. With content-addressed records this is a join, not an investigation:
-- every answer whose prompt held text from a version of the source fetched after the bad edit
SELECT g.request_id, g.model_id, g.created_at, r.user_ref
FROM source_version sv
JOIN chunk c ON c.source_version = sv.hash
JOIN retrieval_selected rs ON rs.chunk_id = c.id
JOIN generation_event g ON g.request_id = rs.request_id
JOIN request r ON r.id = g.request_id
WHERE sv.uri = :bad_uri AND sv.fetched >= :first_bad_fetch
ORDER BY g.created_at;Joining on selected chunks rather than cited chunks matters: injected instructions work by being read, and a model that obeyed them rarely cites them.
Erasure. Deleting a source must reach every derived copy: chunks, embeddings in every index, rerank and semantic caches, cached answers, and fine-tuning datasets built from logs. Walk the lineage graph forward from the source version, write a tombstone that the retrieval filter honours immediately, then purge each store and record a completion entry per store. The tombstone closes the window while physical deletion catches up.
Worked example: a poisoned policy page
Worked example (illustrative). An internal assistant answers HR questions from a wiki. A contractor edits the travel-policy page to add a hidden line telling the model to send expense questions to an external form. Ten days later a user reports the odd link.
- The responder takes the request id from the report and loads its generation and retrieval events. The selected chunks include one whose source_version hash belongs to the travel-policy page, fetched after the contractor's edit.
- Diffing that version against the previous one localises the injected span; its byte range matches the chunk's recorded span, so the chunk is confirmed as the carrier.
- The blast-radius query returns every request whose prompt included any chunk from the bad version: in this scenario a few hundred answers across forty users, listed with timestamps. Users who received the link are contacted.
- A tombstone on the bad source version stops new retrievals within minutes; the purge job removes its chunks from the vector index and the semantic answer cache and logs each step.
- The chain is verified against the hourly checkpoint, showing the log was not altered during the incident window, and the report cites hashes rather than screenshots.
Without content-addressed provenance, step 3 is the hard part: the page has since been reverted, the URL now shows clean text, and citations show only the answers in which the model chose to mention the page.
Failure modes
| Failure | Consequence | Fix |
|---|---|---|
| Mutable chunk ids | Records point at text that has changed | Content-address ids from source hash and span |
| Logging only citations | Blast radius misses injections that were read but not cited | Log selected chunks and filter drops |
| No index snapshots | Replay reproduces today, not the incident | Name a snapshot per retrieval; retain long enough |
| Raw prompts in logs | Lineage store becomes a PII and secrets leak | Hash queries and prompts; separate governed store |
| Caches outside lineage | Deleted or poisoned text keeps being served | Key caches by chunk ids; purge via the graph |
| Log writable by the pipeline | Evidence can be silently edited | Hash chain plus external checkpoints |
| Unversioned pipeline config | Two runs with different chunkers look identical | Hash the config into every ingest_run record |
The last row is the quiet one. If a chunker or embedding model changes without a new ingest_run record, two indexes built a month apart carry indistinguishable lineage, and any comparison between their answers is guesswork.
Operating it: cost, latency, retention
Cost. A retrieval event with 50 candidates, ids and scores, is a few kilobytes of JSON; at a million requests a day that is a few gigabytes daily before compression, small next to the vector index. The larger cost is retaining index snapshots: use append-only index segments so a snapshot is a list of segment ids, not a copy.
Latency. Write lineage asynchronously through a durable queue, but make the request fail closed if the queue is unavailable for regulated workloads; a gap in the log is exactly what an auditor will ask about.
Retention. Lineage records about public documents can live for years; records that reference user queries inherit the retention policy of user data. Decide both explicitly. For data contracts across the wider pipeline, see data lineage contracts for AI knowledge systems.
What to do next
- Switch chunk identifiers to content-addressed hashes of source version, span and chunker version, and re-ingest.
- Log retrieval decisions: candidates, scores, filter drops with reasons, rerank scores and the selected set, plus the index snapshot id.
- Record generation events with prompt hash, model id, parameters and output hash.
- Write everything to a hash-chained log and checkpoint the head to storage the pipeline cannot modify.
- Write and test the blast-radius query against a planted test document before you need it.
- Implement tombstones in the retrieval filter and a purge job that walks the lineage graph into every index and cache.
- Run a replay drill on a random request each week and alert when it does not reproduce retrieval.