A retrieval-augmented system answers questions by pasting retrieved text into a model's prompt. When one of those answers turns out to be wrong, leaked or manipulated, someone has to answer a forensic question quickly: which bytes, from which version of which source, processed by which pipeline, were in the prompt, and why were they chosen over everything else? A list of cited URLs does not answer it. URLs point at documents that have since changed, and citations name what the model chose to mention, not everything it read.

RAG provenance tracking is the engineering that makes that question answerable. This article designs the lineage record end to end: content-addressed chunk identifiers, ingestion runs, retrieval decision logs and generation events, chained so that tampering is evident. It then uses the record for the three jobs that justify its cost: replaying an answer, computing the blast radius of a poisoned document, and propagating deletions. Showing citations to users and verifying them is a separate topic, covered in the citation verification pipeline.

Why provenance is a security control

Provenance is a security control because the retrieval corpus is an input channel. Anyone who can write to a wiki, a ticket system or a shared drive can put text in front of the model, which is the basis of RAG poisoning and indirect prompt injection. Prevention controls such as trust labels and permission-aware retrieval (see RAG defense in depth) reduce the odds. Provenance is the detective and corrective layer: when prevention fails, it tells you what was affected and lets you prove it.

It also serves non-adversarial needs: debugging a bad answer ("was it retrieval or generation?"), honouring erasure requests, and demonstrating to an auditor which data a regulated decision relied on. The same record does all of these, which is the argument for building it once and properly.

Be clear about what provenance does not do. It does not decide whether a document is trustworthy, detect an injection, or check that an answer is supported by its sources. It records facts so that the controls which do those things can be audited and their misses traced afterwards. Output-side provenance, such as citations, watermarks and content credentials, is a related but different problem, surveyed in LLM output provenance architecture. The two connect at one point: every citation shown to a user should resolve to a chunk id in this log.

The lineage model

Borrow the vocabulary of the W3C PROV data model: entities are things (a document version, a chunk, a prompt, an answer), activities transform them (ingestion, retrieval, generation), and agents are responsible (a pipeline service, a user). The relations you need are "used", "was generated by" and "was derived from". You do not have to emit PROV documents; adopting the model keeps the schema honest.

Provenance follows the bytes: every stage writes a record keyed by hashesSourcedoc + versionIngest runparser, chunkerChunk storechunk_id = hashRetrievalscores, filtersGenerationmodel, promptAppend-only lineage log (hash-chained)source_version -> ingest_run -> chunk -> retrieval_event -> generation_eventReplayre-run one answer exactlyBlast radiusanswers that used doc XErasurepropagate deletes everywhere
Every stage writes a record keyed by content hashes into one append-only log; replay, blast-radius and erasure all read it.
RecordKeyMust contain
source_versionsha256 of raw bytessource URI, fetch time, owner, ACL snapshot, trust label
ingest_runrun idparser, chunker and embedding model versions, config hash
chunkhash(source_version, span, chunker)byte span, text hash, embedding model, index id
retrieval_eventrequest idquery hash, index snapshot, candidates with scores, filter drops, rerank scores, selected ids
generation_eventrequest idprompt hash, model id, parameters, output hash, cited chunk ids

The key design decision is content addressing. A chunk's identifier is derived from the hash of the source version it came from, its byte span, and the chunker configuration, so the same text from a different document version gets a different identifier, and nothing can be silently overwritten in place. Mutable identifiers like doc-42#chunk-3 make every downstream record ambiguous the moment doc-42 is edited.

Recording ingestion and retrieval

The ingestion and retrieval sides need only a few lines each. The important property is that the retrieval log records decisions, including candidates that were filtered out and why, because "why did the model not see the right policy?" is as common a question as "why did it see the wrong one?".

import hashlib, json, time

def h(*parts: bytes) -> str:
    d = hashlib.sha256()
    for part in parts:
        d.update(len(part).to_bytes(8, "big")); d.update(part)   # length-prefixed
    return d.hexdigest()

def ingest(uri, raw: bytes, run, chunker, log):
    sv = h(raw)
    log.append("source_version", sv, uri=uri, fetched=time.time(),
               trust=run.trust_for(uri), acl=run.acl_for(uri))
    for start, end in chunker.spans(raw):
        cid = h(sv.encode(), f"{start}:{end}".encode(), chunker.version.encode())
        log.append("chunk", cid, source_version=sv, span=[start, end],
                   text_sha=h(raw[start:end]), ingest_run=run.id,
                   embed_model=run.embed_model)
        yield cid, raw[start:end]

def retrieve(req_id, query, index, filters, reranker, k, log):
    cands = index.search(query, top=50)                # [(cid, score)]
    kept, dropped = [], []
    for cid, s in cands:
        reason = filters.reject_reason(cid, req_id)    # ACL, trust, tombstone
        (dropped if reason else kept).append((cid, s, reason))
    ranked = reranker.score(query, [c for c, _, _ in kept])
    chosen = [cid for cid, _ in ranked[:k]]
    log.append("retrieval_event", req_id, query_sha=h(query.encode()),
               index_snapshot=index.snapshot_id,
               candidates=[[c, round(s, 4)] for c, s, _ in kept],
               dropped=[[c, r] for c, _, r in dropped],
               rerank=[[c, round(s, 4)] for c, s in ranked], selected=chosen)
    return chosen

Notice what is hashed rather than stored. The query text is often personal data; the log keeps its hash, and the raw query lives (if at all) in a separately governed store with shorter retention. The same applies to the assembled prompt. Hashes are enough to prove that a later replay used identical inputs.

Making the log tamper-evident

A lineage log that an attacker or a careless operator can edit is not evidence. Make it tamper-evident by chaining: each entry stores the hash of the previous entry, and a periodic checkpoint of the head hash is written somewhere the pipeline cannot modify (a write-once bucket, a separate account, or a transparency service).

class ChainedLog:
    def __init__(self, sink, head="0" * 64):
        self.sink, self.head = sink, head

    def append(self, kind, key, **fields):
        body = json.dumps({"kind": kind, "key": key, **fields},
                          sort_keys=True, separators=(",", ":")).encode()
        entry_hash = h(self.head.encode(), body)
        self.sink.write({"prev": self.head, "hash": entry_hash, "body": body.decode()})
        self.head = entry_hash

def verify(entries, checkpoint_head):
    head = "0" * 64
    for e in entries:
        if e["prev"] != head or h(head.encode(), e["body"].encode()) != e["hash"]:
            return False, e
        head = e["hash"]
    return head == checkpoint_head, None

Chaining detects edits and deletions between checkpoints; it does not prevent them, and it cannot vouch for an entry that was false when written. Shard the chain per service or per hour so that writes do not serialise across the whole system, and keep the checkpoint interval short enough that the window of undetectable tampering is acceptable to you.

Replay, blast radius and erasure

Replay. Given a request id, load the retrieval event, open the index snapshot it names, and re-run retrieval with the recorded query. If the selected chunk ids match, retrieval is reproduced; then rebuild the prompt from the chunk texts, check its hash against the generation event, and regenerate with the recorded model and parameters. Generation may still differ at non-zero temperature or after a model update, but you have isolated the variable. Without index snapshots, replay degrades to "retrieval today", which answers a different question.

Blast radius. When a document is found to be poisoned, the question is which answers it reached. With content-addressed records this is a join, not an investigation:

-- every answer whose prompt held text from a version of the source fetched after the bad edit
SELECT g.request_id, g.model_id, g.created_at, r.user_ref
FROM   source_version sv
JOIN   chunk c           ON c.source_version = sv.hash
JOIN   retrieval_selected rs ON rs.chunk_id  = c.id
JOIN   generation_event g    ON g.request_id = rs.request_id
JOIN   request r             ON r.id         = g.request_id
WHERE  sv.uri = :bad_uri AND sv.fetched >= :first_bad_fetch
ORDER  BY g.created_at;

Joining on selected chunks rather than cited chunks matters: injected instructions work by being read, and a model that obeyed them rarely cites them.

Erasure. Deleting a source must reach every derived copy: chunks, embeddings in every index, rerank and semantic caches, cached answers, and fine-tuning datasets built from logs. Walk the lineage graph forward from the source version, write a tombstone that the retrieval filter honours immediately, then purge each store and record a completion entry per store. The tombstone closes the window while physical deletion catches up.

Worked example: a poisoned policy page

Worked example (illustrative). An internal assistant answers HR questions from a wiki. A contractor edits the travel-policy page to add a hidden line telling the model to send expense questions to an external form. Ten days later a user reports the odd link.

  1. The responder takes the request id from the report and loads its generation and retrieval events. The selected chunks include one whose source_version hash belongs to the travel-policy page, fetched after the contractor's edit.
  2. Diffing that version against the previous one localises the injected span; its byte range matches the chunk's recorded span, so the chunk is confirmed as the carrier.
  3. The blast-radius query returns every request whose prompt included any chunk from the bad version: in this scenario a few hundred answers across forty users, listed with timestamps. Users who received the link are contacted.
  4. A tombstone on the bad source version stops new retrievals within minutes; the purge job removes its chunks from the vector index and the semantic answer cache and logs each step.
  5. The chain is verified against the hourly checkpoint, showing the log was not altered during the incident window, and the report cites hashes rather than screenshots.

Without content-addressed provenance, step 3 is the hard part: the page has since been reverted, the URL now shows clean text, and citations show only the answers in which the model chose to mention the page.

Failure modes

FailureConsequenceFix
Mutable chunk idsRecords point at text that has changedContent-address ids from source hash and span
Logging only citationsBlast radius misses injections that were read but not citedLog selected chunks and filter drops
No index snapshotsReplay reproduces today, not the incidentName a snapshot per retrieval; retain long enough
Raw prompts in logsLineage store becomes a PII and secrets leakHash queries and prompts; separate governed store
Caches outside lineageDeleted or poisoned text keeps being servedKey caches by chunk ids; purge via the graph
Log writable by the pipelineEvidence can be silently editedHash chain plus external checkpoints
Unversioned pipeline configTwo runs with different chunkers look identicalHash the config into every ingest_run record

The last row is the quiet one. If a chunker or embedding model changes without a new ingest_run record, two indexes built a month apart carry indistinguishable lineage, and any comparison between their answers is guesswork.

Operating it: cost, latency, retention

Cost. A retrieval event with 50 candidates, ids and scores, is a few kilobytes of JSON; at a million requests a day that is a few gigabytes daily before compression, small next to the vector index. The larger cost is retaining index snapshots: use append-only index segments so a snapshot is a list of segment ids, not a copy.

Latency. Write lineage asynchronously through a durable queue, but make the request fail closed if the queue is unavailable for regulated workloads; a gap in the log is exactly what an auditor will ask about.

Retention. Lineage records about public documents can live for years; records that reference user queries inherit the retention policy of user data. Decide both explicitly. For data contracts across the wider pipeline, see data lineage contracts for AI knowledge systems.

What to do next

  1. Switch chunk identifiers to content-addressed hashes of source version, span and chunker version, and re-ingest.
  2. Log retrieval decisions: candidates, scores, filter drops with reasons, rerank scores and the selected set, plus the index snapshot id.
  3. Record generation events with prompt hash, model id, parameters and output hash.
  4. Write everything to a hash-chained log and checkpoint the head to storage the pipeline cannot modify.
  5. Write and test the blast-radius query against a planted test document before you need it.
  6. Implement tombstones in the retrieval filter and a purge job that walks the lineage graph into every index and cache.
  7. Run a replay drill on a random request each week and alert when it does not reproduce retrieval.
Key takeaway: RAG provenance means recording, for every answer, the exact content-addressed chunks the model read, why they were selected, and which source versions and pipeline runs produced them, in a log you can prove was not edited. That record turns poisoning response, replay and erasure into queries instead of investigations.