A retrieval index is a cache of the source systems it was built from, and like every cache it goes stale. A policy page is edited, a price changes, a ticket is resolved, a document is deleted for legal reasons, and until the index catches up the assistant answers from the old version with full confidence and a citation. Stale answers are harder to notice than missing ones, because they look correct and cite a real document.
Rebuilding the entire index on a schedule is the simplest fix and does not scale: re-embedding millions of chunks nightly is slow and expensive, and a nightly rebuild still means answers up to a day out of date. The alternative is an incremental pipeline that detects changes, re-embeds only what changed, applies updates and deletes in order, and proves its freshness with measurements. This article covers each stage, the handling of embedding model upgrades, and how to monitor freshness as a service level objective.
Freshness as a service level objective
Define freshness before building for it. The useful measure is the lag from a change becoming visible in the source system to that change being reflected in retrieval results, measured per change and reported at a percentile, for example 95% of changes queryable within fifteen minutes. Deletes deserve their own, stricter objective, because serving deleted content can be a legal or security problem rather than just a quality one. Different sources can have different targets: a pricing catalogue may need minutes, while an archive of past project documents can tolerate a day.
Writing the objective down changes the design conversation. A nightly batch cannot meet a fifteen-minute target no matter how it is tuned, whereas an event-driven pipeline with a nightly reconciliation sweep can meet it and still catch the events it misses.
Capturing changes
Each source offers different change signals, and most pipelines combine several. Databases support change data capture from the transaction log, for example with Debezium, which delivers inserts, updates and deletes in commit order. SaaS systems offer webhooks, which are fast but can be dropped, duplicated or delivered out of order. Everything else is polled, using modification timestamps, entity tags or change cursors where the API provides them. Polling by timestamp needs an overlap window, because clock skew and in-flight transactions can place a change just before the last cursor.
No change feed is perfect, so pair it with reconciliation: a periodic sweep that lists every document id and version in the source, compares them with what the index holds, and enqueues anything missing, extra or out of date. The event path provides speed; the sweep provides correctness. Report how many discrepancies each sweep finds, since a rising count means the event path is losing changes.
Stable chunks and content-hash diffing
Incremental embedding only saves money if a small edit changes a small number of chunks. Fixed-size chunking with overlap defeats this: inserting one sentence near the start of a document shifts every boundary after it, so every chunk's text changes and every chunk must be re-embedded. Structure-aware chunking, which splits on headings, sections and paragraphs and identifies each chunk by its position in the document structure, keeps boundaries stable, so an edit to one section changes only that section's chunks.
With stable chunk identities, the pipeline can diff. For each chunk, compute a hash of its text together with the embedding model version and the chunking configuration, and store it alongside the vector. On update, re-chunk the document, compare hashes, embed only the chunks whose hash changed, and delete chunks that no longer exist. Including the model version and chunking settings in the hash means that changing either one correctly marks everything as changed.
import hashlib
def chunk_hash(text, model_version, chunker_version):
h = hashlib.sha256()
for part in (model_version, chunker_version, text):
h.update(part.encode("utf-8"))
h.update(b"\x00")
return h.hexdigest()
def sync_document(doc, index, embed, model_version, chunker_version):
new = {c.key: c for c in chunk_by_structure(doc)} # key = doc_id::section_path::ordinal
old = index.stored_hashes(doc.id) # key -> hash from the last sync
changed = []
for key, c in new.items():
h = chunk_hash(c.text, model_version, chunker_version)
if old.get(key) != h:
changed.append((key, c, h))
vectors = embed([c.text for _, c, _ in changed]) if changed else []
index.upsert([
{"key": key, "doc_id": doc.id, "text": c.text, "hash": h,
"vector": v, "source_version": doc.version}
for (key, c, h), v in zip(changed, vectors)
])
index.delete([key for key in old if key not in new]) # orphaned chunks
return len(changed), len(new)In practice most updates touch a small fraction of a document's chunks, so the hash diff typically cuts embedding volume by a large factor compared with re-embedding whole documents. The saving is measurable, so measure it: log changed and total chunks per update, and watch the ratio after any change to the chunker.
Ordering, idempotency and deletes
Change events arrive late, twice or in the wrong order. If an update for version 7 is processed after version 8, the index quietly regresses to old content. Guard against this by storing the source version with each document and ignoring any event whose version is not newer than the one already applied. Route all events for one document through the same queue partition, keyed by document id, so that per-document processing is serialized, and make every step idempotent so that redelivery is harmless.
def apply_event(evt, store, index):
current = store.source_version(evt.doc_id)
if current is not None and evt.source_version <= current:
return "stale" # duplicate or out-of-order delivery
if evt.kind == "delete":
index.delete_by_doc(evt.doc_id)
store.tombstone(evt.doc_id, evt.source_version) # blocks late updates
else:
doc = fetch(evt.doc_id) # fetch current content, not the event payload
sync_document(doc, index, embed, MODEL_VERSION, CHUNKER_VERSION)
store.set_version(evt.doc_id, doc.version)
return "applied"Deletes need tombstones. Without one, a delayed update event for a deleted document looks newer than anything stored and resurrects the content. Keep tombstones for longer than the maximum plausible event delay, and have reconciliation treat documents missing from the source as deletes. Also check what deletion means in your vector store: many engines mark vectors as deleted and reclaim space in a later compaction or segment merge, and graph indexes can lose some recall as deleted nodes accumulate. Monitor the deleted fraction and schedule compaction or rebuilds before it degrades search.
Throughput, batching and backpressure
Embedding is the rate-limited stage. Hosted embedding APIs enforce request and token rate limits, and self-hosted embedding servers have fixed throughput. Batch chunks into requests sized for the provider's limits, and run live updates and bulk backfills in separate queues with separate budgets, so that a large backfill never starves the fresh-change path that the freshness objective depends on. When the embedding stage falls behind, let queues grow and alert on queue age rather than dropping events; the reconciliation sweep is a safety net, not a substitute.
Fetch current content at processing time rather than trusting the event payload. If ten edits to the same document arrive in a minute, coalescing them and fetching once produces the same final state as processing all ten, at a tenth of the cost.
Freshness also belongs on the query side. Store each chunk's source modification time and indexing time as metadata, pass them to the reranker or prompt so that newer versions can win ties between near-duplicate passages, and show the source date in citations. Users who can see that a cited policy was last updated two years ago can judge its reliability themselves, which a silently stale answer never allows.
Embedding model upgrades
Vectors from different embedding models, or even different versions of the same model, live in incompatible spaces, and a query embedded with one cannot be meaningfully compared with documents embedded with another. An embedding model upgrade is therefore a full re-embedding, and it must happen without a period of mixed or partial results. The standard pattern is blue-green: build a new index with the new model in parallel, dual-write live changes to both indexes while the backfill runs, evaluate retrieval quality on the new index with the golden set, then switch queries to it with an alias swap and keep the old index for a rollback window.
Estimate the cost before starting: total tokens in the corpus times the embedding price, or corpus tokens divided by self-hosted throughput. The backfill must also share rate limits with the live path, which is another reason to keep separate queues. Because the model version is part of the chunk hash, the new index's pipeline naturally treats every chunk as changed, and the old index's incremental path keeps working untouched until cutover.
Monitoring freshness
Pipeline metrics such as queue depth, embedding errors and processing time are necessary but indirect. Measure freshness end to end with a canary: at a fixed interval, write or edit a probe document in each source containing a unique marker, then query the retrieval service until the marker is returned, and record the elapsed time. Do the same for deletes, confirming that a deleted probe stops appearing. These probes measure exactly what the objective promises, through every stage, including the ones you did not think to instrument.
Alongside the canary, track the lag distribution for real changes where the source provides a modification time, the discrepancy count from each reconciliation sweep, the age of the oldest queued event, and the share of chunks re-embedded per update. Alert on the canary exceeding its objective and on reconciliation discrepancies rising, since those two signals catch most real failures before users notice stale answers.
Failure modes
- Shifting chunk boundaries. Fixed-size chunking turns every small edit into a full document re-embed.
- Hash without model version. A model or chunker change leaves stale vectors marked as current.
- Out-of-order regression. Without source-version checks, late events roll documents back.
- Resurrected deletes. Missing tombstones let delayed updates bring deleted content back.
- Backfill starvation. A bulk re-embed shares a queue with live changes and blows the freshness objective.
- Mixed embedding spaces. Upgrading the model in place leaves queries comparing vectors from two models.