A retrieval-augmented application has a supply chain that most teams never draw. The language model gets attention: where the weights came from, which container serves them. The retrieval tier gets treated as a database, and databases feel like infrastructure rather than dependencies. But a vector store is assembled from parts you did not write and usually did not review: an engine shipped as a container image, a managed service or a Postgres extension; a client SDK; an embedding model pulled from a hub; connectors and document parsers; and files that carry index state between environments, such as snapshots, backups and saved indexes. Any of these can change what the model is shown, and some of them can execute code on the machine that loads them.
This page treats the vector store as a supply-chain problem. It inventories the components, shows where each one can be compromised and what an attacker gains, then works through controls with code: refusing pickle-based index formats, verifying artifacts against a signed manifest, pinning and fingerprinting the embedding model, and carrying provenance through ingestion. Content poisoning as such, where an attacker writes a malicious document into a source you index, is covered in prompt injection via RAG and RAG document curation; this page is about the machinery around the content.
The retrieval tier's real inventory
Start with an inventory, because a control you cannot attach to a named component does not get maintained. Only two rows below are code in the conventional sense. The others are data that behaves like code: a model whose weights decide which documents are near each other, and index files whose contents decide what comes back for a query.
| Component | How it arrives | If tampered with |
|---|---|---|
| Vector engine | Container image, OS package, Postgres extension, or managed service | Arbitrary behaviour on the data plane: exfiltration, altered results, credential theft |
| Client SDK | Package index (pip, npm, Maven) | Runs inside your application with its credentials |
| Embedding model | Model hub download, sometimes with custom code | Code execution on load if remote code is trusted; silently altered retrieval if weights change |
| Connectors and parsers | Libraries for PDF, HTML, office formats, SaaS APIs | Parser exploits on hostile files; injected or dropped text |
| Index artifacts | Snapshots, backups, saved index directories, exported collections | Code execution if the format is pickle; replaced vectors or payloads |
| Ingestion and query glue | Your own code, plus framework packages | Wrong model or collection used; filters bypassed |
Retrieval quality and retrieval integrity are the same thing seen from two sides: anything that can degrade recall can also be used to steer it.
Index files that execute code
The sharpest risk in the inventory is a file format. Python's pickle is not a data format; it is a small program that the loader executes to rebuild objects, and a crafted pickle can call any importable function during load. Several convenient ways to persist vector indexes use it. LangChain's FAISS integration is the best-known example: FAISS.load_local restores a FAISS index plus a docstore and id mapping, and the docstore is pickled, so the method requires an explicit allow_dangerous_deserialization=True before it will load.
The same pattern shows up elsewhere: torch.load with weights_only=False, numpy.load with allow_pickle=True, joblib dumps of scikit-learn nearest-neighbour models, and ad hoc pickle.dump of vectors and metadata. The rule: an index artifact that crosses a trust boundary must be in a format whose loader cannot execute code. FAISS's own faiss.write_index binary format, Parquet or Arrow for vectors and metadata, .npy loaded with allow_pickle=False, and safetensors for tensors all qualify. Engine-native snapshots are parsed by the engine, not by Python's object loader, but they still need integrity checks, because a well-formed snapshot can carry replaced vectors and payloads.
Enforce the rule in a loader that every service uses, rather than in a code review comment:
import hashlib, json
from pathlib import Path
import faiss
import pyarrow.parquet as pq
PICKLE_SUFFIXES = {".pkl", ".pickle", ".joblib", ".pt", ".pth"}
def sha256(path: Path) -> str:
h = hashlib.sha256()
with path.open("rb") as f:
for block in iter(lambda: f.read(1 << 20), b""):
h.update(block)
return h.hexdigest()
def load_index_bundle(bundle: Path, manifest: dict):
"""manifest has already been signature-verified; see below."""
skip = {"manifest.json", "manifest.json.sig"} # verified separately
files = sorted(p for p in bundle.rglob("*") if p.is_file() and p.name not in skip)
names = {str(p.relative_to(bundle)) for p in files}
if names != set(manifest["files"]):
raise ValueError(f"bundle file set differs from manifest: {names ^ set(manifest['files'])}")
for p in files:
if p.suffix in PICKLE_SUFFIXES:
raise ValueError(f"refusing executable format: {p.name}")
if sha256(p) != manifest["files"][str(p.relative_to(bundle))]:
raise ValueError(f"digest mismatch: {p.name}")
index = faiss.read_index(str(bundle / "vectors.faiss"))
meta = pq.read_table(bundle / "chunks.parquet") # id, text, source_uri, source_digest, model_id
if index.ntotal != meta.num_rows:
raise ValueError("vector count and metadata row count disagree")
if meta.column("model_id").unique().to_pylist() != [manifest["embedding_model"]]:
raise ValueError("bundle was embedded with a different model")
return index, metaSign the manifest, not every file. A JSON manifest listing each file's SHA-256, the embedding model identity and the source snapshot is small, and a detached signature over it covers the bundle. Sigstore's cosign sign-blob and cosign verify-blob work for this, as does any signing service you already use for build artifacts. Verify the signature first, then call the loader with the parsed manifest. The file-set comparison matters: a check that only verifies listed files lets an attacker add an extra file that some other component picks up.
The embedding model is a dependency
The embedding model is a function from text to vectors, and retrieval is geometry on its output. Change the function and the geometry changes everywhere at once. Three ways this goes wrong:
- Code on load. Some embedding models ship custom modelling code and require
trust_remote_code=True. With that flag the hub repository's Python runs in your ingestion job and your query service. A pinned revision makes it reproducible; it does not make it reviewed. - Floating revisions. Loading by repository name pulls whatever the default branch points to today. If the owner pushes new weights, or the account is compromised, the next deployment of your query service embeds queries with a different model from the one that embedded the corpus. Recall collapses quietly, or, worse, shifts in a direction someone chose.
- Targeted weights. A model can be fine-tuned so that ordinary queries behave normally while a trigger phrase maps close to an attacker's document. This is hard to detect from aggregate metrics, which is why the defence is provenance rather than testing alone.
Pin by immutable revision, mirror the artifact into storage you control, and record the digest. Then add a behavioural fingerprint: a fixed set of canary texts whose embeddings you store at the time you approve the model. Both the ingestion job and the query service recompute them at start-up and refuse to run if they drift. This catches a swapped model, a different tokenizer, a changed pooling setting, and a precision change that is larger than you thought.
import numpy as np
from sentence_transformers import SentenceTransformer
MODEL_DIR = "/models/embed/approved-2026-09" # mirrored copy, digest recorded in the release manifest
CANARIES = ["refund policy for annual plans", "rotate the database password",
"error 503 from the payments gateway", "quarterly revenue by region"]
def check_fingerprint(model, expected: np.ndarray, tol: float = 1e-3) -> None:
got = model.encode(CANARIES, normalize_embeddings=True)
cos = np.sum(got * expected, axis=1) # both sides unit-normalised
if (cos < 1 - tol).any():
raise RuntimeError(f"embedding fingerprint drift: min cosine {cos.min():.5f}")
model = SentenceTransformer(MODEL_DIR, trust_remote_code=False)
check_fingerprint(model, np.load("/models/embed/approved-2026-09.canaries.npy", allow_pickle=False))Store the model identity next to every vector, as the loader above does with model_id. Mixed-model collections are the most common silent retrieval failure, and a single column makes them detectable with one query.
Connectors, parsers and provenance
Connectors and parsers sit where hostile input first touches your pipeline. A PDF or office document is a complex format, and parser libraries have a history of memory-safety and resource-exhaustion bugs. Run parsing in a sandboxed worker with no credentials for the vector store, a memory and time limit, and no network access beyond the source it reads. The worker emits plain text and metadata; a separate, credentialed job embeds and upserts.
Record provenance here, because it cannot be reconstructed later: for each chunk keep the source URI, source digest, connector and parser versions, ingestion run id and embedding model id. Then an incident becomes a query: when a parser version is compromised, delete every chunk it produced and re-ingest. The same columns let the query service filter retrieval to sources that passed review, which is the hook the controls in RAG defence architecture build on.
Engine, extension and SDK
The engine itself is the most conventional part of the problem and should get conventional treatment. Pin the container image by digest, not by tag; keep an SBOM; admit only signed images; and treat a Postgres extension such as pgvector as part of the database build rather than something installed by hand on one server. The practices in supply chain for LLM serving containers apply unchanged to vector engines.
Two engine-specific points. First, check the authentication default. In several popular self-hosted engines authentication is optional and off unless configured, which is convenient on a laptop and dangerous on a shared network, because the snapshot and collection-management endpoints are on the same port as search. Turn on API keys or mTLS, and give the query service a credential that can search but not create snapshots or delete collections. Second, for managed services the supply chain is the vendor's, so the questions become contractual and architectural: where are backups kept, who can restore them, can you export your data in an open format, and does the service log administrative actions somewhere you can read. The client SDK belongs in a hash-checked lockfile, since it runs with your application's credentials.
Worked example: one index, three regions
Consider a team that builds a support assistant over 400,000 help-centre chunks. Embedding the corpus takes several GPU-hours, so they build the index once in a staging project and copy it to three production regions. The first version of the pipeline did this with FAISS.save_local on a shared bucket and FAISS.load_local(..., allow_dangerous_deserialization=True) in each region. Anyone with write access to that bucket, including a CI token used by an unrelated job, could have replaced the pickle and executed code in every query pod at the next restart.
The rebuilt flow has five steps, each with a single owner:
- The build job, running from a pinned image, writes
vectors.faisswithfaiss.write_indexandchunks.parquetwith provenance columns. No pickle anywhere. - It writes
manifest.jsonwith per-file SHA-256 digests, the embedding model's mirrored revision and canary fingerprint file digest, the source snapshot id and the build run id. - A signing step, with a key the build job cannot use for anything else, signs the manifest.
- Each region's loader verifies the signature, then runs the bundle check and the embedding fingerprint check before the pod reports ready. A failure keeps the previous bundle serving.
- The bucket grants write access to the build identity only; the regions read. Object versioning keeps prior bundles for rollback.
The added cost is seconds of hashing per load. The bucket stops being a code-execution path, and a swapped embedding model fails start-up instead of degrading answers.
Failure modes
| Failure | Symptom | Control |
|---|---|---|
| Pickled index loaded from shared storage | None until exploited | Ban pickle formats in the loader; scan repos for the dangerous-deserialization flag |
| Embedding model revision floats | Recall drops after a deploy with no code change | Pin revision, mirror, fingerprint at start-up |
| Mixed models in one collection | Some queries never find recent documents | model_id per vector; refuse mixed bundles |
| Snapshot restored from an unverified source | Unexpected documents in answers | Signed manifest, restore only from build identity's bucket |
| Engine exposed without auth | Unknown collections or snapshots appear | Enable auth; separate search and admin credentials |
| Parser exploit on a hostile file | Crashed or hung ingestion workers | Sandboxed parsing with no store credentials |
| Verification only of listed files | Extra file picked up by another tool | Compare the full file set to the manifest |
Most of these have no symptom until exploited, which is why the controls are gates rather than alerts.
Trade-offs
Signed, pickle-free bundles cost engineering time up front and make quick experiments slightly slower, since the convenient save and load helpers in frameworks are often the pickled ones. A reasonable split is to allow anything inside a single notebook or developer machine and enforce the loader at the first boundary: anything written to shared storage, or loaded by a service.
Mirroring embedding models means you own updates. You will fall behind upstream fixes unless re-approving a model is a routine runbook that includes re-embedding the corpus. Fingerprint tolerances need care: measure the cosine identical weights give across your real hardware, and keep the threshold well clear of a different model.
Managed services trade the patching problem for trust in a vendor; keep an open-format export with provenance columns so you can rebuild elsewhere. For how the engines differ, see vector databases compared.
What to do next
- Draw the retrieval tier's dependency inventory: engine, SDK, extension, embedding model, connectors, parsers, artifact stores. Give each an owner.
- Search your repositories for
allow_dangerous_deserialization,allow_pickle=True,weights_only=False,joblib.loadandpickle.loadon index or embedding files. Replace each at a trust boundary. - Write one shared loader that rejects pickle formats, verifies a signed manifest and compares the full file set.
- Pin the embedding model by immutable revision, mirror it, record its digest, and add a canary fingerprint check to ingestion and query start-up.
- Add
model_id, source URI, source digest, parser version and run id to every chunk. - Move document parsing into a sandboxed worker without vector store credentials.
- Pin engine images by digest, enable authentication, and split search credentials from admin and snapshot credentials.
- Restrict write access on snapshot and bundle buckets to the build identity; turn on object versioning.
- Add retrieval canaries and a small labelled query set to the deployment gate.