A retrieval-augmented application has a supply chain that most teams never draw. The language model gets attention: where the weights came from, which container serves them. The retrieval tier gets treated as a database, and databases feel like infrastructure rather than dependencies. But a vector store is assembled from parts you did not write and usually did not review: an engine shipped as a container image, a managed service or a Postgres extension; a client SDK; an embedding model pulled from a hub; connectors and document parsers; and files that carry index state between environments, such as snapshots, backups and saved indexes. Any of these can change what the model is shown, and some of them can execute code on the machine that loads them.

This page treats the vector store as a supply-chain problem. It inventories the components, shows where each one can be compromised and what an attacker gains, then works through controls with code: refusing pickle-based index formats, verifying artifacts against a signed manifest, pinning and fingerprinting the embedding model, and carrying provenance through ingestion. Content poisoning as such, where an attacker writes a malicious document into a source you index, is covered in prompt injection via RAG and RAG document curation; this page is about the machinery around the content.

The retrieval tier's real inventory

Start with an inventory, because a control you cannot attach to a named component does not get maintained. Only two rows below are code in the conventional sense. The others are data that behaves like code: a model whose weights decide which documents are near each other, and index files whose contents decide what comes back for a query.

ComponentHow it arrivesIf tampered with
Vector engineContainer image, OS package, Postgres extension, or managed serviceArbitrary behaviour on the data plane: exfiltration, altered results, credential theft
Client SDKPackage index (pip, npm, Maven)Runs inside your application with its credentials
Embedding modelModel hub download, sometimes with custom codeCode execution on load if remote code is trusted; silently altered retrieval if weights change
Connectors and parsersLibraries for PDF, HTML, office formats, SaaS APIsParser exploits on hostile files; injected or dropped text
Index artifactsSnapshots, backups, saved index directories, exported collectionsCode execution if the format is pickle; replaced vectors or payloads
Ingestion and query glueYour own code, plus framework packagesWrong model or collection used; filters bypassed

Retrieval quality and retrieval integrity are the same thing seen from two sides: anything that can degrade recall can also be used to steer it.

The retrieval tier as a dependency graph: six things you did not writeSource connectorsparsers, crawlers, APIsEmbedding modelweights + tokenizer + codeIngestion jobchunk, embed, upsertdocumentsvectorsVector engineimage / extension / SaaSIndex artifactssnapshots, saved indexesupsert via SDKQuery servicesame embedding modelLLM promptretrieved text = inputtop-k chunkssearchOther envsrestore, copyRed boxes arrive from outside your code review: a model hub, a vendor, a bucket, a parser library.Each can change what is retrieved, and two of them (saved indexes, model code) can run code on load.
The components of a retrieval tier. Red boxes are inputs from outside your review process; arrows show where their influence propagates.

Index files that execute code

The sharpest risk in the inventory is a file format. Python's pickle is not a data format; it is a small program that the loader executes to rebuild objects, and a crafted pickle can call any importable function during load. Several convenient ways to persist vector indexes use it. LangChain's FAISS integration is the best-known example: FAISS.load_local restores a FAISS index plus a docstore and id mapping, and the docstore is pickled, so the method requires an explicit allow_dangerous_deserialization=True before it will load.

The same pattern shows up elsewhere: torch.load with weights_only=False, numpy.load with allow_pickle=True, joblib dumps of scikit-learn nearest-neighbour models, and ad hoc pickle.dump of vectors and metadata. The rule: an index artifact that crosses a trust boundary must be in a format whose loader cannot execute code. FAISS's own faiss.write_index binary format, Parquet or Arrow for vectors and metadata, .npy loaded with allow_pickle=False, and safetensors for tensors all qualify. Engine-native snapshots are parsed by the engine, not by Python's object loader, but they still need integrity checks, because a well-formed snapshot can carry replaced vectors and payloads.

Enforce the rule in a loader that every service uses, rather than in a code review comment:

import hashlib, json
from pathlib import Path

import faiss
import pyarrow.parquet as pq

PICKLE_SUFFIXES = {".pkl", ".pickle", ".joblib", ".pt", ".pth"}

def sha256(path: Path) -> str:
    h = hashlib.sha256()
    with path.open("rb") as f:
        for block in iter(lambda: f.read(1 << 20), b""):
            h.update(block)
    return h.hexdigest()

def load_index_bundle(bundle: Path, manifest: dict):
    """manifest has already been signature-verified; see below."""
    skip = {"manifest.json", "manifest.json.sig"}     # verified separately
    files = sorted(p for p in bundle.rglob("*") if p.is_file() and p.name not in skip)
    names = {str(p.relative_to(bundle)) for p in files}
    if names != set(manifest["files"]):
        raise ValueError(f"bundle file set differs from manifest: {names ^ set(manifest['files'])}")
    for p in files:
        if p.suffix in PICKLE_SUFFIXES:
            raise ValueError(f"refusing executable format: {p.name}")
        if sha256(p) != manifest["files"][str(p.relative_to(bundle))]:
            raise ValueError(f"digest mismatch: {p.name}")
    index = faiss.read_index(str(bundle / "vectors.faiss"))
    meta = pq.read_table(bundle / "chunks.parquet")       # id, text, source_uri, source_digest, model_id
    if index.ntotal != meta.num_rows:
        raise ValueError("vector count and metadata row count disagree")
    if meta.column("model_id").unique().to_pylist() != [manifest["embedding_model"]]:
        raise ValueError("bundle was embedded with a different model")
    return index, meta

Sign the manifest, not every file. A JSON manifest listing each file's SHA-256, the embedding model identity and the source snapshot is small, and a detached signature over it covers the bundle. Sigstore's cosign sign-blob and cosign verify-blob work for this, as does any signing service you already use for build artifacts. Verify the signature first, then call the loader with the parsed manifest. The file-set comparison matters: a check that only verifies listed files lets an attacker add an extra file that some other component picks up.

The embedding model is a dependency

The embedding model is a function from text to vectors, and retrieval is geometry on its output. Change the function and the geometry changes everywhere at once. Three ways this goes wrong:

  1. Code on load. Some embedding models ship custom modelling code and require trust_remote_code=True. With that flag the hub repository's Python runs in your ingestion job and your query service. A pinned revision makes it reproducible; it does not make it reviewed.
  2. Floating revisions. Loading by repository name pulls whatever the default branch points to today. If the owner pushes new weights, or the account is compromised, the next deployment of your query service embeds queries with a different model from the one that embedded the corpus. Recall collapses quietly, or, worse, shifts in a direction someone chose.
  3. Targeted weights. A model can be fine-tuned so that ordinary queries behave normally while a trigger phrase maps close to an attacker's document. This is hard to detect from aggregate metrics, which is why the defence is provenance rather than testing alone.

Pin by immutable revision, mirror the artifact into storage you control, and record the digest. Then add a behavioural fingerprint: a fixed set of canary texts whose embeddings you store at the time you approve the model. Both the ingestion job and the query service recompute them at start-up and refuse to run if they drift. This catches a swapped model, a different tokenizer, a changed pooling setting, and a precision change that is larger than you thought.

import numpy as np
from sentence_transformers import SentenceTransformer

MODEL_DIR = "/models/embed/approved-2026-09"   # mirrored copy, digest recorded in the release manifest
CANARIES = ["refund policy for annual plans", "rotate the database password",
            "error 503 from the payments gateway", "quarterly revenue by region"]

def check_fingerprint(model, expected: np.ndarray, tol: float = 1e-3) -> None:
    got = model.encode(CANARIES, normalize_embeddings=True)
    cos = np.sum(got * expected, axis=1)              # both sides unit-normalised
    if (cos < 1 - tol).any():
        raise RuntimeError(f"embedding fingerprint drift: min cosine {cos.min():.5f}")

model = SentenceTransformer(MODEL_DIR, trust_remote_code=False)
check_fingerprint(model, np.load("/models/embed/approved-2026-09.canaries.npy", allow_pickle=False))

Store the model identity next to every vector, as the loader above does with model_id. Mixed-model collections are the most common silent retrieval failure, and a single column makes them detectable with one query.

Connectors, parsers and provenance

Connectors and parsers sit where hostile input first touches your pipeline. A PDF or office document is a complex format, and parser libraries have a history of memory-safety and resource-exhaustion bugs. Run parsing in a sandboxed worker with no credentials for the vector store, a memory and time limit, and no network access beyond the source it reads. The worker emits plain text and metadata; a separate, credentialed job embeds and upserts.

Record provenance here, because it cannot be reconstructed later: for each chunk keep the source URI, source digest, connector and parser versions, ingestion run id and embedding model id. Then an incident becomes a query: when a parser version is compromised, delete every chunk it produced and re-ingest. The same columns let the query service filter retrieval to sources that passed review, which is the hook the controls in RAG defence architecture build on.

Engine, extension and SDK

The engine itself is the most conventional part of the problem and should get conventional treatment. Pin the container image by digest, not by tag; keep an SBOM; admit only signed images; and treat a Postgres extension such as pgvector as part of the database build rather than something installed by hand on one server. The practices in supply chain for LLM serving containers apply unchanged to vector engines.

Two engine-specific points. First, check the authentication default. In several popular self-hosted engines authentication is optional and off unless configured, which is convenient on a laptop and dangerous on a shared network, because the snapshot and collection-management endpoints are on the same port as search. Turn on API keys or mTLS, and give the query service a credential that can search but not create snapshots or delete collections. Second, for managed services the supply chain is the vendor's, so the questions become contractual and architectural: where are backups kept, who can restore them, can you export your data in an open format, and does the service log administrative actions somewhere you can read. The client SDK belongs in a hash-checked lockfile, since it runs with your application's credentials.

Worked example: one index, three regions

Consider a team that builds a support assistant over 400,000 help-centre chunks. Embedding the corpus takes several GPU-hours, so they build the index once in a staging project and copy it to three production regions. The first version of the pipeline did this with FAISS.save_local on a shared bucket and FAISS.load_local(..., allow_dangerous_deserialization=True) in each region. Anyone with write access to that bucket, including a CI token used by an unrelated job, could have replaced the pickle and executed code in every query pod at the next restart.

The rebuilt flow has five steps, each with a single owner:

  1. The build job, running from a pinned image, writes vectors.faiss with faiss.write_index and chunks.parquet with provenance columns. No pickle anywhere.
  2. It writes manifest.json with per-file SHA-256 digests, the embedding model's mirrored revision and canary fingerprint file digest, the source snapshot id and the build run id.
  3. A signing step, with a key the build job cannot use for anything else, signs the manifest.
  4. Each region's loader verifies the signature, then runs the bundle check and the embedding fingerprint check before the pod reports ready. A failure keeps the previous bundle serving.
  5. The bucket grants write access to the build identity only; the regions read. Object versioning keeps prior bundles for rollback.

The added cost is seconds of hashing per load. The bucket stops being a code-execution path, and a swapped embedding model fails start-up instead of degrading answers.

Failure modes

FailureSymptomControl
Pickled index loaded from shared storageNone until exploitedBan pickle formats in the loader; scan repos for the dangerous-deserialization flag
Embedding model revision floatsRecall drops after a deploy with no code changePin revision, mirror, fingerprint at start-up
Mixed models in one collectionSome queries never find recent documentsmodel_id per vector; refuse mixed bundles
Snapshot restored from an unverified sourceUnexpected documents in answersSigned manifest, restore only from build identity's bucket
Engine exposed without authUnknown collections or snapshots appearEnable auth; separate search and admin credentials
Parser exploit on a hostile fileCrashed or hung ingestion workersSandboxed parsing with no store credentials
Verification only of listed filesExtra file picked up by another toolCompare the full file set to the manifest

Most of these have no symptom until exploited, which is why the controls are gates rather than alerts.

Trade-offs

Signed, pickle-free bundles cost engineering time up front and make quick experiments slightly slower, since the convenient save and load helpers in frameworks are often the pickled ones. A reasonable split is to allow anything inside a single notebook or developer machine and enforce the loader at the first boundary: anything written to shared storage, or loaded by a service.

Mirroring embedding models means you own updates. You will fall behind upstream fixes unless re-approving a model is a routine runbook that includes re-embedding the corpus. Fingerprint tolerances need care: measure the cosine identical weights give across your real hardware, and keep the threshold well clear of a different model.

Managed services trade the patching problem for trust in a vendor; keep an open-format export with provenance columns so you can rebuild elsewhere. For how the engines differ, see vector databases compared.

What to do next

  1. Draw the retrieval tier's dependency inventory: engine, SDK, extension, embedding model, connectors, parsers, artifact stores. Give each an owner.
  2. Search your repositories for allow_dangerous_deserialization, allow_pickle=True, weights_only=False, joblib.load and pickle.load on index or embedding files. Replace each at a trust boundary.
  3. Write one shared loader that rejects pickle formats, verifies a signed manifest and compares the full file set.
  4. Pin the embedding model by immutable revision, mirror it, record its digest, and add a canary fingerprint check to ingestion and query start-up.
  5. Add model_id, source URI, source digest, parser version and run id to every chunk.
  6. Move document parsing into a sandboxed worker without vector store credentials.
  7. Pin engine images by digest, enable authentication, and split search credentials from admin and snapshot credentials.
  8. Restrict write access on snapshot and bundle buckets to the build identity; turn on object versioning.
  9. Add retrieval canaries and a small labelled query set to the deployment gate.
Key takeaway: Treat the vector store as a set of dependencies, not a database you happen to run. Ban pickle-based index formats at every trust boundary, sign a manifest that lists every file and the embedding model, pin and fingerprint that model in both ingestion and query services, parse untrusted documents in a sandbox, and keep provenance on every chunk so an incident becomes a query instead of a rebuild.