Every vector database demo looks the same: embed some text, insert it, ask a question, get the nearest chunks back. The four products people shortlist most often, Pinecone, Weaviate, Qdrant and Chroma, all pass that demo. They stop looking the same when you add a metadata filter that matches 0.5% of the data, a few thousand tenants, a write rate that never stops, a hard memory budget and a requirement that a deleted document disappears from results now rather than eventually.
This article compares the four on the questions that decide production behaviour: how each models data, how each executes a filtered nearest-neighbour query, what hybrid search means in each, how tenants are isolated, what consistency you get after a write, and what it costs to run. Because the only comparison that matters is the one you run on your own data, it ends with a benchmark harness. Product features move quickly; confirm version-specific details against current docs.
What a vector database adds to an ANN index
An approximate-nearest-neighbour (ANN) library such as FAISS or hnswlib answers one question: given a query vector, which stored vectors are closest? A vector database wraps that index with everything a service needs: durable writes through a log, updates and deletes on a structure that was designed to be built once, metadata stored next to each vector, filtering on that metadata during the search, horizontal scale, replication, access control and an API. Most of the engineering difference between products lives in that wrapper, not in the graph algorithm. All four use HNSW, or something close to it, as their main dense index for most workloads.
HNSW builds a layered proximity graph. A search starts at a sparse top layer, greedily walks towards the query, drops a layer, and repeats until the bottom layer holds every vector. Neighbours per node (M) and the build-time candidate list (ef_construction) are fixed at build time; the search-time candidate list (ef) is set per query. Larger values buy recall with memory and latency.
| Pinecone | Weaviate | Qdrant | Chroma | |
|---|---|---|---|---|
| Distribution | Managed service only | Open source (BSD-3-Clause), self-host or Weaviate Cloud | Open source (Apache 2.0), self-host or Qdrant Cloud | Open source (Apache 2.0), embedded, server or Chroma Cloud |
| Unit of data | Index, namespace, record (id, values, sparse values, metadata) | Collection, object (properties, one or more named vectors) | Collection, point (id, named vectors, payload) | Collection, record (id, embedding, document, metadata) |
| Dense index | Managed; not user-tuned on serverless | HNSW, flat, or dynamic (flat that switches to HNSW) | HNSW with payload-aware links | HNSW for local and single-node deployments |
| Keyword or sparse | Sparse-dense vectors | BM25 plus hybrid fusion built in | Sparse vectors plus server-side fusion | Document substring filters in OSS |
| Scale-out | Handled by the service | Shards and replicas | Shards and replicas | Single node in OSS |
The same operation in four clients
The fastest way to feel the data models is to write one operation four times: create a collection for 768-dimensional cosine vectors, insert a chunk with a tenant id and a document type, and run a top-5 query restricted to one tenant and one type. The code below uses the current Python clients. Signatures do change between major versions, so pin client versions in your project.
# --- Pinecone (serverless) ---
from pinecone import Pinecone, ServerlessSpec
pc = Pinecone(api_key=PINECONE_API_KEY)
pc.create_index(name="docs", dimension=768, metric="cosine",
spec=ServerlessSpec(cloud="aws", region="us-east-1"))
idx = pc.Index("docs")
idx.upsert(vectors=[{"id": "c1", "values": vec, "metadata": {"type": "faq"}}],
namespace="tenant-42") # tenant = namespace
res = idx.query(vector=qvec, top_k=5, namespace="tenant-42",
filter={"type": {"$eq": "faq"}}, include_metadata=True)
# --- Weaviate (Python client v4) ---
import weaviate
from weaviate.classes.config import Configure, Property, DataType
from weaviate.classes.query import Filter
client = weaviate.connect_to_local()
client.collections.create(
"Docs",
# bring your own vectors; newer v4 clients replace vectorizer_config with vector_config
vectorizer_config=Configure.Vectorizer.none(),
multi_tenancy_config=Configure.multi_tenancy(enabled=True),
properties=[Property(name="type", data_type=DataType.TEXT),
Property(name="body", data_type=DataType.TEXT)])
docs = client.collections.get("Docs")
docs.tenants.create(["tenant-42"])
t = docs.with_tenant("tenant-42")
t.data.insert(properties={"type": "faq", "body": text}, vector=vec)
res = t.query.near_vector(near_vector=qvec, limit=5,
filters=Filter.by_property("type").equal("faq"))
# --- Qdrant ---
from qdrant_client import QdrantClient, models
qc = QdrantClient(url="http://localhost:6333")
qc.create_collection("docs", vectors_config=models.VectorParams(
size=768, distance=models.Distance.COSINE))
qc.create_payload_index("docs", field_name="tenant", field_schema=models.KeywordIndexParams(
type="keyword", is_tenant=True)) # is_tenant: 1.11+
qc.create_payload_index("docs", field_name="type",
field_schema=models.PayloadSchemaType.KEYWORD)
qc.upsert("docs", points=[models.PointStruct(
id=1, vector=vec, payload={"tenant": "tenant-42", "type": "faq"})])
res = qc.query_points("docs", query=qvec, limit=5, query_filter=models.Filter(must=[
models.FieldCondition(key="tenant", match=models.MatchValue(value="tenant-42")),
models.FieldCondition(key="type", match=models.MatchValue(value="faq"))]))
# --- Chroma (1.x) ---
import chromadb
cc = chromadb.PersistentClient(path="./chroma")
col = cc.get_or_create_collection("docs", configuration={"hnsw": {"space": "cosine"}})
col.add(ids=["c1"], embeddings=[vec], documents=[text],
metadatas=[{"tenant": "tenant-42", "type": "faq"}])
res = col.query(query_embeddings=[qvec], n_results=5,
where={"$and": [{"tenant": "tenant-42"}, {"type": "faq"}]})Notice where the tenant lives. Pinecone and Weaviate make it a structural partition (a namespace, a tenant shard), so a query physically cannot reach another tenant's vectors. Qdrant and Chroma, as written here, make it a metadata field, so isolation depends on every query carrying the filter. Qdrant's is_tenant flag improves locality for tenant-filtered queries, but it is a storage hint, not an access-control rule. Also notice that Qdrant wants payload indexes declared explicitly: filtering on an unindexed field still works, but it cannot use the filter-aware graph, and latency degrades as the collection grows.
Filtering: the part that decides latency
Filtered ANN search is where products differ most. There are three basic strategies. Post-filtering searches the graph for the top k, then drops non-matching results; with a selective filter you get fewer than k results or none, so implementations over-fetch and still fail when only 0.1% of points match. Pre-filtering computes the set of matching ids first, then searches only among them; for very selective filters the best plan is to skip the graph and compute exact distances over the matching ids. Filter-aware traversal walks the graph while skipping non-matching nodes, which works until the matching nodes are so sparse that the graph falls apart into disconnected islands.
Qdrant's answer is to build extra graph links within payload-index values, so that a filtered walk stays connected, and to use a query planner that estimates filter cardinality and switches to a payload-index scan plus exact scoring when the filter is selective. Weaviate builds an allow-list from its inverted index, searches the graph against it, falls back to flat search for small allow-lists, and recent versions add an ACORN-style strategy for restrictive filters. Pinecone plans filtered queries inside the service, so measure it on your own selectivity distribution. Chroma applies the metadata predicate and searches among matching records, which is fine at single-node scale.
Make your most frequent filter a partition key where the product supports one, and benchmark filtered queries at production selectivities: unfiltered recall says little about filtered recall.
Hybrid search and sparse vectors
Dense embeddings miss exact identifiers: product codes, error strings, people's names. Hybrid search combines a lexical or sparse signal with the dense one. Weaviate has the most built-in version: it maintains a BM25 index over text properties, and its hybrid query runs both searches and fuses them with either ranked fusion or relative-score fusion, weighted by an alpha parameter between 0 (pure keyword) and 1 (pure vector). Qdrant stores sparse vectors as first-class named vectors and its query API can prefetch candidates from a dense and a sparse vector and fuse them server-side with reciprocal rank fusion. You produce the sparse vectors yourself, for example from a learned sparse model. Pinecone supports sparse-dense records, where each record carries both dense values and a sparse index-and-value list. Open-source Chroma offers document filters such as substring containment; that is a filter, not a ranked lexical score.
Whichever product you choose, add a reranker on the fused top 50 to 100 if quality matters. A cross-encoder orders candidates far better than either score alone.
Multi-tenancy and isolation
SaaS retrieval almost always means many tenants, each a small fraction of the data. You have three designs. One collection per tenant isolates best and carries the most overhead once tenants number in thousands. One shared collection with a tenant filter is cheap and relies on code never forgetting the filter. Native tenant partitions sit between: Weaviate's multi-tenancy creates a shard per tenant inside one collection, and lets idle tenants be deactivated or offloaded so they stop consuming memory; Pinecone namespaces partition an index so a query addresses exactly one namespace; Chroma's server has a tenant and database hierarchy above collections.
Whatever the product does, route every query through one function that derives the tenant from the authenticated principal, and test that a tenantless query is refused.
Consistency, durability and scale-out
All four acknowledge a write once it is durable in a log, and make it searchable when the index catches up. The gap is the freshness lag, and it bites 'upload, then immediately ask' flows. Pinecone documents that serverless writes are eventually consistent and become visible after a short delay. Weaviate replicates data with a leaderless design and tunable consistency levels per request (ONE, QUORUM, ALL). Qdrant replicates shards with a configurable replication factor, a write consistency factor that sets how many replicas must acknowledge a write, and a read consistency option. Open-source Chroma runs on one node, so durability is that node's disk and your backups.
Deletes deserve a separate test. In graph indexes a delete is usually a tombstone that is cleaned up during later optimisation, so deleted ids must be filtered out of results until then.
Worked example: sizing 40 million chunks
A support-search product holds 40 million chunks across 3,000 tenants, embedded at 1,024 dimensions in float32. Raw vectors take 40,000,000 x 1,024 x 4 bytes = 163.8 GB. The HNSW bottom layer with 32 neighbour links of 4 bytes each adds roughly 128 bytes per vector, about 5.1 GB, plus upper layers and metadata. Keeping everything in RAM therefore needs around 175 GB before replication; with two replicas, about 350 GB.
Quantization changes the picture. Scalar int8 quantization stores one byte per dimension: 41 GB for the vectors, usually with a small recall loss that rescoring against full-precision vectors on disk recovers. Binary quantization stores one bit per dimension: about 5.1 GB, very fast, and acceptable only for embedding models that tolerate it, always with oversampling and rescoring. Qdrant and Weaviate both offer scalar, product and binary quantization. On Pinecone serverless the same arithmetic becomes a cost estimate rather than a capacity plan. On single-node Chroma, 175 GB of in-memory index is past what one machine should hold.
A benchmark harness you can run yourself
Build ground truth from your own data with exact search, then measure each product against it.
import numpy as np, time
def ground_truth(base, queries, masks, k=10):
"""Exact top-k by cosine over only the vectors each query's filter allows."""
base_n = base / np.linalg.norm(base, axis=1, keepdims=True)
truth = []
for q, allowed in zip(queries, masks):
ids = np.flatnonzero(allowed)
sims = base_n[ids] @ (q / np.linalg.norm(q))
truth.append(set(ids[np.argsort(-sims)[:k]]))
return truth
def evaluate(search_fn, queries, filters, truth, k=10):
"""search_fn(query, filter, k) -> list of ids; wraps any of the four clients."""
recalls, lat = [], []
for q, f, t in zip(queries, filters, truth):
t0 = time.perf_counter()
got = search_fn(q, f, k)
lat.append(time.perf_counter() - t0)
recalls.append(len(set(got) & t) / max(1, len(t)))
lat = np.array(lat) * 1000
return {"recall@k": float(np.mean(recalls)),
"p50_ms": float(np.percentile(lat, 50)),
"p99_ms": float(np.percentile(lat, 99))}
# Run evaluate() per selectivity bucket (50%, 5%, 0.5%, 0.05%),
# once idle and once while a writer upserts and deletes in the background.Record recall, p50 and p99 latency at two or three ef settings, and cost per million queries.
Failure modes seen in production
- Empty or short result lists under selective filters. Post-filtering behaviour, or an unindexed payload field in Qdrant. Use payload indexes or partitions.
- Metric mismatch. Inserting unnormalised vectors into a dot-product index built for normalised embeddings, or creating a Chroma collection with the default L2 space for a model trained for cosine.
- Mixed embedding versions. A model upgrade re-embeds half the corpus. Re-embed into a new collection and switch with an alias.
- Tenant leakage. One code path forgot the tenant filter.
- Recall drift after heavy deletes and updates. Tombstones and incremental inserts degrade graph quality. Track recall on a fixed probe set.
Choosing
Pick Pinecone when you want zero operations, accept a managed-only dependency and can live with the service's tuning choices and pricing model. Pick Weaviate when built-in BM25 hybrid search, native multi-tenancy with tenant offloading, and optional vectorizer modules match your application. Pick Qdrant when you need fine control over filtering, quantization, named and sparse vectors,. Pick Chroma for local development, embedded use and modest single-node corpora. If your data already lives in Postgres and the corpus is a few million vectors, pgvector may beat all four on total cost; compare it in the same harness.
Keep a thin retrieval interface in your code and store source text and embedding model version outside the database, so a re-embed or a product move is a batch job.
Related reading on this site
- Vector search infrastructure: index types, sharding and the recall-latency-memory triangle
- Memory and vector store options for agentic systems, including when pgvector is enough
- Pinecone math and architecture: pods, serverless and sizing arithmetic
- Vector database architecture in depth
- RAG pipeline at scale
What to do next
- Write down your corpus size, dimension, write rate, tenant count and the three filters you will run most often, with their selectivity.
- Compute raw and quantized memory with the arithmetic above to see which deployment shapes are even possible.
- Export 100,000 real vectors and 1,000 real queries with filters, and build exact ground truth with the harness.
- Load the same data into your two strongest candidates and measure recall, p50 and p99 per selectivity bucket, idle and under writes.
- Test the operations you will need on a bad day: delete-then-query, tenant isolation without a filter, backup and restore, and a re-embed into a new collection behind an alias.
- Wrap the winner behind a small retrieval interface and record the embedding model version with every vector.