Every vector database demo looks the same: embed some text, insert it, ask a question, get the nearest chunks back. The four products people shortlist most often, Pinecone, Weaviate, Qdrant and Chroma, all pass that demo. They stop looking the same when you add a metadata filter that matches 0.5% of the data, a few thousand tenants, a write rate that never stops, a hard memory budget and a requirement that a deleted document disappears from results now rather than eventually.

This article compares the four on the questions that decide production behaviour: how each models data, how each executes a filtered nearest-neighbour query, what hybrid search means in each, how tenants are isolated, what consistency you get after a write, and what it costs to run. Because the only comparison that matters is the one you run on your own data, it ends with a benchmark harness. Product features move quickly; confirm version-specific details against current docs.

Advertisement

What a vector database adds to an ANN index

An approximate-nearest-neighbour (ANN) library such as FAISS or hnswlib answers one question: given a query vector, which stored vectors are closest? A vector database wraps that index with everything a service needs: durable writes through a log, updates and deletes on a structure that was designed to be built once, metadata stored next to each vector, filtering on that metadata during the search, horizontal scale, replication, access control and an API. Most of the engineering difference between products lives in that wrapper, not in the graph algorithm. All four use HNSW, or something close to it, as their main dense index for most workloads.

HNSW builds a layered proximity graph. A search starts at a sparse top layer, greedily walks towards the query, drops a layer, and repeats until the bottom layer holds every vector. Neighbours per node (M) and the build-time candidate list (ef_construction) are fixed at build time; the search-time candidate list (ef) is set per query. Larger values buy recall with memory and latency.

What every vector database does, and where Pinecone, Weaviate, Qdrant and Chroma differWRITE PATHEmbedyour model or moduleUpsert APIid, vector, metadataWAL / logdurability firstSegmentsmutable then sealedIndex buildHNSW, quantizedfreshness lag lives here: acknowledged is not yet searchableREAD PATHQueryvector + filterFilter plannerselectivity estimateANN searchfilter-aware graphRescorefull precisionFuse / rerankBM25, sparsevery selectiveBrute forceon matching idsPinecone: managed only; serverless, namespaces, the service plans filters for you.Weaviate: open source (Go); HNSW/flat/dynamic; BM25 hybrid and native tenants built in.Qdrant: open source (Rust); filterable HNSW over explicit payload indexes; rich quantization.Chroma: open source; embedded or client-server; smallest API, single node unless you use Chroma Cloud.
The common write and read path. The products differ mainly in the filter planner, how segments are merged and indexed, and where they run.
PineconeWeaviateQdrantChroma
DistributionManaged service onlyOpen source (BSD-3-Clause), self-host or Weaviate CloudOpen source (Apache 2.0), self-host or Qdrant CloudOpen source (Apache 2.0), embedded, server or Chroma Cloud
Unit of dataIndex, namespace, record (id, values, sparse values, metadata)Collection, object (properties, one or more named vectors)Collection, point (id, named vectors, payload)Collection, record (id, embedding, document, metadata)
Dense indexManaged; not user-tuned on serverlessHNSW, flat, or dynamic (flat that switches to HNSW)HNSW with payload-aware linksHNSW for local and single-node deployments
Keyword or sparseSparse-dense vectorsBM25 plus hybrid fusion built inSparse vectors plus server-side fusionDocument substring filters in OSS
Scale-outHandled by the serviceShards and replicasShards and replicasSingle node in OSS

The same operation in four clients

The fastest way to feel the data models is to write one operation four times: create a collection for 768-dimensional cosine vectors, insert a chunk with a tenant id and a document type, and run a top-5 query restricted to one tenant and one type. The code below uses the current Python clients. Signatures do change between major versions, so pin client versions in your project.

# --- Pinecone (serverless) ---
from pinecone import Pinecone, ServerlessSpec
pc = Pinecone(api_key=PINECONE_API_KEY)
pc.create_index(name="docs", dimension=768, metric="cosine",
                spec=ServerlessSpec(cloud="aws", region="us-east-1"))
idx = pc.Index("docs")
idx.upsert(vectors=[{"id": "c1", "values": vec, "metadata": {"type": "faq"}}],
           namespace="tenant-42")                       # tenant = namespace
res = idx.query(vector=qvec, top_k=5, namespace="tenant-42",
                filter={"type": {"$eq": "faq"}}, include_metadata=True)

# --- Weaviate (Python client v4) ---
import weaviate
from weaviate.classes.config import Configure, Property, DataType
from weaviate.classes.query import Filter
client = weaviate.connect_to_local()
client.collections.create(
    "Docs",
    # bring your own vectors; newer v4 clients replace vectorizer_config with vector_config
    vectorizer_config=Configure.Vectorizer.none(),
    multi_tenancy_config=Configure.multi_tenancy(enabled=True),
    properties=[Property(name="type", data_type=DataType.TEXT),
                Property(name="body", data_type=DataType.TEXT)])
docs = client.collections.get("Docs")
docs.tenants.create(["tenant-42"])
t = docs.with_tenant("tenant-42")
t.data.insert(properties={"type": "faq", "body": text}, vector=vec)
res = t.query.near_vector(near_vector=qvec, limit=5,
                          filters=Filter.by_property("type").equal("faq"))

# --- Qdrant ---
from qdrant_client import QdrantClient, models
qc = QdrantClient(url="http://localhost:6333")
qc.create_collection("docs", vectors_config=models.VectorParams(
    size=768, distance=models.Distance.COSINE))
qc.create_payload_index("docs", field_name="tenant", field_schema=models.KeywordIndexParams(
    type="keyword", is_tenant=True))                      # is_tenant: 1.11+
qc.create_payload_index("docs", field_name="type",
                        field_schema=models.PayloadSchemaType.KEYWORD)
qc.upsert("docs", points=[models.PointStruct(
    id=1, vector=vec, payload={"tenant": "tenant-42", "type": "faq"})])
res = qc.query_points("docs", query=qvec, limit=5, query_filter=models.Filter(must=[
    models.FieldCondition(key="tenant", match=models.MatchValue(value="tenant-42")),
    models.FieldCondition(key="type", match=models.MatchValue(value="faq"))]))

# --- Chroma (1.x) ---
import chromadb
cc = chromadb.PersistentClient(path="./chroma")
col = cc.get_or_create_collection("docs", configuration={"hnsw": {"space": "cosine"}})
col.add(ids=["c1"], embeddings=[vec], documents=[text],
        metadatas=[{"tenant": "tenant-42", "type": "faq"}])
res = col.query(query_embeddings=[qvec], n_results=5,
                where={"$and": [{"tenant": "tenant-42"}, {"type": "faq"}]})

Notice where the tenant lives. Pinecone and Weaviate make it a structural partition (a namespace, a tenant shard), so a query physically cannot reach another tenant's vectors. Qdrant and Chroma, as written here, make it a metadata field, so isolation depends on every query carrying the filter. Qdrant's is_tenant flag improves locality for tenant-filtered queries, but it is a storage hint, not an access-control rule. Also notice that Qdrant wants payload indexes declared explicitly: filtering on an unindexed field still works, but it cannot use the filter-aware graph, and latency degrades as the collection grows.

Advertisement

Filtering: the part that decides latency

Filtered ANN search is where products differ most. There are three basic strategies. Post-filtering searches the graph for the top k, then drops non-matching results; with a selective filter you get fewer than k results or none, so implementations over-fetch and still fail when only 0.1% of points match. Pre-filtering computes the set of matching ids first, then searches only among them; for very selective filters the best plan is to skip the graph and compute exact distances over the matching ids. Filter-aware traversal walks the graph while skipping non-matching nodes, which works until the matching nodes are so sparse that the graph falls apart into disconnected islands.

Qdrant's answer is to build extra graph links within payload-index values, so that a filtered walk stays connected, and to use a query planner that estimates filter cardinality and switches to a payload-index scan plus exact scoring when the filter is selective. Weaviate builds an allow-list from its inverted index, searches the graph against it, falls back to flat search for small allow-lists, and recent versions add an ACORN-style strategy for restrictive filters. Pinecone plans filtered queries inside the service, so measure it on your own selectivity distribution. Chroma applies the metadata predicate and searches among matching records, which is fine at single-node scale.

Make your most frequent filter a partition key where the product supports one, and benchmark filtered queries at production selectivities: unfiltered recall says little about filtered recall.

Hybrid search and sparse vectors

Dense embeddings miss exact identifiers: product codes, error strings, people's names. Hybrid search combines a lexical or sparse signal with the dense one. Weaviate has the most built-in version: it maintains a BM25 index over text properties, and its hybrid query runs both searches and fuses them with either ranked fusion or relative-score fusion, weighted by an alpha parameter between 0 (pure keyword) and 1 (pure vector). Qdrant stores sparse vectors as first-class named vectors and its query API can prefetch candidates from a dense and a sparse vector and fuse them server-side with reciprocal rank fusion. You produce the sparse vectors yourself, for example from a learned sparse model. Pinecone supports sparse-dense records, where each record carries both dense values and a sparse index-and-value list. Open-source Chroma offers document filters such as substring containment; that is a filter, not a ranked lexical score.

Whichever product you choose, add a reranker on the fused top 50 to 100 if quality matters. A cross-encoder orders candidates far better than either score alone.

Multi-tenancy and isolation

SaaS retrieval almost always means many tenants, each a small fraction of the data. You have three designs. One collection per tenant isolates best and carries the most overhead once tenants number in thousands. One shared collection with a tenant filter is cheap and relies on code never forgetting the filter. Native tenant partitions sit between: Weaviate's multi-tenancy creates a shard per tenant inside one collection, and lets idle tenants be deactivated or offloaded so they stop consuming memory; Pinecone namespaces partition an index so a query addresses exactly one namespace; Chroma's server has a tenant and database hierarchy above collections.

Whatever the product does, route every query through one function that derives the tenant from the authenticated principal, and test that a tenantless query is refused.

Consistency, durability and scale-out

All four acknowledge a write once it is durable in a log, and make it searchable when the index catches up. The gap is the freshness lag, and it bites 'upload, then immediately ask' flows. Pinecone documents that serverless writes are eventually consistent and become visible after a short delay. Weaviate replicates data with a leaderless design and tunable consistency levels per request (ONE, QUORUM, ALL). Qdrant replicates shards with a configurable replication factor, a write consistency factor that sets how many replicas must acknowledge a write, and a read consistency option. Open-source Chroma runs on one node, so durability is that node's disk and your backups.

Deletes deserve a separate test. In graph indexes a delete is usually a tombstone that is cleaned up during later optimisation, so deleted ids must be filtered out of results until then.

Worked example: sizing 40 million chunks

A support-search product holds 40 million chunks across 3,000 tenants, embedded at 1,024 dimensions in float32. Raw vectors take 40,000,000 x 1,024 x 4 bytes = 163.8 GB. The HNSW bottom layer with 32 neighbour links of 4 bytes each adds roughly 128 bytes per vector, about 5.1 GB, plus upper layers and metadata. Keeping everything in RAM therefore needs around 175 GB before replication; with two replicas, about 350 GB.

Quantization changes the picture. Scalar int8 quantization stores one byte per dimension: 41 GB for the vectors, usually with a small recall loss that rescoring against full-precision vectors on disk recovers. Binary quantization stores one bit per dimension: about 5.1 GB, very fast, and acceptable only for embedding models that tolerate it, always with oversampling and rescoring. Qdrant and Weaviate both offer scalar, product and binary quantization. On Pinecone serverless the same arithmetic becomes a cost estimate rather than a capacity plan. On single-node Chroma, 175 GB of in-memory index is past what one machine should hold.

A benchmark harness you can run yourself

Build ground truth from your own data with exact search, then measure each product against it.

import numpy as np, time

def ground_truth(base, queries, masks, k=10):
    """Exact top-k by cosine over only the vectors each query's filter allows."""
    base_n = base / np.linalg.norm(base, axis=1, keepdims=True)
    truth = []
    for q, allowed in zip(queries, masks):
        ids = np.flatnonzero(allowed)
        sims = base_n[ids] @ (q / np.linalg.norm(q))
        truth.append(set(ids[np.argsort(-sims)[:k]]))
    return truth

def evaluate(search_fn, queries, filters, truth, k=10):
    """search_fn(query, filter, k) -> list of ids; wraps any of the four clients."""
    recalls, lat = [], []
    for q, f, t in zip(queries, filters, truth):
        t0 = time.perf_counter()
        got = search_fn(q, f, k)
        lat.append(time.perf_counter() - t0)
        recalls.append(len(set(got) & t) / max(1, len(t)))
    lat = np.array(lat) * 1000
    return {"recall@k": float(np.mean(recalls)),
            "p50_ms": float(np.percentile(lat, 50)),
            "p99_ms": float(np.percentile(lat, 99))}

# Run evaluate() per selectivity bucket (50%, 5%, 0.5%, 0.05%),
# once idle and once while a writer upserts and deletes in the background.

Record recall, p50 and p99 latency at two or three ef settings, and cost per million queries.

Failure modes seen in production

  • Empty or short result lists under selective filters. Post-filtering behaviour, or an unindexed payload field in Qdrant. Use payload indexes or partitions.
  • Metric mismatch. Inserting unnormalised vectors into a dot-product index built for normalised embeddings, or creating a Chroma collection with the default L2 space for a model trained for cosine.
  • Mixed embedding versions. A model upgrade re-embeds half the corpus. Re-embed into a new collection and switch with an alias.
  • Tenant leakage. One code path forgot the tenant filter.
  • Recall drift after heavy deletes and updates. Tombstones and incremental inserts degrade graph quality. Track recall on a fixed probe set.

Choosing

Pick Pinecone when you want zero operations, accept a managed-only dependency and can live with the service's tuning choices and pricing model. Pick Weaviate when built-in BM25 hybrid search, native multi-tenancy with tenant offloading, and optional vectorizer modules match your application. Pick Qdrant when you need fine control over filtering, quantization, named and sparse vectors,. Pick Chroma for local development, embedded use and modest single-node corpora. If your data already lives in Postgres and the corpus is a few million vectors, pgvector may beat all four on total cost; compare it in the same harness.

Keep a thin retrieval interface in your code and store source text and embedding model version outside the database, so a re-embed or a product move is a batch job.

What to do next

  1. Write down your corpus size, dimension, write rate, tenant count and the three filters you will run most often, with their selectivity.
  2. Compute raw and quantized memory with the arithmetic above to see which deployment shapes are even possible.
  3. Export 100,000 real vectors and 1,000 real queries with filters, and build exact ground truth with the harness.
  4. Load the same data into your two strongest candidates and measure recall, p50 and p99 per selectivity bucket, idle and under writes.
  5. Test the operations you will need on a bad day: delete-then-query, tenant isolation without a filter, backup and restore, and a re-embed into a new collection behind an alias.
  6. Wrap the winner behind a small retrieval interface and record the embedding model version with every vector.
Key takeaway: Pinecone, Weaviate, Qdrant and Chroma share an HNSW-style core and differ in everything around it: managed versus self-hosted, how filters are planned, whether hybrid search is built in, how tenants are partitioned and how replication behaves. Decide with your own data: size the memory, then benchmark filtered recall and tail latency at real selectivities and under write load. Put the tenant boundary and the embedding version in your own code so that changing products later is a migration rather than a rewrite.