An embedding index is mostly a pile of floating-point numbers. One hundred million documents embedded at 1,024 dimensions in float32 occupy 409.6 GB before any index structure is added, and that memory is what you pay for every month. Quantization stores the same vectors in 2, 1 or even one-eighth of a byte per dimension. Done carelessly, it quietly drops relevant documents from the results. Done well, it cuts memory by 4 to 32 times while the final ranking barely changes.

This article explains the options from first principles, shows the code for each, and gives you a harness to measure exactly what your data loses. It is engine-neutral; if you run Postgres, the pgvector at 100M scale guide shows the same ideas with its halfvec and bit types.

Advertisement

Why embeddings are a different quantization problem

Quantizing model weights, covered in INT8 quantization in depth, aims to keep each layer's output close to the original so that the model's next-token predictions do not change. Quantizing stored embeddings has a narrower goal: keep the order of similarity scores. If document A scored higher than document B against a query before quantization, it should still score higher afterwards, at least near the top of the list. Absolute score values can drift as long as the ranking survives.

That goal has two useful consequences. First, errors that are the same for every document cancel out, since only comparisons matter. Second, you can tolerate a coarse first pass if a second pass re-ranks a short list with better precision. Those two facts are why one-bit codes work at all.

There are also two separate things you can quantize: the stored vectors, which decides memory and scan cost, and the embedding model, which decides how fast you produce vectors. They have different risks and are tested differently, so treat them as separate decisions.

The representations and what they cost

FormatBytes for 1,024 dimsCompressionSimilarity computed asTypical role
float324,0961xfloat dot productReference and rescoring
float16 or bfloat162,0482xfloat dot productSafe default for storage
int8 scalar1,0244xinteger dot product plus scaleMain index or rescoring tier
binary (1 bit)12832xHamming distance via popcountFirst-pass candidate search
product quantizationfor example 64about 64xtable lookups per sub-vectorVery large indexes

Float16 almost never changes rankings for normalized embeddings and should be the first thing you try. Product quantization splits a vector into sub-vectors and replaces each with the id of its nearest centroid; it is covered in product quantization maths. The rest of this article focuses on int8 and binary, which need no training and fit any engine.

Before any of these, normalize vectors to unit length if your similarity is cosine. Quantization ranges are then comparable across documents, and dot product equals cosine similarity, which keeps the later arithmetic simple.

Advertisement

Scalar int8: calibrate ranges, then map to 256 levels

Scalar quantization maps each dimension's values from a range [lo, hi] onto 256 integer levels. The range is calibrated per dimension on a sample of real document vectors. Using the minimum and maximum lets one outlier stretch the range and waste most of the levels on values that never occur, so use percentiles, such as the 0.1st and 99.9th, and clip values outside them.

Queries do not need to be quantized. In the asymmetric form, the query stays in float and the score is computed against integer codes by folding the scale into the query, so each document costs one integer-by-float dot product plus a constant.

import numpy as np

def calibrate(sample, lo_pct=0.1, hi_pct=99.9):
    lo = np.percentile(sample, lo_pct, axis=0)
    hi = np.percentile(sample, hi_pct, axis=0)
    step = np.maximum(hi - lo, 1e-9) / 255.0
    return lo.astype(np.float32), step.astype(np.float32)

def to_int8(x, lo, step):
    q = np.rint((np.clip(x, lo, lo + 255 * step) - lo) / step) - 128
    return q.astype(np.int8)

def asym_scores(query, codes, lo, step):
    # x is approximately lo + (code + 128) * step, so q.x = q.lo + 128 * sum(q*step) + (q*step).code
    qs = query * step
    const = float(query @ lo + 128.0 * qs.sum())
    return const + codes.astype(np.float32) @ qs

Recalibrate when you change the embedding model, and check periodically that the share of clipped values in new documents has not grown. A rising clip rate means the data has drifted away from the calibration sample.

Binary codes: one bit per dimension

Binary quantization keeps only the sign of each dimension: 1 if the value is positive, 0 otherwise. A 1,024-dimension vector becomes 128 bytes. The distance between two codes is the Hamming distance, the number of differing bits, computed with XOR and a population count instruction that modern CPUs execute on 64 bits at once. A full scan over millions of codes is therefore fast even without an index.

Why does a sign keep any meaning? When dimensions carry roughly balanced, spread-out information, two vectors that point in similar directions agree in sign on most dimensions, so Hamming distance tracks the angle between them. The approximation is coarse, and it works better with more dimensions and worse with models whose values are lopsided, for example where one dimension is almost always positive. Centering each dimension by subtracting its mean before taking the sign can help such models; measure it rather than assuming.

def to_bits(x, center=None):
    if center is not None:
        x = x - center
    return np.packbits(x > 0, axis=-1)                # (n, dims/8) uint8

POP = np.array([bin(i).count("1") for i in range(256)], dtype=np.uint8)

def hamming(q_bits, doc_bits):
    return POP[np.bitwise_xor(doc_bits, q_bits)].sum(axis=1, dtype=np.int32)

def binary_candidates(query, doc_bits, n, center=None):
    d = hamming(to_bits(query[None, :], center), doc_bits)
    return np.argpartition(d, n)[:n]

Oversample and rescore: the pipeline that makes it work

Two-stage search: cheap codes in memory, better vectors for the shortlistQuery textembed once, floatBinarize querysign of each valueHamming scan or index1 bit per dim, all docs in RAMShortlist k x oversamplefor example 10 x 4 = 40 idsRescore with int8 or floatvectors on SSD or object storeTop k resultsranked by rescored dot productfloat query used for rescoringMemory holds only the compact codes. Precision is spent only on the few candidates that might be returned.
Figure 1. Binary codes find a shortlist cheaply; int8 or float vectors, which can live on slower storage, decide the final order.

Coarse codes are good at finding the neighbourhood and bad at ordering within it. So ask the first stage for more candidates than you need, k times an oversampling factor, then rescore that shortlist with better vectors and keep the top k. The rescoring vectors can sit on SSD or in object storage because only a few dozen are read per query.

There is also a cheap middle step. The dot product of the float query with a binary document code, reading each bit as +1 or -1, equals twice the sum of query values where the bit is 1, minus the sum of all query values. It needs no extra storage and often reorders the shortlist better than Hamming distance alone.

def search(query, doc_bits, int8_codes, lo, step, k=10, oversample=4):
    cand = binary_candidates(query, doc_bits, k * oversample)
    scores = asym_scores(query, int8_codes[cand], lo, step)   # reads only len(cand) vectors
    order = np.argsort(-scores)[:k]
    return cand[order], scores[order]

Measuring what you lost: a recall harness

Published retention figures for binary and int8 embeddings vary widely by model, dimension and dataset, so measure your own. The harness compares each configuration against exact float32 search on a held-out set of real queries. Recall@10 answers the question: of the ten documents exact search returns, how many does the quantized pipeline also return?

def recall_at_k(exact_ids, approx_ids, k=10):
    hits = [len(set(e[:k]) & set(a[:k])) for e, a in zip(exact_ids, approx_ids)]
    return sum(hits) / (k * len(hits))

def evaluate(queries, docs_f32, doc_bits, int8_codes, lo, step, k=10):
    exact = [np.argsort(-(docs_f32 @ q))[:k] for q in queries]
    for m in (1, 2, 4, 8, 16):
        approx = [search(q, doc_bits, int8_codes, lo, step, k, m)[0] for q in queries]
        print(f"oversample {m:2d}: recall@{k} = {recall_at_k(exact, approx, k):.3f}")

Use at least a few hundred real queries, not documents reused as queries, because real queries are shorter and distributed differently. If you have relevance labels, also compute nDCG on them: recall against exact search measures fidelity to float32, while labelled metrics measure whether users get good answers, and a small fidelity loss sometimes costs nothing in relevance.

Worked example: 20 million documents at 768 dimensions

A support-search service holds 20 million passages embedded at 768 dimensions. In float32 that is 20M x 768 x 4 bytes = 61.44 GB, which needs a large memory instance per replica. Int8 brings it to 15.36 GB and binary codes to 20M x 96 bytes = 1.92 GB.

The team keeps binary codes in memory for the first pass and int8 vectors on local NVMe for rescoring. With k = 10 and oversampling of 4, each query reads 40 int8 vectors of 768 bytes, about 30 KB, which local SSD serves in well under a millisecond. They run the harness with oversampling from 1 to 16 and choose the smallest factor whose recall@10 against float32 meets the threshold they agreed with the product owner. They then confirm on their labelled set that nDCG did not drop. Memory per replica falls by roughly thirty times, and the float32 vectors stay in object storage so that any future re-quantization or recalibration does not require re-embedding.

Truncation and quantizing the encoder

Dimension truncation compounds with quantization: models trained with a Matryoshka objective keep useful information in their leading dimensions, so you can keep the first 256 of 1,024 and then binarize. Only do this with models trained for it; truncating an ordinary embedding model degrades it sharply. See Matryoshka embeddings for how the training works.

Quantizing the encoder model to int8 speeds up embedding but changes the vectors slightly. Never mix vectors from the quantized and unquantized encoder in one index without measuring: re-embed your evaluation queries and documents with the new encoder, rerun the recall harness against the old float32 results, and plan a full re-embedding if the drift is material.

Choosing a configuration: the trade-offs

Each step down in precision buys memory and scan speed with recall, and oversampling buys some of that recall back with latency and storage reads. A useful way to decide is to fix the recall threshold first and then search for the cheapest configuration that meets it, rather than picking a format and hoping.

At a few million vectors, float16 in memory is usually the whole answer: the savings from going further are small next to the engineering effort. Between tens and hundreds of millions, int8 in memory is a strong default, and binary-plus-rescore wins when memory cost dominates or when a brute-force scan must stay fast without an index. Beyond that, codes are usually combined with an approximate index rather than scanned, and product quantization inside an inverted-file index, described in IVF-PQ index architecture, becomes the main option.

Remember the operational costs too. Every format change requires re-encoding the whole corpus from the stored originals, every new model requires recalibration, and two-tier storage adds a second system that must stay consistent with the first. Write the chosen configuration, its calibration date and the measured recall into the index metadata so that the next engineer can reproduce the decision.

Failure modes

SymptomCauseFix
Recall collapses after binarizingLopsided dimensions, or too few dimensionsCenter before sign, raise oversampling, or use int8 for the first pass
Int8 recall worse than expectedRanges from min and max with outliersCalibrate with percentiles and clip
Quality drifts over monthsNew content outside calibration rangesTrack clip rate; recalibrate and requantize
Recall good, users unhappyMeasured fidelity, not relevanceAdd a labelled set and nDCG
Latency rises with oversamplingRescoring vectors on slow storageKeep the rescoring tier on local SSD or in memory
Scores inconsistent after upgradeEncoder changed, old vectors keptVersion vectors by model; re-embed on change

What to do next

  1. Normalize your embeddings and switch storage to float16; confirm the ranking is unchanged on a sample.
  2. Collect a few hundred real queries and compute exact float32 top-10 results as your reference.
  3. Calibrate int8 ranges with percentiles and measure recall@10 using asymmetric scoring.
  4. Try binary codes with oversampling 1 to 16 plus int8 rescoring, and plot recall against oversampling.
  5. Pick the cheapest configuration that meets a recall threshold you have written down, then check nDCG on labelled data.
  6. Keep float32 originals in cheap storage and record the model version alongside every vector.
Key takeaway: Embedding quantization only has to preserve ranking, not exact values, which is why float16, int8 and even one-bit codes work. Normalize first, calibrate int8 ranges with percentiles, and use asymmetric scoring so queries stay in float. Binary codes with popcount make a fast first pass, and oversample-and-rescore with int8 or float vectors restores the order. Retention varies by model, so measure recall@k against exact search on real queries, confirm relevance on labelled data, and keep the float originals so you can change your mind.