Embedding models turn text into fixed-length vectors that power semantic search, retrieval-augmented generation, deduplication and clustering. Serving them looks easy next to serving a chat model: there is no token-by-token decoding, no KV cache to manage, and each request returns a few kilobytes. Yet embedding services fail in their own characteristic ways. GPUs sit half idle computing padding, query latency spikes behind bulk indexing jobs, and the most expensive bug of all produces no error: vectors from a slightly different model configuration land in the same index and quietly degrade retrieval.

This page explains what the GPU actually does for an embedding request, how pooling, prefixes and normalisation define the vector, how token-budget batching recovers throughput, how to configure Hugging Face Text Embeddings Inference (TEI), how to size a re-embedding job from first principles, and how to version embeddings so an index never mixes incompatible vectors.

What the GPU does for an embedding

An embedding request is a single forward pass through an encoder, the equivalent of the prefill phase of a language model with nothing after it. All tokens are processed in parallel, so the work is dense matrix multiplication and the GPU is compute bound once batches are reasonably large. A useful estimate is about 2 x parameters FLOPs per token for the linear layers, plus attention, which grows with the square of sequence length and starts to matter for long chunks. Memory is dominated by weights and by activations for the batch in flight; there is nothing to keep between requests.

Two consequences follow. First, cost is proportional to tokens processed, including padding, so the batching strategy decides efficiency far more than kernel choice. Second, the workload naturally splits in two: short queries arriving one at a time from users who are waiting, and large volumes of document chunks that nobody is waiting for. Serving both through one undifferentiated queue serves both badly.

Architecture: a query path and a document path

Embedding serving: two paths, one model contractSearch queriesshort, latency boundDocument chunkslong, throughput boundPrefix + tokenizequery promptPrefix + tokenizepassage promptSmall batcheswait under 5 msToken-budget batchessorted by lengthGPU encoderforward pass onlyPool + normalisethen truncate or quantiseVector indextagged with contract hashEmbedding cachekey: contract + text hashBoth paths must use the same model revision, pooling, prefixes and normalisation as the index.
Queries and documents take separate paths with different batching policies, then share the encoder, pooling and normalisation that define the vector contract.

The query path optimises tail latency: tiny batches, a wait of a few milliseconds at most, and capacity reserved so bulk work cannot starve it. The document path optimises tokens per second per GPU: large batches packed by token count and sorted by length. In small deployments both paths can be two priority classes in one server; at scale they are usually separate deployments of the same model image so they can be scaled and paused independently.

Pooling, prefixes and normalisation

The encoder outputs one hidden vector per token. Pooling reduces those to one vector. Encoder models trained with a classification token often use CLS pooling, taking the first token. Many sentence-transformer models use mean pooling, averaging token vectors while ignoring padding. Embedding models built on decoder LLMs typically use last-token pooling, because with causal attention only the final token has seen the whole input. The model card and its pooling configuration decide which one is correct; using a different one still produces vectors, just worse ones.

import torch
import torch.nn.functional as F

def pool(hidden, mask, method):
    # hidden: [batch, seq, dim], mask: [batch, seq] with 1 for real tokens
    if method == "cls":
        v = hidden[:, 0]
    elif method == "mean":
        m = mask.unsqueeze(-1).to(hidden.dtype)
        v = (hidden * m).sum(1) / m.sum(1).clamp(min=1e-6)
    elif method == "last":   # assumes right padding
        idx = mask.sum(1) - 1
        v = hidden[torch.arange(hidden.size(0)), idx]
    else:
        raise ValueError(method)
    return F.normalize(v.float(), dim=-1)   # unit length: dot product equals cosine

Two more details belong to the vector's definition. Many retrieval models are trained with instruction prefixes, for example a query: and a passage: prefix, or a task instruction prepended to queries only. Omitting them, or using the query prefix for documents, measurably hurts retrieval without raising any error. And normalisation to unit length lets the index use inner product as cosine similarity; do it once, in one place, and record that you did. Pool and normalise in fp32 even when the encoder runs in fp16, which costs almost nothing.

Token-budget batching

A batch is a rectangle: every sequence is padded to the longest one in the batch, and the GPU computes the padding. Chunk lengths in real corpora vary widely, so batching by count in arrival order wastes a large share of compute. The fix is to batch by token budget and to group similar lengths together.

Padding waste: the same eight inputs batched two waysArrival order, one batch of 8Sorted, two batches of 4useful 1,705 of 4,080 padded tokens: 42%useful 1,705 of 2,400 padded tokens: 71%Blue is real tokens; red is padding the GPU still computes.
Eight inputs from 25 to 510 tokens. Padding to the batch maximum computes 4,080 token slots for 1,705 real tokens; sorting and splitting computes 2,400.
import asyncio

class TokenBudgetBatcher:
    """Collect requests until the padded batch would exceed max_tokens or max_wait passes."""
    def __init__(self, encode_fn, max_tokens=16384, max_wait=0.005):
        self.encode_fn, self.max_tokens, self.max_wait = encode_fn, max_tokens, max_wait
        self.queue = asyncio.Queue()

    async def embed(self, ids):
        fut = asyncio.get_running_loop().create_future()
        await self.queue.put((ids, fut))
        return await fut

    async def run(self):
        while True:
            batch = [await self.queue.get()]
            deadline = asyncio.get_running_loop().time() + self.max_wait
            while True:
                longest = max(len(x[0]) for x in batch)
                timeout = deadline - asyncio.get_running_loop().time()
                if timeout <= 0:
                    break
                try:
                    nxt = await asyncio.wait_for(self.queue.get(), timeout)
                except asyncio.TimeoutError:
                    break
                if max(longest, len(nxt[0])) * (len(batch) + 1) > self.max_tokens:
                    await self.queue.put(nxt)   # defer to the next batch (reorders slightly)
                    break
                batch.append(nxt)
            batch.sort(key=lambda x: len(x[0]))          # neighbours share lengths
            vectors = await asyncio.to_thread(self.encode_fn, [x[0] for x in batch])
            for (_, fut), v in zip(batch, vectors):
                fut.set_result(v)

The budget counts padded tokens, longest times batch size, because that is what the GPU computes and what bounds activation memory. For bulk jobs, sort the whole corpus shard by length before batching; that one line often improves throughput more than any kernel flag. The general family of batching policies is compared in GPU batching strategies.

Serving with Text Embeddings Inference

You rarely need to write the batcher yourself. TEI is an inference server built for embedding, reranking and sequence-classification models, with token-based dynamic batching built in. Its own documentation describes --max-batch-tokens as the critical control, defaulting to 16384, and advises making it as large as possible until the model is compute bound. A deployment looks like this, with the image tag chosen for your GPU architecture from the TEI documentation. This is the query-path deployment; the document path runs the same command with the passage prompt:

# query-path deployment: the default prompt is the query prefix
docker run --gpus all -p 8080:80 -v $PWD/data:/data \
  ghcr.io/huggingface/text-embeddings-inference:<tag-for-your-gpu> \
  --model-id <org/model> \
  --revision <commit-sha> \
  --dtype float16 \
  --pooling mean \
  --default-prompt-name query \
  --max-batch-tokens 65536 \
  --max-client-batch-size 64 \
  --max-concurrent-requests 512

curl -s localhost:8080/embed -H 'Content-Type: application/json' \
  -d '{"inputs": ["why is the sky blue"]}'
FlagDocumented defaultWhat to decide
--revisionnonePin a commit so the model cannot change under the index
--poolingfrom the model pooling configOverride only if the model card says so; cls, mean, splade, last-token
--dtypenonefloat16 or float32 are the documented values
--max-batch-tokens16384Raise until throughput stops improving or memory runs out
--max-client-batch-size32Largest inputs array a client may send per request
--max-concurrent-requests512Lower it to shed load early instead of queueing
--auto-truncatetrue in current docsDecide explicitly; truncation is silent data loss
--default-prompt-namenoneApplies a prefix from the model prompts config to every input

Defaults have changed between releases, so read the --help output of the exact image you deploy, and set the truncation policy explicitly using the syntax that help output shows. Note that a default prompt applies to every input, so a single server with --default-prompt-name query is wrong for documents; either run the two paths as separate deployments or have clients send the prefix explicitly. TEI exposes Prometheus metrics on port 9000 by default. If you serve through ONNX instead, the provider and IOBinding issues in ONNX Runtime in depth apply directly.

Worked example: sizing a re-embed

Suppose you must re-embed a corpus of 50 million chunks averaging 300 tokens with an encoder of about 335 million parameters. The figures below are arithmetic from those assumptions, not benchmarks; measure your own throughput before committing to a plan.

QuantityCalculationResult
Tokens50M x 30015 billion
FLOPs per token2 x 335M0.67 GFLOP
Total compute15e9 x 0.67e9about 1.0e19 FLOP
At 100 TFLOP/s sustained1.0e19 / 1e14about 28 GPU-hours
With 42% padding efficiency28 / 0.42about 66 GPU-hours
With 90% after length sorting28 / 0.90about 31 GPU-hours

The sustained rate is an assumption you replace with a measurement: run a few thousand real chunks at your chosen batch budget and divide tokens by seconds. The ratio between the padded and sorted rows is the robust lesson. On the query path, the arithmetic is different: a 20-token query is so little compute that latency is dominated by tokenisation, network and queueing, so keep the query deployment lightly loaded and watch its p99, as discussed in GPU inference latency.

The embedding contract

A vector is only comparable with vectors produced under the same configuration. Define an embedding contract: model id, revision, pooling, prefixes for each side, maximum length and truncation policy, normalisation, output dimension and any post-processing such as dimension truncation or quantisation. Hash it, store the hash with every vector and on the index, and make the query path refuse to search an index whose hash differs from its own.

import hashlib, json

CONTRACT = {
    "model": "org/model", "revision": "commit-sha", "pooling": "mean",
    "query_prefix": "query: ", "doc_prefix": "passage: ",
    "max_tokens": 512, "truncate": "reject", "normalize": True,
    "dim": 1024, "post": "none",
}
CONTRACT_ID = hashlib.sha256(json.dumps(CONTRACT, sort_keys=True).encode()).hexdigest()[:16]

def search(index, query_vec):
    if index.meta["contract_id"] != CONTRACT_ID:
        raise RuntimeError("query and index were embedded under different contracts")
    return index.search(query_vec)

Changing any field means a new index built in parallel and a cut-over, never in-place mixing. The same hash makes a safe cache key: content hash plus contract id lets you skip unchanged chunks during incremental re-embedding, the pattern described in incremental embedding pipelines. Shrinking stored vectors with Matryoshka truncation or int8 and binary quantisation is a contract field too, because it changes what is comparable.

Failure modes

  • Silent truncation. Chunks longer than the model limit are cut and their tails never become searchable. Count truncations per job, or reject and re-chunk.
  • Prefix mismatch. Documents embedded with the query prefix, or none. Retrieval quality drops with no error. Test with a labelled query set on every deploy.
  • Revision drift. An unpinned model id pulls a new revision on restart, and new vectors join an old index. Pin revisions and enforce the contract hash.
  • Bulk starves queries. An indexing job fills the queue and search p99 jumps. Separate deployments or strict priority with reserved capacity.
  • Tokenizer bottleneck. The GPU waits on CPU tokenisation. Watch GPU utilisation against request rate and add tokeniser workers or CPU cores.
  • Expecting bitwise equality. fp16 results differ slightly across batch sizes and hardware. Compare with cosine tolerance, never exact equality.

Trade-offs

Larger token budgets raise throughput and also raise latency and memory per batch; the query path wants small budgets and the document path wants large ones. Bigger embedding models usually retrieve better, but every chunk in the corpus pays for that quality on every re-embed, so evaluate on your own queries before upgrading. fp16 halves memory against fp32 with small numerical differences that rarely matter for retrieval. Truncating dimensions or quantising vectors cuts index memory and search cost at some recall loss, measured as in quantization for embeddings. Running embeddings on CPU can be reasonable for low query volumes; bulk re-embedding is where GPUs pay off.

What to do next

  1. Write down the embedding contract for your current index and check every field against how vectors were actually produced.
  2. Pin the model revision in the serving config and store the contract hash on the index.
  3. Split query and document traffic into separate deployments or priority classes.
  4. Measure padding efficiency on a sample of real chunks, then sort by length for bulk jobs.
  5. Raise the token budget step by step while watching throughput, p99 latency and GPU memory.
  6. Decide the truncation policy explicitly and count truncated inputs per job.
  7. Build a labelled query set and run a retrieval check after every model or config change.
  8. Size the next full re-embed from a measured tokens-per-second figure, not a datasheet.
Key takeaway: Embedding serving is a forward pass whose cost is set by tokens processed, padding included, so token-budget batching and length sorting matter more than kernels. Separate latency-bound queries from throughput-bound documents, and pin every setting that defines the vector in a hashed contract so an index never mixes incompatible embeddings.