Cloudflare Vectorize is a managed vector database that you query from a Cloudflare Worker through a binding, or from anywhere through an HTTP API. You put embeddings in with an id and some metadata, and you ask for the nearest neighbours of a query vector, optionally filtered. That is the whole surface, and it is small on purpose. The design decisions that matter are the ones Vectorize makes for you: index configuration is immutable, writes are asynchronous, metadata filtering only works on properties you indexed in advance, and result sizes are capped.
This article explains how those choices shape a real application, builds a complete multi-tenant retrieval Worker, and lists the failure modes teams hit in their first month. API names, limits and behaviour were checked against Cloudflare's Vectorize documentation on 2026-10-02 and describe the current (V2) indexes; limits change, so re-check the limits page before you design around a number. Pricing is deliberately left out.
What an index is, and what it fixes forever
An index is a named collection of vectors of one dimensionality, compared with one distance metric. You choose both at creation, with --dimensions and --metric (cosine, euclidean or dot-product), and neither can be changed afterwards. The dimensionality must match your embedding model exactly: 768 for Workers AI's @cf/baai/bge-base-en-v1.5, for example.
The practical consequence is that an index belongs to an embedding model. Changing models, or even changing how you normalise vectors, means a new index and a full re-embed. Name indexes after the model (docs-bge-768) so the coupling is visible, and plan the migration path before you need it. Vectors from two different models are not comparable even when they have the same length.
The metric decides what a score means. For cosine, a higher score is more similar; Cloudflare's own example returns about 0.897 for a close match. For Euclidean, a smaller distance is closer. Check the ordering in a small test against your own data before you write threshold logic, and never reuse a threshold across models.
The architecture at a glance
Cloudflare documents the write path plainly. A mutation is written immediately to a write-ahead log, so it is durable. To make it visible to queries, an asynchronous job reads the current index files from R2, builds an updated index and writes it back. Jobs batch work, up to 200,000 vectors or 1,000 individual updates per job, which means many small writes are much slower to become visible than a few large ones. Cloudflare's example: 250,000 vectors written one at a time need about 250 jobs; the same vectors in 100 requests need two or three, and are queryable within minutes.
Reads are synchronous: a query from a Worker goes through the binding and returns matches in the request. The index is not your system of record. Keep the chunk text and document state in D1, R2 or KV and treat Vectorize as a derived structure you can rebuild, as the R2 article and the Durable Objects article describe for the storage side.
The data model and its limits
A vector has an id, its values (32-bit floats), an optional namespace and optional metadata, a JSON object whose keys cannot contain dots or double quotes or start with a dollar sign. The limits below shape the design more than any API call.
| Limit | Value |
|---|---|
| Dimensions per vector | 1,536 (float32) |
| Vectors per index | 20,000,000 |
| Vector id length | 64 bytes |
| Metadata per vector | 10 KiB |
| Metadata indexes per index | 10, indexing up to 64 bytes per value |
| Namespaces per index | 50,000 on Workers Paid, 1,000 on Free |
| Indexes per account | 50,000 on Workers Paid, 100 on Free |
| topK | 100 without values or metadata, 50 with values or all metadata |
| Upsert batch | 1,000 from Workers, 5,000 over the HTTP API |
| Filter size | compact JSON under 2,048 bytes |
Three of these drive design. The 64-byte id is exactly the length of a hex SHA-256, which makes content-derived ids convenient. The 10 KiB metadata cap means chunk text belongs elsewhere; store a pointer. The 20 million vector cap means a large corpus needs several indexes and a fan-out query.
Creating an index and its metadata indexes
Setup is a handful of Wrangler commands plus a binding. The order matters: metadata indexes must exist before you insert, because vectors written earlier are not added to a metadata index created later. If you forget, the fix is to re-upsert every affected vector.
# One index per embedding model: dimensions and metric are fixed at creation.
npx wrangler vectorize create docs-bge-768 --dimensions=768 --metric=cosine
# Create metadata indexes BEFORE inserting; earlier vectors are not back-filled.
npx wrangler vectorize create-metadata-index docs-bge-768 --property-name=product --type=string
npx wrangler vectorize create-metadata-index docs-bge-768 --property-name=updated --type=number
npx wrangler vectorize info docs-bge-768
# wrangler.toml
[[vectorize]]
binding = "DOCS"
index_name = "docs-bge-768"
[ai]
binding = "AI"Metadata indexes accept string, number and boolean properties. A string index covers only the first 64 bytes of each value, truncated on a UTF-8 boundary, so two long values that share a 64-byte prefix are indistinguishable to a filter. Use short codes (product: "billing"), not titles or URLs, as filter keys.
Writing: insert, upsert and mutation ids
There are two write calls with different conflict rules. insert keeps the first vector written for an id; a second insert of the same id does not replace it. upsert keeps the last. For pipelines that re-embed changed documents, upsert with deterministic ids is the safe default: re-running the job converges to the same state instead of creating duplicates or keeping stale vectors.
Both calls return a mutation id that identifies the asynchronous change. Log it with the document id, so that when someone asks why a page is not searchable yet you can tell whether the write was accepted and is still being indexed. deleteByIds is asynchronous too, so deleted content can still be returned for a short time; filter results against your source of truth if that matters, as the worked example does.
For bulk loads, wrangler vectorize insert accepts an NDJSON file with --file; Cloudflare recommends at most 5,000 vectors per file. Batch writes from Workers into groups of up to 1,000, not one call per chunk: the async job model rewards large batches.
Querying, filters and namespaces
query(vector, options) returns the nearest matches with ids and scores. queryById does the same using a vector already in the index, which is handy for related-item features. The options are topK (default 5), returnValues (default false), returnMetadata (none, indexed or all) and namespace and filter. Ask for indexed metadata rather than all when you only need filterable fields: it is lighter, and all lowers the topK ceiling to 50.
Filters use $eq, $ne, $in, $nin, $lt, $lte, $gt and $gte. An upper bound can be combined with a lower bound on the same property to form a range; other combinations are rejected. Cloudflare notes that range queries over very large datasets can lose some accuracy, so test recall on your own data if you filter by date ranges.
A namespace is a partition label on each vector, and Cloudflare applies the namespace filter before any metadata filter. That makes it the natural tenant boundary: one namespace per customer, with metadata for finer facets inside it.
Worked example: a multi-tenant documentation search Worker
The Worker below ingests chunked documents for a tenant and answers queries. It embeds with Workers AI, stores chunk text in D1 under a SHA-256 id, upserts vectors in batches with the tenant as namespace, and at query time joins matches back to D1. The join both supplies the text and drops any match whose row has been deleted, which hides the asynchronous delete window.
export interface Env { DOCS: Vectorize; AI: Ai; DB: D1Database; }
const MODEL = "@cf/baai/bge-base-en-v1.5"; // 768 dimensions
const BATCH = 1000; // Workers upsert limit per call
async function chunkId(docId: string, n: number): Promise<string> {
const bytes = new TextEncoder().encode(`${docId}#${n}`);
const hash = await crypto.subtle.digest("SHA-256", bytes);
return [...new Uint8Array(hash)].map(b => b.toString(16).padStart(2, "0")).join(""); // 64 chars
}
async function ingest(env: Env, tenant: string, doc: {id: string; product: string; chunks: string[]}) {
const vectors: VectorizeVector[] = [];
for (let i = 0; i < doc.chunks.length; i += 50) {
const slice = doc.chunks.slice(i, i + 50);
const emb = await env.AI.run(MODEL, { text: slice });
for (let j = 0; j < slice.length; j++) {
const id = await chunkId(doc.id, i + j);
await env.DB.prepare("INSERT OR REPLACE INTO chunks (id, doc_id, body) VALUES (?, ?, ?)")
.bind(id, doc.id, slice[j]).run();
vectors.push({ id, values: emb.data[j], namespace: tenant,
metadata: { product: doc.product, updated: Date.now(), doc: doc.id } });
}
}
const mutations: string[] = [];
for (let i = 0; i < vectors.length; i += BATCH) {
const res = await env.DOCS.upsert(vectors.slice(i, i + BATCH));
mutations.push(res.mutationId);
}
return mutations; // log these; they identify the async index jobs
}
async function search(env: Env, tenant: string, q: string, product?: string) {
const emb = await env.AI.run(MODEL, { text: [q] });
const res = await env.DOCS.query(emb.data[0], {
topK: 8,
namespace: tenant,
returnMetadata: "indexed",
filter: product ? { product: { $eq: product } } : undefined,
});
const ids = res.matches.map(m => m.id);
if (ids.length === 0) return [];
const rows = await env.DB.prepare(
`SELECT id, doc_id, body FROM chunks WHERE id IN (${ids.map(() => "?").join(",")})`
).bind(...ids).all();
const byId = new Map(rows.results.map((r: any) => [r.id, r]));
return res.matches.filter(m => byId.has(m.id)).map(m => ({ score: m.score, ...byId.get(m.id) }));
}Walk through one request. A support page in the billing product is split into twelve chunks. ingest embeds them in one Workers AI call, writes twelve D1 rows and one upsert of twelve vectors, and returns one mutation id. A few minutes later, after the index job, a query for "refund an annual plan" with product = billing embeds the question, searches only that tenant's namespace with the metadata filter, takes the top eight ids, and fetches their text from D1 for the prompt. If you re-ingest the page after an edit, the same ids are upserted and the old vectors are replaced, not duplicated. For keeping the index fresh as documents change, seeincremental embedding pipelines.
Multi-tenancy: namespace, filter or index
| Approach | Isolation | Limits to watch | Use when |
|---|---|---|---|
| Namespace per tenant | Logical; applied before filters | 50,000 namespaces; 20M vectors shared | Many small and medium tenants |
| Metadata field per tenant | Weakest; a bug in the filter leaks data | 10 metadata indexes; 64-byte strings | Avoid for tenancy |
| Index per tenant | Strongest; separate config and capacity | 50,000 indexes per account | Few large tenants, or per-tenant models |
A hybrid is common: namespaces for the long tail, dedicated indexes for the handful of tenants that approach the per-index vector cap or need a different embedding model.
Failure modes
- Read-after-write tests fail. A test that upserts then immediately queries sees nothing. Poll with a timeout in tests, and in production tell users that new content is searchable within minutes.
- Filters return nothing. The metadata index was created after the vectors were written. Re-upsert the affected vectors.
- Dimension mismatch. A model upgrade produces 1,024-dimensional vectors for a 768-dimensional index and every write fails. Version indexes by model and switch the binding only after the new index is fully populated.
- Truncated filter keys. Long string values collide after 64 bytes. Filter on short codes.
- topK caps. Results are capped at 100, or 50 with values or all metadata. Ask for ids only, then hydrate from your own store.
- Capacity cliff at 20M vectors. Shard large corpora across indexes by a stable key and merge results by score, only when indexes share a model and metric.
When to choose Vectorize
Vectorize fits well when your application already runs on Workers, the corpus is millions rather than billions of vectors, and minutes of indexing latency are acceptable. It fits less well when you need hybrid lexical-plus-vector ranking in one query, rich boolean filters, or immediate read-after-write. The vector database comparison covers the alternatives.
What to do next
- Pick the embedding model, then create an index named after it with matching dimensions and metric.
- List every property you will ever filter on, keep it to ten or fewer short values, and create those metadata indexes before the first insert.
- Store chunk text and document state in D1 or R2 under deterministic 64-byte ids; store only pointers and filter fields in metadata.
- Write with upsert in batches of up to 1,000 and log every mutation id.
- Use one namespace per tenant and join results back to your source of truth.
- Write a recall test with known query-answer pairs, and rerun it after every model or chunking change.