Cloudflare Vectorize is a managed vector database that you query from a Cloudflare Worker through a binding, or from anywhere through an HTTP API. You put embeddings in with an id and some metadata, and you ask for the nearest neighbours of a query vector, optionally filtered. That is the whole surface, and it is small on purpose. The design decisions that matter are the ones Vectorize makes for you: index configuration is immutable, writes are asynchronous, metadata filtering only works on properties you indexed in advance, and result sizes are capped.

This article explains how those choices shape a real application, builds a complete multi-tenant retrieval Worker, and lists the failure modes teams hit in their first month. API names, limits and behaviour were checked against Cloudflare's Vectorize documentation on 2026-10-02 and describe the current (V2) indexes; limits change, so re-check the limits page before you design around a number. Pricing is deliberately left out.

Advertisement

What an index is, and what it fixes forever

An index is a named collection of vectors of one dimensionality, compared with one distance metric. You choose both at creation, with --dimensions and --metric (cosine, euclidean or dot-product), and neither can be changed afterwards. The dimensionality must match your embedding model exactly: 768 for Workers AI's @cf/baai/bge-base-en-v1.5, for example.

The practical consequence is that an index belongs to an embedding model. Changing models, or even changing how you normalise vectors, means a new index and a full re-embed. Name indexes after the model (docs-bge-768) so the coupling is visible, and plan the migration path before you need it. Vectors from two different models are not comparable even when they have the same length.

The metric decides what a score means. For cosine, a higher score is more similar; Cloudflare's own example returns about 0.897 for a close match. For Euclidean, a smaller distance is closer. Check the ordering in a small test against your own data before you write threshold logic, and never reuse a threshold across models.

The architecture at a glance

Vectorize: a synchronous query path and an asynchronous write pathClientHTTP requestWorkeryour code at the edgeWorkers AIembedding modelVectorize bindingquery / upsert / deleteWrite-ahead logdurable on writeAsync index jobbatches mutationsIndex files in R2read by queriesD1 / R2 / KVsource text, truthtextvectorupsert returns mutationIdnew indexquery readsids to rowsWrites are durable at once but visible to queries only after an index job runs.Keep the authoritative copy of each chunk elsewhere; the index stores vectors and small metadata.
Queries read the current index files. Writes land in a write-ahead log and become visible when an asynchronous job rebuilds the index from R2.

Cloudflare documents the write path plainly. A mutation is written immediately to a write-ahead log, so it is durable. To make it visible to queries, an asynchronous job reads the current index files from R2, builds an updated index and writes it back. Jobs batch work, up to 200,000 vectors or 1,000 individual updates per job, which means many small writes are much slower to become visible than a few large ones. Cloudflare's example: 250,000 vectors written one at a time need about 250 jobs; the same vectors in 100 requests need two or three, and are queryable within minutes.

Reads are synchronous: a query from a Worker goes through the binding and returns matches in the request. The index is not your system of record. Keep the chunk text and document state in D1, R2 or KV and treat Vectorize as a derived structure you can rebuild, as the R2 article and the Durable Objects article describe for the storage side.

Advertisement

The data model and its limits

A vector has an id, its values (32-bit floats), an optional namespace and optional metadata, a JSON object whose keys cannot contain dots or double quotes or start with a dollar sign. The limits below shape the design more than any API call.

LimitValue
Dimensions per vector1,536 (float32)
Vectors per index20,000,000
Vector id length64 bytes
Metadata per vector10 KiB
Metadata indexes per index10, indexing up to 64 bytes per value
Namespaces per index50,000 on Workers Paid, 1,000 on Free
Indexes per account50,000 on Workers Paid, 100 on Free
topK100 without values or metadata, 50 with values or all metadata
Upsert batch1,000 from Workers, 5,000 over the HTTP API
Filter sizecompact JSON under 2,048 bytes

Three of these drive design. The 64-byte id is exactly the length of a hex SHA-256, which makes content-derived ids convenient. The 10 KiB metadata cap means chunk text belongs elsewhere; store a pointer. The 20 million vector cap means a large corpus needs several indexes and a fan-out query.

Creating an index and its metadata indexes

Setup is a handful of Wrangler commands plus a binding. The order matters: metadata indexes must exist before you insert, because vectors written earlier are not added to a metadata index created later. If you forget, the fix is to re-upsert every affected vector.

# One index per embedding model: dimensions and metric are fixed at creation.
npx wrangler vectorize create docs-bge-768 --dimensions=768 --metric=cosine

# Create metadata indexes BEFORE inserting; earlier vectors are not back-filled.
npx wrangler vectorize create-metadata-index docs-bge-768 --property-name=product --type=string
npx wrangler vectorize create-metadata-index docs-bge-768 --property-name=updated --type=number

npx wrangler vectorize info docs-bge-768

# wrangler.toml
[[vectorize]]
binding = "DOCS"
index_name = "docs-bge-768"

[ai]
binding = "AI"

Metadata indexes accept string, number and boolean properties. A string index covers only the first 64 bytes of each value, truncated on a UTF-8 boundary, so two long values that share a 64-byte prefix are indistinguishable to a filter. Use short codes (product: "billing"), not titles or URLs, as filter keys.

Writing: insert, upsert and mutation ids

There are two write calls with different conflict rules. insert keeps the first vector written for an id; a second insert of the same id does not replace it. upsert keeps the last. For pipelines that re-embed changed documents, upsert with deterministic ids is the safe default: re-running the job converges to the same state instead of creating duplicates or keeping stale vectors.

Both calls return a mutation id that identifies the asynchronous change. Log it with the document id, so that when someone asks why a page is not searchable yet you can tell whether the write was accepted and is still being indexed. deleteByIds is asynchronous too, so deleted content can still be returned for a short time; filter results against your source of truth if that matters, as the worked example does.

For bulk loads, wrangler vectorize insert accepts an NDJSON file with --file; Cloudflare recommends at most 5,000 vectors per file. Batch writes from Workers into groups of up to 1,000, not one call per chunk: the async job model rewards large batches.

Querying, filters and namespaces

query(vector, options) returns the nearest matches with ids and scores. queryById does the same using a vector already in the index, which is handy for related-item features. The options are topK (default 5), returnValues (default false), returnMetadata (none, indexed or all) and namespace and filter. Ask for indexed metadata rather than all when you only need filterable fields: it is lighter, and all lowers the topK ceiling to 50.

Filters use $eq, $ne, $in, $nin, $lt, $lte, $gt and $gte. An upper bound can be combined with a lower bound on the same property to form a range; other combinations are rejected. Cloudflare notes that range queries over very large datasets can lose some accuracy, so test recall on your own data if you filter by date ranges.

A namespace is a partition label on each vector, and Cloudflare applies the namespace filter before any metadata filter. That makes it the natural tenant boundary: one namespace per customer, with metadata for finer facets inside it.

Worked example: a multi-tenant documentation search Worker

The Worker below ingests chunked documents for a tenant and answers queries. It embeds with Workers AI, stores chunk text in D1 under a SHA-256 id, upserts vectors in batches with the tenant as namespace, and at query time joins matches back to D1. The join both supplies the text and drops any match whose row has been deleted, which hides the asynchronous delete window.

export interface Env { DOCS: Vectorize; AI: Ai; DB: D1Database; }

const MODEL = "@cf/baai/bge-base-en-v1.5";   // 768 dimensions
const BATCH = 1000;                            // Workers upsert limit per call

async function chunkId(docId: string, n: number): Promise<string> {
  const bytes = new TextEncoder().encode(`${docId}#${n}`);
  const hash = await crypto.subtle.digest("SHA-256", bytes);
  return [...new Uint8Array(hash)].map(b => b.toString(16).padStart(2, "0")).join(""); // 64 chars
}

async function ingest(env: Env, tenant: string, doc: {id: string; product: string; chunks: string[]}) {
  const vectors: VectorizeVector[] = [];
  for (let i = 0; i < doc.chunks.length; i += 50) {
    const slice = doc.chunks.slice(i, i + 50);
    const emb = await env.AI.run(MODEL, { text: slice });
    for (let j = 0; j < slice.length; j++) {
      const id = await chunkId(doc.id, i + j);
      await env.DB.prepare("INSERT OR REPLACE INTO chunks (id, doc_id, body) VALUES (?, ?, ?)")
        .bind(id, doc.id, slice[j]).run();
      vectors.push({ id, values: emb.data[j], namespace: tenant,
                     metadata: { product: doc.product, updated: Date.now(), doc: doc.id } });
    }
  }
  const mutations: string[] = [];
  for (let i = 0; i < vectors.length; i += BATCH) {
    const res = await env.DOCS.upsert(vectors.slice(i, i + BATCH));
    mutations.push(res.mutationId);
  }
  return mutations;            // log these; they identify the async index jobs
}

async function search(env: Env, tenant: string, q: string, product?: string) {
  const emb = await env.AI.run(MODEL, { text: [q] });
  const res = await env.DOCS.query(emb.data[0], {
    topK: 8,
    namespace: tenant,
    returnMetadata: "indexed",
    filter: product ? { product: { $eq: product } } : undefined,
  });
  const ids = res.matches.map(m => m.id);
  if (ids.length === 0) return [];
  const rows = await env.DB.prepare(
    `SELECT id, doc_id, body FROM chunks WHERE id IN (${ids.map(() => "?").join(",")})`
  ).bind(...ids).all();
  const byId = new Map(rows.results.map((r: any) => [r.id, r]));
  return res.matches.filter(m => byId.has(m.id)).map(m => ({ score: m.score, ...byId.get(m.id) }));
}

Walk through one request. A support page in the billing product is split into twelve chunks. ingest embeds them in one Workers AI call, writes twelve D1 rows and one upsert of twelve vectors, and returns one mutation id. A few minutes later, after the index job, a query for "refund an annual plan" with product = billing embeds the question, searches only that tenant's namespace with the metadata filter, takes the top eight ids, and fetches their text from D1 for the prompt. If you re-ingest the page after an edit, the same ids are upserted and the old vectors are replaced, not duplicated. For keeping the index fresh as documents change, seeincremental embedding pipelines.

Multi-tenancy: namespace, filter or index

ApproachIsolationLimits to watchUse when
Namespace per tenantLogical; applied before filters50,000 namespaces; 20M vectors sharedMany small and medium tenants
Metadata field per tenantWeakest; a bug in the filter leaks data10 metadata indexes; 64-byte stringsAvoid for tenancy
Index per tenantStrongest; separate config and capacity50,000 indexes per accountFew large tenants, or per-tenant models

A hybrid is common: namespaces for the long tail, dedicated indexes for the handful of tenants that approach the per-index vector cap or need a different embedding model.

Failure modes

  • Read-after-write tests fail. A test that upserts then immediately queries sees nothing. Poll with a timeout in tests, and in production tell users that new content is searchable within minutes.
  • Filters return nothing. The metadata index was created after the vectors were written. Re-upsert the affected vectors.
  • Dimension mismatch. A model upgrade produces 1,024-dimensional vectors for a 768-dimensional index and every write fails. Version indexes by model and switch the binding only after the new index is fully populated.
  • Truncated filter keys. Long string values collide after 64 bytes. Filter on short codes.
  • topK caps. Results are capped at 100, or 50 with values or all metadata. Ask for ids only, then hydrate from your own store.
  • Capacity cliff at 20M vectors. Shard large corpora across indexes by a stable key and merge results by score, only when indexes share a model and metric.

When to choose Vectorize

Vectorize fits well when your application already runs on Workers, the corpus is millions rather than billions of vectors, and minutes of indexing latency are acceptable. It fits less well when you need hybrid lexical-plus-vector ranking in one query, rich boolean filters, or immediate read-after-write. The vector database comparison covers the alternatives.

What to do next

  1. Pick the embedding model, then create an index named after it with matching dimensions and metric.
  2. List every property you will ever filter on, keep it to ten or fewer short values, and create those metadata indexes before the first insert.
  3. Store chunk text and document state in D1 or R2 under deterministic 64-byte ids; store only pointers and filter fields in metadata.
  4. Write with upsert in batches of up to 1,000 and log every mutation id.
  5. Use one namespace per tenant and join results back to your source of truth.
  6. Write a recall test with known query-answer pairs, and rerun it after every model or chunking change.
Key takeaway: Vectorize is a deliberately small vector database: an index is fixed to one model's dimensionality and metric, writes are durable at once but visible only after an asynchronous index job, metadata filters work only on up to ten properties indexed before insertion, and results are capped at 100. Design around that: upsert large batches with deterministic ids, keep text in D1 or R2, use namespaces for tenants, filter on short codes, and treat the index as a rebuildable derivative of your data.