Cloudflare Workers KV is a global key-value store you call from a Worker with a single line: await env.FLAGS.get("flags:prod"). Reads are fast because they are usually served from a cache in the same data centre as the Worker. That speed comes from a specific design, and the design has consequences. KV is eventually consistent, one key accepts at most one write per second, and a change can take a minute or more to reach every location. Applications that fit this shape get near-local reads everywhere for almost no effort. Applications that do not, such as counters, locks or inventory, get bugs that only show up under load and across regions.

This article explains how KV moves data, states the consistency model precisely, walks through the API with working TypeScript, and builds a feature-flag service as a worked example. Limits and behaviours were checked against Cloudflare's KV documentation on 2026-10-03; Cloudflare changes limits from time to time, so re-check the limits page before you design around a number.

Advertisement

What KV is, and what it is not

A KV namespace is a bucket of keys. You bind it to a Worker under a variable name and call methods on that binding. Keys are strings up to 512 bytes; values can be text, JSON, binary or a stream, up to 25 MiB; each key can also carry up to 1,024 bytes of JSON metadata and an expiry.

KV is a read-optimised, globally cached store. It is not a database with transactions, it has no compare-and-set, and it does not promise that a read sees the latest write. Cloudflare's own guidance names the fit: static assets, application configuration, user preferences, allow-lists and deny-lists. It names the misfit too: write-heavy workloads that update the same key tens or hundreds of times per second, for which it points to Durable Objects.

How a read and a write travel

Writes go to a small set of central data stores. Reads go the other way, through layers of cache. A Worker's read first checks the cache in its own location. On a miss, the request falls back to a regional tier, then a central tier, and only then to the central stores. A read served locally is a hot read; one that goes all the way back is a cold read and takes noticeably longer, especially far from the central stores.

Each layer keeps what it fetched for a time-to-live. That is the whole source of KV's speed and of its staleness: a cached value is served without asking anyone whether it is still current.

Workers KV: writes go to central stores, reads are served from tiered cachesWorker in Lisbonenv.FLAGS.get(...)Worker in Tokyoenv.FLAGS.put(...)Local edge cachehot read, cacheTtlRegional tiershared cacheCentral tiershared cacheCentral storessource of truthreadmissmisscold readwrite: lands centrallyEach cache keeps what it fetched, including "key not found", until its TTL expires.A new value is invisible in Lisbon until Lisbon's cached copy ages out: up to 60 s or more.Design consequencemany reads, few writes per key, staleness tolerated: config, flags, allow-lists, rendered fragments
A read walks outwards through local, regional and central caches to the central stores; a write lands centrally and reaches other locations only as their cached copies expire.
Advertisement

The consistency model, precisely

  • Changes may take up to 60 seconds or more to be visible elsewhere, or as long as the cacheTtl a reader chose. Treat 60 seconds as typical, not as a guarantee.
  • Not-found is cached too. If a location read a key before it existed, it caches the miss, so a newly created key can be invisible there for the same delay as an update.
  • Recent readers see changes later. Cloudflare notes that visibility takes longer in locations that recently read the previous version, which are precisely your busiest locations.
  • Read-your-writes is likely but not promised, even in the location that wrote.
  • list() is eventually consistent as well, so a just-written key may be missing from a listing and a just-deleted one may still appear.

Two practical rules follow. Never read a value, change it and write it back expecting to be the only writer: concurrent Workers in different cities will overwrite each other with stale copies. And never use KV to coordinate, for example as a lock or a once-only marker.

Binding a namespace

# Create the namespace; the command prints its id
npx wrangler kv namespace create FLAGS

# wrangler.toml
[[kv_namespaces]]
binding = "FLAGS"
id = "<id printed above>"

# Seed or inspect values from the command line
# --binding picks the namespace; --remote targets the deployed one, not local dev storage
npx wrangler kv key put "flags:prod" '{"checkoutV2":10}' --binding FLAGS --remote
npx wrangler kv key get "flags:prod" --binding FLAGS --remote
npx wrangler kv bulk put seed.json --binding FLAGS --remote

Older tutorials show wrangler kv:namespace with a colon; current Wrangler uses the space-separated form shown here.

The Workers API

export interface Env { FLAGS: KVNamespace; PREFS: KVNamespace; }

// Single read: type "text" (default), "json", "arrayBuffer" or "stream"
const flags = await env.FLAGS.get("flags:prod", { type: "json", cacheTtl: 60 });

// Value plus metadata in one round trip
const { value, metadata } = await env.FLAGS.getWithMetadata("banner:en", "text");

// Bulk read: up to 100 keys, returns a Map; missing keys map to null
const prefs = await env.PREFS.get(["user:17", "user:18", "user:19"], "json");
const p17 = prefs.get("user:17");

// Write with a relative expiry (minimum 60 s) and metadata (max 1,024 bytes)
await env.PREFS.put("session:ab12", JSON.stringify(state), {
  expirationTtl: 3600,
  metadata: { plan: "pro", v: 3 },
});

await env.PREFS.delete("session:ab12");

// Listing: sorted by UTF-8 bytes, up to 1,000 keys per page, paginate with cursor
async function listAll(ns: KVNamespace, prefix: string): Promise<string[]> {
  const names: string[] = [];
  let cursor: string | undefined;
  do {
    const page = await ns.list({ prefix, cursor });
    for (const k of page.keys) names.push(k.name);   // k.metadata, k.expiration if set
    cursor = page.list_complete ? undefined : page.cursor;
  } while (cursor);
  return names;
}

Note the loop condition in listAll. A page can come back with an empty keys array while list_complete is still false, because expired and deleted keys are iterated but not returned. Code that stops at the first empty page silently truncates the listing.

Limits that shape the design

LimitValueDesign consequence
Writes to the same key1 per secondNo hot counters; shard or use Durable Objects
Operations per Worker invocation1,000Use bulk get, which counts as one operation
Keys per bulk get100Batch larger reads in chunks of 100
Key size512 bytesKeep IDs, not content, in keys
Value size25 MiBLarge files belong in object storage
Metadata per key1,024 bytesEnough for small labels returned by list()
cacheTtlminimum 30 s, default 60 sLower bound on how fresh hot reads can be
expirationTtlminimum 60 sNot suitable for very short-lived tokens
Free plan100,000 reads and 1,000 writes per day, 1 GBPrototype only

Designing keys and values

Use prefixes as namespaces within a namespace. Keys like user:17:prefs and tenant:acme:theme let list({ prefix }) enumerate one group, since listing is a sorted prefix scan.

Put small attributes in metadata. list() returns each key's metadata, so a page of 1,000 keys can render an index without 1,000 reads.

Choose the value granularity deliberately. One JSON blob for all flags is one read and one atomic replacement, but every edit rewrites the whole blob and is bound by the one-write-per-second rule. Many small keys allow independent edits but cost more reads; bulk get softens that.

Make content immutable and point at it. Write a new configuration under config:v42 and then update a small pointer key config:current to v42. Versioned keys are never overwritten, so they are never stale; only the pointer is eventually consistent. A reader therefore sees either the whole old configuration or the whole new one, never a mixture, and rolling back is one pointer write.

Tuning cacheTtl

cacheTtl sets how long the reading location may keep a value, with a minimum of 30 seconds and a default of 60. Raising it makes more reads hot and reduces latency and cold-read cost, at the price of staleness: a reader that cached a value for an hour may serve it for an hour after it changed. Values that never change, like the versioned keys above, can use a long TTL safely. Pointer keys and kill switches should stay near the default. You cannot go below 30 seconds, so KV alone cannot give you faster global propagation than that.

Worked example: feature flags at the edge

A storefront wants percentage rollouts and a kill switch, evaluated in a Worker before the request reaches the origin. Flags change a few times a day; requests arrive thousands of times a second worldwide. That is the KV shape: many reads, rare writes, and a minute of staleness acceptable.

type Flags = { checkoutV2: number; killSwitch: boolean };
const DEFAULTS: Flags = { checkoutV2: 0, killSwitch: false };

function bucket(userId: string): number {           // stable 0-99 per user
  let h = 0;
  for (const ch of userId) h = (h * 31 + ch.charCodeAt(0)) >>> 0;
  return h % 100;
}

export default {
  async fetch(req: Request, env: Env): Promise<Response> {
    let flags = DEFAULTS;
    try {
      const version = await env.FLAGS.get("flags:current", { cacheTtl: 60 });
      if (version) {
        flags = (await env.FLAGS.get<Flags>(`flags:${version}`, { type: "json", cacheTtl: 86400 })) ?? DEFAULTS;
      }
    } catch {
      // KV unavailable: fall back to safe defaults rather than failing the request
    }
    const user = req.headers.get("x-user-id") ?? "anon";
    const useV2 = !flags.killSwitch && bucket(user) < flags.checkoutV2;
    const url = new URL(req.url);
    if (useV2) url.pathname = "/v2" + url.pathname;
    return fetch(new Request(url, req));
  },
};

The release process writes flags:v7 with the new percentages, waits a few seconds, then points flags:current at v7. Over the next minute or so locations pick up the change. Because the hash per user is stable, a user does not flip between versions on successive requests, except during that propagation window, when two locations can disagree.

The team writes down the honest limit: the kill switch takes effect within roughly a minute, not instantly. For incidents that need a faster stop, they keep a second path that does not depend on KV propagation: a Durable Object holding the switch, read only on the routes that need it.

Failure modes and how to handle them

  • 429 on rapid writes to one key. More than one write per second to the same key is rate limited. Retry with exponential backoff, as Cloudflare's documentation recommends, and redesign if it happens routinely.
  • Lost updates from read-modify-write. Two Workers increment a stale copy and the second write wins. There is no fix within KV; move the counter to a Durable Object.
  • A new key that stays missing. A cached negative lookup hides a key created after a location checked for it. Write keys before any code path reads them.
  • Truncated listings from stopping on an empty page, as described above.
  • Cold-read latency for keys that are rarely read in a location; warm critical keys or accept the occasional slow request.
  • Operation budget. Loops that call get per item hit the 1,000-operations ceiling; use bulk get.
async function putWithRetry(ns: KVNamespace, key: string, value: string, attempts = 5) {
  let delay = 1000;
  for (let i = 1; ; i++) {
    try { return await ns.put(key, value); }
    catch (err) {
      if (!String(err).includes("429") || i >= attempts) throw err;
      await new Promise(r => setTimeout(r, delay));
      delay *= 2;
    }
  }
}

When to reach for something else

Pick the store by consistency and write pattern, not by habit. Durable Objects give a single-threaded, strongly consistent home for each object: counters, rate limiters, sessions, coordination. R2 holds large objects and files. Vectorize serves embedding similarity search. KV sits in front of all of them as the cheap, fast, global read cache for small, slowly changing values. For how these pieces fit into an edge architecture generally, see edge compute.

What to do next

  1. List every key your application writes and how often; anything written more than about once a second, or read-modified-written, moves out of KV.
  2. Decide the staleness each value can tolerate and set cacheTtl per read accordingly.
  3. Adopt immutable versioned keys plus a pointer for configuration, so updates and rollbacks are all-or-nothing.
  4. Replace per-item reads in loops with bulk get in chunks of 100.
  5. Fix every list() loop to continue until list_complete is true.
  6. Wrap writes in retry with backoff and alert on sustained 429s.
  7. Give every KV read a safe default so a KV error degrades a feature instead of failing the request.
Key takeaway: Workers KV writes to central stores and serves reads from tiered caches, so reads are fast everywhere and changes take up to a minute or more to arrive, including cached not-found results. Use it for small, read-heavy, slowly changing data such as configuration, flags and allow-lists; respect the one-write-per-second-per-key limit; use bulk get, metadata and versioned keys with a pointer; paginate list() to completion; and move counters, locks and anything needing strong consistency to Durable Objects.