During prefill an LLM turns every prompt token into a key and a value vector per layer. That KV cache is what lets decode run one token at a time instead of re-reading the whole prompt. It is also expensive to rebuild: a long system prompt, a retrieved document set or a multi-turn conversation costs the same prefill FLOPs every time it arrives, unless the KV produced last time is still somewhere you can read it. GPU memory holds minutes of it at best, CPU DRAM holds more, and local NVMe holds terabytes. This article is about that last tier: when reading KV from disk beats recomputing it, how to lay it out and name it so the wrong cache is never loaded, what LMCache and similar systems expose today, and the failure modes that only appear once a disk is involved.
Two neighbouring pages cover the surrounding ground. KV-cache offloading architecture describes the whole HBM-to-DRAM-to-NVMe hierarchy, and NVMe offload maths works through disk bandwidth for training state. Here the disk tier for inference KV is the subject.
How big the KV cache really is
Start from the bytes. For each token the cache stores one key and one value vector per layer per KV head. The size per token is 2 x layers x kv_heads x head_dim x bytes_per_element. Take an 8B model with grouped-query attention: 32 layers, 8 KV heads, head dimension 128, stored in bf16 (2 bytes). That is 2 x 32 x 8 x 128 x 2 = 131,072 bytes, exactly 128 KiB per token. A 256-token chunk is 32 MiB. A 32,768-token document context is 4 GiB. A 70B-class model with 80 layers and the same KV head layout needs 320 KiB per token, so the same document is 10 GiB.
Those numbers explain the tiering. An 80 GB GPU running the 8B model has perhaps 50 GB left for KV, about 400,000 tokens. 500 GB of host DRAM holds about four million tokens; a pair of 7.68 TB NVMe drives holds over a hundred million. Prefixes reused hours apart, such as per-customer system prompts or a conversation resumed after lunch, survive only on disk.
Precision matters directly. Storing KV in FP8 halves every figure above and halves read time, at whatever quality cost FP8 KV has for your model. Whatever format you store, it must match what the attention kernel expects when the KV is loaded back; that becomes part of the cache key below. KV cache sizing walks through the formula for other architectures.
When disk beats recompute
Loading from disk only helps if it is faster than recomputing. Recomputing a prefix of n tokens costs roughly 2 x params x n FLOPs for the dense layers plus an attention term that grows with n squared. Loading costs n x bytes_per_token / disk_bandwidth plus a fixed overhead for lookup, file opens and the host-to-device copy.
Work it through for the 8B model on one modern data-centre GPU. Assume the GPU sustains about 400 TFLOP/s of useful bf16 throughput during prefill, which is a realistic fraction of peak rather than the datasheet figure. The dense part costs 2 x 8e9 = 1.6e10 FLOPs per token, so about 40 microseconds per token. A single PCIe Gen4 NVMe drive reads around 7 GB/s sequentially, so 131,072 bytes take about 19 microseconds. Disk is roughly twice as fast as recompute even before attention, and the attention term makes recompute worse as the prefix grows: at 32k tokens it adds about half again to the FLOP count. The 32k document takes about 0.6 seconds to read from one drive versus about 2 seconds to recompute.
Recompute wins for short prefixes, where fixed lookup and I/O overhead exceeds the compute, and when the disk is slow or shared with checkpoint writes. Larger models favour disk, because FLOPs per token grow with total parameters while KV bytes grow only with layers and KV heads. Measure with the calculator below.
def kv_bytes_per_token(layers, kv_heads, head_dim, elem_bytes=2):
return 2 * layers * kv_heads * head_dim * elem_bytes
def prefill_seconds(params, n, layers, d_model, gpu_flops):
dense = 2 * params * n
attn = 2 * layers * d_model * n * n # causal QK^T and AV, rough
return (dense + attn) / gpu_flops
def disk_seconds(n, bpt, disk_bw, fixed=0.002):
return fixed + n * bpt / disk_bw
bpt = kv_bytes_per_token(32, 8, 128) # 131072 bytes
for n in (64, 256, 2048, 32768):
r = prefill_seconds(8e9, n, 32, 4096, 4e14)
d = disk_seconds(n, bpt, 7e9)
print(n, f"recompute {r:.3f}s disk {d:.3f}s -> {'disk' if d < r else 'recompute'}")
Architecture of a disk tier
The usual design has four parts. A chunker splits the prompt into fixed-size token chunks; LMCache defaults to 256 tokens. A key function names each chunk. An index maps keys to where the bytes live and how big they are. Storage backends hold the bytes: pinned CPU memory first, then local disk, sometimes a remote store after that.
On a new request the engine walks the chunks in order and asks the index for each key. It stops at the first miss, because KV for chunk k is only valid if every earlier token is identical. Hits are read into pinned host buffers and copied into the engine's paged KV blocks; the engine then prefills only the remaining tokens. After prefill, newly computed full chunks are written back asynchronously: to DRAM immediately, and to disk either straight away or when DRAM evicts them. A partial final chunk is usually not stored because it will not match next time.
Prefetching is what makes disk usable in latency terms. If a request is queued behind others, the engine can look up its prefix while it waits and start moving chunks from disk to DRAM, so that by the time the request is scheduled the reads are done. Prefix caching covers the in-GPU version of the same lookup.
Cache keys that cannot lie
The key is the most important correctness decision in the system. A KV chunk is a function of the model weights, the tokenizer, the exact token ids of the whole prefix up to the end of the chunk, the positional encoding scheme, the stored dtype and layout, and, under tensor parallelism, which shard of the heads this worker owns. Loading a chunk computed under any different value of any of these produces fluent, wrong output with no error. That is the worst kind of bug: nothing crashes.
Use a chained hash so a chunk's key covers every token before it, and put a namespace in front that changes whenever anything that affects the bytes changes:
import hashlib, struct
def namespace(model_id, weights_sha, tokenizer_sha, kv_dtype, layout, tp_rank, tp_size):
s = f"{model_id}|{weights_sha}|{tokenizer_sha}|{kv_dtype}|{layout}|{tp_rank}/{tp_size}"
return hashlib.sha256(s.encode()).digest()
def chunk_keys(token_ids, ns, chunk=256):
keys, prev = [], ns
full = len(token_ids) // chunk * chunk # never key a partial chunk
for i in range(0, full, chunk):
h = hashlib.sha256(prev)
h.update(struct.pack(f"<{chunk}I", *token_ids[i:i + chunk]))
prev = h.digest()
keys.append(prev.hex())
return keysTwo details are easy to miss. Hash token ids, not text, because the same text can tokenize differently after a tokenizer update. And include the weights digest, not only the model name, because a fine-tune pushed under the same name silently invalidates every chunk.
Laying bytes out on NVMe
Disks reward large, aligned, sequential reads and punish small random ones. A 32 MiB chunk is a good unit: big enough that per-I/O overhead vanishes, small enough that a cache hit on half a document does not drag in the other half. Store one chunk per file named by its key, or pack chunks into larger segment files with an offset index if your filesystem struggles with millions of files.
Bypass the page cache. Buffered reads copy every byte into kernel memory first, which competes with the DRAM tier you are already managing and evicts it unpredictably. LMCache exposes this as extra_config.use_odirect; with O_DIRECT, buffers and offsets must be aligned to the filesystem block size. Use several I/O threads or an asynchronous interface so the drive sees a queue depth above one; a single synchronous reader rarely reaches rated bandwidth. NVIDIA GPUDirect Storage can move data from NVMe to GPU memory without a host bounce buffer, but it needs supported drivers, filesystems and hardware, so treat it as an optimisation once the plain path works.
Write atomically. Write to a temporary name, fsync if you need the chunk to survive a crash, then rename into place. A reader must never see a half-written file under a valid key. Store a checksum alongside the payload and verify it on read; a torn write or a bad sector should become a cache miss, never corrupt attention state.
A reference implementation
A minimal disk tier with those properties; production systems add asynchronous I/O and pinned buffers.
import os, zlib, time, threading
class DiskKVTier:
def __init__(self, root, max_bytes):
self.root, self.max_bytes = root, max_bytes
self.index = {} # key -> (size, last_used)
self.used = 0
self.lock = threading.Lock()
os.makedirs(root, exist_ok=True)
def _path(self, key):
return os.path.join(self.root, key[:2], key)
def put(self, key, payload: bytes):
if key in self.index:
return
self._make_room(len(payload) + 4)
path = self._path(key)
os.makedirs(os.path.dirname(path), exist_ok=True)
tmp = path + ".tmp"
with open(tmp, "wb") as f:
f.write(zlib.crc32(payload).to_bytes(4, "little"))
f.write(payload)
os.replace(tmp, path) # atomic on POSIX filesystems
with self.lock:
self.index[key] = (len(payload) + 4, time.monotonic())
self.used += len(payload) + 4
def get(self, key):
if key not in self.index:
return None
try:
with open(self._path(key), "rb") as f:
crc, data = int.from_bytes(f.read(4), "little"), f.read()
except FileNotFoundError:
self._drop(key); return None
if zlib.crc32(data) != crc:
self._drop(key); return None # corruption becomes a miss
with self.lock:
self.index[key] = (self.index[key][0], time.monotonic())
return data
def _make_room(self, need):
with self.lock:
victims = sorted(self.index.items(), key=lambda kv: kv[1][1])
for key, _ in victims:
if self.used + need <= self.max_bytes:
break
self._drop(key)
def _drop(self, key):
with self.lock:
size, _ = self.index.pop(key, (0, 0))
self.used -= size
try: os.remove(self._path(key))
except FileNotFoundError: passThe lookup loop on the engine side walks chunk_keys() in order, calls get() until the first None, and hands the hits to the KV allocator. Rebuild the index at startup by scanning the directory, deleting any .tmp leftovers.
Using LMCache with vLLM
You rarely need to write the tier yourself. LMCache plugs into vLLM through the KV connector interface. A configuration that enables CPU and disk tiers looks like this:
# lmcache.yaml
chunk_size: 256
local_cpu: true
max_local_cpu_size: 40.0 # GB of pinned host memory
local_disk: "file:///mnt/nvme0/lmcache/"
max_local_disk_size: 1500.0 # GB
extra_config:
use_odirect: true
disk_io_threads: 8
# launch
LMCACHE_CONFIG_FILE=lmcache.yaml vllm serve meta-llama/Llama-3.1-8B-Instruct \
--kv-transfer-config '{"kv_connector":"LMCacheConnectorV1","kv_role":"kv_both"}'Disk offload is off by default in LMCache, and the same settings exist as LMCACHE_* environment variables. With several drives, local_disk accepts a comma-separated list of paths and shards them by GPU index. SGLang offers a comparable hierarchical cache behind --enable-hierarchical-cache, and NVIDIA Dynamo has its own KV block manager; option names change between releases, so check the documentation for the version you deploy rather than copying flags from older posts.
Failure modes
These are the failure modes that appear only once a disk tier exists.
- Stale or foreign KV. A model update without a namespace change serves old attention state. Symptom: quality drops only on cached prefixes. Test by toggling the cache and diffing greedy outputs on the same prompts; they should match token for token.
- Tail latency from the drive. NVMe garbage collection and contention with checkpoint or log writes can turn a 5 ms read into 200 ms. Put the cache on dedicated drives and time out slow reads by falling back to recompute.
- Write endurance. Every stored chunk is a write. A server writing 2 GB/s of fresh KV writes about 170 TB a day, which exhausts consumer drives in weeks. Only write chunks that have been hit at least once in DRAM, or that belong to prefixes you know repeat, and check the drive's rated drive writes per day.
- Disk full or slow startup. Enforce a byte quota below the filesystem size and rebuild the index lazily.
- Privacy. Cached KV encodes the prompt. Treat the directory as user data: encrypt at rest, scope namespaces per tenant, and delete chunks on deletion requests. A shared cache can also leak through timing, since hits return faster.
Operating it and trade-offs
Track hit rate by tier in tokens, not requests; read latency percentiles; bytes written per day against drive rating; checksum failures; recompute fallbacks; and time to first token split by cache outcome.
| Choice | Gains | Costs |
|---|---|---|
| Disk tier on | Huge capacity, reuse across hours | Endurance, tail latency, more code paths |
| Larger chunks | Fewer I/Os, higher bandwidth | Coarser reuse, more partial-chunk waste |
| FP8 stored KV | Half the bytes and read time | Possible quality loss; format coupling |
| Write on first compute | Maximum hit rate | Writes every prompt, drives wear |
| Write after DRAM hit | Writes only proven reuse | First reuse after eviction misses |
| Per-tenant namespaces | Isolation, clean deletion | Lower hit rate on shared prompts |
Routing completes the picture: a disk cache on one server does nothing for a request routed to another. Use prefix-aware routing so repeat prefixes land where their chunks live. Multi-turn KV discusses routing for returning conversations.
What to do next
- Log prompt prefixes for a day (as hashes) and measure how many tokens repeat, and how far apart in time. If reuse distance is mostly under a few minutes, DRAM is enough.
- Run the break-even calculator with your model, GPU throughput and measured drive bandwidth.
- Enable LMCache with CPU only, confirm greedy outputs are identical with and without the cache, then add the disk path on a dedicated NVMe drive with O_DIRECT.
- Put the weights digest, tokenizer digest, KV dtype and tensor-parallel rank in the cache namespace.
- Add dashboards for token hit rate by tier, read p99, bytes written per day and recompute fallbacks.
- Decide the write policy from endurance arithmetic, and document retention and deletion for the cache directory.