S-LoRA is a serving system, published by researchers at UC Berkeley and Stanford in late 2023 (arXiv 2311.03285, later at MLSys 2024), whose title states its goal: serve thousands of concurrent LoRA adapters on top of one base model. S-LoRA keeps one copy of the base weights on the GPU, keeps every adapter in host memory, copies in only the adapters that the current batch needs, and computes the low-rank part for each request separately while the base matmul is shared by the whole batch.
This page explains the three ideas that make that work, in the order you would meet them when operating such a system: a memory pool that pages adapters and KV cache together, a scheduler that prefetches adapters and decides who to admit, and a tensor-parallel layout for the adapter matrices. The batched gather kernels themselves (the SGMV and BGMV family) are covered in Multi-LoRA serving in depth; here they appear only where they constrain the memory design. If LoRA itself is new to you, read LoRA fine-tuning on GPUs first.
Merge or keep adapters separate
A LoRA adapter replaces a weight matrix W of shape h by h with W + AB, where A has shape h by r, B has shape r by h, and the rank r is small, typically 8 to 64. For one token with activation x the layer output is xW + (xA)B. There are two ways to serve that.
Merge. Compute W' = W + AB once and serve W' as an ordinary model. Inference has zero extra cost, but you now have one full model per adapter. A batch can only contain requests for the same merged weights, so with many adapters you either keep many full copies in memory or swap weights between batches and leave the GPU idle during each swap.
Unmerged. Keep W shared and compute the small (xA)B term per request. The expensive xW product is batched across every request regardless of adapter, and the extra work is two skinny matmuls. The S-LoRA authors measured both: merging only won when a single adapter was in use, because it avoids the extra low-rank work; from two adapters upward, switching merged weights left the GPU under-utilised and the unmerged approach was faster. Unmerged serving is the starting point, and S-LoRA's contribution is what happens to memory and scheduling once there are thousands of adapters.
The architecture
The components, following a request: it arrives with a prompt and an adapter id and waits in a queue. The scheduler builds the next batch under a memory budget. Any adapter in that batch that is not already resident is copied from host memory into pages from the pool; the prefetcher tries to start that copy while the previous batch is still computing. Each transformer layer then runs the base projection for the whole batch and a gather kernel that, for each request, reads that request's A and B pages wherever they sit in memory. When an adapter has no active requests left, its pages can be reclaimed for KV cache or for another adapter.
S-LoRA was built on LightLLM, a PyTorch and Triton serving system, so the base model path is a normal continuous-batching engine of the kind described in continuous batching on GPUs.
Unified paging: adapters and KV cache in one pool
Two kinds of memory grow and shrink while the server runs. KV cache grows by one entry per token per request and is freed when the request finishes. Adapter weights appear when a request for a new adapter is admitted and can disappear when its last request ends. Ranks differ between adapters, so adapter tensors are not one size. If each kind had its own reserved region, you would have to guess the split in advance; guess too much for adapters and KV cache runs out, too little and adapters thrash.
S-LoRA's answer, called Unified Paging, extends the paged KV cache idea to adapters. The pool is made of pages, and each page holds one vector of length H, the hidden size. A KV cache tensor for a sequence of length S uses S pages. A LoRA weight tensor of rank R uses R pages, because each of A and B is just r vectors of length h. Both kinds come from the same free list, so the split between adapters and KV cache is decided by the workload at each moment rather than by a config file.
Worked example. Take a 7B Llama-shaped model: hidden size 4096, 32 layers, fp16. The KV cache for one token is 2 vectors (K and V) per layer, so 64 vectors of 4096 values, which is 64 x 4096 x 2 bytes = 512 KiB per token. A rank-16 adapter on the q, k, v and o projections has, per layer, 4 projections x 2 matrices x 16 vectors = 128 vectors, so 4096 vectors over 32 layers, which is 32 MiB. One rank-16 adapter therefore costs the same memory as 64 tokens of KV cache. A request with a 2,000-token context costs as much as about 31 such adapters. The same arithmetic says 2,000 adapters would need about 62 GiB, which is why they live in host memory and only the active ones are paged in.
# Sketch of a unified page pool. One page = one H-length vector.
class UnifiedPool:
def __init__(self, num_pages):
self.free = list(range(num_pages))
self.adapter_pages = {} # adapter_id -> list of page ids
self.adapter_refs = {} # adapter_id -> active request count
self.kv_pages = {} # request_id -> list of page ids
def pages_for_adapter(self, rank, layers, projections):
return rank * 2 * projections * layers # A and B per projection
def ensure_adapter(self, aid, rank, layers=32, projections=4):
if aid in self.adapter_pages:
self.adapter_refs[aid] += 1
return True
need = self.pages_for_adapter(rank, layers, projections)
if len(self.free) < need and not self.evict_idle_adapters(need):
return False
self.adapter_pages[aid] = [self.free.pop() for _ in range(need)]
self.adapter_refs[aid] = 1
return True # caller now copies host weights into these pages
def grow_kv(self, rid, layers=32):
need = 2 * layers # K and V vectors for one new token
if len(self.free) < need:
return False # caller must preempt or abort a request
self.kv_pages.setdefault(rid, []).extend(self.free.pop() for _ in range(need))
return True
def evict_idle_adapters(self, need):
for aid in [a for a, n in self.adapter_refs.items() if n == 0]:
self.free.extend(self.adapter_pages.pop(aid))
del self.adapter_refs[aid]
if len(self.free) >= need:
return True
return len(self.free) >= need
Kernels for a batch of mixed ranks
Paging solves fragmentation but scatters each adapter's vectors through memory, and different requests in one batch have different ranks. Off-the-shelf batched GEMM wants contiguous, equally shaped operands, so S-LoRA wrote its own kernels. For prefill, where each request contributes many tokens, it uses a kernel the paper calls MBGMM (multi-size batched gather matrix-matrix multiplication), written in Triton with tiling. For decode, where each request contributes one token, it uses MBGMV (the matrix-vector version), in one Triton version and one derived from an early version of Punica's kernels, extended to handle non-contiguous memory, several ranks in one batch, and finer-grained gathering.
The practical consequence for capacity planning is that the low-rank work scales with the number of tokens and their ranks, not with the number of distinct adapters, while memory scales with the number of distinct resident adapters. The kernel-level cost model is worked through in the multi-LoRA article linked above.
Scheduling: prefetch, clustering and early abort
Three scheduling ideas sit on top of the pool.
Prefetching. While the current decode batch runs, the scheduler looks at the waiting queue, predicts which adapters the next batch will need, and starts copying them from host memory. A 32 MiB adapter over a PCIe link that delivers, say, 20 GB/s in practice takes roughly 1.6 ms, a noticeable share of a decode step, and several new adapters in one batch add up to a whole step. Pinned host memory and a separate copy stream are what make the overlap possible.
Adapter clustering. Preferring to batch requests that share an adapter reduces the number of resident adapters, which frees pages for KV cache and allows a bigger batch.
Early abort. Under overload a first-come-first-served queue serves everyone late. S-LoRA's admission policy estimates which of the most recent requests can still meet the latency objective and serves those in arrival order, dropping requests that would miss it anyway.
def next_batch(queue, pool, slo_ms, now_ms, est_ms, max_batch):
"""One scheduling round: drop hopeless requests, cluster by adapter, admit under memory."""
alive = [r for r in queue if now_ms - r.arrival_ms + est_ms(r) <= slo_ms]
aborted = [r for r in queue if r not in alive] # reply 429/503 immediately
resident = set(pool.adapter_pages)
# Prefer adapters already on the GPU, then larger groups, then oldest request.
alive.sort(key=lambda r: (r.adapter not in resident, -group_size(alive, r.adapter), r.arrival_ms))
batch = []
for r in alive:
if len(batch) == max_batch:
break
if pool.ensure_adapter(r.adapter, r.rank) and pool.grow_kv(r.id):
batch.append(r)
prefetch([r.adapter for r in alive[len(batch):len(batch) + 16]]) # async host-to-device copies
return batch, aborted
Tensor parallelism for adapters
Models too large for one GPU are split with Megatron-style tensor parallelism: in each attention block the q, k and v projections are column-partitioned and the output projection is row-partitioned, so a block needs one all-reduce. The adapters have to be partitioned to match, and the naive choice (partition A and B the same way as the base weight) adds communication of the same size as the base model's.
S-LoRA's layout column-partitions both A and B for the adapters on column-partitioned base weights and uses an all-gather on the small intermediate xA; for the row-partitioned output projection, A is row-partitioned and B column-partitioned. The result is three extra all-gathers for q, k and v and one extra all-reduce for the output projection, on tensors of width r rather than h. The paper gives the added cost as 5(N-1)Br/N against 2(N-1)Bh/N for the base model, where N is the number of GPUs and B the number of tokens. With h = 4096 and r = 16 the ratio is 5 x 16 / (2 x 4096), under 1 percent.
What the paper measured and where the ideas live now
The paper evaluates Llama 7B, 13B, 30B and 70B on A10G (24 GB) and A100 (40 GB and 80 GB) GPUs with up to 2,000 adapters and synthetic traffic in which a few adapters are popular and most are rarely used. Its headline numbers: up to 4x the throughput of vLLM with adapters packed as separate merged models (which ran out of memory beyond a handful of adapters), and up to 30x the throughput of HuggingFace PEFT. On a single A100 80 GB with 2,000 adapters, S-LoRA still sustained several requests per second, a setting in which the baselines could not run at all.
Treat these as results for the paper's workloads and 2023 baselines. Production engines have since added their own multi-LoRA paths. vLLM, for example, serves LoRA adapters with flags such as --enable-lora, --max-loras, --max-lora-rank and --max-cpu-loras; it reserves GPU slots for a fixed number of adapters sized at the maximum rank and keeps more in a CPU cache, rather than paging adapters into the KV pool. That is simpler to operate and wastes memory when ranks vary, which is exactly the trade S-LoRA's unified pool addresses. Measure your own engine version before assuming either behaviour.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| p99 latency spikes when a new customer appears | Adapter copied from host on the critical path | Pinned memory, prefetch from the queue, warm popular adapters at start |
| Throughput drops as adapter count grows | Resident adapters eat pages, batches shrink | Cluster by adapter, cap distinct adapters per batch, lower ranks |
| Preemption storms | KV growth finds the pool full of idle adapters | Evict adapters with zero references before preempting requests |
| Wrong answers for one tenant | Request mapped to the wrong adapter id or stale version | Version adapter ids, log the id per request, canary new versions |
| OOM at startup with many ranks | Static slot sizing at the maximum rank | Group adapters by rank onto different replicas |
Operating guidance
- Track resident adapters and their page share next to KV cache utilisation. When adapters hold more than a few percent of the pool, batch size is what you are paying with.
- Standardise ranks per deployment. Mixed ranks work in S-LoRA's design, but most engines size slots at the maximum rank, so one rank-128 adapter can multiply memory per slot by eight.
- Merge the one adapter that dominates traffic into its own deployment if it carries most requests; the paper's own measurement says merged weights win when only one adapter is in play.
What to do next
- Compute the per-token KV size and per-adapter size for your model with the arithmetic above, and write down how many tokens one adapter is worth.
- Pull a week of request logs and plot requests per adapter; the shape of that curve decides how many adapters need to be resident.
- Load-test your engine's multi-LoRA mode with 1, 10, 100 and 1,000 adapters at your real rank mix and record throughput and p99.
- Add metrics for adapter load time, resident adapter count and requests aborted at admission.
- Decide per adapter whether it should be merged, resident, or swapped, and encode that in routing.
- Read the multi-LoRA kernel article to understand the compute side of the same batch.