A LoRA adapter is a pair of thin matrices that nudges a frozen base model towards one task, one customer or one tone. Because the base model is frozen, a hundred fine-tunes built on the same base share almost all of their weights. Multi-LoRA serving turns that into money: load the base once, keep many adapters next to it, and let one batch contain requests for different adapters at the same time.
The serving architecture, meaning routing, the adapter registry and the cache tiers, is covered in multi-LoRA serving architecture and multi-LoRA serving for small models. This article goes one level down, to the GPU: the arithmetic each row of a mixed batch runs, why it needs special kernels, what an adapter costs in memory, and how decode cost grows with adapter diversity. By the end you should be able to size a deployment on paper and predict when adding tenants will slow it down.
The computation, unmerged
LoRA replaces a weight update with a low-rank product. For a linear layer with weight W of shape d_in x d_out, the adapter stores A (d_in x r) and B (r x d_out) with a small rank r, usually 8 to 64, and a scale s = alpha / r. The adapted layer computes y = x W + s * (x A) B. The background on how these matrices are trained is in LoRA and QLoRA.
With one adapter you can merge it: compute W' = W + s A B once and serve W' as an ordinary model. That costs nothing at inference time, but it creates a full copy of every adapted weight, which is exactly what multi-tenant serving is trying to avoid. So multi-LoRA serving keeps the adapter unmerged and computes the two products at run time: a shrink v = x A that projects the activation down to r numbers, and an expand s * v B that projects it back up and is added to the base output.
The order matters: (x A) B costs 2 * r * (d_in + d_out) operations per token, while forming A B first would rebuild a full-size matrix.
Why a mixed batch needs special kernels
GPUs are fast at large matrix multiplications and slow at many small ones. The base product x W is the same for every row of the batch, so it stays one big GEMM no matter how many adapters are involved. The trouble is the low-rank path: row 0 wants adapter A, row 1 wants adapter B, row 2 wants A again. There are two naive options, and both are bad.
- Batch by adapter. Only put requests for the same adapter together. With fifty tenants each sending one request, you run fifty batches of size one and lose the batching that makes GPU inference cheap.
- Loop over adapters inside the step. Slice the batch by adapter and launch a small GEMM pair for each group. The work is correct, but each launch is tiny, the launch overhead and poor occupancy dominate, and the step time grows with the number of groups.
The fix is a gathered product. Stack the adapters that are resident on the GPU into tensors indexed by slot, give each row an integer telling it which slot to use, and write one kernel that, for each row, loads the right A or B slice and multiplies. The whole batch runs one shrink launch and one expand launch per adapted layer, whatever the mix of adapters.
import numpy as np
def lora_batched(x, W, A_stack, B_stack, scale, idx):
"""Reference semantics of a gathered multi-LoRA layer.
x: (n, d_in) one row per token in the step
W: (d_in, d_out) shared base weight
A_stack: (slots, d_in, r) resident adapters, padded to the max rank
B_stack: (slots, r, d_out)
scale: (slots,) alpha / r for each adapter
idx: (n,) slot per row, -1 means no adapter
"""
y = x @ W # one shared GEMM
for i in range(x.shape[0]): # a real kernel does this in parallel
s = idx[i]
if s < 0:
continue
v = x[i] @ A_stack[s] # shrink: d_in -> r
y[i] += scale[s] * (v @ B_stack[s]) # expand: r -> d_out
return yThis loop is the specification; a production kernel parallelises it across rows and output columns.
Two shapes: SGMV for prefill, BGMV for decode
The Punica work (Chen et al., 2023) named the two shapes this gather takes. In prefill, a request contributes many consecutive tokens, so the rows that share an adapter form contiguous segments. Its kernel, SGMV (segmented gather matrix-vector multiplication), treats each segment as a small matrix product against its adapter. The segments are short and fat, so the work is closer to a GEMM and uses the tensor cores reasonably well.
In decode, each request contributes exactly one row per step. The batch is a set of single vectors, each pointing at its own adapter. This is the batched gather matrix-vector shape usually called BGMV. Each row does very little arithmetic against the adapter weights it loads, so decode LoRA work is bound by memory bandwidth, just as the base model's decode is.
S-LoRA (Sheng et al., 2023) added the memory side. It keeps adapter weights and KV cache in one paged pool, so that adapters of different ranks and the growing KV cache share GPU memory without fragmenting it. The idea is the same one described in the paged KV cache. Serving engines now ship their own kernels but keep the same split: shrink into rank space, expand back out, each gathering by a per-token index. Details change often, so this article describes the shapes and quotes no kernel names or speed-ups.
Worked example: what one adapter costs in memory
Take Llama-3-8B: hidden size 4096, 32 layers, 32 query heads and 8 key-value heads of dimension 128, and an MLP width of 14336. Suppose the adapter has rank 16 and targets all seven projections in each layer. Each LoRA matrix pair costs r * (d_in + d_out) parameters.
| Projection | d_in x d_out | LoRA parameters at r = 16 |
|---|---|---|
| q_proj | 4096 x 4096 | 131,072 |
| k_proj | 4096 x 1024 | 81,920 |
| v_proj | 4096 x 1024 | 81,920 |
| o_proj | 4096 x 4096 | 131,072 |
| gate_proj | 4096 x 14336 | 294,912 |
| up_proj | 4096 x 14336 | 294,912 |
| down_proj | 14336 x 4096 | 294,912 |
| per layer | 1,310,720 | |
| 32 layers | 41,943,040 (about 84 MB in 16-bit) |
The base model is about 8 billion parameters, about 16 GB in 16-bit, so this adapter is roughly half a percent of the base. An adapter that only targets q_proj and v_proj is 6,815,744 parameters, about 14 MB. Rank scales everything linearly, so a rank-64 all-projection adapter is about 336 MB.
Now size the GPU. On an 80 GB card, with 16 GB of base weights and, say, 10 GB kept for activations and workspace, the rest is shared between KV cache and adapter slots. Sixteen resident rank-16 adapters take about 1.3 GB, which is modest. Sixteen rank-64 adapters take about 5.4 GB, which is thousands of tokens of KV cache you no longer have. The KV cost per token is worked out in vLLM on GPU. Adapter rank is a capacity decision, not only a quality decision.
The decode cost model: distinct adapters, not requests
A decode step is limited by bytes read from GPU memory. Each step reads the whole base model once, whatever the batch size; that sharing is why batching works. It also reads the KV cache of every sequence, and the weights of every adapter used in the step. A good kernel reads each adapter once per step no matter how many rows use it. So the extra traffic depends on how many distinct adapters are in the batch, not on how many requests there are.
BASE_BYTES = 16.06e9 # Llama-3-8B weights, 16-bit
ADAPTER_BYTES = 83.9e6 # rank 16, all seven projections, 16-bit
for distinct in (1, 8, 32, 64):
extra = distinct * ADAPTER_BYTES
print(f"{distinct:>3} distinct adapters: +{extra/1e9:5.2f} GB per step, "
f"{100*extra/BASE_BYTES:5.1f}% of base weight traffic")
# -> 1 distinct adapters: + 0.08 GB per step, 0.5% of base weight traffic
# -> 8 distinct adapters: + 0.67 GB per step, 4.2% of base weight traffic
# -> 32 distinct adapters: + 2.68 GB per step, 16.7% of base weight traffic
# -> 64 distinct adapters: + 5.37 GB per step, 33.4% of base weight trafficThis is a lower bound on the extra work and ignores KV traffic, kernel efficiency and launch overhead. The shape is still the lesson. A batch of 64 requests that all use one adapter costs almost the same as the base model. A batch of 64 requests spread over 64 adapters can add a third to the weight traffic of every step. Prefill is different: it is compute-bound, and the adapter adds about 2 * 41.9M operations per token against about 2 * 8B for the base, again about half a percent.
So routing a tenant's requests to the same replica improves both cache hits and step time, and a cap on distinct adapters per step is a latency control, not only a memory limit.
Running it in vLLM
vLLM exposes the pieces above as server flags. --enable-lora turns the feature on. --lora-modules name=path registers adapters at start-up. --max-loras is the number of adapters that can be served at the same time, which means the number of GPU slots. --max-lora-rank must be at least the highest rank you will load, and the documentation advises setting it to exactly that. Setting it higher than you need makes the slots bigger than the adapters in them. --max-cpu-loras sizes the host-memory cache that sits behind the GPU slots.
# Start a server with two adapters registered and room for eight on the GPU.
vllm serve meta-llama/Meta-Llama-3-8B-Instruct \
--enable-lora \
--lora-modules support=/adapters/support sql=/adapters/sql \
--max-loras 8 --max-lora-rank 16 --max-cpu-loras 32
# A request selects an adapter through the "model" field.
curl http://localhost:8000/v1/completions -H 'Content-Type: application/json' \
-d '{"model": "sql", "prompt": "List the ten largest orders:", "max_tokens": 64}'
# Runtime loading needs VLLM_ALLOW_RUNTIME_LORA_UPDATING=True in the server environment.
curl -X POST http://localhost:8000/v1/load_lora_adapter -H 'Content-Type: application/json' \
-d '{"lora_name": "billing", "lora_path": "/adapters/billing"}'# Offline, the same engine takes a LoRARequest(name, globally unique int id, path).
from vllm import LLM, SamplingParams
from vllm.lora.request import LoRARequest
llm = LLM(model="meta-llama/Meta-Llama-3-8B-Instruct", enable_lora=True,
max_loras=8, max_lora_rank=16)
params = SamplingParams(max_tokens=64)
out = llm.generate(["List the ten largest orders:"], params,
lora_request=LoRARequest("sql", 1, "/adapters/sql"))Treat the slot count as a scheduling constraint. When the step already uses as many adapters as there are slots, a request for another adapter has to wait for a slot to free up, even if there is KV space for it. Flags change between releases, so check vllm serve --help for your version before copying these lines.
Failure modes
- Rank mismatch. An adapter with a rank above
--max-lora-rankfails to load. Standardise ranks per base model, or run a second pool for the high-rank tenants. - Wrong base. An adapter trained on a different base or revision will often load, because the shapes match, and then produce quietly worse output. Record the base model's hash in adapter metadata and check it at registration.
- Cold-load stalls. The first request for an adapter pays for a disk or network read and a copy to the GPU. At p99 this looks like random slow requests. Preload the hot adapters and track the adapter-load rate.
- Slot thrashing. More active tenants than slots means constant eviction and reloading. Add slots or route tenants to fixed replicas.
- Diversity-driven latency. Time per output token rises as the number of distinct adapters per step rises, which is the byte model above. Dashboards that only show request rate will miss it, so export distinct adapters per step.
Choosing a deployment shape
| Shape | GPU memory | Step cost | Use when |
|---|---|---|---|
| Merged model per tenant | a full copy each | base only | a few tenants with steady, high traffic each |
| One base, unmerged adapters | base + slots | base + distinct adapters | many tenants with bursty or long-tail traffic |
| Unmerged, tenant-affine replicas | base + small slot set | close to base | many tenants and strict latency targets |
| Full fine-tunes | a full copy each | base only | when LoRA quality is not enough |
Start with one shared base; pin the heaviest tenants to their own replicas or merged models once their traffic justifies it.
What to do next
- List your adapters with their base-model hash, rank and target modules, and settle on one standard rank per base model.
- Compute each adapter's size with
r * (d_in + d_out)per targeted projection, and budget slots against KV cache on your GPU. - Set
--max-lora-rankto your highest rank and--max-lorasfrom that budget, then load-test with a realistic mix of tenants rather than one adapter. - Plot time per output token against the number of distinct adapters per step, and compare it with the byte model on this page.
- Add tenant-affine routing if the curve bends, and preload the adapters that produce most of your traffic.
- For every new adapter, compare its output with a merged reference before routing real traffic to it.