A LoRA adapter is a pair of thin matrices that nudges a frozen base model towards one task, one customer or one tone. Because the base model is frozen, a hundred fine-tunes built on the same base share almost all of their weights. Multi-LoRA serving turns that into money: load the base once, keep many adapters next to it, and let one batch contain requests for different adapters at the same time.

The serving architecture, meaning routing, the adapter registry and the cache tiers, is covered in multi-LoRA serving architecture and multi-LoRA serving for small models. This article goes one level down, to the GPU: the arithmetic each row of a mixed batch runs, why it needs special kernels, what an adapter costs in memory, and how decode cost grows with adapter diversity. By the end you should be able to size a deployment on paper and predict when adding tenants will slow it down.

Advertisement

The computation, unmerged

LoRA replaces a weight update with a low-rank product. For a linear layer with weight W of shape d_in x d_out, the adapter stores A (d_in x r) and B (r x d_out) with a small rank r, usually 8 to 64, and a scale s = alpha / r. The adapted layer computes y = x W + s * (x A) B. The background on how these matrices are trained is in LoRA and QLoRA.

With one adapter you can merge it: compute W' = W + s A B once and serve W' as an ordinary model. That costs nothing at inference time, but it creates a full copy of every adapted weight, which is exactly what multi-tenant serving is trying to avoid. So multi-LoRA serving keeps the adapter unmerged and computes the two products at run time: a shrink v = x A that projects the activation down to r numbers, and an expand s * v B that projects it back up and is added to the base output.

The order matters: (x A) B costs 2 * r * (d_in + d_out) operations per token, while forming A B first would rebuild a full-size matrix.

Why a mixed batch needs special kernels

GPUs are fast at large matrix multiplications and slow at many small ones. The base product x W is the same for every row of the batch, so it stays one big GEMM no matter how many adapters are involved. The trouble is the low-rank path: row 0 wants adapter A, row 1 wants adapter B, row 2 wants A again. There are two naive options, and both are bad.

  • Batch by adapter. Only put requests for the same adapter together. With fifty tenants each sending one request, you run fifty batches of size one and lose the batching that makes GPU inference cheap.
  • Loop over adapters inside the step. Slice the batch by adapter and launch a small GEMM pair for each group. The work is correct, but each launch is tiny, the launch overhead and poor occupancy dominate, and the step time grows with the number of groups.

The fix is a gathered product. Stack the adapters that are resident on the GPU into tensors indexed by slot, give each row an integer telling it which slot to use, and write one kernel that, for each row, loads the right A or B slice and multiplies. The whole batch runs one shrink launch and one expand launch per adapted layer, whatever the mix of adapters.

import numpy as np

def lora_batched(x, W, A_stack, B_stack, scale, idx):
    """Reference semantics of a gathered multi-LoRA layer.

    x:        (n, d_in)            one row per token in the step
    W:        (d_in, d_out)        shared base weight
    A_stack:  (slots, d_in, r)     resident adapters, padded to the max rank
    B_stack:  (slots, r, d_out)
    scale:    (slots,)             alpha / r for each adapter
    idx:      (n,)                 slot per row, -1 means no adapter
    """
    y = x @ W                                   # one shared GEMM
    for i in range(x.shape[0]):                 # a real kernel does this in parallel
        s = idx[i]
        if s < 0:
            continue
        v = x[i] @ A_stack[s]                   # shrink: d_in -> r
        y[i] += scale[s] * (v @ B_stack[s])     # expand: r -> d_out
    return y

This loop is the specification; a production kernel parallelises it across rows and output columns.

One decode step, four requests, three adapters: the base product is shared, the low-rank path is gathered per rowbatch rows x (4 x 4096)row 0 -> adapter Arow 1 -> adapter Brow 2 -> adapter Arow 3 -> adapter Cbase GEMM: x Wone product for all rows, W read onceshrink: v = x A[idx]gather A per row, 4096 -> rexpand: s * v B[idx]gather B per row, r -> 4096addy = x W + s (x A) Bper row, its own A, B and sGPU memorybase weights | KV cache pages | adapter slots (--max-loras)host memory and diskCPU adapter cache (--max-cpu-loras) | adapter filesPrefill: rows sharing an adapter form contiguous segments (SGMV shape).Decode: one row per request, each with its own index (BGMV shape).Extra bytes per decode step grow with the number of DISTINCT adapters, not the number of requests.
A gathered multi-LoRA layer. The base product runs once for the whole batch; the shrink and expand products read each row's own adapter from the slot tensors.
Advertisement

Two shapes: SGMV for prefill, BGMV for decode

The Punica work (Chen et al., 2023) named the two shapes this gather takes. In prefill, a request contributes many consecutive tokens, so the rows that share an adapter form contiguous segments. Its kernel, SGMV (segmented gather matrix-vector multiplication), treats each segment as a small matrix product against its adapter. The segments are short and fat, so the work is closer to a GEMM and uses the tensor cores reasonably well.

In decode, each request contributes exactly one row per step. The batch is a set of single vectors, each pointing at its own adapter. This is the batched gather matrix-vector shape usually called BGMV. Each row does very little arithmetic against the adapter weights it loads, so decode LoRA work is bound by memory bandwidth, just as the base model's decode is.

S-LoRA (Sheng et al., 2023) added the memory side. It keeps adapter weights and KV cache in one paged pool, so that adapters of different ranks and the growing KV cache share GPU memory without fragmenting it. The idea is the same one described in the paged KV cache. Serving engines now ship their own kernels but keep the same split: shrink into rank space, expand back out, each gathering by a per-token index. Details change often, so this article describes the shapes and quotes no kernel names or speed-ups.

Worked example: what one adapter costs in memory

Take Llama-3-8B: hidden size 4096, 32 layers, 32 query heads and 8 key-value heads of dimension 128, and an MLP width of 14336. Suppose the adapter has rank 16 and targets all seven projections in each layer. Each LoRA matrix pair costs r * (d_in + d_out) parameters.

Projectiond_in x d_outLoRA parameters at r = 16
q_proj4096 x 4096131,072
k_proj4096 x 102481,920
v_proj4096 x 102481,920
o_proj4096 x 4096131,072
gate_proj4096 x 14336294,912
up_proj4096 x 14336294,912
down_proj14336 x 4096294,912
per layer1,310,720
32 layers41,943,040 (about 84 MB in 16-bit)

The base model is about 8 billion parameters, about 16 GB in 16-bit, so this adapter is roughly half a percent of the base. An adapter that only targets q_proj and v_proj is 6,815,744 parameters, about 14 MB. Rank scales everything linearly, so a rank-64 all-projection adapter is about 336 MB.

Now size the GPU. On an 80 GB card, with 16 GB of base weights and, say, 10 GB kept for activations and workspace, the rest is shared between KV cache and adapter slots. Sixteen resident rank-16 adapters take about 1.3 GB, which is modest. Sixteen rank-64 adapters take about 5.4 GB, which is thousands of tokens of KV cache you no longer have. The KV cost per token is worked out in vLLM on GPU. Adapter rank is a capacity decision, not only a quality decision.

The decode cost model: distinct adapters, not requests

A decode step is limited by bytes read from GPU memory. Each step reads the whole base model once, whatever the batch size; that sharing is why batching works. It also reads the KV cache of every sequence, and the weights of every adapter used in the step. A good kernel reads each adapter once per step no matter how many rows use it. So the extra traffic depends on how many distinct adapters are in the batch, not on how many requests there are.

BASE_BYTES = 16.06e9          # Llama-3-8B weights, 16-bit
ADAPTER_BYTES = 83.9e6        # rank 16, all seven projections, 16-bit

for distinct in (1, 8, 32, 64):
    extra = distinct * ADAPTER_BYTES
    print(f"{distinct:>3} distinct adapters: +{extra/1e9:5.2f} GB per step, "
          f"{100*extra/BASE_BYTES:5.1f}% of base weight traffic")
# ->   1 distinct adapters: + 0.08 GB per step,   0.5% of base weight traffic
# ->   8 distinct adapters: + 0.67 GB per step,   4.2% of base weight traffic
# ->  32 distinct adapters: + 2.68 GB per step,  16.7% of base weight traffic
# ->  64 distinct adapters: + 5.37 GB per step,  33.4% of base weight traffic

This is a lower bound on the extra work and ignores KV traffic, kernel efficiency and launch overhead. The shape is still the lesson. A batch of 64 requests that all use one adapter costs almost the same as the base model. A batch of 64 requests spread over 64 adapters can add a third to the weight traffic of every step. Prefill is different: it is compute-bound, and the adapter adds about 2 * 41.9M operations per token against about 2 * 8B for the base, again about half a percent.

So routing a tenant's requests to the same replica improves both cache hits and step time, and a cap on distinct adapters per step is a latency control, not only a memory limit.

Running it in vLLM

vLLM exposes the pieces above as server flags. --enable-lora turns the feature on. --lora-modules name=path registers adapters at start-up. --max-loras is the number of adapters that can be served at the same time, which means the number of GPU slots. --max-lora-rank must be at least the highest rank you will load, and the documentation advises setting it to exactly that. Setting it higher than you need makes the slots bigger than the adapters in them. --max-cpu-loras sizes the host-memory cache that sits behind the GPU slots.

# Start a server with two adapters registered and room for eight on the GPU.
vllm serve meta-llama/Meta-Llama-3-8B-Instruct \
    --enable-lora \
    --lora-modules support=/adapters/support sql=/adapters/sql \
    --max-loras 8 --max-lora-rank 16 --max-cpu-loras 32

# A request selects an adapter through the "model" field.
curl http://localhost:8000/v1/completions -H 'Content-Type: application/json' \
  -d '{"model": "sql", "prompt": "List the ten largest orders:", "max_tokens": 64}'

# Runtime loading needs VLLM_ALLOW_RUNTIME_LORA_UPDATING=True in the server environment.
curl -X POST http://localhost:8000/v1/load_lora_adapter -H 'Content-Type: application/json' \
  -d '{"lora_name": "billing", "lora_path": "/adapters/billing"}'
# Offline, the same engine takes a LoRARequest(name, globally unique int id, path).
from vllm import LLM, SamplingParams
from vllm.lora.request import LoRARequest

llm = LLM(model="meta-llama/Meta-Llama-3-8B-Instruct", enable_lora=True,
          max_loras=8, max_lora_rank=16)
params = SamplingParams(max_tokens=64)
out = llm.generate(["List the ten largest orders:"], params,
                   lora_request=LoRARequest("sql", 1, "/adapters/sql"))

Treat the slot count as a scheduling constraint. When the step already uses as many adapters as there are slots, a request for another adapter has to wait for a slot to free up, even if there is KV space for it. Flags change between releases, so check vllm serve --help for your version before copying these lines.

Failure modes

  • Rank mismatch. An adapter with a rank above --max-lora-rank fails to load. Standardise ranks per base model, or run a second pool for the high-rank tenants.
  • Wrong base. An adapter trained on a different base or revision will often load, because the shapes match, and then produce quietly worse output. Record the base model's hash in adapter metadata and check it at registration.
  • Cold-load stalls. The first request for an adapter pays for a disk or network read and a copy to the GPU. At p99 this looks like random slow requests. Preload the hot adapters and track the adapter-load rate.
  • Slot thrashing. More active tenants than slots means constant eviction and reloading. Add slots or route tenants to fixed replicas.
  • Diversity-driven latency. Time per output token rises as the number of distinct adapters per step rises, which is the byte model above. Dashboards that only show request rate will miss it, so export distinct adapters per step.

Choosing a deployment shape

ShapeGPU memoryStep costUse when
Merged model per tenanta full copy eachbase onlya few tenants with steady, high traffic each
One base, unmerged adaptersbase + slotsbase + distinct adaptersmany tenants with bursty or long-tail traffic
Unmerged, tenant-affine replicasbase + small slot setclose to basemany tenants and strict latency targets
Full fine-tunesa full copy eachbase onlywhen LoRA quality is not enough

Start with one shared base; pin the heaviest tenants to their own replicas or merged models once their traffic justifies it.

What to do next

  1. List your adapters with their base-model hash, rank and target modules, and settle on one standard rank per base model.
  2. Compute each adapter's size with r * (d_in + d_out) per targeted projection, and budget slots against KV cache on your GPU.
  3. Set --max-lora-rank to your highest rank and --max-loras from that budget, then load-test with a realistic mix of tenants rather than one adapter.
  4. Plot time per output token against the number of distinct adapters per step, and compare it with the byte model on this page.
  5. Add tenant-affine routing if the curve bends, and preload the adapters that produce most of your traffic.
  6. For every new adapter, compare its output with a merged reference before routing real traffic to it.
Key takeaway: Multi-LoRA serving keeps adapters unmerged and adds a shrink and an expand product to every adapted layer. Gathered kernels do this for a whole mixed batch in one launch each: segmented for prefill, one row per request for decode. A rank-16 adapter on all projections of Llama-3-8B is about 84 MB, so memory rarely stops you. The number of distinct adapters in each decode step does, because every distinct adapter adds bytes to a memory-bound step. Size the slots, standardise the rank, measure latency against adapter diversity, and route tenants to the same replicas when the curve bends.