"The model is slow" is not a diagnosis. An LLM job that underperforms is limited by one resource at a time: arithmetic throughput, memory bandwidth, the interconnect, the host CPU that launches work, or memory capacity that caps how much work can be in flight. Each of those has a different fix, and the fixes conflict. Quantising weights rescues a bandwidth-bound decode loop and does nothing for a host-bound one. Raising the batch rescues an under-filled GPU and makes a capacity-bound server preempt more.
This page is a procedure for naming the bound before touching anything. You will compute two utilisation numbers, MFU and MBU, from figures you already log. You will learn which hardware counters agree with them and which mislead. Then you will run small perturbation experiments that confirm the answer, because a single profile can be read several ways. The per-operator byte and FLOP ledger is in decode compute math and the latency model is in GPU inference latency; this page uses both and re-derives neither.
Five bounds, and why only one matters at a time
Start from the five candidates. Each one leaves a recognisable signature in throughput, latency and counters. A real workload usually sits close to two of them, but at any given batch size and context length one dominates, and that one sets the ceiling.
| Bound | What is saturated | Typical LLM situation | Signature |
|---|---|---|---|
| Compute | Tensor-core arithmetic | Prefill of long prompts; training with large micro-batches | Time scales with tokens processed; high tensor-pipe activity; MFU of 40% or more |
| Memory bandwidth | HBM reads | Decode at small or moderate batch; long-context decode | Step time barely moves with batch; DRAM busy, tensor pipe mostly idle; MBU of 60% or more |
| Communication | NVLink, PCIe or network | Tensor parallel decode across many GPUs; data-parallel gradient all-reduce | Gaps where GPUs wait on collectives; scaling efficiency falls as GPUs are added |
| Host / launch | The CPU thread issuing kernels | Small models, tiny batches, Python-heavy schedulers | Idle gaps between short kernels; one CPU core pinned; low SM activity |
| Capacity | HBM space for KV cache or activations | Many long concurrent requests; long-sequence training | Batch cannot grow; requests queue or are preempted; out-of-memory without the fix |
Capacity is the odd one out. It does not make any single step slow; it stops you reaching the batch size at which the other resources would be well used. That is why a capacity-bound server often shows a bandwidth-bound profile: the GPU is doing small-batch decode because it has no room for a bigger batch.
MFU and MBU: two ratios that name the bound
Two ratios turn a throughput number into a statement about the hardware. Model FLOPs Utilisation (MFU) is the arithmetic the model requires per second divided by the GPU's peak. Model Bandwidth Utilisation (MBU) is the bytes the model must read per second divided by peak memory bandwidth. Both use the model's requirement, not what the kernels actually executed, so wasted work never inflates them.
For a dense transformer with N parameters, a forward pass costs about 2N FLOPs per token and training about 6N (forward plus a backward pass twice as expensive). Attention adds a term that grows with context; at short context it is small enough to ignore for a first pass. A decode step must read every weight once, plus the KV cache of every sequence in the batch. The calculator below is deliberately small so you can check every line.
from dataclasses import dataclass
@dataclass
class Gpu:
peak_flops: float # dense FLOP/s at your dtype, from the datasheet
peak_bw: float # bytes/s of HBM bandwidth
def decode_step(gpu, params, bytes_per_param, batch, ctx, kv_bytes_per_token, step_s):
weight_bytes = params * bytes_per_param
kv_bytes = batch * ctx * kv_bytes_per_token # read once per step
flops = 2 * params * batch # attention term ignored
mbu = (weight_bytes + kv_bytes) / step_s / gpu.peak_bw
mfu = flops / step_s / gpu.peak_flops
floor_s = (weight_bytes + kv_bytes) / gpu.peak_bw # cannot beat this
return dict(mbu=mbu, mfu=mfu, floor_ms=floor_s * 1e3,
kv_share=kv_bytes / (weight_bytes + kv_bytes))
def prefill(gpu, params, prompt_tokens, ttft_s):
return 2 * params * prompt_tokens / ttft_s / gpu.peak_flops # MFU
def train_mfu(gpu, params, tokens_per_s):
return 6 * params * tokens_per_s / gpu.peak_flops
H100 = Gpu(peak_flops=989e12, peak_bw=3.35e12) # SXM, BF16 denseThe ratio of the two peaks is the ridge point of the roofline: 989e12 / 3.35e12 is about 295 FLOPs per byte. A kernel that performs fewer FLOPs per byte read than that cannot be compute-bound on this GPU, however well written. Batch-1 decode performs roughly 1 FLOP per byte of BF16 weight (2 FLOPs per 2-byte parameter), which is why it is bandwidth-bound everywhere.
Worked example: an 8B model on one H100
Take an 8-billion-parameter model in BF16 on one H100 SXM. The weights are 16 GB. Assume grouped-query attention with 32 layers, 8 KV heads and a head dimension of 128. The KV cache then costs 2 (K and V) x 32 x 8 x 128 x 2 bytes = 131,072 bytes, 128 KiB, per token of context.
- Batch 1, short context, 7.0 ms per token. Bytes per step are about 16 GB, so achieved bandwidth is 16e9 / 0.007 = 2.29 TB/s and MBU is 68%. The floor is 16e9 / 3.35e12 = 4.8 ms. FLOPs are 16 GFLOP per step, 2.3 TFLOP/s, an MFU of 0.2%. Verdict: bandwidth-bound and reasonably efficient. Better kernels might reach 5.5 ms; only fewer bytes (quantisation, a smaller model, speculative decoding) can go below 4.8 ms.
- Batch 64, 2,000 tokens of context each, 12 ms per step. KV bytes are 64 x 2,000 x 131,072 = 16.8 GB, now as large as the weights. Total 32.8 GB per step gives 2.73 TB/s, MBU 82%. Throughput is 64 / 0.012 = 5,333 tokens/s and MFU is 8.6%. Still bandwidth-bound, but half the bytes are KV cache, so weight quantisation alone buys at most about half the step time. KV-cache quantisation or shorter context now matters as much.
- Prefill of a 2,000-token prompt with a 60 ms time to first token. 2 x 8e9 x 2,000 = 32 TFLOP in 60 ms is 533 TFLOP/s, MFU 54%. This phase is compute-bound, and the fixes are different: FP8 matmuls, better attention kernels, or chunking the prefill so it does not stall decode (see chunked prefill).
The arithmetic alone shows one deployment with a compute-bound phase and a bandwidth-bound phase whose character changes with context. When both ratios are low, say MBU of 25% in decode, something else is eating the time: usually the host, the interconnect or the scheduler.
Which counters to trust
Hardware counters are useful as corroboration, and treacherous as a starting point. The figure most dashboards show, the utilisation column of nvidia-smi, reports the fraction of time in which at least one kernel was running. A GPU executing a single tiny kernel continuously reads 100%. It says nothing about how much of the chip did work.
The DCGM profiling fields are closer to what you need. DCGM_FI_PROF_SM_ACTIVE is the fraction of time streaming multiprocessors had work assigned. DCGM_FI_PROF_PIPE_TENSOR_ACTIVE is the fraction of cycles the tensor pipe was busy. DCGM_FI_PROF_DRAM_ACTIVE is the fraction of cycles the memory interface was moving data. Read them as a triple: high DRAM with low tensor activity agrees with a bandwidth bound; high tensor activity agrees with a compute bound; low SM activity with low everything means the GPU is starved and the problem is upstream.
For kernel-level confirmation, Nsight Compute's speed-of-light section (ncu --section SpeedOfLight) reports each kernel's compute and memory throughput as a percentage of peak. A decode GEMV at 80% of memory throughput and 5% of compute is behaving as it must. A GEMM in prefill at 30% of both is a kernel-selection or shape problem. For the timeline view that exposes host gaps and collective waits, use Nsight Systems; how both tools work, and how their overhead distorts what you see, is covered in GPU profiling with Nsight.
Perturbation tests: change one thing, predict the response
A profile is a single observation, and the same observation fits several stories. A perturbation test changes one variable and predicts how each bound would respond. The bound whose prediction matches is the one you have. These tests are cheap enough to script and rerun after every change.
| Change | Compute-bound | Bandwidth-bound | Host-bound | Communication-bound |
|---|---|---|---|---|
| Double the decode batch | Step time roughly doubles | Step time rises a little (KV bytes only) | Step time barely moves; throughput nearly doubles | Rises with message size |
| Double context at fixed batch | Prefill time more than doubles (attention) | Decode step grows with the KV share | Little change | Little change |
| Lock the SM clock lower | Slows in proportion | Barely slows | Barely slows | Barely slows |
| Capture the decode loop in CUDA graphs | No change | No change | Large improvement | Small change |
| Halve tensor-parallel degree (if it fits) | Per-GPU time doubles | Per-GPU time doubles | Little change | Per-token latency may improve |
Lowering the clock is the most discriminating test and needs administrator rights: nvidia-smi -lgc <min>,<max> locks the graphics clock and nvidia-smi -rgc resets it. Memory clocks are unaffected, so a bandwidth-bound step hardly changes while a compute-bound one slows by the clock ratio. Run it on a test node, never on a shared production host.
# Batch sweep: the shape of the curve names the bound.
import time, torch
def step_time(run_step, batch, warmup=5, iters=20):
for _ in range(warmup):
run_step(batch)
torch.cuda.synchronize()
t0 = time.perf_counter()
for _ in range(iters):
run_step(batch)
torch.cuda.synchronize() # without this you time the launch, not the work
return (time.perf_counter() - t0) / iters
for b in (1, 2, 4, 8, 16, 32, 64, 128):
s = step_time(run_decode_step, b)
print(f"batch={b:4d} step={s*1e3:7.2f} ms tok/s={b/s:9.0f}")Read the output as a curve. A flat region means fixed per-step cost dominates (weights or launch overhead; the clock test separates them), a linear region means a compute or KV-bandwidth bound, and a curve that stops early means capacity.
The same method for training jobs
Training uses the same method with different suspects. Compute MFU from tokens per second with the 6N rule and compare against the peak at your precision. A well-tuned dense run sits far above single digits; a low figure means one of four things.
- Input pipeline. The GPU idles at the start of each step while the data loader catches up. The timeline shows a gap before the first forward kernel. Fixes: more loader workers, pinned memory, pre-tokenised shards, prefetching.
- Exposed communication. Gradient all-reduce or sharded-parameter all-gathers that do not overlap with compute. Scaling from 8 to 64 GPUs costs more than the extra data parallelism buys. Fixes: larger gradient buckets, overlap settings in your framework, keeping tensor parallelism inside one NVLink domain.
- Recomputation. Activation checkpointing re-runs forward passes. Hardware FLOPs rise while model FLOPs do not, so MFU falls even though the tensor cores look busy. That is a deliberate capacity-for-compute trade, not a defect; report hardware FLOPs utilisation alongside MFU so nobody chases it.
- Small kernels. Element-wise operations, normalisation and optimizer steps that are each bandwidth-bound and too short to fill the GPU. Fixes: fused kernels,
torch.compile, fused optimizers.
How the analysis goes wrong
The analysis itself has failure modes, and most of them come from measuring the wrong thing.
- Mixing phases. Averaging prefill and decode into one tokens-per-second figure produces an MFU and MBU that describe neither. Measure them separately, or use the engine's per-phase timings.
- Timing without synchronising. CUDA calls are asynchronous; a Python timer around them measures launch, not execution. Synchronise, or use CUDA events.
- Wrong peak. Using a sparsity figure (twice the dense number) or an FP8 peak while running BF16 halves your apparent MFU. Use the dense peak for the dtype the matmuls actually run in, and the HBM figure for the exact SKU, since PCIe and SXM variants differ.
- Forgetting KV bytes. Computing MBU from weights alone at long context makes a healthy server look half-idle, and leads to buying faster GPUs when shorter context or a quantised KV cache would do.
- Benchmark versus production. A fixed-batch benchmark hides queueing and preemption; confirm with production traces.
Trade-offs: bounds to accept and bounds that move
Some bounds should be accepted rather than removed. An interactive chat service that keeps batches small for latency will always be bandwidth-bound in decode, with MFU in single digits, and that is correct. The right target there is MBU close to the hardware limit and a latency objective met, not a high MFU. An offline batch job has the opposite priority: raise the batch until capacity binds, accept higher per-request latency, and buy throughput.
Fixes also move the bound rather than eliminating it: 4-bit weights may expose the host or the KV cache next, and a bigger batch spends KV capacity. Rerun the measurement after every change. For a ranked list of levers once you know the bound, see the LLM inference optimisation overview.
What to do next
- Log prefill and decode times separately, with batch size and total context per step, for one representative model.
- Compute MFU for prefill and MBU for decode with the calculator above, using the dense datasheet peaks for your exact GPU and dtype.
- Run a batch sweep and save the curve; label the flat and linear regions.
- If neither ratio is above about 50%, capture a short Nsight Systems trace and look for gaps between kernels and waits on collectives.
- On a test node, run the clock-lock test to separate compute from bandwidth or host bounds.
- Export the three DCGM profiling fields to your dashboards next to nvidia-smi utilisation, and stop alerting on the latter alone.
- Write the verdict down per phase ("decode: bandwidth-bound, MBU 78%, KV is 45% of bytes") and pick the one lever that moves that term.
- Re-measure after each change and record which bound you moved to.