"The model is slow" is not a diagnosis. An LLM job that underperforms is limited by one resource at a time: arithmetic throughput, memory bandwidth, the interconnect, the host CPU that launches work, or memory capacity that caps how much work can be in flight. Each of those has a different fix, and the fixes conflict. Quantising weights rescues a bandwidth-bound decode loop and does nothing for a host-bound one. Raising the batch rescues an under-filled GPU and makes a capacity-bound server preempt more.

This page is a procedure for naming the bound before touching anything. You will compute two utilisation numbers, MFU and MBU, from figures you already log. You will learn which hardware counters agree with them and which mislead. Then you will run small perturbation experiments that confirm the answer, because a single profile can be read several ways. The per-operator byte and FLOP ledger is in decode compute math and the latency model is in GPU inference latency; this page uses both and re-derives neither.

Five bounds, and why only one matters at a time

Start from the five candidates. Each one leaves a recognisable signature in throughput, latency and counters. A real workload usually sits close to two of them, but at any given batch size and context length one dominates, and that one sets the ceiling.

BoundWhat is saturatedTypical LLM situationSignature
ComputeTensor-core arithmeticPrefill of long prompts; training with large micro-batchesTime scales with tokens processed; high tensor-pipe activity; MFU of 40% or more
Memory bandwidthHBM readsDecode at small or moderate batch; long-context decodeStep time barely moves with batch; DRAM busy, tensor pipe mostly idle; MBU of 60% or more
CommunicationNVLink, PCIe or networkTensor parallel decode across many GPUs; data-parallel gradient all-reduceGaps where GPUs wait on collectives; scaling efficiency falls as GPUs are added
Host / launchThe CPU thread issuing kernelsSmall models, tiny batches, Python-heavy schedulersIdle gaps between short kernels; one CPU core pinned; low SM activity
CapacityHBM space for KV cache or activationsMany long concurrent requests; long-sequence trainingBatch cannot grow; requests queue or are preempted; out-of-memory without the fix

Capacity is the odd one out. It does not make any single step slow; it stops you reaching the batch size at which the other resources would be well used. That is why a capacity-bound server often shows a bandwidth-bound profile: the GPU is doing small-batch decode because it has no room for a bigger batch.

MFU and MBU: two ratios that name the bound

Two ratios turn a throughput number into a statement about the hardware. Model FLOPs Utilisation (MFU) is the arithmetic the model requires per second divided by the GPU's peak. Model Bandwidth Utilisation (MBU) is the bytes the model must read per second divided by peak memory bandwidth. Both use the model's requirement, not what the kernels actually executed, so wasted work never inflates them.

For a dense transformer with N parameters, a forward pass costs about 2N FLOPs per token and training about 6N (forward plus a backward pass twice as expensive). Attention adds a term that grows with context; at short context it is small enough to ignore for a first pass. A decode step must read every weight once, plus the KV cache of every sequence in the batch. The calculator below is deliberately small so you can check every line.

from dataclasses import dataclass

@dataclass
class Gpu:
    peak_flops: float      # dense FLOP/s at your dtype, from the datasheet
    peak_bw: float         # bytes/s of HBM bandwidth

def decode_step(gpu, params, bytes_per_param, batch, ctx, kv_bytes_per_token, step_s):
    weight_bytes = params * bytes_per_param
    kv_bytes = batch * ctx * kv_bytes_per_token       # read once per step
    flops = 2 * params * batch                         # attention term ignored
    mbu = (weight_bytes + kv_bytes) / step_s / gpu.peak_bw
    mfu = flops / step_s / gpu.peak_flops
    floor_s = (weight_bytes + kv_bytes) / gpu.peak_bw  # cannot beat this
    return dict(mbu=mbu, mfu=mfu, floor_ms=floor_s * 1e3,
                kv_share=kv_bytes / (weight_bytes + kv_bytes))

def prefill(gpu, params, prompt_tokens, ttft_s):
    return 2 * params * prompt_tokens / ttft_s / gpu.peak_flops   # MFU

def train_mfu(gpu, params, tokens_per_s):
    return 6 * params * tokens_per_s / gpu.peak_flops

H100 = Gpu(peak_flops=989e12, peak_bw=3.35e12)   # SXM, BF16 dense

The ratio of the two peaks is the ridge point of the roofline: 989e12 / 3.35e12 is about 295 FLOPs per byte. A kernel that performs fewer FLOPs per byte read than that cannot be compute-bound on this GPU, however well written. Batch-1 decode performs roughly 1 FLOP per byte of BF16 weight (2 FLOPs per 2-byte parameter), which is why it is bandwidth-bound everywhere.

Worked example: an 8B model on one H100

Take an 8-billion-parameter model in BF16 on one H100 SXM. The weights are 16 GB. Assume grouped-query attention with 32 layers, 8 KV heads and a head dimension of 128. The KV cache then costs 2 (K and V) x 32 x 8 x 128 x 2 bytes = 131,072 bytes, 128 KiB, per token of context.

  1. Batch 1, short context, 7.0 ms per token. Bytes per step are about 16 GB, so achieved bandwidth is 16e9 / 0.007 = 2.29 TB/s and MBU is 68%. The floor is 16e9 / 3.35e12 = 4.8 ms. FLOPs are 16 GFLOP per step, 2.3 TFLOP/s, an MFU of 0.2%. Verdict: bandwidth-bound and reasonably efficient. Better kernels might reach 5.5 ms; only fewer bytes (quantisation, a smaller model, speculative decoding) can go below 4.8 ms.
  2. Batch 64, 2,000 tokens of context each, 12 ms per step. KV bytes are 64 x 2,000 x 131,072 = 16.8 GB, now as large as the weights. Total 32.8 GB per step gives 2.73 TB/s, MBU 82%. Throughput is 64 / 0.012 = 5,333 tokens/s and MFU is 8.6%. Still bandwidth-bound, but half the bytes are KV cache, so weight quantisation alone buys at most about half the step time. KV-cache quantisation or shorter context now matters as much.
  3. Prefill of a 2,000-token prompt with a 60 ms time to first token. 2 x 8e9 x 2,000 = 32 TFLOP in 60 ms is 533 TFLOP/s, MFU 54%. This phase is compute-bound, and the fixes are different: FP8 matmuls, better attention kernels, or chunking the prefill so it does not stall decode (see chunked prefill).

The arithmetic alone shows one deployment with a compute-bound phase and a bandwidth-bound phase whose character changes with context. When both ratios are low, say MBU of 25% in decode, something else is eating the time: usually the host, the interconnect or the scheduler.

Which counters to trust

Hardware counters are useful as corroboration, and treacherous as a starting point. The figure most dashboards show, the utilisation column of nvidia-smi, reports the fraction of time in which at least one kernel was running. A GPU executing a single tiny kernel continuously reads 100%. It says nothing about how much of the chip did work.

The DCGM profiling fields are closer to what you need. DCGM_FI_PROF_SM_ACTIVE is the fraction of time streaming multiprocessors had work assigned. DCGM_FI_PROF_PIPE_TENSOR_ACTIVE is the fraction of cycles the tensor pipe was busy. DCGM_FI_PROF_DRAM_ACTIVE is the fraction of cycles the memory interface was moving data. Read them as a triple: high DRAM with low tensor activity agrees with a bandwidth bound; high tensor activity agrees with a compute bound; low SM activity with low everything means the GPU is starved and the problem is upstream.

For kernel-level confirmation, Nsight Compute's speed-of-light section (ncu --section SpeedOfLight) reports each kernel's compute and memory throughput as a percentage of peak. A decode GEMV at 80% of memory throughput and 5% of compute is behaving as it must. A GEMM in prefill at 30% of both is a kernel-selection or shape problem. For the timeline view that exposes host gaps and collective waits, use Nsight Systems; how both tools work, and how their overhead distorts what you see, is covered in GPU profiling with Nsight.

Measure step timeper phase: prefill, decodeCompute MFU and MBUrequirement / peakMFU highMBU highboth lowCompute-boundFP8, kernels, less workBandwidth-boundfewer bytes, bigger batchLook at the timelinewhere is the idle time?gaps between kernelswaits on collectivesHost / launch-boundCUDA graphs, batchingCommunicationless TP, overlapCapacity-boundbatch capped by KV memoryThen: can the batch grow?if not, check capacity
Bottleneck triage: compute the two utilisation ratios per phase first, and only open a timeline when neither ratio is high.

Perturbation tests: change one thing, predict the response

A profile is a single observation, and the same observation fits several stories. A perturbation test changes one variable and predicts how each bound would respond. The bound whose prediction matches is the one you have. These tests are cheap enough to script and rerun after every change.

ChangeCompute-boundBandwidth-boundHost-boundCommunication-bound
Double the decode batchStep time roughly doublesStep time rises a little (KV bytes only)Step time barely moves; throughput nearly doublesRises with message size
Double context at fixed batchPrefill time more than doubles (attention)Decode step grows with the KV shareLittle changeLittle change
Lock the SM clock lowerSlows in proportionBarely slowsBarely slowsBarely slows
Capture the decode loop in CUDA graphsNo changeNo changeLarge improvementSmall change
Halve tensor-parallel degree (if it fits)Per-GPU time doublesPer-GPU time doublesLittle changePer-token latency may improve

Lowering the clock is the most discriminating test and needs administrator rights: nvidia-smi -lgc <min>,<max> locks the graphics clock and nvidia-smi -rgc resets it. Memory clocks are unaffected, so a bandwidth-bound step hardly changes while a compute-bound one slows by the clock ratio. Run it on a test node, never on a shared production host.

# Batch sweep: the shape of the curve names the bound.
import time, torch

def step_time(run_step, batch, warmup=5, iters=20):
    for _ in range(warmup):
        run_step(batch)
    torch.cuda.synchronize()
    t0 = time.perf_counter()
    for _ in range(iters):
        run_step(batch)
    torch.cuda.synchronize()               # without this you time the launch, not the work
    return (time.perf_counter() - t0) / iters

for b in (1, 2, 4, 8, 16, 32, 64, 128):
    s = step_time(run_decode_step, b)
    print(f"batch={b:4d} step={s*1e3:7.2f} ms tok/s={b/s:9.0f}")

Read the output as a curve. A flat region means fixed per-step cost dominates (weights or launch overhead; the clock test separates them), a linear region means a compute or KV-bandwidth bound, and a curve that stops early means capacity.

The same method for training jobs

Training uses the same method with different suspects. Compute MFU from tokens per second with the 6N rule and compare against the peak at your precision. A well-tuned dense run sits far above single digits; a low figure means one of four things.

  • Input pipeline. The GPU idles at the start of each step while the data loader catches up. The timeline shows a gap before the first forward kernel. Fixes: more loader workers, pinned memory, pre-tokenised shards, prefetching.
  • Exposed communication. Gradient all-reduce or sharded-parameter all-gathers that do not overlap with compute. Scaling from 8 to 64 GPUs costs more than the extra data parallelism buys. Fixes: larger gradient buckets, overlap settings in your framework, keeping tensor parallelism inside one NVLink domain.
  • Recomputation. Activation checkpointing re-runs forward passes. Hardware FLOPs rise while model FLOPs do not, so MFU falls even though the tensor cores look busy. That is a deliberate capacity-for-compute trade, not a defect; report hardware FLOPs utilisation alongside MFU so nobody chases it.
  • Small kernels. Element-wise operations, normalisation and optimizer steps that are each bandwidth-bound and too short to fill the GPU. Fixes: fused kernels, torch.compile, fused optimizers.

How the analysis goes wrong

The analysis itself has failure modes, and most of them come from measuring the wrong thing.

  • Mixing phases. Averaging prefill and decode into one tokens-per-second figure produces an MFU and MBU that describe neither. Measure them separately, or use the engine's per-phase timings.
  • Timing without synchronising. CUDA calls are asynchronous; a Python timer around them measures launch, not execution. Synchronise, or use CUDA events.
  • Wrong peak. Using a sparsity figure (twice the dense number) or an FP8 peak while running BF16 halves your apparent MFU. Use the dense peak for the dtype the matmuls actually run in, and the HBM figure for the exact SKU, since PCIe and SXM variants differ.
  • Forgetting KV bytes. Computing MBU from weights alone at long context makes a healthy server look half-idle, and leads to buying faster GPUs when shorter context or a quantised KV cache would do.
  • Benchmark versus production. A fixed-batch benchmark hides queueing and preemption; confirm with production traces.

Trade-offs: bounds to accept and bounds that move

Some bounds should be accepted rather than removed. An interactive chat service that keeps batches small for latency will always be bandwidth-bound in decode, with MFU in single digits, and that is correct. The right target there is MBU close to the hardware limit and a latency objective met, not a high MFU. An offline batch job has the opposite priority: raise the batch until capacity binds, accept higher per-request latency, and buy throughput.

Fixes also move the bound rather than eliminating it: 4-bit weights may expose the host or the KV cache next, and a bigger batch spends KV capacity. Rerun the measurement after every change. For a ranked list of levers once you know the bound, see the LLM inference optimisation overview.

What to do next

  1. Log prefill and decode times separately, with batch size and total context per step, for one representative model.
  2. Compute MFU for prefill and MBU for decode with the calculator above, using the dense datasheet peaks for your exact GPU and dtype.
  3. Run a batch sweep and save the curve; label the flat and linear regions.
  4. If neither ratio is above about 50%, capture a short Nsight Systems trace and look for gaps between kernels and waits on collectives.
  5. On a test node, run the clock-lock test to separate compute from bandwidth or host bounds.
  6. Export the three DCGM profiling fields to your dashboards next to nvidia-smi utilisation, and stop alerting on the latter alone.
  7. Write the verdict down per phase ("decode: bandwidth-bound, MBU 78%, KV is 45% of bytes") and pick the one lever that moves that term.
  8. Re-measure after each change and record which bound you moved to.
Key takeaway: Every slow LLM job is limited by one resource at a time: compute, memory bandwidth, communication, host launch overhead or memory capacity. Compute MFU for prefill and training and MBU for decode, counting KV-cache bytes, against the dense datasheet peaks; a high ratio names the bound and a low pair sends you to the timeline. Corroborate with DCGM tensor, DRAM and SM activity rather than nvidia-smi utilisation, confirm with batch sweeps and clock locking, and re-measure after every fix because fixes move the bound rather than removing it.