Every capacity plan for an LLM service starts with the same question: how many tokens per second does one GPU produce for this model and this traffic? Vendors quote peak FLOP/s, benchmark posts quote one number from one configuration, and neither predicts what your deployment will do. The good news is that a first-order answer needs only a few inputs and a few lines of arithmetic, and once calibrated against two measurements it is close enough to plan capacity.

This article builds that arithmetic as a system-level model. It treats a GPU as three ceilings (compute, memory bandwidth and memory capacity), shows which ceiling each phase of inference hits, turns that into tokens per second, latency per token and dollars per million tokens, and then explains how to calibrate the model against measurements. The per-operator byte and FLOP ledger for a single decode step is covered in Decode Compute Math; here we stay one level up, where capacity decisions are made.

Three ceilings

A GPU can be described, for this purpose, by three numbers from its data sheet. For an H100 SXM, NVIDIA lists 1,979 TFLOP/s of BF16 tensor throughput with sparsity, which means 989 TFLOP/s for the dense matrices LLMs actually use, 3.35 TB/s of HBM bandwidth and 80 GB of HBM. Every step of inference must do some arithmetic and move some bytes, so its time is bounded below by both of these:

t_compute   = FLOPs / (peak_flops * MFU)
t_bandwidth = bytes_moved / (peak_bw * bw_efficiency)
t_step      ~ max(t_compute, t_bandwidth)

MFU (model FLOPs utilisation) and bandwidth efficiency are the fudge factors that turn the data sheet into reality. Well-tuned prefill kernels commonly sustain 40-60% of peak FLOP/s, and decode kernels commonly sustain 60-80% of peak bandwidth. Treat both as measured inputs, not constants: the model below defaults to 0.5 and 0.7 and tells you which one matters for your workload.

The third ceiling is different in kind. HBM capacity does not slow a step down; it caps how many sequences can be in flight at once, because every active sequence keeps its KV cache resident. Since decode throughput grows with batch size, capacity is often the ceiling that decides throughput in practice, even though it never appears in a timing formula.

Three ceilings: a serving step takes as long as the slowest of them allowsModel shapeparams, layers, KV headsTraffic shapeprompt, output, batchGPU sheet + measuredFLOP/s, HBM BW, HBM sizeCompute ceilingFLOPs / (peak x MFU)Bandwidth ceilingbytes / (peak BW x eff)Capacity ceilingmax batch that fits in HBMStep timemax(compute, bandwidth)Tokens/s and $/Mtokbatch / step timecaps batchPrefill sits on the compute ceiling; decode sits on the bandwidth ceiling until the batch is very large.
The three ceilings. Model and traffic shape determine FLOPs and bytes; the GPU determines how fast each is served and how large a batch fits.

Prefill is compute-bound

A forward pass costs about 2 FLOPs per parameter per token (one multiply and one add in each weight matrix), plus attention, which adds roughly 4 x layers x context x d_model FLOPs per token. During prefill the whole prompt is processed in one pass, so every weight read from HBM is reused across hundreds or thousands of tokens. The arithmetic intensity is enormous and the phase is compute-bound:

prefill_tokens_per_s = peak_flops * MFU / (2 * N + 4 * L * ctx * d)

8B model, 500-token average context, MFU 0.5, H100 SXM:
  989e12 * 0.5 / (2 * 8.03e9 + 4 * 32 * 500 * 4096)  ~ 30,000 tokens/s

Thirty thousand prompt tokens per second sounds like plenty, but in a chat service with 1,000-token prompts that is only 30 requests per second of prefill, and as the worked example shows, prefill ends up consuming a large share of the GPU's time. Long prompts make prefill worse twice over: more tokens, and a larger attention term per token.

Decode is bandwidth-bound

Decode generates one token per sequence per step. Each step must stream every weight from HBM once, no matter how many sequences are in the batch, and must also read each sequence's whole KV cache. The bytes moved per step are therefore:

bytes_per_step = N * bytes_per_weight + B * ctx * kv_bytes_per_token
kv_bytes_per_token = 2 (K and V) * layers * kv_heads * head_dim * bytes_per_element

8B model with grouped-query attention (32 layers, 8 KV heads, head_dim 128, fp16 KV):
  kv_bytes_per_token = 2 * 32 * 8 * 128 * 2 = 131,072 bytes (128 KiB)

The FLOPs per step are B x (2N + attention), so the arithmetic intensity of the weight traffic is about B FLOPs per byte for a bf16 model. The H100's ratio of compute to bandwidth is about 989e12 / 3.35e12, or roughly 295 FLOPs per byte, so on weight traffic alone decode would stay bandwidth-bound until the batch reached a few hundred. The KV reads make it worse: they scale with B and are used for only a couple of FLOPs per byte, so with long contexts decode never becomes compute-bound at any batch that fits.

The chart shows the model's output for the 8B shape at a 2,048-token context. At B = 1 the step is almost entirely weight streaming, about 7 ms, so a single user sees about 144 tokens per second and the GPU is nearly idle on arithmetic. Batching amortises that weight read, and throughput rises steeply until KV traffic, which grows with each added sequence, takes over.

Modelled decode throughput, 8B model on one H100 SXM, 2,048-token context, 70% of peak bandwidthB = 1144 tok/s, step 6.96 msB = 81,030 tok/s, step 7.76 msB = 323,044 tok/s, step 10.51 msB = 644,515 tok/s, step 14.17 msB = 1285,953 tok/s, step 21.5 msB = 2006,724 tok/s, step 29.74 msThroughput climbs about 47x from B = 1 to B = 200 while per-token latency only grows about 4x.The KV reads that grow with the batch, not compute, are what bend the curve.
Modelled decode throughput against batch size. Compute time at B = 200 is about 7 ms against about 30 ms of memory time, so the step is still bandwidth-bound.

Capacity sets the batch

How large can B be? Subtract the weights and a safety reserve for activations, the CUDA context and fragmentation from HBM, then divide by the KV bytes one sequence needs:

max_batch = (HBM * (1 - reserve) - N * bytes_per_weight) / (ctx * kv_bytes_per_token)

80e9 * 0.9 - 16.06e9 = 55.9e9 bytes free
55.9e9 / (2048 * 131072) = 208 sequences at 2,048 tokens

Double the context and the batch halves, which on the chart above costs a large part of the throughput. This is why KV-cache precision, grouped-query attention, paged allocation and prefix sharing show up as throughput features: they raise the capacity ceiling, which lets B grow. KV Cache Sizing for Deployments works through the capacity side in detail, including tensor parallelism, which divides both weights and KV heads across GPUs.

A runnable estimator

The whole model fits in a short script. It is deliberately first-order: it ignores kernel launch overhead, sampling, communication and scheduler gaps, and it time-shares prefill and decode on one GPU the way chunked-prefill servers do. Its value is that every output can be traced to an input you can measure.

from dataclasses import dataclass

@dataclass
class Gpu:
    flops: float          # dense FLOP/s at serving precision
    bw: float             # HBM bytes/s
    hbm: float            # HBM bytes
    eff_bw: float = 0.7   # measured fraction of peak bandwidth in decode
    mfu: float = 0.5      # measured fraction of peak FLOP/s in prefill

@dataclass
class Model:
    params: float
    layers: int
    kv_heads: int
    head_dim: int
    q_dim: int            # n_heads * head_dim
    w_bytes: float = 2
    kv_bytes: float = 2

    def kv_per_token(self):
        return 2 * self.layers * self.kv_heads * self.head_dim * self.kv_bytes

    def flops_per_token(self, ctx):
        return 2 * self.params + 4 * self.layers * ctx * self.q_dim

def decode_step(g, m, batch, ctx):
    bytes_moved = m.params * m.w_bytes + batch * ctx * m.kv_per_token()
    t_mem = bytes_moved / (g.bw * g.eff_bw)
    t_cmp = batch * m.flops_per_token(ctx) / (g.flops * g.mfu)
    return max(t_mem, t_cmp)

def serve(g, m, batch, prompt, output, usd_per_hour):
    step = decode_step(g, m, batch, prompt + output // 2)
    prefill_s = prompt * m.flops_per_token(prompt / 2) / (g.flops * g.mfu)
    gpu_s_per_req = prefill_s + output * step / batch
    req_s = 1 / gpu_s_per_req
    share = req_s * prefill_s                  # fraction of time spent in prefill
    out_tok_s = req_s * output
    return dict(req_s=req_s, out_tok_s=out_tok_s,
                tpot_ms=1e3 * step / (1 - share), prefill_share=share,
                usd_per_mtok=usd_per_hour / 3600 / out_tok_s * 1e6)

h100 = Gpu(flops=989e12, bw=3.35e12, hbm=80e9)
m8b = Model(params=8.03e9, layers=32, kv_heads=8, head_dim=128, q_dim=4096)
print(serve(h100, m8b, batch=64, prompt=1000, output=250, usd_per_hour=3.0))

Worked example: a chat workload

Take a chat workload with 1,000-token prompts and 250-token answers on the 8B model, one H100 SXM, and an assumed price of $3 per GPU-hour (substitute your own). Running the script at three decode concurrencies gives:

Concurrent sequencesRequests/sOutput tok/sTime per output tokenPrefill share of GPU time$ per M output tokens
166.41,60510.0 ms21%0.52
6413.23,31219.3 ms44%0.25
12816.14,02631.8 ms53%0.21

Three things stand out. First, raising concurrency from 16 to 128 cuts cost per token by about 60% but triples time per output token; the right point is set by your latency objective, not by the hardware. Second, prefill takes 21-53% of the GPU even though prompts are cheap per token, which is why splitting the phases across pools is attractive at scale (see Prefill vs Decode Split). Third, the output token rate is well below the pure-decode chart, because time spent on prefill is time not spent decoding.

Little's law ties these together: sequences in flight = request rate x time in system. At 13.2 requests per second and 250 tokens at about 19 ms each, a request spends about 4.8 s decoding, so about 64 requests are in flight, which matches the concurrency we assumed. Use the same identity backwards when planning: given a target request rate and a latency budget, it tells you the concurrency each replica must sustain, and the capacity ceiling tells you whether it can.

Training throughput

Training throughput uses the same compute ceiling with a different constant. A training step does a forward and a backward pass, about 6 FLOPs per parameter per token, and with large batches it is compute-bound:

train_tokens_per_s_per_gpu = peak_flops * MFU / (6 * N)
8B model, MFU 0.4, H100 SXM: 989e12 * 0.4 / (6 * 8.03e9) ~ 8,200 tokens/s per GPU

Multiply by GPU count and by goodput (the fraction of wall time that produces committed steps) to get a schedule. FLOPS Budget for LLM Training covers the attention correction, the difference between MFU and hardware FLOPs utilisation, and goodput losses from checkpoints and restarts.

Calibrating against measurements

The model is only useful once it is calibrated. Calibrate the two efficiency factors separately, because they fail for different reasons:

  1. Run decode at a fixed batch and context with prefill disabled or excluded, time many steps, and compute achieved bandwidth as bytes_per_step / step_time. If it is far below 60% of peak, profile the attention kernel; Nsight Compute reports DRAM throughput as a percentage of peak per kernel.
  2. Run prefill only, at the prompt lengths you serve, and compute MFU as prompt tokens x FLOPs per token / (elapsed time x peak FLOP/s).
  3. Plug the measured factors back into the model and predict a mixed workload. Then run the mixed workload with a load generator at a fixed arrival rate, not a closed loop, and compare throughput and per-token latency. A gap above about 20% means something the model ignores (scheduler gaps, preemption, sampling, CPU-side tokenisation) is significant and worth finding.
  4. Record the inputs next to the result. A throughput number without model shape, prompt and output lengths, concurrency and precision cannot be compared with anything.

Failure modes

  • Using the sparse FLOP/s figure. Data sheets often headline throughput with 2:4 structured sparsity. Dense LLM inference gets half of it.
  • Forgetting the KV term. A model that counts only weight bytes predicts decode throughput that keeps rising with batch size, and plans that rely on it overshoot badly at long contexts.
  • Benchmarking with short contexts. A benchmark at 128-token prompts and 128-token outputs measures a different regime from a production mix with 4,000-token prompts; capacity and KV traffic both change by an order of magnitude.
  • Closed-loop load tests. A fixed number of clients sending back-to-back requests hides queueing, so latency looks flat right up to the point where production traffic falls over.
  • Ignoring the output length distribution. Mean output length sets throughput, but the tail sets how long sequences hold KV memory, and therefore the effective batch.
  • Treating tensor parallelism as free. Splitting a model over GPUs adds aggregate bandwidth and capacity but also all-reduce traffic per layer, and the per-GPU throughput usually drops.

Trade-offs

The model makes the core trade-off explicit: batch size buys throughput and costs latency, and memory capacity decides how much batch you can buy. Quantising weights to 8 bits halves the weight term, which helps most at small batch; quantising the KV cache helps most at large batch and long context, because it raises the capacity ceiling and cuts the KV term at once. Prefill and decode want different hardware balances, which is the argument for disaggregation, while colocating them with chunked prefill keeps utilisation high on a small fleet. Continuous batching, described in Continuous batching, is what lets a server actually hold the batch the arithmetic says is optimal.

What to do next

  1. Write down your model shape (parameters, layers, KV heads, head dimension) and precision for weights and KV.
  2. Pull prompt and output length percentiles from production logs, not from a benchmark default.
  3. Run the script with data-sheet peaks and default efficiencies to get a first estimate.
  4. Measure decode bandwidth efficiency and prefill MFU on your server, and replace the defaults.
  5. Pick the concurrency that meets your per-token latency objective, check it fits under the capacity ceiling, and read off cost per million tokens.
  6. Validate with an open-loop load test at the planned request rate and keep the inputs with the result.
Key takeaway: Model a GPU as three ceilings. Prefill is limited by compute (about 2N FLOPs per token), decode by bandwidth (all weights plus every sequence's KV cache per step), and HBM capacity limits the batch that decode needs to be efficient. A twenty-line estimator with measured efficiency factors predicts tokens per second, per-token latency and cost per million tokens well enough to plan capacity, provided you feed it your real prompt and output lengths.