Most serving decisions can be made, or ruled out, with arithmetic you can do in a notebook before renting a GPU. Will the model fit? How long will the first token take? How fast will tokens stream? How many requests run at once, and how many GPUs does a forecast need? Each has a first-order answer from five numbers: weight bytes, KV-cache bytes per token, prefill FLOPs, bytes moved per decode step, and Little's law. This primer derives each, wires them into one short Python calculator, and runs a complete sizing of Llama 3.1 70B on H100 SXM GPUs, with what-ifs, cost per token and an honest list of what the math leaves out.
The inputs and the chain
From the model's config.json: parameter count, layers, KV heads (num_key_value_heads, not the query-head count) and head dimension. From the datasheet: memory, bandwidth and dense tensor throughput at your weight precision. From traffic: arrival rate, prompt length and output length, ideally as distributions.
This article uses NVIDIA's published H100 SXM figures: 80 GB HBM3, 3.35 TB/s, 989 dense BF16 TFLOPS and 1,979 dense FP8 TFLOPS (the larger headline numbers assume structured sparsity, which serving generally does not use). The PCIe H100 is slower, so check which variant you rent. Llama 3.1 70B has about 70.6 billion parameters, 80 layers and 8 KV heads of dimension 128.
Datasheet peaks are ceilings, so the calculator carries two explicit assumptions: 80 percent of peak bandwidth in decode and 50 percent of peak FLOPS (model FLOPs utilisation, MFU) in prefill. Replace both with measurements.
Memory: weights first, cache second
Number 1, weight bytes. Parameters times bytes per parameter. In BF16, 70.6 billion parameters is about 141 GB. Split across two GPUs with tensor parallelism (TP=2), each holds 70.6 GB, leaving nothing for the cache: BF16 70B on two H100s does not work. Quantised to FP8 (1 byte per weight) it is 70.6 GB in total, 35.3 GB per GPU at TP=2.
Number 2, KV bytes per token. For each token, each layer caches one key and one value per KV head: 2 x layers x kv_heads x head_dim x bytes per element. For Llama 3.1 70B with a BF16 cache that is 2 x 80 x 8 x 128 x 2 = 327,680 bytes, 320 KiB per token, halved per GPU at TP=2. Cache precision is chosen separately from weight precision: quantising weights to FP8 does not shrink a BF16 cache.
Combine them into a memory check. Let the engine use 90 percent of each 80 GB GPU (72 GB), subtract 35.3 GB of weights and an assumed 4 GB of runtime overhead (activation workspace and similar), and 32.7 GB remains for the cache. At 163,840 bytes per token per GPU that is about 200,000 tokens. A request with a 1,500-token prompt and 300 output tokens occupies 1,800 tokens at its peak, so at most about 110 such requests fit at once. KV cache sizing covers the budget in depth.
Prefill is a FLOPs problem
Number 3, prefill FLOPs. Each token costs roughly two floating-point operations per parameter in a dense model, one multiply and one add. Prefill processes the whole prompt in large matrix multiplications that keep tensor cores busy, so it is compute bound: time is about 2 x params x prompt tokens / (TP x peak FLOPS x MFU). For a 1,500-token prompt that is 2 x 70.6e9 x 1,500 = 2.1e14 FLOPs, split over two GPUs at 1,979 TFLOPS and 50 percent MFU: about 107 ms. That is the floor under time to first token (TTFT) before any queueing. Attention adds work that grows with the square of the prompt length; it is small at 1,500 tokens and significant at tens of thousands, so long-context workloads need the fuller model in GPU inference latency.
Decode is a bytes problem
Number 4, decode step time. Each decode step produces one token for every sequence in the batch, and to do so it reads every weight and every live sequence's KV cache from HBM. Decode does 2 FLOPs per weight per sequence, so at batch 1 it performs about 2 FLOPs per byte of FP8 weights. The H100 can do about 1,979 / 3.35, roughly 590, FP8 FLOPs in the time it reads one byte (for BF16, 989 / 3.35 is about 295 FLOPs per byte against 1 FLOP per byte of two-byte weights, the same crossover batch). Until the batch approaches a few hundred, the tensor cores wait on memory, and step time is bytes divided by achieved bandwidth.
At batch 1 with 1,650 live tokens (the average during a 1,500 + 300 request), each GPU moves 35.6 GB per step: 13.3 ms, or 75 tokens per second. At batch 64 the KV term adds 17.3 GB and the step takes 19.6 ms, but it yields 64 tokens: about 3,260 tokens per second per pair of GPUs, 43 times the throughput for 1.5 times the latency. But KV reads are per sequence and never shared, so at long contexts decode stays memory bound at any batch size. See decode compute math.
Little's law: the batch you actually get
Number 5, Little's law. In any stable system, the average number of requests inside equals the arrival rate times the average time each spends inside. For an LLM replica the number decoding at once is the batch. So the batch is not a knob you set; it is arrival rate x output tokens x TPOT, and TPOT itself grows with the batch. That loop has a fixed point, which the calculator finds by iteration, or no fixed point, when the replica is overloaded.
Prefill and decode share the GPU. With continuous batching and chunked prefill, prompt chunks ride in the same forward passes as decode tokens. To first order, prefill occupies a share rho = rate x prefill time of each second and stretches decode steps by 1 / (1 - rho). Queueing in front of prefill is approximated as an M/D/1 queue, adding rho x prefill / (2 x (1 - rho)) to TTFT.
The calculator
import math
from dataclasses import dataclass
@dataclass
class Model:
params: float # parameters
layers: int
kv_heads: int
head_dim: int
w_bytes: float # bytes per weight: 2 = BF16, 1 = FP8
kv_bytes: float # bytes per cached K or V element
@dataclass
class GPU:
mem: float # bytes of HBM
bw: float # datasheet bytes/s
flops: float # datasheet dense FLOP/s at the weight dtype
BW_EFF, MFU = 0.8, 0.5 # achieved fractions: assumptions, replace with measurements
def kv_per_token(m):
return 2 * m.layers * m.kv_heads * m.head_dim * m.kv_bytes
def decode_step(m, g, tp, batch, ctx):
"""Slower of moving the bytes and doing the FLOPs, per GPU."""
bytes_ = (m.params * m.w_bytes + batch * ctx * kv_per_token(m)) / tp
flops = 2 * m.params * batch / tp
return max(bytes_ / (g.bw * BW_EFF), flops / (g.flops * MFU))
def replica(m, g, tp, prompt, output, rps):
"""Steady state of one replica fed `rps` requests per second."""
prefill = 2 * m.params * prompt / tp / (g.flops * MFU)
rho = rps * prefill # share of each second spent on prefill
if rho >= 1:
return None
ctx = prompt + output / 2 # mean live context while decoding
batch = 1.0
for _ in range(200): # Little's law: batch = rate x time decoding
tpot = decode_step(m, g, tp, batch, ctx) / (1 - rho)
new = rps * output * tpot
if new > 10_000:
return None # KV term outruns bandwidth: unstable
if abs(new - batch) < 1e-6:
break
batch = new
ttft = prefill + rho * prefill / (2 * (1 - rho)) # M/D/1 queue in front of prefill
return dict(prefill_ms=prefill * 1e3, rho=rho, batch=batch,
tpot_ms=tpot * 1e3, ttft_ms=ttft * 1e3)
def plan(m, g, tp, prompt, output, rps, tpot_slo_ms, ttft_slo_ms,
util=0.9, overhead=4e9, max_rho=0.7):
pool = g.mem * util - m.params * m.w_bytes / tp - overhead
if pool <= 0:
raise ValueError("weights do not fit: raise TP or quantise")
max_seqs = pool / (kv_per_token(m) / tp * (prompt + output))
for n in range(1, 1000):
s = replica(m, g, tp, prompt, output, rps / n)
if (s and s["rho"] <= max_rho and s["batch"] <= max_seqs
and s["tpot_ms"] <= tpot_slo_ms and s["ttft_ms"] <= ttft_slo_ms):
return dict(replicas=n, gpus=n * tp, pool_gb=round(pool / 1e9, 1),
max_seqs=int(max_seqs), **{k: round(v, 2) for k, v in s.items()})
raise ValueError("no replica count meets the targets")
llama70b = Model(70.6e9, 80, 8, 128, w_bytes=1, kv_bytes=2) # FP8 weights, BF16 KV
h100_sxm = GPU(80e9, 3.35e12, 1979e12) # FP8 dense peak
print(plan(llama70b, h100_sxm, tp=2, prompt=1500, output=300, rps=20,
tpot_slo_ms=30, ttft_slo_ms=500))
Worked example: 20 requests per second on H100s
The forecast: 20 requests per second, 1,500-token prompts, 300-token answers, a streaming target of 30 ms per token (about 33 tokens per second, faster than people read) and TTFT under 500 ms. The calculator walks the replica count up until every check passes. Here is what it finds at each step, for Llama 3.1 70B in FP8 at TP=2:
| Replicas | Prefill share (rho) | Batch per replica | TPOT | TTFT | Verdict |
|---|---|---|---|---|---|
| 2 | 1.07 | - | - | - | saturated: prefill alone exceeds the GPU |
| 3 | 0.71 | 311 | 155 ms | 240 ms | fails: batch exceeds the 110 that fit |
| 4 | 0.54 | 63 | 42 ms | 169 ms | fails TPOT |
| 5 | 0.43 | 35 | 29.2 ms | 147 ms | passes: 10 GPUs |
| 6 | 0.36 | 24 | 24.3 ms | 137 ms | passes with headroom: 12 GPUs |
| 7 | 0.31 | 19 | 21.7 ms | 131 ms | passes |
Three lessons. This workload is prefill heavy: 30,000 prompt tokens per second against 6,000 output tokens. The transition is a cliff: four replicas run at 42 ms, three collapse to 155 ms, because the Little's law loop amplifies itself as TPOT rises. And five replicas pass with almost no margin, so a production plan would take six.
What-ifs cost one line each. Raising TP to 4 gives a 50.4 GB cache pool per GPU and halves both prefill time and step time; the calculator says two replicas (8 GPUs) with 21 ms TPOT and 84 ms TTFT. Treat that with suspicion: the model ignores the all-reduce traffic tensor parallelism adds on every layer, so measure before believing that 8 GPUs beat 10. 6,000-token prompts (with a 1,500 ms TTFT target) quadruple prefill and cut the sequences that fit to 31: 20 replicas, 40 GPUs.
Cost follows directly. The fleet serves 20 x 300 x 3,600 = 21.6 million output tokens per hour, so a million output tokens costs about 0.46 GPU-hours at the five-replica minimum and 0.56 at the recommended six, prompts included. Multiply by your contract price.
What first-order math leaves out
First-order math rules options in or out. It ignores real costs that push the answer up:
- Achieved efficiency. BW_EFF and MFU vary by kernel, batch shape and engine version. Measure batch-1 decode step time and a single prefill at your prompt length, then solve for both.
- Tensor-parallel communication. Every layer ends in an all-reduce. On NVLink it is modest; over PCIe it can dominate decode.
- Length variance. Averages hide tails. A few 30,000-token prompts cause TTFT spikes and preemptions that a mean of 1,500 never shows. Run the calculator on your p95 lengths too.
- Host overhead. Scheduling, sampling and detokenisation add time per step.
- Prefix caching and speculative decoding. Both help and need their own terms.
Failure modes
- Using query heads for KV size. Counting Llama 3.1 70B's 64 attention heads instead of its 8 KV heads overestimates the cache eightfold.
- Sparse datasheet numbers. Doubling FLOPS by quoting the sparsity figure halves predicted prefill time for no reason.
- Treating batch as a setting. Max batch is a cap; the batch you get comes from Little's law. A plan that assumes batch 128 at 20 requests per second over-promises.
- Sizing at the average. Capacity planned at mean load runs out at peak. Size at the peak hour with a utilisation target, as in the max_rho check above.
Trade-offs
| Lever | Moves | Cost |
|---|---|---|
| FP8 or INT4 weights | weight bytes, decode floor, fits on fewer GPUs | accuracy, must be evaluated |
| FP8 KV cache | KV bytes per token, concurrency | accuracy at long context |
| Higher TP | per-GPU bytes and FLOPs | all-reduce latency, more GPUs per replica |
| Shorter prompts or prefix caching | prefill FLOPs, TTFT, rho | product or cache-hit dependence |
| Looser TPOT target | larger batch, cheaper tokens | slower streaming |
What to do next
- Copy the calculator and replace the model and GPU with yours, reading KV heads from
config.json. - Pull prompt and output length distributions from real logs; run the plan at mean and p95.
- Measure batch-1 decode step time and one prefill on your engine; solve for BW_EFF and MFU.
- Compare the engine's logged KV cache size with the calculator's pool; investigate gaps over 10 percent.
- Load test at the planned replica count with open-loop traffic and confirm TTFT and TPOT at p99.
- Record which term dominates your workload (prefill share or decode bytes) and pick levers from that row.
- Recompute cost per million tokens whenever prompt length, model or GPU price changes.