Most serving decisions can be made, or ruled out, with arithmetic you can do in a notebook before renting a GPU. Will the model fit? How long will the first token take? How fast will tokens stream? How many requests run at once, and how many GPUs does a forecast need? Each has a first-order answer from five numbers: weight bytes, KV-cache bytes per token, prefill FLOPs, bytes moved per decode step, and Little's law. This primer derives each, wires them into one short Python calculator, and runs a complete sizing of Llama 3.1 70B on H100 SXM GPUs, with what-ifs, cost per token and an honest list of what the math leaves out.

The inputs and the chain

Five numbers turn a model card and a traffic forecast into a GPU countmodel configparams, layers, KV headsGPU datasheetHBM, TB/s, TFLOPStrafficreq/s, prompt, output1. weight bytesparams x bytes / TP2. KV bytes per token2 x L x kv_heads x d x dtype3. prefill time2 x params x prompt / FLOPS4. decode stepbytes moved / bandwidth5. Little's lawin flight = rate x timememory checkpool / KV per sequencelatency checkTTFT and TPOT vs SLOreplicas x TP = GPUsthen cost per token
The calculation chain: three inputs, five derived numbers, three checks.

From the model's config.json: parameter count, layers, KV heads (num_key_value_heads, not the query-head count) and head dimension. From the datasheet: memory, bandwidth and dense tensor throughput at your weight precision. From traffic: arrival rate, prompt length and output length, ideally as distributions.

This article uses NVIDIA's published H100 SXM figures: 80 GB HBM3, 3.35 TB/s, 989 dense BF16 TFLOPS and 1,979 dense FP8 TFLOPS (the larger headline numbers assume structured sparsity, which serving generally does not use). The PCIe H100 is slower, so check which variant you rent. Llama 3.1 70B has about 70.6 billion parameters, 80 layers and 8 KV heads of dimension 128.

Datasheet peaks are ceilings, so the calculator carries two explicit assumptions: 80 percent of peak bandwidth in decode and 50 percent of peak FLOPS (model FLOPs utilisation, MFU) in prefill. Replace both with measurements.

Memory: weights first, cache second

Number 1, weight bytes. Parameters times bytes per parameter. In BF16, 70.6 billion parameters is about 141 GB. Split across two GPUs with tensor parallelism (TP=2), each holds 70.6 GB, leaving nothing for the cache: BF16 70B on two H100s does not work. Quantised to FP8 (1 byte per weight) it is 70.6 GB in total, 35.3 GB per GPU at TP=2.

Number 2, KV bytes per token. For each token, each layer caches one key and one value per KV head: 2 x layers x kv_heads x head_dim x bytes per element. For Llama 3.1 70B with a BF16 cache that is 2 x 80 x 8 x 128 x 2 = 327,680 bytes, 320 KiB per token, halved per GPU at TP=2. Cache precision is chosen separately from weight precision: quantising weights to FP8 does not shrink a BF16 cache.

Combine them into a memory check. Let the engine use 90 percent of each 80 GB GPU (72 GB), subtract 35.3 GB of weights and an assumed 4 GB of runtime overhead (activation workspace and similar), and 32.7 GB remains for the cache. At 163,840 bytes per token per GPU that is about 200,000 tokens. A request with a 1,500-token prompt and 300 output tokens occupies 1,800 tokens at its peak, so at most about 110 such requests fit at once. KV cache sizing covers the budget in depth.

Prefill is a FLOPs problem

Number 3, prefill FLOPs. Each token costs roughly two floating-point operations per parameter in a dense model, one multiply and one add. Prefill processes the whole prompt in large matrix multiplications that keep tensor cores busy, so it is compute bound: time is about 2 x params x prompt tokens / (TP x peak FLOPS x MFU). For a 1,500-token prompt that is 2 x 70.6e9 x 1,500 = 2.1e14 FLOPs, split over two GPUs at 1,979 TFLOPS and 50 percent MFU: about 107 ms. That is the floor under time to first token (TTFT) before any queueing. Attention adds work that grows with the square of the prompt length; it is small at 1,500 tokens and significant at tens of thousands, so long-context workloads need the fuller model in GPU inference latency.

Decode is a bytes problem

Decode step time per GPU: weights are read once per step, KV once per sequenceLlama 3.1 70B, FP8 weights, BF16 KV, TP=2 on H100 SXM, 1,650 live tokens per sequence, 80% of peak bandwidthbatch 113.3 msbatch 3216.4 msbatch 6419.6 msbatch 11024.3 msweights: 35.3 GB per GPU, shared by the whole batchKV cache: about 0.27 GB per sequence per GPU, never shared
Why batching is cheap until it is not: the weight term is constant per step, the KV term grows with batch and context.

Number 4, decode step time. Each decode step produces one token for every sequence in the batch, and to do so it reads every weight and every live sequence's KV cache from HBM. Decode does 2 FLOPs per weight per sequence, so at batch 1 it performs about 2 FLOPs per byte of FP8 weights. The H100 can do about 1,979 / 3.35, roughly 590, FP8 FLOPs in the time it reads one byte (for BF16, 989 / 3.35 is about 295 FLOPs per byte against 1 FLOP per byte of two-byte weights, the same crossover batch). Until the batch approaches a few hundred, the tensor cores wait on memory, and step time is bytes divided by achieved bandwidth.

At batch 1 with 1,650 live tokens (the average during a 1,500 + 300 request), each GPU moves 35.6 GB per step: 13.3 ms, or 75 tokens per second. At batch 64 the KV term adds 17.3 GB and the step takes 19.6 ms, but it yields 64 tokens: about 3,260 tokens per second per pair of GPUs, 43 times the throughput for 1.5 times the latency. But KV reads are per sequence and never shared, so at long contexts decode stays memory bound at any batch size. See decode compute math.

Little's law: the batch you actually get

Number 5, Little's law. In any stable system, the average number of requests inside equals the arrival rate times the average time each spends inside. For an LLM replica the number decoding at once is the batch. So the batch is not a knob you set; it is arrival rate x output tokens x TPOT, and TPOT itself grows with the batch. That loop has a fixed point, which the calculator finds by iteration, or no fixed point, when the replica is overloaded.

Prefill and decode share the GPU. With continuous batching and chunked prefill, prompt chunks ride in the same forward passes as decode tokens. To first order, prefill occupies a share rho = rate x prefill time of each second and stretches decode steps by 1 / (1 - rho). Queueing in front of prefill is approximated as an M/D/1 queue, adding rho x prefill / (2 x (1 - rho)) to TTFT.

The calculator

import math
from dataclasses import dataclass

@dataclass
class Model:
    params: float      # parameters
    layers: int
    kv_heads: int
    head_dim: int
    w_bytes: float     # bytes per weight: 2 = BF16, 1 = FP8
    kv_bytes: float    # bytes per cached K or V element

@dataclass
class GPU:
    mem: float         # bytes of HBM
    bw: float          # datasheet bytes/s
    flops: float       # datasheet dense FLOP/s at the weight dtype

BW_EFF, MFU = 0.8, 0.5   # achieved fractions: assumptions, replace with measurements

def kv_per_token(m):
    return 2 * m.layers * m.kv_heads * m.head_dim * m.kv_bytes

def decode_step(m, g, tp, batch, ctx):
    """Slower of moving the bytes and doing the FLOPs, per GPU."""
    bytes_ = (m.params * m.w_bytes + batch * ctx * kv_per_token(m)) / tp
    flops = 2 * m.params * batch / tp
    return max(bytes_ / (g.bw * BW_EFF), flops / (g.flops * MFU))

def replica(m, g, tp, prompt, output, rps):
    """Steady state of one replica fed `rps` requests per second."""
    prefill = 2 * m.params * prompt / tp / (g.flops * MFU)
    rho = rps * prefill                     # share of each second spent on prefill
    if rho >= 1:
        return None
    ctx = prompt + output / 2               # mean live context while decoding
    batch = 1.0
    for _ in range(200):                    # Little's law: batch = rate x time decoding
        tpot = decode_step(m, g, tp, batch, ctx) / (1 - rho)
        new = rps * output * tpot
        if new > 10_000:
            return None                     # KV term outruns bandwidth: unstable
        if abs(new - batch) < 1e-6:
            break
        batch = new
    ttft = prefill + rho * prefill / (2 * (1 - rho))   # M/D/1 queue in front of prefill
    return dict(prefill_ms=prefill * 1e3, rho=rho, batch=batch,
                tpot_ms=tpot * 1e3, ttft_ms=ttft * 1e3)

def plan(m, g, tp, prompt, output, rps, tpot_slo_ms, ttft_slo_ms,
         util=0.9, overhead=4e9, max_rho=0.7):
    pool = g.mem * util - m.params * m.w_bytes / tp - overhead
    if pool <= 0:
        raise ValueError("weights do not fit: raise TP or quantise")
    max_seqs = pool / (kv_per_token(m) / tp * (prompt + output))
    for n in range(1, 1000):
        s = replica(m, g, tp, prompt, output, rps / n)
        if (s and s["rho"] <= max_rho and s["batch"] <= max_seqs
                and s["tpot_ms"] <= tpot_slo_ms and s["ttft_ms"] <= ttft_slo_ms):
            return dict(replicas=n, gpus=n * tp, pool_gb=round(pool / 1e9, 1),
                        max_seqs=int(max_seqs), **{k: round(v, 2) for k, v in s.items()})
    raise ValueError("no replica count meets the targets")

llama70b = Model(70.6e9, 80, 8, 128, w_bytes=1, kv_bytes=2)   # FP8 weights, BF16 KV
h100_sxm = GPU(80e9, 3.35e12, 1979e12)                        # FP8 dense peak
print(plan(llama70b, h100_sxm, tp=2, prompt=1500, output=300, rps=20,
           tpot_slo_ms=30, ttft_slo_ms=500))

Worked example: 20 requests per second on H100s

The forecast: 20 requests per second, 1,500-token prompts, 300-token answers, a streaming target of 30 ms per token (about 33 tokens per second, faster than people read) and TTFT under 500 ms. The calculator walks the replica count up until every check passes. Here is what it finds at each step, for Llama 3.1 70B in FP8 at TP=2:

ReplicasPrefill share (rho)Batch per replicaTPOTTTFTVerdict
21.07---saturated: prefill alone exceeds the GPU
30.71311155 ms240 msfails: batch exceeds the 110 that fit
40.546342 ms169 msfails TPOT
50.433529.2 ms147 mspasses: 10 GPUs
60.362424.3 ms137 mspasses with headroom: 12 GPUs
70.311921.7 ms131 mspasses

Three lessons. This workload is prefill heavy: 30,000 prompt tokens per second against 6,000 output tokens. The transition is a cliff: four replicas run at 42 ms, three collapse to 155 ms, because the Little's law loop amplifies itself as TPOT rises. And five replicas pass with almost no margin, so a production plan would take six.

What-ifs cost one line each. Raising TP to 4 gives a 50.4 GB cache pool per GPU and halves both prefill time and step time; the calculator says two replicas (8 GPUs) with 21 ms TPOT and 84 ms TTFT. Treat that with suspicion: the model ignores the all-reduce traffic tensor parallelism adds on every layer, so measure before believing that 8 GPUs beat 10. 6,000-token prompts (with a 1,500 ms TTFT target) quadruple prefill and cut the sequences that fit to 31: 20 replicas, 40 GPUs.

Cost follows directly. The fleet serves 20 x 300 x 3,600 = 21.6 million output tokens per hour, so a million output tokens costs about 0.46 GPU-hours at the five-replica minimum and 0.56 at the recommended six, prompts included. Multiply by your contract price.

What first-order math leaves out

First-order math rules options in or out. It ignores real costs that push the answer up:

  • Achieved efficiency. BW_EFF and MFU vary by kernel, batch shape and engine version. Measure batch-1 decode step time and a single prefill at your prompt length, then solve for both.
  • Tensor-parallel communication. Every layer ends in an all-reduce. On NVLink it is modest; over PCIe it can dominate decode.
  • Length variance. Averages hide tails. A few 30,000-token prompts cause TTFT spikes and preemptions that a mean of 1,500 never shows. Run the calculator on your p95 lengths too.
  • Host overhead. Scheduling, sampling and detokenisation add time per step.
  • Prefix caching and speculative decoding. Both help and need their own terms.

Failure modes

  • Using query heads for KV size. Counting Llama 3.1 70B's 64 attention heads instead of its 8 KV heads overestimates the cache eightfold.
  • Sparse datasheet numbers. Doubling FLOPS by quoting the sparsity figure halves predicted prefill time for no reason.
  • Treating batch as a setting. Max batch is a cap; the batch you get comes from Little's law. A plan that assumes batch 128 at 20 requests per second over-promises.
  • Sizing at the average. Capacity planned at mean load runs out at peak. Size at the peak hour with a utilisation target, as in the max_rho check above.

Trade-offs

LeverMovesCost
FP8 or INT4 weightsweight bytes, decode floor, fits on fewer GPUsaccuracy, must be evaluated
FP8 KV cacheKV bytes per token, concurrencyaccuracy at long context
Higher TPper-GPU bytes and FLOPsall-reduce latency, more GPUs per replica
Shorter prompts or prefix cachingprefill FLOPs, TTFT, rhoproduct or cache-hit dependence
Looser TPOT targetlarger batch, cheaper tokensslower streaming

What to do next

  1. Copy the calculator and replace the model and GPU with yours, reading KV heads from config.json.
  2. Pull prompt and output length distributions from real logs; run the plan at mean and p95.
  3. Measure batch-1 decode step time and one prefill on your engine; solve for BW_EFF and MFU.
  4. Compare the engine's logged KV cache size with the calculator's pool; investigate gaps over 10 percent.
  5. Load test at the planned replica count with open-loop traffic and confirm TTFT and TPOT at p99.
  6. Record which term dominates your workload (prefill share or decode bytes) and pick levers from that row.
  7. Recompute cost per million tokens whenever prompt length, model or GPU price changes.
Key takeaway: Weights must fit, the cache gets what is left, prefill costs about 2 x params FLOPs per prompt token, decode costs the bytes of the weights plus every live sequence's cache per step, and Little's law decides the batch you really run. Chain those five numbers, check memory, TTFT and TPOT, then multiply replicas by TP. Calibrate the two efficiency factors on your own stack before trusting the GPU count.