A CPU core is a latency machine: it spends its transistors on branch prediction, out-of-order execution and large caches so that one instruction stream finishes as soon as possible. A GPU is a throughput machine: it spends them on thousands of arithmetic lanes and enough resident threads that the lanes stay busy while individual threads wait on memory. GPU Architecture Overview covers that hardware contrast in depth, so this article starts from the other end, the workload. What does a training step actually ask a machine to do, how do the two machines score on those demands, and what is left for the CPU once the GPU takes the heavy work?
The short answer is that training is dominated by large dense matrix multiplications, which have enough arithmetic per byte to exploit whatever compute a chip offers, and a modern data-center GPU offers roughly two orders of magnitude more of it than a server CPU at the precisions training uses. The longer answer explains why the gap is not just FLOPs, and why a GPU training job still fails in interesting ways when its CPU side is neglected.
What a training step computes
For a dense transformer with N parameters, one training step over T tokens costs about 6·N·T floating-point operations: 2·N·T for the forward pass, where every parameter takes part in one multiply-add per token, and about 4·N·T for the backward pass, which computes gradients with respect to both activations and weights. Attention adds a term that grows with sequence length, but for typical models and contexts the parameter matmuls dominate.
Those FLOPs are not evenly spread across operations. The projections in attention and the feed-forward block are general matrix multiplications (GEMMs), and they account for the overwhelming majority of the arithmetic. Around them sit many cheap operations: normalization, activation functions, residual adds, dropout, softmax, and the optimizer update. These perform a handful of operations per element and are limited by how fast data moves, not how fast it is multiplied. A good mental model is that the GEMMs set the speed limit of training, and the memory-bound operations set how close you get to it.
Arithmetic intensity: the number that decides
Arithmetic intensity is the FLOPs performed per byte moved to or from main memory. For a GEMM multiplying an M×K matrix by a K×N matrix, the work is 2·M·N·K FLOPs and the minimum traffic is reading both inputs and writing the output. With square 4096 matrices in a 2-byte format, that is 2 × 4096³ FLOPs over 2 × 3 × 4096² bytes, about 1,365 FLOPs per byte. A LayerNorm over the same activations reads and writes each element once and does a few operations on it, which comes to around one FLOP per byte or less.
A machine's ridge point is its peak compute divided by its memory bandwidth. Operations with intensity above the ridge are compute-bound; below it they are bandwidth-bound. This single comparison explains most of the performance behavior of deep-learning workloads on either kind of hardware.
Putting numbers on both machines
For the GPU, take an NVIDIA H100 SXM, whose datasheet lists 67 TFLOP/s of FP32, about 989 TFLOP/s of dense BF16 on tensor cores (the headline 1,979 figure assumes structured sparsity), and 3.35 TB/s of HBM3 bandwidth. Its BF16 ridge point is therefore about 295 FLOPs per byte.
For the CPU, build an illustrative server socket from arithmetic rather than a specific product: 64 cores at 2.5 GHz, each with two 512-bit fused multiply-add units, so 64 × 2 units × 16 FP32 lanes × 2 FLOPs × 2.5 GHz ≈ 10.2 TFLOP/s of FP32. Give it twelve channels of DDR5-4800 at 38.4 GB/s each, about 461 GB/s. Its ridge point is about 22 FLOPs per byte. Real CPUs vary widely around these figures, some have matrix extensions that raise low-precision throughput for inference, and sustained vector clocks are often below nominal, but the order of magnitude holds.
Now compare. A large GEMM is compute-bound on both chips, so the gap is the ratio of compute roofs: about 97 times for GPU BF16 tensor cores against CPU FP32, or about 6.6 times if you forced the GPU to use plain FP32. A LayerNorm is bandwidth-bound on both, so the gap is the bandwidth ratio, about 7 times. Training time is dominated by the first case, which is why precision support matters as much as core counts.
A worked estimate: one run, two machines
Take a 350M-parameter model trained on 100B tokens with one million tokens per step, so 100,000 steps. Each step needs 6 × 3.5e8 × 1e6 = 2.1e15 FLOPs.
- CPU socket at 50% of its 10.2 TF peak (optimistic for a full training step): 2.1e15 / 5.1e12 ≈ 412 seconds per step, about 477 days for the run.
- One H100 at 40% model FLOPs utilization in BF16, a realistic figure for a well-tuned small-model run: 2.1e15 / 3.96e14 ≈ 5.3 seconds per step, about 6.1 days.
The ratio, around 78 times, understates the practical difference. The GPU run can be split across eight GPUs in a node with fast interconnect and finish in under a day; spreading the CPU run across many sockets hits network bandwidth long before it approaches linear scaling, because gradient synchronization moves the full gradient every step. Memory capacity matters too: optimizer state for mixed-precision Adam is roughly 16 bytes per parameter, which fits comfortably in 80 GB for this model, and the GPU's bandwidth keeps the optimizer update itself from becoming a bottleneck.
Measure it on your own hardware
Datasheet numbers set ceilings; measurements tell you where you are. The script below times a large matmul and a LayerNorm on each device and reports achieved TFLOP/s and GB/s. Two details make or break it: warm-up iterations, so the timing excludes one-time costs, and torch.cuda.synchronize(), because GPU calls return before the work finishes and unsynchronized timing measures only the launch.
import time, torch
def bench(fn, device, iters=20):
for _ in range(3): # warm-up: allocator, kernel selection
fn()
if device == "cuda":
torch.cuda.synchronize() # GPU calls are async; wait before timing
t0 = time.perf_counter()
for _ in range(iters):
fn()
if device == "cuda":
torch.cuda.synchronize()
return (time.perf_counter() - t0) / iters
def run(device, dtype):
n = 4096
a = torch.randn(n, n, device=device, dtype=dtype)
b = torch.randn(n, n, device=device, dtype=dtype)
x = torch.randn(8192, 4096, device=device, dtype=dtype)
ln = torch.nn.LayerNorm(4096, device=device, dtype=dtype)
t_mm = bench(lambda: a @ b, device)
t_ln = bench(lambda: ln(x), device)
flops = 2 * n**3
moved = 2 * x.numel() * x.element_size() # read x, write y (ignores weights)
print(f"{device:4s} {str(dtype):15s} matmul {flops / t_mm / 1e12:7.2f} TFLOP/s"
f" layernorm {moved / t_ln / 1e9:8.1f} GB/s")
torch.set_num_threads(torch.get_num_threads()) # pin explicitly; see failure modes
run("cpu", torch.float32)
if torch.cuda.is_available():
torch.backends.cuda.matmul.allow_tf32 = True
run("cuda", torch.float32)
run("cuda", torch.bfloat16)Expect the CPU matmul to land well below its theoretical peak and the GPU BF16 matmul to land at a large fraction of the tensor-core peak for this shape. The LayerNorm lines show the bandwidth story, and the FP32-with-TF32 line on the GPU shows how much a single precision setting changes the result.
Why the gap is more than FLOPs
Four further properties widen the practical gap. Low-precision matrix units: tensor cores execute BF16 and FP8 matrix tiles at many times the rate of scalar FP32 lanes, and mixed-precision training keeps accuracy by accumulating in FP32 and holding master weights in FP32. Memory bandwidth: HBM stacked beside the die delivers several times the bandwidth of socketed DRAM, which speeds every memory-bound operation. Scale-out interconnect: GPUs in a node exchange gradients over dedicated links much faster than PCIe or a data-center network, which is what makes data, tensor and pipeline parallelism practical. Software: cuBLAS, cuDNN, fused attention kernels, NCCL collectives and compiler stacks have been tuned for training for more than a decade, and frameworks assume them.
What the CPU does inside a GPU training node
Moving the math to the GPU does not make the CPU irrelevant; it changes its job to feeding and orchestrating. The CPU reads and decodes data, tokenizes text or decodes and augments images, collates batches, copies them to the device, launches every kernel, runs the Python training loop, writes checkpoints and, in offloading setups, may even run the optimizer (see Optimizer Offload to CPU). A GPU that computes a step in 300 ms is wasted if the input pipeline needs 400 ms to produce the batch.
loader = torch.utils.data.DataLoader(
dataset,
batch_size=per_gpu_batch,
num_workers=8, # CPU processes doing decode/tokenize/augment
pin_memory=True, # page-locked buffers so the copy can be async DMA
prefetch_factor=4, # batches queued ahead per worker
persistent_workers=True, # do not re-fork workers every epoch
)
for step, (x, y) in enumerate(loader):
x = x.cuda(non_blocking=True) # overlaps with queued GPU work
y = y.cuda(non_blocking=True)
with torch.autocast("cuda", dtype=torch.bfloat16):
loss = model(x, y)
loss.backward()
opt.step(); opt.zero_grad(set_to_none=True)
if step % 100 == 0:
print(step, loss.item()) # .item() syncs: keep it off the hot pathMultiple loader workers parallelize preprocessing across cores, pinned memory lets the host-to-device copy run as asynchronous DMA, and non_blocking=True lets that copy overlap with queued GPU work. On the launch side, each kernel launch costs the host a few microseconds, which matters for small models and small batches where kernels are short; fusion, CUDA graphs and compilation reduce the count. The GPU Dataloader Bottleneck and GPU Kernel Launch cover each side in depth. The practical rule when buying or renting nodes is to size CPU cores and host memory to the input pipeline, not as an afterthought.
Where CPUs still win or are good enough
- Small-model, batch-one inference. Decoding one token at a time is memory-bound, a quantized small model fits in CPU caches and DRAM, and there is no PCIe hop. For many edge and cost-sensitive deployments a CPU is the right answer; see CPU Inference Pipelines.
- Classical machine learning. Gradient-boosted trees, linear models and most tabular pipelines are branchy and irregular, and run well on CPUs.
- Huge sparse embedding tables. Recommendation models with terabyte-scale embedding tables often keep them in host memory, with lookups on the CPU and dense layers on the GPU.
- Data preparation. Parsing, filtering, deduplication and tokenization at corpus scale are CPU work.
- Debugging and tiny experiments. A unit test on a two-layer model runs fine on a laptop CPU, and deterministic CPU execution is a useful reference when chasing numerical bugs.
Failure modes
- Starved GPU. Utilization looks high in coarse monitoring, but profiles show gaps between steps while the loader catches up.
- Timing without synchronization. Benchmarks report impossible speeds because they measured kernel launches, not execution.
- Hidden sync points. Calling
.item(),.cpu()or printing a tensor every step forces the host to wait for the device and empties the launch queue. - Plain FP32 on tensor-core hardware. Leaving TF32 and mixed precision off forfeits most of the GPU's matrix throughput.
- Launch-bound small models. Tiny batches produce microsecond kernels, and the host cannot launch them fast enough.
- CPU thread oversubscription. Many loader workers each running multi-threaded math libraries contend for cores and slow everything down.
Choosing, and the trade-offs
Train on GPUs, or on other accelerators with matrix units and high-bandwidth memory, whenever the model is a neural network of meaningful size; the arithmetic above leaves little room for debate. Spend real design effort on the CPU side of those nodes, because an under-provisioned input pipeline wastes the most expensive component in the system. For inference, decide per workload: large models, large batches and long contexts favor GPUs, while small quantized models serving one request at a time can be cheaper and simpler on CPUs. The costs to weigh are hardware or rental price, energy per useful FLOP, utilization you can actually sustain, and the engineering effort of each software stack.