A convolutional neural network (CNN) is a network built mainly from convolution layers: small learned filters slid across an input, computing the same weighted sum at every position. That one idea, reusing the same weights everywhere, made it practical to learn directly from pixels. LeNet-5 read handwritten digits in 1998; AlexNet won the 2012 ImageNet challenge with a top-5 error of 15.3% against 26.2% for the runner-up; ResNet's residual connections made networks of 150 layers trainable in 2015. Vision transformers have since taken much of the frontier, but CNNs still run most on-device vision, many detection and segmentation backbones, audio models and the image encoders inside larger systems.
This article builds CNNs up from the operation: shape and compute arithmetic, receptive fields, how a convolution becomes a GPU matrix multiply, a PyTorch training loop, failure modes and the CNN versus transformer decision.
Why convolution: locality, sharing, equivariance
Feed a 224 x 224 colour image to a fully connected layer with 1,000 outputs and you need 150,528 x 1,000, about 150 million, weights for one layer, and the layer has no idea that neighbouring pixels are related or that a cat in the corner is still a cat in the centre. Convolution builds those two facts in.
- Locality. Each output depends only on a small window, typically 3 x 3, so a layer learns local patterns such as edges and textures. Deeper layers combine them into larger ones.
- Weight sharing. The same filter is applied at every position, so the parameter count depends on filter size and channels, not on image size.
- Translation equivariance. Shift the input and the output of a stride-1 convolution shifts the same way, away from the borders. Pooling and striding later trade some of this for invariance, imperfectly: Zhang (2019) showed that standard downsampling aliases, so small shifts can change predictions.
These inductive biases make CNNs data-efficient on images; a transformer gives them up for flexibility.
The operation, precisely
A 2D convolution layer takes an input of shape Cin x H x W and has Cout filters, each of shape Cin x K x K, plus one bias per filter. Output channel o at position (i, j) is the bias plus the sum, over every input channel and every kernel offset, of input times weight. Strictly this is cross-correlation (the kernel is not flipped), but every framework calls it convolution and, because the weights are learned, the difference is irrelevant.
Four hyperparameters set the output size. Padding P adds zeros around the border; stride S moves the window S pixels at a time; dilation D spaces kernel taps D apart to see further without more weights; groups G split channels into independent groups (G = Cin gives a depthwise convolution). The output height is
H_out = floor((H + 2P - D(K - 1) - 1) / S) + 1and the same for width. A 3 x 3 kernel with P = 1, S = 1 preserves size; S = 2 halves it. Parameters are Cout x (Cin / G) x K x K, plus Cout if there is a bias, and the compute is that weight count times Hout x Wout multiply-accumulates (MACs) per image.
Backpropagation needs two more convolutions. The gradient with respect to the input is a convolution of the output gradient with the spatially flipped kernel (a transposed convolution), and the gradient with respect to the weights correlates the input with the output gradient. Training therefore costs roughly three times the forward pass, the same rule of thumb as for matrix multiplies. Gradient-based training itself is covered in the gradient descent analysis.
Worked example: shapes, parameters and compute
Take a small classifier for 32 x 32 CIFAR-10 images: three 3 x 3 convolutions with padding 1, each followed by batch normalisation and ReLU, with 2 x 2 max pooling after the first two, then global average pooling and a linear layer to 10 classes. Convolutions have no bias because batch normalisation adds its own shift.
| Layer | Output shape | Parameters | MACs per image |
|---|---|---|---|
| conv1 3 to 32 | 32 x 32 x 32 | 3 x 32 x 9 = 864 | 32 x 32 x 864 = 884,736 |
| bn1 | same | 64 | negligible |
| pool | 16 x 16 x 32 | 0 | negligible |
| conv2 32 to 64 | 16 x 16 x 64 | 18,432 | 16 x 16 x 18,432 = 4,718,592 |
| bn2 | same | 128 | negligible |
| pool | 8 x 8 x 64 | 0 | negligible |
| conv3 64 to 128 | 8 x 8 x 128 | 73,728 | 8 x 8 x 73,728 = 4,718,592 |
| bn3 | same | 256 | negligible |
| GAP + fc 128 to 10 | 10 | 1,290 | 1,280 |
| Total | 94,762 | about 10.3 million |
Two patterns are typical. Parameters concentrate in the deep layers, where channels are wide (conv3 holds 78% of them), while compute spreads across layers because early layers work on large spatial maps. And because global average pooling replaced flattening, the classifier has 1,290 weights instead of the 81,930 a flatten of the 8 x 8 x 128 map into a 10-way layer would need, and the network accepts other input sizes.
Receptive fields
The receptive field of a unit is the input region that can influence it. Track two numbers through the layers: the receptive field r and the jump j, the input distance between adjacent units. Start with r = 1, j = 1; each layer with kernel k and stride s updates r = r + (k - 1) j, then j = j s.
For the example network: conv1 gives r = 3; the pool gives r = 4 and j = 2; conv2 gives r = 8; the second pool gives r = 10 and j = 4; conv3 gives r = 18. So each final feature sees an 18 x 18 window of a 32 x 32 image, so no unit can relate opposite corners. That is why deep CNNs downsample aggressively and why dilated convolutions exist. The effective receptive field is smaller still, because central pixels reach the output through more paths.
From loops to matrix multiply on a GPU
Seven nested loops implement convolution directly, and they are slow. The standard trick, im2col, copies every input patch into a column of a matrix so the whole convolution becomes one matrix multiply, the operation GPUs and their Tensor Cores are built for (see how Tensor Cores work and parallel matrix multiplication). The NumPy below implements both forms and checks that they agree.
import numpy as np
def out_size(n, k, s, p):
return (n + 2 * p - k) // s + 1
def conv2d_naive(x, w, b, s=1, p=0): # x: (C_in,H,W) w: (C_out,C_in,K,K)
C_out, _, K, _ = w.shape
xp = np.pad(x, ((0, 0), (p, p), (p, p)))
Ho, Wo = out_size(x.shape[1], K, s, p), out_size(x.shape[2], K, s, p)
y = np.zeros((C_out, Ho, Wo))
for o in range(C_out):
for i in range(Ho):
for j in range(Wo):
patch = xp[:, i*s:i*s+K, j*s:j*s+K]
y[o, i, j] = np.sum(patch * w[o]) + b[o]
return y
def im2col(x, K, s=1, p=0):
C = x.shape[0]
xp = np.pad(x, ((0, 0), (p, p), (p, p)))
Ho, Wo = out_size(x.shape[1], K, s, p), out_size(x.shape[2], K, s, p)
cols = np.empty((C * K * K, Ho * Wo))
for i in range(Ho):
for j in range(Wo):
cols[:, i * Wo + j] = xp[:, i*s:i*s+K, j*s:j*s+K].ravel()
return cols, Ho, Wo
def conv2d_gemm(x, w, b, s=1, p=0):
C_out, _, K, _ = w.shape
cols, Ho, Wo = im2col(x, K, s, p)
return (w.reshape(C_out, -1) @ cols + b[:, None]).reshape(C_out, Ho, Wo)
rng = np.random.default_rng(0)
x, w, b = rng.standard_normal((3, 8, 8)), rng.standard_normal((4, 3, 3, 3)), rng.standard_normal(4)
assert np.allclose(conv2d_naive(x, w, b, s=2, p=1), conv2d_gemm(x, w, b, s=2, p=1))The price is memory. For conv2 in the example, the patch matrix is 288 x 256 = 73,728 values against an input of 8,192: nine times larger, as the diagram shows.
Production libraries avoid materialising it. cuDNN offers several algorithms per layer: implicit GEMM, which forms patch tiles in on-chip memory; Winograd transforms, which cut the multiplies for 3 x 3 kernels at some cost in numerical precision; and FFT-based convolution, which pays off for large kernels (the transform is explained in the FFT article). In PyTorch, setting torch.backends.cudnn.benchmark = True times the candidates on the first batch of each shape and caches the fastest. Memory layout matters too: Tensor Core kernels prefer channels-last (NHWC), so converting model and inputs with memory_format=torch.channels_last often speeds up mixed-precision training.
The building blocks of modern CNNs
- Batch normalisation normalises each channel over the batch and spatial positions, then applies a learned scale and shift. It stabilises training at high learning rates, but behaves differently in training and evaluation mode, and small batches make its statistics noisy; group normalisation is the usual fix.
- Residual connections compute y = x + F(x), so each block learns a correction and gradients have a direct path back. They are what made very deep CNNs train.
- Downsampling by max pooling or by a stride-2 convolution. Strided convolutions are learned and now common; either can alias, which blur-then-subsample layers reduce.
- Depthwise separable convolutions (MobileNet and successors) split a K x K convolution into a per-channel K x K depthwise step and a 1 x 1 pointwise step, cutting parameters and MACs by close to a factor of K2 for wide layers, at the cost of lower arithmetic intensity on GPUs.
Training one end to end
The loop below trains the example network on CIFAR-10 with standard augmentation, SGD with Nesterov momentum, a one-cycle learning-rate schedule, bfloat16 autocast and channels-last layout. It is deliberately small; the point is a correct, fast loop you can extend.
import torch, torch.nn as nn, torch.nn.functional as F
import torchvision, torchvision.transforms as T
class SmallCNN(nn.Module):
def __init__(self, n_classes=10):
super().__init__()
def block(cin, cout):
return nn.Sequential(nn.Conv2d(cin, cout, 3, padding=1, bias=False),
nn.BatchNorm2d(cout), nn.ReLU(inplace=True))
self.features = nn.Sequential(block(3, 32), nn.MaxPool2d(2),
block(32, 64), nn.MaxPool2d(2), block(64, 128))
self.head = nn.Linear(128, n_classes)
def forward(self, x):
return self.head(self.features(x).mean(dim=(2, 3))) # global average pooling
def train(epochs=15, bs=128, lr=0.1, dev="cuda"):
torch.backends.cudnn.benchmark = True
norm = T.Normalize((0.4914, 0.4822, 0.4465), (0.2470, 0.2435, 0.2616))
aug = T.Compose([T.RandomCrop(32, padding=4), T.RandomHorizontalFlip(), T.ToTensor(), norm])
tr = torchvision.datasets.CIFAR10("data", train=True, download=True, transform=aug)
te = torchvision.datasets.CIFAR10("data", train=False, download=True,
transform=T.Compose([T.ToTensor(), norm]))
tl = torch.utils.data.DataLoader(tr, bs, shuffle=True, num_workers=4, drop_last=True)
vl = torch.utils.data.DataLoader(te, 512, num_workers=4)
model = SmallCNN().to(dev, memory_format=torch.channels_last)
print(sum(p.numel() for p in model.parameters())) # 94762
opt = torch.optim.SGD(model.parameters(), lr=lr, momentum=0.9,
nesterov=True, weight_decay=5e-4)
sched = torch.optim.lr_scheduler.OneCycleLR(opt, lr, total_steps=epochs * len(tl))
for ep in range(epochs):
model.train()
for x, y in tl:
x = x.to(dev, non_blocking=True, memory_format=torch.channels_last)
with torch.autocast("cuda", dtype=torch.bfloat16):
loss = F.cross_entropy(model(x), y.to(dev, non_blocking=True))
opt.zero_grad(set_to_none=True)
loss.backward()
opt.step(); sched.step()
model.eval(); correct = 0
with torch.no_grad(), torch.autocast("cuda", dtype=torch.bfloat16):
for x, y in vl:
pred = model(x.to(dev, memory_format=torch.channels_last)).argmax(1)
correct += (pred.cpu() == y).sum().item()
print(f"epoch {ep} loss {loss.item():.3f} test acc {correct / len(te):.3f}")The printed parameter count should match the table. Note model.eval() before testing: without it batch normalisation uses the test batch's statistics. A network this small plateaus well short of state of the art; widen it and add residual blocks to go further.
Failure modes
- Forgetting eval mode: batch-norm statistics and dropout behave as in training, so accuracy at inference varies with batch composition.
- Preprocessing mismatch: different normalisation constants, channel order (RGB versus BGR) or resize method between training and serving silently costs accuracy; the serving side is covered in image classification serving.
- Receptive field too small for the task, so the model cannot relate distant parts of the image; compute it before training.
- Shortcut learning: CNNs readily latch onto texture, backgrounds or watermarks that correlate with labels. Check errors by slice, not just overall accuracy.
- Shift sensitivity from aliasing in strided layers: predictions flip under one-pixel shifts. Test with shifted inputs.
- Slow kernels: odd channel counts, NCHW layout, or input shapes that change every batch and defeat
cudnn.benchmark.
Trade-offs: CNN or vision transformer
Vision transformers (2020) split an image into patches and use attention, which lets every patch attend to every other from the first layer but brings no locality bias, so they need more data or stronger augmentation to match CNNs, and attention cost grows quadratically with the number of patches. ConvNeXt (2022) showed that a CNN modernised with transformer-era recipes is competitive with similar-sized vision transformers. Many vision-language systems, described in vision language models, use transformer encoders; many phone and embedded deployments still use CNNs.
| Situation | Prefer | Why |
|---|---|---|
| Small or medium labelled dataset | CNN, often pretrained | locality bias is free data efficiency |
| Edge or mobile latency budget | CNN with depthwise blocks | mature, quantisation-friendly kernels |
| Very large pretraining corpus | Vision transformer | scales well, shares tooling with LLMs |
| Image input to a language model | Transformer encoder | tokens fit the LLM interface |
What to do next
- Run the NumPy code and extend
conv2d_naivewith dilation and groups; keep the assert passing againsttorch.nn.functional.conv2d. - Recompute the worked example's table for your own architecture before training it: shapes, parameters, MACs and receptive field.
- Train the example on CIFAR-10, confirm the parameter count prints 94,762, then add residual blocks and compare.
- Profile one epoch with channels-last on and off, and with
cudnn.benchmarkon and off, to see what layout and algorithm selection buy on your GPU. - Test robustness: one-pixel shifts, different resize methods, and per-class error slices.
- Before choosing an architecture for a new project, fine-tune a pretrained CNN and a pretrained vision transformer of similar size on your data and compare.