cuDNN is NVIDIA's library of deep-learning primitives: convolutions, normalisations, attention, matrix multiplies with fused epilogues, and the backward passes of all of them. You rarely call it directly. When PyTorch runs torch.nn.functional.conv2d on a CUDA tensor, or picks the cuDNN backend for scaled dot-product attention, the work is handed to cuDNN, which chooses one of many kernels for that exact shape, dtype, layout and GPU. Most of what goes right or wrong with cuDNN in production comes from that choice: a slow first iteration, a run that is fast on one batch size and slow on the next, a gradient that differs bit-for-bit between two identical runs, or an error saying a graph is not supported.
This article explains the library from the point of view of the software that drives it. It covers the two APIs, how the cuDNN 9 library is packaged, how a graph of operations becomes an execution plan, and a worked example that fuses a convolution, a bias and a ReLU in the Python frontend and counts the memory traffic the fusion saves. It ends with the PyTorch switches that control cuDNN, the failure modes worth recognising on sight, and a checklist. Hardware detail appears only where it changes what the software should do.
Two APIs and the cuDNN 9 library layout
cuDNN has two programming models. The legacy API is imperative: you create descriptors for tensors, filters and a convolution, ask for an algorithm, and call a fixed entry point. It exposes a fixed set of operations and a handful of hard-wired fusions. The graph API, introduced in cuDNN 8, is declarative: you describe a small graph of operations (a convolution feeding a bias, an activation and a scale, say) and cuDNN finds engines that can execute the whole graph, often as a single kernel.
cuDNN 9 reorganised the shipped library around that split. The graph API lives in libcudnn_graph, and the implementations live in separately loaded sub-libraries: cudnn_engines_precompiled for kernels compiled ahead of time, cudnn_engines_runtime_compiled for kernels generated when a plan is built, and cudnn_heuristics for the models that rank candidate engines. The graph library loads these on demand. Legacy functionality sits in its own sub-libraries. The practical consequence is packaging: a container that copies only one shared object, or a wheel that bundles a different cuDNN from the system one, can fail when a sub-library is missing, not when the process starts.
Graphs, engines and execution plans
The graph API has a small vocabulary, and every frontend call maps onto it.
- Tensor: dimensions, strides, a data type and a unique id. Strides encode layout, so an NHWC tensor is a 4-D tensor in NCHW dimension order with strides
[H*W*C, 1, W*C, C]. There is no separate layout flag. - Operation: a node such as convolution forward, data-gradient, weight-gradient, matmul, pointwise (add, multiply, ReLU and others), reduction, normalisation or attention.
- Operation graph: the nodes and the virtual tensors that connect them. A virtual tensor never needs to reach global memory, which is where fusion savings come from.
- Engine: one implementation strategy able to run some class of graphs. An engine has knobs (tile sizes, split factors, kernel configurations), and an engine with every knob set is an engine config.
- Execution plan: a finalised engine config for one specific graph, with a known workspace size. Building a plan can cost milliseconds, and much more when it involves runtime compilation.
- Variant pack: the device pointers for the graph's real tensors, plus the workspace, supplied on every execution.
Heuristics turn a graph into a ranked list of engine configs. The frontend exposes heuristic modes: mode A is the fast default, and heur_mode.FALLBACK returns engines that trade speed for broad support. Passing both, as in [heur_mode.A, heur_mode.FALLBACK], is the usual way to say "the best ranked plan, but never fail just because the ranked ones refuse this graph". Autotuning replaces the ranking with measurement: build several plans, time each on real data, keep the fastest. The ranking is a prediction; the timing is ground truth for the GPU the process is running on.
Worked example: fusing a convolution, bias and ReLU
Take a ResNet-style layer: a batch of 8 images, 64 input channels, 56 by 56 pixels, 32 output channels, a 3 by 3 filter, padding 1, half precision, followed by a per-channel bias and a ReLU. The Python frontend (pip install nvidia-cudnn-frontend, imported as cudnn) describes all three operations as one graph:
import cudnn
import torch
handle = cudnn.create_handle()
N, C, H, W, K = 8, 64, 56, 56, 32
graph = cudnn.pygraph(
handle=handle,
io_data_type=cudnn.data_type.HALF,
intermediate_data_type=cudnn.data_type.FLOAT, # virtual tensors stay fp32
compute_data_type=cudnn.data_type.FLOAT,
)
# NHWC expressed through strides; dims are always listed N, C, H, W.
X = graph.tensor(name="X", dim=[N, C, H, W], stride=[H * W * C, 1, W * C, C])
Wt = graph.tensor(name="W", dim=[K, C, 3, 3], stride=[3 * 3 * C, 1, 3 * C, C])
B = graph.tensor(name="B", dim=[1, K, 1, 1], stride=[K, 1, K, K])
conv = graph.conv_fprop(X, Wt, padding=[1, 1], stride=[1, 1], dilation=[1, 1],
compute_data_type=cudnn.data_type.FLOAT)
biased = graph.bias(name="bias", input=conv, bias=B)
Y = graph.relu(name="relu", input=biased)
Y.set_output(True) # only Y is materialised; conv and biased are virtual
graph.build([cudnn.heur_mode.A, cudnn.heur_mode.FALLBACK])
cl = torch.channels_last
x = torch.randn(N, C, H, W, device="cuda", dtype=torch.float16).to(memory_format=cl)
w = torch.randn(K, C, 3, 3, device="cuda", dtype=torch.float16).to(memory_format=cl)
b = torch.randn(1, K, 1, 1, device="cuda", dtype=torch.float16)
y = torch.empty(N, K, H, W, device="cuda", dtype=torch.float16).to(memory_format=cl)
workspace = torch.empty(graph.get_workspace_size(), device="cuda", dtype=torch.uint8)
graph.execute({X: x, Wt: w, B: b, Y: y}, workspace, handle=handle)
torch.cuda.synchronize()
ref = torch.relu(torch.nn.functional.conv2d(x, w, padding=1) + b)
torch.testing.assert_close(y, ref, atol=5e-3, rtol=3e-3)graph.build is shorthand for the explicit sequence the frontend also exposes: validate the graph, build the backend operation graph, create candidate execution plans from the heuristic modes, check support, and build plans. Writing the steps out is useful when you want to filter engines or autotune yourself.
Now count what fusion buys. The convolution performs 2 x 8 x 32 x 56 x 56 x 64 x 9 = 924,844,032 floating-point operations, about 0.92 GFLOP. It reads 3,211,264 bytes of input and 36,864 bytes of weights and writes 1,605,632 bytes of output: 4,853,760 bytes in total, roughly 190 FLOP per byte, comfortably compute-bound on a modern GPU. Run bias and ReLU as two separate elementwise kernels and each one reads and writes the whole output again: 4 x 1,605,632 = 6,422,528 extra bytes, more traffic than the convolution itself needed, spent on two operations that do almost no arithmetic. In the fused graph those intermediate tensors are virtual: the epilogue applies bias and ReLU to values still in registers, and the output is written once.
Controlling cuDNN from PyTorch
Inside PyTorch you control cuDNN through a few process-wide switches in torch.backends.cudnn:
| Setting | What it does | When to change it |
|---|---|---|
benchmark | Times candidate algorithms the first time each new input configuration is seen and caches the winner | Turn on for fixed shapes (vision training); leave off when shapes vary every step |
benchmark_limit | Caps how many candidates the benchmark tries | Lower it if first-step autotuning is too slow |
deterministic | Restricts cuDNN to deterministic algorithms | Reproducibility work, debugging divergence; expect some slowdown |
allow_tf32 | Lets fp32 convolutions use TF32 Tensor Core math on GPUs that support it | Turn off when a model is sensitive to the reduced mantissa |
import torch
torch.backends.cudnn.benchmark = True # autotune per new shape, then reuse
torch.backends.cudnn.deterministic = False # set True when chasing nondeterminism
torch.backends.cudnn.allow_tf32 = True # fp32 convs may use TF32
# Attention: ask SDPA for the cuDNN backend explicitly when comparing backends.
from torch.nn.attention import sdpa_kernel, SDPBackend
with sdpa_kernel(SDPBackend.CUDNN_ATTENTION):
out = torch.nn.functional.scaled_dot_product_attention(q, k, v, is_causal=True)The benchmark flag is the one that bites. Its cache is keyed by input configuration, so a model fed variable sequence lengths or ragged image sizes autotunes again for every new shape, and each of those steps pays for timing several kernels. Bucket or pad inputs to a small set of shapes, or leave the flag off. Attention is covered in more depth in the FlashAttention article; for cuDNN, the point is that its fused attention is one more backend SDPA can select, so benchmark it against the others on your own shapes rather than assuming a winner.
Failure modes and how to diagnose them
Most cuDNN incidents fall into a few recognisable shapes.
- Graph not supported. Fused graphs are supported only for particular combinations of operation pattern, data type, layout, alignment and GPU architecture. A graph that builds on one GPU can be refused on another. Always include a fallback mode, and keep an unfused path you can switch to with a flag.
- Slow first iterations. Plan building, runtime compilation and autotuning all happen on first use of a shape. Warm up every shape before timing or serving traffic, and cache built graphs keyed by shape and dtype in your own code if you drive the frontend directly.
- Layout conversions. Tensor Core convolution kernels generally prefer channels-last data. Feed NCHW tensors and you can pay for transposes around every convolution. Convert the model and inputs to
torch.channels_lastonce and profile to confirm the transposes disappear. - Nondeterminism. Some backward algorithms accumulate with atomics, so summation order and the last bits of gradients vary between runs. That is expected; set
deterministiconly when you need bitwise repeatability. - Precision surprises. TF32 keeps a 10-bit mantissa. Models with large dynamic range in fp32 layers can drift; compare with
allow_tf32off before blaming the optimiser. The Tensor Core article explains the number formats. - Version and packaging mismatch. A framework wheel ships its own cuDNN; a system install or a second wheel can shadow it. Log
torch.backends.cudnn.version()and, for the frontend,cudnn.backend_version()at startup. - Workspace pressure. The fastest plan may want a large workspace. Under memory pressure frameworks fall back to slower plans that fit, which looks like an unexplained slowdown when batch size grows.
For diagnosis, the backend logs through CUDNN_LOGLEVEL_DBG and CUDNN_LOGDEST_DBG, and the frontend through CUDNN_FRONTEND_LOG_INFO and CUDNN_FRONTEND_LOG_FILE. The frontend's full logging level also dumps tensors, so use it on a reproducer, not a production job. Pair logs with a profiler trace: the kernel names it records tell you which engine actually ran.
Trade-offs
| Choice | You gain | You give up |
|---|---|---|
| Framework defaults | Zero code, tested paths | Fusion only where the framework already uses it |
| cuDNN graph API via the frontend | Vendor-tuned fused kernels for conv, norm, attention | Support depends on pattern, dtype and architecture; you own fallbacks |
| Compiler-generated kernels (torch.compile, Triton) | Arbitrary elementwise fusion, portable code | May not match cuDNN on convolutions |
| Hand-written kernels (CUTLASS) | Full control of tiling and epilogues | Engineering and maintenance cost per GPU generation |
benchmark=True | Measured best kernel per shape | Warm-up cost, repeated for every new shape |
A sound default: let the framework call cuDNN, turn on channels-last and benchmarking for fixed-shape convolutional work, and reach for the graph API only when a profile shows time going to unfused elementwise kernels around convolutions or normalisations that cuDNN can fuse. The kernel fusion article covers how to find those kernels in a trace.
What to do next
- Log the cuDNN version your framework actually loaded, on every node, at startup.
- Profile one training step and list the convolution, normalisation and attention kernels with their times; note any layout transposes and standalone elementwise kernels.
- Convert a convolutional model to channels-last and compare step time with the profile.
- Try
benchmark=Trueon fixed shapes and measure the warm-up cost against the steady-state gain. - Rebuild the conv, bias and ReLU example above, check it against PyTorch, then time it against the unfused three-kernel version on your GPU.
- Write down your fallback: which flag disables the fused path when a graph is not supported on new hardware.