cuDNN is NVIDIA's library of deep-learning primitives: convolutions, normalisations, attention, matrix multiplies with fused epilogues, and the backward passes of all of them. You rarely call it directly. When PyTorch runs torch.nn.functional.conv2d on a CUDA tensor, or picks the cuDNN backend for scaled dot-product attention, the work is handed to cuDNN, which chooses one of many kernels for that exact shape, dtype, layout and GPU. Most of what goes right or wrong with cuDNN in production comes from that choice: a slow first iteration, a run that is fast on one batch size and slow on the next, a gradient that differs bit-for-bit between two identical runs, or an error saying a graph is not supported.

This article explains the library from the point of view of the software that drives it. It covers the two APIs, how the cuDNN 9 library is packaged, how a graph of operations becomes an execution plan, and a worked example that fuses a convolution, a bias and a ReLU in the Python frontend and counts the memory traffic the fusion saves. It ends with the PyTorch switches that control cuDNN, the failure modes worth recognising on sight, and a checklist. Hardware detail appears only where it changes what the software should do.

Two APIs and the cuDNN 9 library layout

cuDNN has two programming models. The legacy API is imperative: you create descriptors for tensors, filters and a convolution, ask for an algorithm, and call a fixed entry point. It exposes a fixed set of operations and a handful of hard-wired fusions. The graph API, introduced in cuDNN 8, is declarative: you describe a small graph of operations (a convolution feeding a bias, an activation and a scale, say) and cuDNN finds engines that can execute the whole graph, often as a single kernel.

cuDNN 9 reorganised the shipped library around that split. The graph API lives in libcudnn_graph, and the implementations live in separately loaded sub-libraries: cudnn_engines_precompiled for kernels compiled ahead of time, cudnn_engines_runtime_compiled for kernels generated when a plan is built, and cudnn_heuristics for the models that rank candidate engines. The graph library loads these on demand. Legacy functionality sits in its own sub-libraries. The practical consequence is packaging: a container that copies only one shared object, or a wheel that bundles a different cuDNN from the system one, can fail when a sub-library is missing, not when the process starts.

Where cuDNN sits: from a framework call to a GPU kernelFramework opconv2d, SDPA, normcuDNN frontendC++ header / PythonPlan cachekeyed by shape+dtypehitlibcudnn_graphoperation graphmiss: describecudnn_heuristicsrank engine configsengines_precompiledshipped kernelsengines_runtime_compiledgenerated at buildqueryExecution planengine + knobs + workspaceordered listexecute(variant pack)store plan
The request path. A plan-cache hit skips everything below the frontend; a miss describes the graph, asks the heuristics for an ordered list of engine configs, builds the first one that supports the graph, and stores the resulting plan.

Graphs, engines and execution plans

The graph API has a small vocabulary, and every frontend call maps onto it.

  • Tensor: dimensions, strides, a data type and a unique id. Strides encode layout, so an NHWC tensor is a 4-D tensor in NCHW dimension order with strides [H*W*C, 1, W*C, C]. There is no separate layout flag.
  • Operation: a node such as convolution forward, data-gradient, weight-gradient, matmul, pointwise (add, multiply, ReLU and others), reduction, normalisation or attention.
  • Operation graph: the nodes and the virtual tensors that connect them. A virtual tensor never needs to reach global memory, which is where fusion savings come from.
  • Engine: one implementation strategy able to run some class of graphs. An engine has knobs (tile sizes, split factors, kernel configurations), and an engine with every knob set is an engine config.
  • Execution plan: a finalised engine config for one specific graph, with a known workspace size. Building a plan can cost milliseconds, and much more when it involves runtime compilation.
  • Variant pack: the device pointers for the graph's real tensors, plus the workspace, supplied on every execution.

Heuristics turn a graph into a ranked list of engine configs. The frontend exposes heuristic modes: mode A is the fast default, and heur_mode.FALLBACK returns engines that trade speed for broad support. Passing both, as in [heur_mode.A, heur_mode.FALLBACK], is the usual way to say "the best ranked plan, but never fail just because the ranked ones refuse this graph". Autotuning replaces the ranking with measurement: build several plans, time each on real data, keep the fastest. The ranking is a prediction; the timing is ground truth for the GPU the process is running on.

Worked example: fusing a convolution, bias and ReLU

Take a ResNet-style layer: a batch of 8 images, 64 input channels, 56 by 56 pixels, 32 output channels, a 3 by 3 filter, padding 1, half precision, followed by a per-channel bias and a ReLU. The Python frontend (pip install nvidia-cudnn-frontend, imported as cudnn) describes all three operations as one graph:

import cudnn
import torch

handle = cudnn.create_handle()
N, C, H, W, K = 8, 64, 56, 56, 32

graph = cudnn.pygraph(
    handle=handle,
    io_data_type=cudnn.data_type.HALF,
    intermediate_data_type=cudnn.data_type.FLOAT,  # virtual tensors stay fp32
    compute_data_type=cudnn.data_type.FLOAT,
)
# NHWC expressed through strides; dims are always listed N, C, H, W.
X = graph.tensor(name="X", dim=[N, C, H, W], stride=[H * W * C, 1, W * C, C])
Wt = graph.tensor(name="W", dim=[K, C, 3, 3], stride=[3 * 3 * C, 1, 3 * C, C])
B = graph.tensor(name="B", dim=[1, K, 1, 1], stride=[K, 1, K, K])

conv = graph.conv_fprop(X, Wt, padding=[1, 1], stride=[1, 1], dilation=[1, 1],
                        compute_data_type=cudnn.data_type.FLOAT)
biased = graph.bias(name="bias", input=conv, bias=B)
Y = graph.relu(name="relu", input=biased)
Y.set_output(True)  # only Y is materialised; conv and biased are virtual

graph.build([cudnn.heur_mode.A, cudnn.heur_mode.FALLBACK])

cl = torch.channels_last
x = torch.randn(N, C, H, W, device="cuda", dtype=torch.float16).to(memory_format=cl)
w = torch.randn(K, C, 3, 3, device="cuda", dtype=torch.float16).to(memory_format=cl)
b = torch.randn(1, K, 1, 1, device="cuda", dtype=torch.float16)
y = torch.empty(N, K, H, W, device="cuda", dtype=torch.float16).to(memory_format=cl)

workspace = torch.empty(graph.get_workspace_size(), device="cuda", dtype=torch.uint8)
graph.execute({X: x, Wt: w, B: b, Y: y}, workspace, handle=handle)
torch.cuda.synchronize()

ref = torch.relu(torch.nn.functional.conv2d(x, w, padding=1) + b)
torch.testing.assert_close(y, ref, atol=5e-3, rtol=3e-3)

graph.build is shorthand for the explicit sequence the frontend also exposes: validate the graph, build the backend operation graph, create candidate execution plans from the heuristic modes, check support, and build plans. Writing the steps out is useful when you want to filter engines or autotune yourself.

Now count what fusion buys. The convolution performs 2 x 8 x 32 x 56 x 56 x 64 x 9 = 924,844,032 floating-point operations, about 0.92 GFLOP. It reads 3,211,264 bytes of input and 36,864 bytes of weights and writes 1,605,632 bytes of output: 4,853,760 bytes in total, roughly 190 FLOP per byte, comfortably compute-bound on a modern GPU. Run bias and ReLU as two separate elementwise kernels and each one reads and writes the whole output again: 4 x 1,605,632 = 6,422,528 extra bytes, more traffic than the convolution itself needed, spent on two operations that do almost no arithmetic. In the fused graph those intermediate tensors are virtual: the epilogue applies bias and ReLU to values still in registers, and the output is written once.

Controlling cuDNN from PyTorch

Inside PyTorch you control cuDNN through a few process-wide switches in torch.backends.cudnn:

SettingWhat it doesWhen to change it
benchmarkTimes candidate algorithms the first time each new input configuration is seen and caches the winnerTurn on for fixed shapes (vision training); leave off when shapes vary every step
benchmark_limitCaps how many candidates the benchmark triesLower it if first-step autotuning is too slow
deterministicRestricts cuDNN to deterministic algorithmsReproducibility work, debugging divergence; expect some slowdown
allow_tf32Lets fp32 convolutions use TF32 Tensor Core math on GPUs that support itTurn off when a model is sensitive to the reduced mantissa
import torch
torch.backends.cudnn.benchmark = True        # autotune per new shape, then reuse
torch.backends.cudnn.deterministic = False   # set True when chasing nondeterminism
torch.backends.cudnn.allow_tf32 = True       # fp32 convs may use TF32

# Attention: ask SDPA for the cuDNN backend explicitly when comparing backends.
from torch.nn.attention import sdpa_kernel, SDPBackend
with sdpa_kernel(SDPBackend.CUDNN_ATTENTION):
    out = torch.nn.functional.scaled_dot_product_attention(q, k, v, is_causal=True)

The benchmark flag is the one that bites. Its cache is keyed by input configuration, so a model fed variable sequence lengths or ragged image sizes autotunes again for every new shape, and each of those steps pays for timing several kernels. Bucket or pad inputs to a small set of shapes, or leave the flag off. Attention is covered in more depth in the FlashAttention article; for cuDNN, the point is that its fused attention is one more backend SDPA can select, so benchmark it against the others on your own shapes rather than assuming a winner.

Failure modes and how to diagnose them

Most cuDNN incidents fall into a few recognisable shapes.

  • Graph not supported. Fused graphs are supported only for particular combinations of operation pattern, data type, layout, alignment and GPU architecture. A graph that builds on one GPU can be refused on another. Always include a fallback mode, and keep an unfused path you can switch to with a flag.
  • Slow first iterations. Plan building, runtime compilation and autotuning all happen on first use of a shape. Warm up every shape before timing or serving traffic, and cache built graphs keyed by shape and dtype in your own code if you drive the frontend directly.
  • Layout conversions. Tensor Core convolution kernels generally prefer channels-last data. Feed NCHW tensors and you can pay for transposes around every convolution. Convert the model and inputs to torch.channels_last once and profile to confirm the transposes disappear.
  • Nondeterminism. Some backward algorithms accumulate with atomics, so summation order and the last bits of gradients vary between runs. That is expected; set deterministic only when you need bitwise repeatability.
  • Precision surprises. TF32 keeps a 10-bit mantissa. Models with large dynamic range in fp32 layers can drift; compare with allow_tf32 off before blaming the optimiser. The Tensor Core article explains the number formats.
  • Version and packaging mismatch. A framework wheel ships its own cuDNN; a system install or a second wheel can shadow it. Log torch.backends.cudnn.version() and, for the frontend, cudnn.backend_version() at startup.
  • Workspace pressure. The fastest plan may want a large workspace. Under memory pressure frameworks fall back to slower plans that fit, which looks like an unexplained slowdown when batch size grows.

For diagnosis, the backend logs through CUDNN_LOGLEVEL_DBG and CUDNN_LOGDEST_DBG, and the frontend through CUDNN_FRONTEND_LOG_INFO and CUDNN_FRONTEND_LOG_FILE. The frontend's full logging level also dumps tensors, so use it on a reproducer, not a production job. Pair logs with a profiler trace: the kernel names it records tell you which engine actually ran.

Trade-offs

ChoiceYou gainYou give up
Framework defaultsZero code, tested pathsFusion only where the framework already uses it
cuDNN graph API via the frontendVendor-tuned fused kernels for conv, norm, attentionSupport depends on pattern, dtype and architecture; you own fallbacks
Compiler-generated kernels (torch.compile, Triton)Arbitrary elementwise fusion, portable codeMay not match cuDNN on convolutions
Hand-written kernels (CUTLASS)Full control of tiling and epiloguesEngineering and maintenance cost per GPU generation
benchmark=TrueMeasured best kernel per shapeWarm-up cost, repeated for every new shape

A sound default: let the framework call cuDNN, turn on channels-last and benchmarking for fixed-shape convolutional work, and reach for the graph API only when a profile shows time going to unfused elementwise kernels around convolutions or normalisations that cuDNN can fuse. The kernel fusion article covers how to find those kernels in a trace.

What to do next

  1. Log the cuDNN version your framework actually loaded, on every node, at startup.
  2. Profile one training step and list the convolution, normalisation and attention kernels with their times; note any layout transposes and standalone elementwise kernels.
  3. Convert a convolutional model to channels-last and compare step time with the profile.
  4. Try benchmark=True on fixed shapes and measure the warm-up cost against the steady-state gain.
  5. Rebuild the conv, bias and ReLU example above, check it against PyTorch, then time it against the unfused three-kernel version on your GPU.
  6. Write down your fallback: which flag disables the fused path when a graph is not supported on new hardware.
Key takeaway: cuDNN chooses a kernel per shape, dtype, layout and GPU, and the graph API lets it fuse whole subgraphs so intermediate tensors never touch global memory; in the worked example that avoids more traffic than the convolution itself needs. Feed it channels-last data, warm up every shape, keep a fallback for unsupported graphs, and use the benchmark and determinism switches deliberately.