A GPU that gets too hot does not usually crash. It slows down, quietly, by lowering its clock. For a single inference server that costs a few percent of throughput. For a synchronous training job across hundreds of GPUs it can cost far more, because every step waits for the slowest GPU, and a single poorly cooled card sets the pace for the whole cluster while every dashboard shows 100% utilisation.
This article explains the chip-level loop that sits between cooling and performance: how power becomes heat, how heat flows out, how the GPU's firmware chooses a clock, which counters tell you why it chose that clock, and what to do about it. Facility cooling, from air to cold plates and immersion, is covered in GPU datacenter cooling, and electrical sizing in GPU datacenter power. Here the focus is the individual GPU and the job running on it.
Heat is power: where the watts come from
Essentially every watt a GPU draws ends up as heat in the package, so thermal management starts with power. Power has two parts. Dynamic power is spent switching transistors and scales roughly as capacitance times voltage squared times frequency. Because higher clocks need higher voltage, the top of the clock range is disproportionately expensive: the last few hundred megahertz can cost a large share of the power. Leakage power flows even when transistors are not switching and rises with temperature, so a hotter chip draws more power at the same clock, which makes it hotter still. That feedback is why cooling affects performance even before any temperature limit is reached: leakage eats into the power budget the firmware would otherwise spend on clock speed.
Workload matters as much as the chip. Dense matrix multiplies on tensor cores draw far more power than memory-bound kernels or kernels waiting on communication. A training step therefore swings between high-power and lower-power phases many times a second, and so does the heat.
The heat path and its time constants
Heat flows from the transistors (the junction) through the die, a thermal interface material, the package lid, another interface and into a heatsink or cold plate, and from there into air or liquid. Each layer has a thermal resistance, measured in degrees per watt. In steady state the junction sits at roughly the coolant inlet temperature plus power times the total resistance. With an illustrative total of 0.06 °C per watt, a 700 W load runs 42 °C above the inlet: with 25 °C inlet air that is 67 °C, with 35 °C inlet air 77 °C. Every degree of inlet temperature shows up at the junction, which is why a hot aisle leaking into a cold aisle matters.
The layers also respond at different speeds. The die and package heat up in milliseconds to seconds, a heatsink or cold plate over tens of seconds, and the air in a rack or the water in a loop over minutes. Two practical consequences follow. A short benchmark measures a cool GPU and overstates sustained throughput; run anything you intend to compare for at least fifteen to twenty minutes so the system reaches steady state. And a node can be fine for the first quarter hour of a job and then start throttling as it heat-soaks, which looks like a mysterious slowdown unless you keep time series.
The clock-management loop and its reasons
The GPU does not run at a fixed clock. Firmware and driver continually choose the highest clock and voltage that keep several quantities under their limits: board power under the enforced power limit, GPU temperature under its maximum operating temperature, memory temperature under its own limit, and hard protection thresholds well above those. When the clock is lower than it could be, the GPU records why, and NVIDIA exposes those flags as clock event reasons (older tools and docs call them clock throttle reasons; the old names still work as aliases).
| Reason | What it means | Usual cause |
|---|---|---|
| SW power cap | Power draw has reached the enforced limit; clocks are reduced to hold it | Heavy tensor-core work; normal under full load, but also a hotter chip leaking more |
| SW thermal slowdown | Clocks are reduced to keep GPU or memory temperature from exceeding its max operating temperature | Insufficient cooling: inlet too hot, airflow blocked, fan or pump degraded |
| HW thermal slowdown | Hardware protection reacting to temperature; large, abrupt clock cuts | Severe cooling failure; investigate immediately |
| HW power brake slowdown | An external signal from the platform asked the GPU to slow down | Power supply or chassis-level event, not the GPU itself |
| HW slowdown | Any hardware-initiated slowdown, including the two above | Read together with the specific flags |
The actual thresholds depend on the board, so do not copy numbers from a forum. Run nvidia-smi -q -d TEMPERATURE on the machine to see its shutdown, slowdown and maximum operating temperatures, including a memory maximum where the board reports one. High-bandwidth memory sits on the same package as the GPU and has its own limit; on many training GPUs the memory reaches its limit before the core does, so watch temperature.memory as well. Some boards report N/A for it.
SW power cap being active is not by itself a problem: a well-cooled GPU running dense matrix multiplies will often sit at its power limit. It becomes a thermal problem when one GPU hits the power cap at a lower clock than its neighbours doing the same work, because it is hotter and leaking more.
Why one hot GPU slows the whole job
In synchronous data-parallel or tensor-parallel training, each step ends with a collective operation that every rank must reach before any can continue. The step time is the time of the slowest rank plus communication. Suppose a step has 400 ms of compute and 100 ms of non-overlapped communication on 64 GPUs. If one GPU throttles to 80% of the others' clock on compute-bound kernels, its compute takes about 500 ms, and so does everyone's step: 600 ms instead of 500, a 20% loss for the whole job from one card (illustrative numbers, but the arithmetic is general). The other 63 GPUs spend that time waiting inside the collective, and their utilisation counters still read 100%, because a kernel waiting on the network counts as busy. nvidia-smi explains why GPU-Util is a duty cycle and not a measure of useful work.
This is why thermal health in training clusters is judged per GPU and against peers, not against an absolute temperature. A GPU running at 83 °C with full clocks is fine; one at 75 °C with clocks 15% below its neighbours is the problem.
Measuring it
Three sources give you what you need. nvidia-smi in query mode gives a quick per-second view on one node:
# One line per GPU per second: temperatures, power against the enforced limit,
# SM clock against its maximum, and the clock event reasons that matter.
nvidia-smi --query-gpu=timestamp,index,temperature.gpu,temperature.memory,power.draw,enforced.power.limit,clocks.sm,clocks.max.sm,clocks_event_reasons.sw_power_cap,clocks_event_reasons.sw_thermal_slowdown,clocks_event_reasons.hw_thermal_slowdown,clocks_event_reasons.hw_power_brake_slowdown --format=csv -l 1
# The per-SKU limits for THIS board (shutdown, slowdown, max operating temps)
nvidia-smi -q -d TEMPERATURE
# Current clock event reasons in readable form
nvidia-smi -q -d PERFORMANCEFor a script you can schedule or run beside a job, query NVML directly. The snippet below samples every GPU once a second for ten minutes and reports, per GPU, the share of samples in which each reason was active and the mean SM clock as a fraction of its maximum. It looks up the reason constants by name, so it works with both the current and the older binding names, and it asserts no bit values.
import time
from collections import Counter
import pynvml as N # pip install nvidia-ml-py
N.nvmlInit()
def reason(name):
# Current "ClocksEvent" names first, deprecated "ClocksThrottle" aliases for old bindings.
return getattr(N, "nvmlClocksEventReason" + name, None) or getattr(N, "nvmlClocksThrottleReason" + name)
REASONS = {k: reason(k) for k in
["SwPowerCap", "SwThermalSlowdown", "HwThermalSlowdown", "HwPowerBrakeSlowdown", "HwSlowdown"]}
get_reasons = (getattr(N, "nvmlDeviceGetCurrentClocksEventReasons", None)
or N.nvmlDeviceGetCurrentClocksThrottleReasons)
handles = [N.nvmlDeviceGetHandleByIndex(i) for i in range(N.nvmlDeviceGetCount())]
seen = [Counter() for _ in handles]
clock_ratio = [[] for _ in handles]
SAMPLES = 600 # ten minutes at 1 Hz, under real load
for _ in range(SAMPLES):
for i, h in enumerate(handles):
mask = get_reasons(h)
seen[i].update(k for k, bit in REASONS.items() if mask & bit)
clock_ratio[i].append(N.nvmlDeviceGetClockInfo(h, N.NVML_CLOCK_SM)
/ N.nvmlDeviceGetMaxClockInfo(h, N.NVML_CLOCK_SM))
time.sleep(1)
for i, h in enumerate(handles):
temp = N.nvmlDeviceGetTemperature(h, N.NVML_TEMPERATURE_GPU)
mean_clock = sum(clock_ratio[i]) / SAMPLES
shares = {k: round(v / SAMPLES, 2) for k, v in seen[i].items()}
print(f"gpu{i} temp={temp}C mean_sm_clock={mean_clock:.0%} of max reasons={shares}")The flags are instantaneous, so one sample per second misses short events; report shares over many samples, not single readings. For clusters, DCGM and dcgm-exporter collect temperature, power and clocks centrally into Prometheus. Check the exporter's counters file to confirm the clock event reasons field is enabled, since not every field is collected by default.
Worked example: the node that slowed down after twenty minutes
A team notices that a 256-GPU training run is about 9% slower than the same configuration last week. Per-rank timing shows the slowdown starts about twenty minutes into each restart and that the slowest rank is always on one node. On that node, the NVML script reports mean SM clocks around 98% of maximum on GPUs 0 to 5 and around 85% on GPUs 6 and 7, with SW thermal slowdown active in roughly 40% of samples on those two, and memory temperature at its maximum operating value. GPU temperatures on the other six are a few degrees cooler and show only the SW power cap reason. (Numbers illustrative.)
The pattern, two adjacent GPUs at the rear of the chassis, onset after heat soak, thermal rather than power reasons, points to airflow, not silicon. The baseboard management controller shows one chassis fan running well below its siblings, and a blanking panel above the node is missing, letting hot exhaust recirculate into the front. The team drains the node, replaces the fan and fits the panel, and the reasons return to SW power cap only. Before the repair, there was a useful stopgap: capping power on all eight GPUs of that node slightly lower made them throttle evenly and cooler, which cost less than having two stragglers. The node health check now includes a twenty-minute burn-in with the same script, and it fails any GPU whose mean clock falls more than 5% below the node median.
Operational levers and their trade-offs
| Lever | Effect | Trade-off |
|---|---|---|
| Lower inlet temperature, fix airflow | Every degree removed at the inlet comes off the junction | Facility cost; blanking panels and containment are cheap |
Power limit (nvidia-smi -pl, needs root) | Less heat and steadier clocks; often better work per joule | Peak throughput drops; setting typically resets on reboot, so apply it from provisioning |
Lock clocks (nvidia-smi -lgc) | Identical clocks across GPUs; removes jitter in benchmarks and stragglers | Leaves headroom unused on cool GPUs |
| Fan or pump policy via the BMC | More cooling at the cost of noise and fan power | Datacenter GPUs are often passively cooled; the chassis controls the fans |
| Coolant supply temperature and flow (liquid) | Directly sets the base of the heat path | Warmer water saves facility energy but raises junction temperatures |
| Scheduling | Drain or down-rank nodes that fail thermal checks | Capacity loss; needs automation to be practical |
Power capping deserves emphasis. Because the last increments of clock cost disproportionate power, a modest cap often costs much less throughput than it saves in power and heat. The exact curve depends on the GPU and the workload, so measure it: run the same training step at three or four power limits for twenty minutes each and plot throughput against average power.
Failure modes
- Averaging away the signal. Cluster-wide average temperature looks healthy while one GPU throttles. Alert on per-GPU clock against the node median.
- Confusing power cap with thermal. SW power cap under heavy load is normal; SW thermal and HW thermal are cooling problems.
- Short benchmarks. Five-minute tests miss heat soak; burn in for at least fifteen to twenty minutes.
- Trusting GPU-Util. It reads 100% while ranks wait for a straggler.
- Ignoring memory temperature. HBM can hit its limit first; collect it where the board reports it.
- Power limits that silently reset. A reboot restores the default; enforce it from provisioning and verify with
enforced.power.limit. - Physical decay. Dust, failing fans, degraded thermal interface material and pump wear make a node slowly worse; trend clocks per GPU over weeks.
What to do next
- Run
nvidia-smi -q -d TEMPERATUREon each GPU model you operate and record its limits. - Collect per-GPU temperature, memory temperature, power, SM clock and the clock event reasons at one-second resolution into your metrics system.
- Alert on SW or HW thermal slowdown and on any GPU whose clock is more than 5% below its node median under load.
- Add a twenty-minute burn-in with the NVML script to node acceptance and periodic health checks.
- Measure throughput against power limit for your main workload and choose a cap deliberately.
- Walk the racks: blanking panels, fan health, inlet temperatures, and for liquid loops, supply temperature and flow.