A cold plate looks simple: a copper block with channels, a lid, two fittings. Yet every step of how it was made shows up later in your telemetry. A braze void, a bent fin, a base warped in the furnace or debris left in a channel raises one plate's thermal resistance. That GPU runs hotter, reaches its clock limits sooner, and in a synchronous training job sets the pace for every other rank.
This article covers how plates are made, what each process can get wrong, what to test, and how to measure plate quality from software so a bad batch is found in burn-in. Direct liquid cooling explains how a finished plate works; this page looks upstream.
Why manufacturing shows up in telemetry
The heat path from die to coolant is a chain of resistances in series: the die and its lid, the thermal interface material (TIM), the plate base, the fin surface into the fluid, and the coolant's own temperature rise. A useful single number for one GPU is its effective thermal resistance:
R_eff = (T_gpu - T_coolant_in) / P_gpu # kelvin per watt, at steady stateManufacturing sets several of those links: base thickness and material set spreading, fin geometry sets convection, and contact-face flatness sets the TIM thickness after mounting. A partly blocked channel also cuts the plate's share of flow on a shared manifold.
None of this is visible on the part. Two plates from the same drawing can differ by a few kelvin at full power, which can separate a GPU that holds its clocks from one that throttles.
Why manufacturing shows up in telemetry
The heat path from die to coolant is a chain of resistances in series: the die and its lid, the thermal interface material (TIM), the plate base, the fin surface into the fluid, and the coolant's own temperature rise. A useful single number for one GPU is its effective thermal resistance:
R_eff = (T_gpu - T_coolant_in) / P_gpu # kelvin per watt, at steady stateManufacturing sets several of those links: base thickness and material set spreading, fin geometry sets convection, and contact-face flatness sets the TIM thickness after mounting. A partly blocked channel also cuts the plate's share of flow on a shared manifold.
None of this is visible on the part. Two plates from the same drawing can differ by a few kelvin at full power, which can separate a GPU that holds its clocks from one that throttles.
Materials and the base
Most GPU cold plates are copper: its thermal conductivity, close to 400 W/m·K, is about twice that of common aluminium alloys, which matters in the base, where heat from a small die must spread sideways to a much larger fin area. Oxygen-free grades are preferred for brazing and welding, because oxygen in copper can cause embrittlement at joining temperatures.
Aluminium plates are lighter and cheaper, but mixing metals in one loop invites galvanic corrosion. Check every plate, hose and fitting against the CDU vendor's wetted-materials list; many operators simply keep aluminium out of the loop.
The copper's temper changes during manufacture. Brazing and diffusion bonding anneal the whole part, so hard copper leaves the furnace soft, and a soft base can bow under the mounting load and thicken the TIM at the edges. A change of joining process can change a plate's mechanical behaviour even when the drawing is identical.
Making the fins
The fin structure does most of the work, and it is made in one of several ways. Each leaves its own signature in quality data.
| Process | How it works | Strengths | Typical defects |
|---|---|---|---|
| Skiving | A blade slices thin layers from a copper block and bends each up into a fin | One piece with the base, no joint resistance; fine pitch | Bent or torn fins, burrs |
| CNC milling or gang-saw slitting | Channels are cut with end mills or stacked slitting saws | Flexible geometry | Tool wear drifts channel width; chips left behind |
| Stamped or folded fins | A fin strip is formed, then brazed to the base | Cheap at volume | Incomplete fin-to-base joints |
| Etched sheets, diffusion bonded | Etched copper sheets are bonded in a stack under heat and pressure | Fine, complex channels | Unbonded regions |
| Additive manufacturing | Laser powder bed fusion builds channels layer by layer | Shapes no cutter can reach | Porosity, rough walls, trapped powder |
Copper reflects most infrared laser light, so many printers use green or blue lasers for it, and loose powder left inside is a cleanliness problem as much as a geometry one.
Closing the plate
A lid closes the fins into a sealed volume with an inlet and outlet. The joint must stay leak-free through years of pressure and thermal cycling.
| Method | What it is | Watch for |
|---|---|---|
| Vacuum brazing | Filler alloy melts in a vacuum furnace and wicks into the joint; no flux | Joint voids, filler flowing into fins, annealed base |
| Friction stir welding | A rotating tool stirs the parts together below melting point; no filler | Incomplete penetration, seam placement near channels |
| Diffusion bonding | Flat surfaces held at heat and pressure until they fuse | Unbonded patches visible only by ultrasound or CT |
| Gasketed lid | An O-ring and screws seal the lid | Every gasket is a potential leak; depends on torque |
Flux residue inside a sealed channel cannot be inspected and slowly attacks the coolant's inhibitors, which is one reason flux-free brazing and welding dominate. Whatever the method, the joint is the first suspect when a plate leaks after thermal cycling.
The contact face
The face that touches the GPU package is finished after joining, because brazing and welding distort the part. It is machined or lapped flat and usually plated, commonly with nickel, to resist corrosion and keep the surface stable. Flatness is specified across the contact area, and roughness is specified too, because both decide how thin the TIM can get.
The TIM layer is often the largest single resistance in the chain, and its thickness is set by the gap between two imperfect surfaces. A face that is slightly concave leaves a thick layer in the middle, exactly over the hottest part of the die. A face that is convex leaves the edges thin and the corners of the package, where HBM stacks often sit, poorly covered. Memory temperature then rises before core temperature does, which looks like a memory problem rather than a plate problem unless you record both. Uneven plating can also undo the flatness achieved by machining.
Cleanliness
Fine channels are excellent filters. Machining chips, skiving burrs, braze spatter, printing powder and plating flakes all end up in the narrowest passage, usually inside a cold plate. A partly blocked plate gets less flow and runs hotter, so the symptom is one hot GPU rather than a facility alarm.
Factories use ultrasonic cleaning, high-velocity flushing and a particle count of the flush fluid, and ship plates capped. At your end, the CDU filter is the last line of defence; its differential pressure is telemetry worth graphing: a filter whose pressure drop climbs quickly after commissioning is catching manufacturing debris, which tells you something about the batch of plates just installed. Cleanliness limits differ by supplier and loop design, so take them from your CDU and server vendors and write them into the purchase specification.
Factory tests
A plate should not leave the factory without a record of these tests, and you should be able to see the record for every serial number you receive.
| Test | What it catches | Notes |
|---|---|---|
| Pressure decay or helium leak test | Leaks, porosity | Helium finds far smaller leaks |
| Proof pressure test | Weak joints | Above maximum working pressure; burst tests on samples |
| Flow versus pressure drop | Blocked or narrowed channels | Compared with the design's golden curve |
| Thermal test with a heater | High resistance from any cause | A thermal test vehicle mimics the die |
| Flatness measurement | Warped contact face | After final machining and plating |
| CT, X-ray or cross-section | Voids, unbonded areas, braze in fins | Sampled per furnace run or lot |
The flow test is the cheapest way to find internal blockage; repeat it on a sample at incoming inspection. High pressure drop at a fixed flow means something is in the channels, whatever the leak test said.
Traceability from serial to slot
Traceability turns a hot GPU into evidence. Each plate should carry a permanent serial linked to its copper lot, machining batch, joining run and test results. At server assembly, record the plate serial against the server serial and GPU slot, plus the TIM lot and torque record if the line captures them.
Without that join you can see that some GPUs run hot but not whether they share a plate lot, a TIM lot or a chassis position. Ask for the data in the purchase contract and store it beside the GPU-to-host map, so one query goes from a hot GPU to the furnace run that made its plate.
Measuring plates from software
Measure each plate's effective resistance during burn-in: run a steady load on every GPU, let temperatures settle, then sample GPU temperature, power and coolant inlet temperature. GPU temperature and power come from nvidia-smi or DCGM (field DCGM_FI_DEV_GPU_TEMP and DCGM_FI_DEV_POWER_USAGE). Inlet temperature does not come from the GPU; read it from the CDU, the rack manifold sensor or the server BMC. Log the memory temperature too where the GPU reports it, for the edge-contact case described above.
# one sample per second per GPU while the burn-in load runs
nvidia-smi --query-gpu=timestamp,serial,temperature.gpu,power.draw \
--format=csv,noheader,nounits -l 1 >> gpu_$(hostname).csvJoin the samples with inlet temperatures and the inventory map into one CSV with columns host, gpu_serial, slot, plate_lot, t, gpu_temp_c, power_w, inlet_c, where t is seconds since the load started, then compute a steady-state resistance per GPU, and find outliers with the median and median absolute deviation, so a few bad plates cannot hide by inflating the spread.
import csv, statistics, sys
from collections import defaultdict
def steady_resistance(rows, settle_s=600, min_w=300):
"""Median (T_gpu - T_inlet) / P per GPU, skipping warm-up and idle samples."""
per_gpu, meta = defaultdict(list), {}
for r in rows:
t, p = float(r["t"]), float(r["power_w"])
if t >= settle_s and p >= min_w:
rise = float(r["gpu_temp_c"]) - float(r["inlet_c"])
per_gpu[r["gpu_serial"]].append(rise / p)
meta[r["gpu_serial"]] = (r["host"], r["plate_lot"], r["slot"])
return {g: (statistics.median(v), *meta[g]) for g, v in per_gpu.items() if len(v) >= 30}
def report(res, z_limit=3.5):
med = statistics.median(v[0] for v in res.values())
mad = statistics.median(abs(v[0] - med) for v in res.values()) or 1e-9
z = {g: (res[g][0] - med) / (1.4826 * mad) for g in res}
flagged = sorted((g for g in res if z[g] > z_limit), key=lambda g: -z[g])
print(f"fleet median R = {med * 1000:.1f} mK/W over {len(res)} GPUs")
for g in flagged:
r, host, lot, slot = res[g]
print(f" FLAG {host} slot={slot} {g} lot={lot} R={r * 1000:.1f} mK/W z={z[g]:.1f}")
for lot in sorted({v[2] for v in res.values()}):
gs = [g for g in res if res[g][2] == lot]
bad = sum(g in flagged for g in gs) / len(gs)
lot_med = statistics.median(res[g][0] for g in gs)
print(f" lot {lot}: n={len(gs)} median={lot_med * 1000:.1f} mK/W flagged={bad:.0%}")
with open(sys.argv[1], newline="") as f:
report(steady_resistance(list(csv.DictReader(f))))Use a fixed load, not a training job: resistance is only comparable when every GPU carries similar power with a similar power map across the die.
Worked example: a burn-in with two lots
Here is the script's output on a synthetic burn-in of eight nodes with eight GPUs each: five nodes with plates from lot L2407 and three from lot L2411.
fleet median R = 45.3 mK/W over 64 GPUs
FLAG node01 slot=3 GPU-13 lot=L2407 R=59.8 mK/W z=8.7
FLAG node05 slot=5 GPU-55 lot=L2411 R=53.1 mK/W z=4.7
FLAG node06 slot=2 GPU-62 lot=L2411 R=52.9 mK/W z=4.5
FLAG node07 slot=5 GPU-75 lot=L2411 R=52.5 mK/W z=4.3
FLAG node07 slot=2 GPU-72 lot=L2411 R=52.2 mK/W z=4.1
FLAG node06 slot=5 GPU-65 lot=L2411 R=51.6 mK/W z=3.7
lot L2407: n=40 median=45.2 mK/W flagged=2%
lot L2411: n=24 median=45.5 mK/W flagged=21%First, node01 GPU-13 is a lone outlier, about a third above the fleet median: at 700 W, roughly 10 K hotter for the same inlet. A single outlier in a good lot points to assembly, so reseat the plate and check the TIM spread and torque before blaming the plate.
Second, lot L2411 has almost the same median as L2407, yet a fifth of its GPUs are flagged, all in slots 2 and 5. A lot median hides a sub-population, so always read the slot column too. A slot pattern within one lot suggests a plate variant or fixture used for those positions. Keep those nodes away from straggler-sensitive jobs, and send the serials and measurements to the supplier with a request for that lot's joining and flow-test records.
Failure modes
- Braze voids: local extra resistance; the GPU runs hot but passes every leak test.
- Channel narrowing: filler or debris raises pressure drop and starves the plate on a parallel manifold.
- Warped or soft base: thick TIM at the centre or edges; memory temperature may rise before core temperature.
- Leaks after thermal cycling: joints that passed at the factory open after many heat cycles; a leak detection cable or tray sensor gives the first signal.
- Galvanic corrosion: mixed metals in one loop slowly dissolve the less noble metal and shed particles into every plate downstream.
- Erosion: high local velocity at an impingement jet wears the base; resistance creeps up fleet-wide rather than on one plate.
Trade-offs
Skived fins are cheap and joint-free but limited in geometry; bonded or printed channels allow complex shapes at higher cost. Vacuum brazing gives strong flux-free joints but anneals the copper; friction stir welding keeps the base hard but constrains the seam. Per-plate thermal testing catches more than sampling but costs line time. A burn-in resistance check costs hours of idle hardware, small next to a straggling GPU in a large training job.
What to do next
- Ask suppliers for the fin and joining process, per-plate test list and sample records.
- Record plate serial, lot, server serial, GPU slot and TIM lot at receipt.
- Add the fixed-load resistance check to node acceptance, grouped by lot and slot; re-run it after every reseat.
- Graph CDU filter pressure drop from commissioning onward; a fast rise means debris.
- Re-measure resistance quarterly to catch fouling or erosion trends; read liquid cooling loop telemetry and GPU thermal management for the alarms and throttle behaviour that follow.
- If you are still choosing a cooling approach, compare with immersion cooling and check the rack design in the NVL72 rack.