A cold plate looks simple: a copper block with channels, a lid, two fittings. Yet every step of how it was made shows up later in your telemetry. A braze void, a bent fin, a base warped in the furnace or debris left in a channel raises one plate's thermal resistance. That GPU runs hotter, reaches its clock limits sooner, and in a synchronous training job sets the pace for every other rank.

This article covers how plates are made, what each process can get wrong, what to test, and how to measure plate quality from software so a bad batch is found in burn-in. Direct liquid cooling explains how a finished plate works; this page looks upstream.

From copper bar to fleet telemetry, and the feedback loop back to the factoryCopper stocklot and temper recordedFin formingskive, mill, etch, printLid joiningbraze, FSW, bondFinishflatten, plateCleanflush, particle countFactory testsleak, proof, flow, thermalSerial recordevery result per plateServer buildplate serial to GPU slotBurn-insteady load, sample T and PR = (T_gpu - T_in) / Pper GPU, steady stateFleet analyticsoutliers grouped by lot and slota bad lot found in the fleet goes back to the supplier as evidence
The manufacturing chain of a cold plate and the software measurement that closes the loop. Each plate serial is joined to its GPU slot so fleet outliers can be traced to a lot.

Why manufacturing shows up in telemetry

The heat path from die to coolant is a chain of resistances in series: the die and its lid, the thermal interface material (TIM), the plate base, the fin surface into the fluid, and the coolant's own temperature rise. A useful single number for one GPU is its effective thermal resistance:

R_eff = (T_gpu - T_coolant_in) / P_gpu        # kelvin per watt, at steady state

Manufacturing sets several of those links: base thickness and material set spreading, fin geometry sets convection, and contact-face flatness sets the TIM thickness after mounting. A partly blocked channel also cuts the plate's share of flow on a shared manifold.

None of this is visible on the part. Two plates from the same drawing can differ by a few kelvin at full power, which can separate a GPU that holds its clocks from one that throttles.

Why manufacturing shows up in telemetry

The heat path from die to coolant is a chain of resistances in series: the die and its lid, the thermal interface material (TIM), the plate base, the fin surface into the fluid, and the coolant's own temperature rise. A useful single number for one GPU is its effective thermal resistance:

R_eff = (T_gpu - T_coolant_in) / P_gpu        # kelvin per watt, at steady state

Manufacturing sets several of those links: base thickness and material set spreading, fin geometry sets convection, and contact-face flatness sets the TIM thickness after mounting. A partly blocked channel also cuts the plate's share of flow on a shared manifold.

None of this is visible on the part. Two plates from the same drawing can differ by a few kelvin at full power, which can separate a GPU that holds its clocks from one that throttles.

Materials and the base

Most GPU cold plates are copper: its thermal conductivity, close to 400 W/m·K, is about twice that of common aluminium alloys, which matters in the base, where heat from a small die must spread sideways to a much larger fin area. Oxygen-free grades are preferred for brazing and welding, because oxygen in copper can cause embrittlement at joining temperatures.

Aluminium plates are lighter and cheaper, but mixing metals in one loop invites galvanic corrosion. Check every plate, hose and fitting against the CDU vendor's wetted-materials list; many operators simply keep aluminium out of the loop.

The copper's temper changes during manufacture. Brazing and diffusion bonding anneal the whole part, so hard copper leaves the furnace soft, and a soft base can bow under the mounting load and thicken the TIM at the edges. A change of joining process can change a plate's mechanical behaviour even when the drawing is identical.

Making the fins

The fin structure does most of the work, and it is made in one of several ways. Each leaves its own signature in quality data.

ProcessHow it worksStrengthsTypical defects
SkivingA blade slices thin layers from a copper block and bends each up into a finOne piece with the base, no joint resistance; fine pitchBent or torn fins, burrs
CNC milling or gang-saw slittingChannels are cut with end mills or stacked slitting sawsFlexible geometryTool wear drifts channel width; chips left behind
Stamped or folded finsA fin strip is formed, then brazed to the baseCheap at volumeIncomplete fin-to-base joints
Etched sheets, diffusion bondedEtched copper sheets are bonded in a stack under heat and pressureFine, complex channelsUnbonded regions
Additive manufacturingLaser powder bed fusion builds channels layer by layerShapes no cutter can reachPorosity, rough walls, trapped powder

Copper reflects most infrared laser light, so many printers use green or blue lasers for it, and loose powder left inside is a cleanliness problem as much as a geometry one.

Closing the plate

A lid closes the fins into a sealed volume with an inlet and outlet. The joint must stay leak-free through years of pressure and thermal cycling.

MethodWhat it isWatch for
Vacuum brazingFiller alloy melts in a vacuum furnace and wicks into the joint; no fluxJoint voids, filler flowing into fins, annealed base
Friction stir weldingA rotating tool stirs the parts together below melting point; no fillerIncomplete penetration, seam placement near channels
Diffusion bondingFlat surfaces held at heat and pressure until they fuseUnbonded patches visible only by ultrasound or CT
Gasketed lidAn O-ring and screws seal the lidEvery gasket is a potential leak; depends on torque

Flux residue inside a sealed channel cannot be inspected and slowly attacks the coolant's inhibitors, which is one reason flux-free brazing and welding dominate. Whatever the method, the joint is the first suspect when a plate leaks after thermal cycling.

The contact face

The face that touches the GPU package is finished after joining, because brazing and welding distort the part. It is machined or lapped flat and usually plated, commonly with nickel, to resist corrosion and keep the surface stable. Flatness is specified across the contact area, and roughness is specified too, because both decide how thin the TIM can get.

The TIM layer is often the largest single resistance in the chain, and its thickness is set by the gap between two imperfect surfaces. A face that is slightly concave leaves a thick layer in the middle, exactly over the hottest part of the die. A face that is convex leaves the edges thin and the corners of the package, where HBM stacks often sit, poorly covered. Memory temperature then rises before core temperature does, which looks like a memory problem rather than a plate problem unless you record both. Uneven plating can also undo the flatness achieved by machining.

Cleanliness

Fine channels are excellent filters. Machining chips, skiving burrs, braze spatter, printing powder and plating flakes all end up in the narrowest passage, usually inside a cold plate. A partly blocked plate gets less flow and runs hotter, so the symptom is one hot GPU rather than a facility alarm.

Factories use ultrasonic cleaning, high-velocity flushing and a particle count of the flush fluid, and ship plates capped. At your end, the CDU filter is the last line of defence; its differential pressure is telemetry worth graphing: a filter whose pressure drop climbs quickly after commissioning is catching manufacturing debris, which tells you something about the batch of plates just installed. Cleanliness limits differ by supplier and loop design, so take them from your CDU and server vendors and write them into the purchase specification.

Factory tests

A plate should not leave the factory without a record of these tests, and you should be able to see the record for every serial number you receive.

TestWhat it catchesNotes
Pressure decay or helium leak testLeaks, porosityHelium finds far smaller leaks
Proof pressure testWeak jointsAbove maximum working pressure; burst tests on samples
Flow versus pressure dropBlocked or narrowed channelsCompared with the design's golden curve
Thermal test with a heaterHigh resistance from any causeA thermal test vehicle mimics the die
Flatness measurementWarped contact faceAfter final machining and plating
CT, X-ray or cross-sectionVoids, unbonded areas, braze in finsSampled per furnace run or lot

The flow test is the cheapest way to find internal blockage; repeat it on a sample at incoming inspection. High pressure drop at a fixed flow means something is in the channels, whatever the leak test said.

Traceability from serial to slot

Traceability turns a hot GPU into evidence. Each plate should carry a permanent serial linked to its copper lot, machining batch, joining run and test results. At server assembly, record the plate serial against the server serial and GPU slot, plus the TIM lot and torque record if the line captures them.

Without that join you can see that some GPUs run hot but not whether they share a plate lot, a TIM lot or a chassis position. Ask for the data in the purchase contract and store it beside the GPU-to-host map, so one query goes from a hot GPU to the furnace run that made its plate.

Measuring plates from software

Measure each plate's effective resistance during burn-in: run a steady load on every GPU, let temperatures settle, then sample GPU temperature, power and coolant inlet temperature. GPU temperature and power come from nvidia-smi or DCGM (field DCGM_FI_DEV_GPU_TEMP and DCGM_FI_DEV_POWER_USAGE). Inlet temperature does not come from the GPU; read it from the CDU, the rack manifold sensor or the server BMC. Log the memory temperature too where the GPU reports it, for the edge-contact case described above.

# one sample per second per GPU while the burn-in load runs
nvidia-smi --query-gpu=timestamp,serial,temperature.gpu,power.draw \
           --format=csv,noheader,nounits -l 1 >> gpu_$(hostname).csv

Join the samples with inlet temperatures and the inventory map into one CSV with columns host, gpu_serial, slot, plate_lot, t, gpu_temp_c, power_w, inlet_c, where t is seconds since the load started, then compute a steady-state resistance per GPU, and find outliers with the median and median absolute deviation, so a few bad plates cannot hide by inflating the spread.

import csv, statistics, sys
from collections import defaultdict

def steady_resistance(rows, settle_s=600, min_w=300):
    """Median (T_gpu - T_inlet) / P per GPU, skipping warm-up and idle samples."""
    per_gpu, meta = defaultdict(list), {}
    for r in rows:
        t, p = float(r["t"]), float(r["power_w"])
        if t >= settle_s and p >= min_w:
            rise = float(r["gpu_temp_c"]) - float(r["inlet_c"])
            per_gpu[r["gpu_serial"]].append(rise / p)
            meta[r["gpu_serial"]] = (r["host"], r["plate_lot"], r["slot"])
    return {g: (statistics.median(v), *meta[g]) for g, v in per_gpu.items() if len(v) >= 30}

def report(res, z_limit=3.5):
    med = statistics.median(v[0] for v in res.values())
    mad = statistics.median(abs(v[0] - med) for v in res.values()) or 1e-9
    z = {g: (res[g][0] - med) / (1.4826 * mad) for g in res}
    flagged = sorted((g for g in res if z[g] > z_limit), key=lambda g: -z[g])
    print(f"fleet median R = {med * 1000:.1f} mK/W over {len(res)} GPUs")
    for g in flagged:
        r, host, lot, slot = res[g]
        print(f"  FLAG {host} slot={slot} {g} lot={lot} R={r * 1000:.1f} mK/W z={z[g]:.1f}")
    for lot in sorted({v[2] for v in res.values()}):
        gs = [g for g in res if res[g][2] == lot]
        bad = sum(g in flagged for g in gs) / len(gs)
        lot_med = statistics.median(res[g][0] for g in gs)
        print(f"  lot {lot}: n={len(gs)} median={lot_med * 1000:.1f} mK/W flagged={bad:.0%}")

with open(sys.argv[1], newline="") as f:
    report(steady_resistance(list(csv.DictReader(f))))

Use a fixed load, not a training job: resistance is only comparable when every GPU carries similar power with a similar power map across the die.

Worked example: a burn-in with two lots

Here is the script's output on a synthetic burn-in of eight nodes with eight GPUs each: five nodes with plates from lot L2407 and three from lot L2411.

fleet median R = 45.3 mK/W over 64 GPUs
  FLAG node01 slot=3 GPU-13 lot=L2407 R=59.8 mK/W z=8.7
  FLAG node05 slot=5 GPU-55 lot=L2411 R=53.1 mK/W z=4.7
  FLAG node06 slot=2 GPU-62 lot=L2411 R=52.9 mK/W z=4.5
  FLAG node07 slot=5 GPU-75 lot=L2411 R=52.5 mK/W z=4.3
  FLAG node07 slot=2 GPU-72 lot=L2411 R=52.2 mK/W z=4.1
  FLAG node06 slot=5 GPU-65 lot=L2411 R=51.6 mK/W z=3.7
  lot L2407: n=40 median=45.2 mK/W flagged=2%
  lot L2411: n=24 median=45.5 mK/W flagged=21%

First, node01 GPU-13 is a lone outlier, about a third above the fleet median: at 700 W, roughly 10 K hotter for the same inlet. A single outlier in a good lot points to assembly, so reseat the plate and check the TIM spread and torque before blaming the plate.

Second, lot L2411 has almost the same median as L2407, yet a fifth of its GPUs are flagged, all in slots 2 and 5. A lot median hides a sub-population, so always read the slot column too. A slot pattern within one lot suggests a plate variant or fixture used for those positions. Keep those nodes away from straggler-sensitive jobs, and send the serials and measurements to the supplier with a request for that lot's joining and flow-test records.

Failure modes

  • Braze voids: local extra resistance; the GPU runs hot but passes every leak test.
  • Channel narrowing: filler or debris raises pressure drop and starves the plate on a parallel manifold.
  • Warped or soft base: thick TIM at the centre or edges; memory temperature may rise before core temperature.
  • Leaks after thermal cycling: joints that passed at the factory open after many heat cycles; a leak detection cable or tray sensor gives the first signal.
  • Galvanic corrosion: mixed metals in one loop slowly dissolve the less noble metal and shed particles into every plate downstream.
  • Erosion: high local velocity at an impingement jet wears the base; resistance creeps up fleet-wide rather than on one plate.

Trade-offs

Skived fins are cheap and joint-free but limited in geometry; bonded or printed channels allow complex shapes at higher cost. Vacuum brazing gives strong flux-free joints but anneals the copper; friction stir welding keeps the base hard but constrains the seam. Per-plate thermal testing catches more than sampling but costs line time. A burn-in resistance check costs hours of idle hardware, small next to a straggling GPU in a large training job.

What to do next

  1. Ask suppliers for the fin and joining process, per-plate test list and sample records.
  2. Record plate serial, lot, server serial, GPU slot and TIM lot at receipt.
  3. Add the fixed-load resistance check to node acceptance, grouped by lot and slot; re-run it after every reseat.
  4. Graph CDU filter pressure drop from commissioning onward; a fast rise means debris.
  5. Re-measure resistance quarterly to catch fouling or erosion trends; read liquid cooling loop telemetry and GPU thermal management for the alarms and throttle behaviour that follow.
  6. If you are still choosing a cooling approach, compare with immersion cooling and check the rack design in the NVL72 rack.
Key takeaway: A cold plate's thermal resistance is set by its manufacture: fin forming, lid joining, flatness, plating and cleanliness. Require per-plate test records and serial-to-slot traceability, measure (T_gpu - T_inlet) / P for every GPU in burn-in, and group outliers by lot and slot so a bad batch is caught before training starts.