Most writing about immersion cooling stops at the moment the servers go into the tank. The physics is the easy part: a dielectric liquid touching every component removes heat far better than air, and Immersion Cooling for GPU, in depth already covers single- and two-phase heat transfer, sizing the fluid loop, the GPU telemetry and the scheduler hooks. This article covers everything around it: choosing a fluid you can still buy in five years, proving that your cables, optics and thermal paste survive in it, carrying the weight, keeping the fluid healthy, and repairing a server that is submerged.

These are the steps where immersion projects actually stall. A tank that cools beautifully is a failure if the optics fog after six months, if the floor was not rated for it, if the OEM voids the warranty or if the fluid supplier stops making the fluid. By the end you will have a qualification plan, a fluid-monitoring script and a worked example for a small pilot. Numbers marked illustrative are there to show the method; use your vendor's datasheets and your structural engineer's figures.

When immersion is on the table

Immersion is one answer to rack density. Air cooling becomes impractical as racks climb into the tens of kilowatts, because the airflow and fan power grow out of proportion. Direct-to-chip cold plates, covered in liquid cooling for GPU datacenters, capture most of the heat at the GPU and CPU and leave the rest to air. Immersion captures essentially all of it into the fluid, removes server fans and tolerates warm facility water, which can mean dry coolers instead of chillers in many climates.

The cost is that immersion changes everything physical: rack form, floor, service procedures, supply chain and safety case. Current flagship GPU systems such as rack-scale NVLink designs are engineered and sold for cold plates, so immersion often means qualifying hardware the vendor did not design for it. A sensible screen is: choose immersion when the site cannot reject heat to air at the density you need, when you control the hardware choice for years, and when you can staff a different maintenance workflow. Otherwise cold plates are the lower-risk path. GPU datacenter cooling overview places all the options side by side.

The programme as a lifecycle

An immersion programme is a lifecycle, not a purchase1. Decidedensity, site, fluid2. Qualify fluidcompatibility soak3. Qualify serversoptics, TIM, warranty4. Build sitefloor, bunds, hoist5. Commissionburn-in, baseline6. Operatesample fluid, top up, pull servers7. Retiredrain, dispose fluidFluid health loop (monthly to quarterly)sample -> lab: breakdown voltage, water, acidinline: level, temp, flow, filter pressure dropcompare with the commissioning baselinetrend breach -> filter, dry, or replace fluidService loop (per failed server)drain jobs and cordon node in the schedulerpower off, hoist, drip over tank, swap partreinsert, burn-in, uncordon, log drag-outMTTR is longer than air: hold spare servers
The seven stages of an immersion programme and the two loops that run during operation.

Choosing a fluid

The fluid decision is a ten-year decision, because the tanks, seals and materials are qualified against one fluid and switching later means re-qualification. There are three broad families:

FamilyTypical chemistryStrengthsWatch for
Single-phase hydrocarbonSynthetic oils such as polyalphaolefinsLow cost per litre, widely available, not PFASCombustible: flash point drives fire code; viscosity raises pump power; leaches some plastics
Single-phase fluorinatedEngineered fluorinated fluidsNon-flammable, low viscosityCost; PFAS regulation; global warming potential of some fluids
Two-phase fluorinatedLow-boiling engineered fluidsVery high heat flux at near-constant temperatureVapour loss whenever the lid opens; PFAS regulation; sealed tanks

Regulation is now a primary selection criterion. 3M, historically a major supplier of engineered fluorinated fluids, announced in 2022 that it would exit PFAS manufacturing by the end of 2025, and a broad European Union PFAS restriction was proposed in 2023. A programme built on a fluid that could become unavailable or restricted is a stranded-asset risk, which is a large part of why most new GPU immersion deployments are single-phase hydrocarbon.

When you request datasheets, ask for: dielectric breakdown voltage and the test method; viscosity at your operating temperature, not only at 40 degrees C; flash and fire points; pour point if tanks sit cold; maximum water content; expected service life and the vendor's reconditioning or take-back service; and a material compatibility list with test conditions. Ask whether the vendor will guarantee supply of the same formulation, because topping up with a different batch chemistry is a failure mode.

Qualifying the hardware

A server built for air contains materials that were never meant to sit in oil at 45 degrees C for five years. Qualification finds them before production does.

  • Cable jackets and plastics. Plasticisers in some PVC can leach into the fluid, stiffening the cable and contaminating the bath. Prefer jackets on the fluid vendor's compatible list.
  • Labels and adhesives. Paper labels and some adhesives detach and clog filters. Asset tags must be immersion-rated or moved above the fluid line.
  • Thermal interface material. Some thermal pastes wash out over months, and the GPU temperature creeps upward. Use a TIM the fluid vendor or OEM has validated.
  • Optics. Fluid in a pluggable optical transceiver's light path degrades the signal. Keep optics above the fluid line or use transceivers rated for immersion. This is the most common surprise in GPU tanks, because every GPU has a fast NIC.
  • Fans and BMC. Fans are removed, so the BMC must be configured not to alarm or throttle on missing fan tachometers, and its temperature sensors re-baselined.
  • Storage. Solid-state drives are generally fine; spinning disks need the vendor's confirmation.
  • Warranty. Get written OEM support for the specific model in the specific fluid. Without it, every hardware fault becomes a negotiation.

The standard test is a soak: coupons of every material, plus a full sample server, held in the fluid at or above operating temperature for several weeks, with the fluid measured before and after. Accept a material if the fluid's properties stay within the vendor's limits and the material keeps its mass, hardness and function. Then run a burn-in under full GPU load and compare GPU temperatures, clock-event reasons and link error counters against the same server in air.

Preparing the site

Tanks change the building. A horizontal tank holds the servers vertically, top-loaded, so the aisle layout becomes rows of long tubs with a hoist path above them rather than rows of racks with hot and cold aisles. Four facility questions decide whether a room can take them:

  1. Floor loading. A full tank concentrates fluid, steel and servers on a small footprint; check the slab or raised floor rating with a structural engineer.
  2. Containment. Spill bunds or drip trays sized for a tank's fluid volume, and drainage that does not lead fluid into building drains.
  3. Fire safety. Hydrocarbon fluids are combustible, so the local fire authority will ask about flash point, suppression and fluid volume per room.
  4. Heat rejection. Warm return fluid lets the heat exchanger run on warm facility water, which can enable dry coolers and heat reuse. Size it from the loop calculation in the physics article.

Keeping the fluid healthy

The fluid is now part of your hardware, and it ages. Water ingress from humid air lowers dielectric strength. Leached plastics and oxidation raise the acid number. Particles from filters, labels and wear accumulate. Each of these is slow and invisible until a sample is tested, so the programme needs two loops: inline sensors for level, temperature, flow and pressure drop across the filter, which catch leaks and clogged filters within minutes; and laboratory samples monthly to quarterly for breakdown voltage, water content and acid number, which catch degradation over months.

Absolute limits differ by fluid, so the useful signal is the trend against the commissioning baseline. This script evaluates a new sample against both the vendor limits and the baseline, and is easy to wire into an existing alerting pipeline:

from dataclasses import dataclass

@dataclass
class Sample:
    tank: str
    breakdown_kv: float   # dielectric breakdown voltage
    water_ppm: float
    acid_mg_koh_g: float
    particles_ml: float   # particles per mL above the counter's size threshold

# Fill these from your fluid vendor's datasheet; values here are placeholders.
LIMITS = dict(breakdown_kv_min=None, water_ppm_max=None,
              acid_max=None, particles_max=None)
DRIFT = dict(breakdown_drop=0.20, water_rise=2.0, acid_rise=2.0)

def evaluate(s, base, limits=LIMITS):
    issues = []
    if limits["breakdown_kv_min"] and s.breakdown_kv < limits["breakdown_kv_min"]:
        issues.append(("critical", "breakdown below vendor minimum"))
    if limits["water_ppm_max"] and s.water_ppm > limits["water_ppm_max"]:
        issues.append(("critical", "water above vendor maximum"))
    if limits["acid_max"] and s.acid_mg_koh_g > limits["acid_max"]:
        issues.append(("critical", "acid number above vendor maximum"))
    if s.breakdown_kv < base.breakdown_kv * (1 - DRIFT["breakdown_drop"]):
        issues.append(("warn", "breakdown voltage falling: check for water"))
    if s.water_ppm > base.water_ppm * DRIFT["water_rise"]:
        issues.append(("warn", "water rising: inspect seals, lid, humidity"))
    if s.acid_mg_koh_g > base.acid_mg_koh_g * DRIFT["acid_rise"]:
        issues.append(("warn", "acid rising: look for leaching material"))
    if limits["particles_max"] and s.particles_ml > limits["particles_max"]:
        issues.append(("warn", "particles high: change filter"))
    return issues or [("ok", s.tank)]

The responses are graded. A water warning usually means drying the fluid and finding the ingress, often a lid seal or a humid service area. A rising acid number points at something new in the bath, frequently a cable or label added without qualification. A critical breakdown result means draining the affected servers before a fault, and replacing or reconditioning the fluid.

Servicing a submerged server

Repairing a submerged server is the workflow operators notice most. The steps: drain jobs and cordon the node in the scheduler; power it off; lift it with the tank's hoist; let it drip over the tank for several minutes so fluid returns to the bath; move it to a drip tray or service bench; replace the part; reinsert it, letting trapped air escape; then run a burn-in before uncordoning. Single-phase tanks let you pull one server without disturbing its neighbours; two-phase tanks lose vapour whenever the lid opens, which pushes teams to batch repairs.

Two numbers belong on a dashboard. Drag-out is the fluid that leaves with each pulled server, on its surfaces and in its cavities; measure it by logging top-up volume against pulls. Mean time to repair is longer than in air, so availability depends on spares: hot spare servers that slot straight in turn a slow repair into a fast swap, with the actual repair done on the bench later.

Worked example: a four-tank pilot

Size a four-tank pilot. Each tank holds six 8-GPU servers drawing 11 kW, so 66 kW per tank and 264 kW for the pilot. Illustratively, each tank holds 1,400 litres of a hydrocarbon fluid at 0.80 kg per litre, so 1,120 kg of fluid, plus a 500 kg empty tank and six 60 kg servers. The full tank weighs about 1,980 kg on a 2.6 m by 1.1 m footprint, roughly 2.86 square metres, which is about 690 kg per square metre spread evenly, more at the feet. Compare that with the rated load of the floor, not with an air rack's weight; many raised floors were not designed for it.

Fluid: four tanks hold 5,600 litres. If drag-out is 0.5 litres per pull and the pilot pulls 10 servers a month across failures, upgrades and inspections, that is 5 litres a month, small next to the fill but large enough to need a stocked supply of the same formulation. Availability: with 24 servers, two hardware failures a month and a repair time of three hours each in the tank, against one hour in air, you lose 6 server-hours a month instead of 2. With one hot spare server the loss falls to the swap time. The conclusion is typical: weight and service, not heat, decide the design.

Failure modes

  • Fogged optics. Link errors climb weeks after commissioning because transceivers sit below the fluid line.
  • Unqualified material added later. A replacement cable or label leaches into the bath and the acid number climbs.
  • Water ingress. Open lids in humid service areas lower breakdown voltage.
  • Mismatched top-up. Fluid from a different formulation or supplier changes the bath's properties.
  • Regulatory stranding. A PFAS fluid becomes unavailable and the programme cannot expand.
  • Warranty dispute. The OEM declines a failed GPU because the fluid was not on its approved list.
  • Silent BMC throttling. Firmware treats missing fans as a fault and caps power; check clock-event reasons after every firmware update.

Trade-offs

Against cold plates, immersion removes all server fans, captures all the heat and tolerates warmer water, and it can cool components cold plates ignore. It costs you standard rack form, vendor-native designs for the newest GPU systems, quick service and a simple supply chain. For a site whose density is driven by power density and whose hardware is the latest rack-scale GPU design, cold plates are usually the path the vendor supports. Immersion makes most sense for fleets of standard servers in sites where air or chillers cannot reject the heat, and where you can run the lifecycle in this article.

What to do next

  1. Decide the fluid family first, with regulation and long-term supply as explicit criteria, and get the vendor's compatibility list.
  2. Inventory every material in your target server, including cables, labels, TIM and optics, and soak-test anything not on the list.
  3. Obtain written OEM warranty support for the exact server and fluid.
  4. Have a structural engineer approve the full tank load, and agree containment and fire measures with the local authority.
  5. Record a commissioning baseline for fluid properties and GPU temperatures, then sample on a schedule and alert on trends.
  6. Write the service runbook, including scheduler cordon steps, and measure drag-out and repair time from the first pull.
  7. Run a pilot of a few tanks for at least two quarters before committing a hall.
Key takeaway: Immersion cooling succeeds or fails outside the tank. Choose a fluid you can buy for a decade, qualify every material and the optics, get the warranty in writing, carry the weight, sample the fluid against a baseline, and design the service workflow around spares. Pilot before committing a hall.