We use cookies to make your experience better. To comply with the new e-Privacy directive, we need to ask for your consent to set the cookies. Learn more.
Liquid Cooling for AI Clusters: Keeping Cool at the Cutting Edge
Liquid Cooling for AI Clusters: Keeping Cool at the Cutting Edge
Many modern data centers now use pipes and manifolds to carry warm water away from high-powered chips. It’s no accident: per unit volume, water can carry roughly 3,000–3,500× more heat than air, which is why liquid-cooled AI racks are engineered for around 120–130 kW per rack, whereas typical air-cooled designs tend to stay near ~20–30 kW and most facilities still have few (if any) racks above 30 kW.
To put it in perspective: a single current-gen 8-GPU AI server such as NVIDIA DGX B200 is about 14.3 kW max. Under air cooling you’re typically limited to one (maybe two) such systems in a 20–30 kW rack. With liquid cooling you can place ~8 systems per rack (~8 × 14 kW ≈ 112 kW), which fits comfortably inside a 120–130 kW liquid-cooled envelope; for a 10,000-GPU build (8 GPUs per server), that’s ~1,250 servers drawing ~17.5 MW, or roughly 150–160 liquid-cooled racks versus ~1,250 racks if constrained to ~1 server per air-cooled rack.
For a turnkey implementation, explore Server Simply’s Liquid-Cooled AI SuperCluster Solution to see how direct-to-chip cooling can transform high-density data center performance.
How Liquid Cooling Works
Liquid cooling replaces much of the air-handling with fluid circuits. In a direct-to-chip system, non-conductive coolant (often water-glycol) is pumped through cold plates bolted to the CPUs/GPUs. As shown in the figure below, the coolant absorbs heat from the silicon and returns to a coolant distribution unit (CDU) where an internal heat exchanger dumps the heat to a secondary loop.
Figure: Typical direct-to-chip liquid cooling loop. Coolant (blue) flows through cold plates on CPUs/GPUs, then returns (red) to a CDU with a liquid-to-liquid heat exchanger. In this example, each server’s hot chips are fitted with cold plates. The warmed coolant flows via quick-disconnect hoses back to an in-rack or in-row CDU. The CDU contains redundant pumps and a liquid-to-liquid exchanger: it circulates cold fluid to the server manifolds and absorbs the warm return fluid, using facility chilled water to cool it. The CDU also regulates flow and keeps the coolant above the dew point to prevent condensation in the racks.

In effect, the system creates two isolated loops: (1) the IT coolant loop (cold plates → servers → CDU) and (2) the facility water loop (chiller/cooling tower → CDU heat exchanger → return). The CDU’s heat exchangers couple these loops without mixing fluids. This separation lets data-center engineers run the facility loop at higher temperature (saving chiller energy) while precisely controlling the cold plates’ temperature for reliability.
Aside from direct-to-chip, other liquid-cooled approaches exist. In rear-door heat exchangers, each rack has a radiator door that removes heat from exhaust air. In immersion cooling, entire server assemblies are submerged in a dielectric fluid: this captures 100% of the heat directly from all components. Immersion tanks can achieve extremely high capacity (e.g. up to ~100 kW in a single 42U tank), making them suitable for the most intense GPU/AI loads. In any case, liquid systems eliminate hot air recirculation, enabling AI servers to sustain full clock speeds without throttling.
The Benefits of Liquid Cooling
At its core, liquid cooling lets data centers keep their cool by whisking away heat much more efficiently. Here are some of the key gains:
- Energy savings: Liquid-cooled racks can be about 40% more energy-efficient than air-cooled design. (That extra power can go straight to computation instead of running fans and chillers.)
- Sky-high density: Liquid cooling allows about 3–5× more processors per rack, freeing space to stack dozens of GPUs per unit. (Air cooling simply can’t handle that much heat.)
- Lower total cost: Shrinking rack counts and power use can cut 18–25% off total data-center costs in dense setups. ABI Research analyst Rithika Thomas puts it bluntly: "Efficient thermal management is critical to performance, stability, and equipment lifespan".
These benefits help explain why companies are moving fast. ABI Research forecasts that liquid-cooling installations will quadruple from 2023 to 2030, reaching about $3.7 billion in value (around 22% CAGR). A few big trends are driving the change:
- Massive AI models. Modern LLM training rigs gobble up 10–100× more power than legacy workloads, so data centers are thirstier for cooling.
- Regulations & ESG. The EU’s new data-center rules mandate 40% fewer power losses by 2030, and roughly 78% of Fortune 500 firms have cooling efficiency on their sustainability scorecards.
- Competitive edge. Over 70% of operators plan to test liquid cooling by 2026, with 35% eyeing full liquid-immersion for their densest clusters. Jumping in early can give companies a serious leg up — sticking to old cooling methods could mean falling behind.
Real-World Racks and Solutions
Real-world racks and turnkey solutions. Major vendors are already shipping liquid-cooled AI racks. For example, NVIDIA’s GB200 NVL72—integrated by multiple OEMs including Supermicro and HPE—combines 36 Grace CPUs and 72 Blackwell B200 GPUs in a single rack with up to ~13.5 TB of HBM3e. Depending on the vendor build you’ll see 48U or 50U integrations, paired with in-rack or in-row CDUs (Supermicro specifies in-row options around 250 kW) to move the heat. That density is exactly why liquid cooling shrinks footprints: Supermicro’s Generative AI SuperCluster shows 256 GPUs using 32 × 4U, 8-GPU liquid-cooled systems across 5 racks, versus air-cooled 32 × 10U across 9 racks—nearly 2× denser—and claims up to ~40% lower electricity cost with DLC. (It is not “five 4U servers”; it’s thirty-two 4U systems distributed across five racks.)
Beyond the flagships, the market already offers the full toolkit: direct-to-chip cold plates for CPUs/GPUs from CoolIT (including GB200-specific loops) and Asetek; rear-door heat exchangers such as Motivair’s ChilledDoor rated up to ~75 kW per rack; and immersion systems—two-phase tanks from LiquidStack rated up to ~252 kW for a 48U unit and single-phase solutions from Asperitas typically in the ~44–60 kW range. Many turnkey racks also include redundant hot-swap pumps/PSUs and integrated monitoring to simplify deployment.
Challenges and the Road Ahead
Of course, switching to liquid cooling comes with a learning curve. Key hurdles include:
- Infrastructure retrofit: An existing server room may need new plumbing — chillers, pipes, or specialized in-row units. Many vendors now offer modular racks to let facilities phase in liquid cooling alongside air-cooled systems, and for greenfield builds it just means planning a coolant loop from day one.
- Coolant maintenance: Working with liquids means guarding against leaks, filtering the coolant, and often using dielectric fluids (so they won’t short electronics). Staff have to learn to handle pumps, valves, and coolant chemistry instead of just fans. With proper sensors (flow meters, leak detectors, etc.) and maintenance, these issues can be managed safely.
- Upfront cost: Liquid cooling gear (like chillers and distribution units) is more expensive up front than plain fans and vents. Many operators treat it like an investment: the higher initial cost pays off over time through energy savings, smaller footprints, and better reliability. In fact, studies show that in dense AI clusters, advanced cooling can recoup its cost by cutting total data-center TCO by nearly 25%.
As AI and HPC workloads continue to grow, liquid cooling is becoming a de facto requirement for cutting-edge data centers. By coupling coolant directly to chips, these systems unlock performance and efficiency unattainable with air. Industry leaders report PUE values approaching 1.05, multi-fold throughput gains, and major cost savings when moving to liquid-cooled GPUs. Integrated solutions now exist – for example, ServerSimply’s Liquid-Cooled AI Super Cluster – that package direct-to-chip cold-plate cooling with optimized CDUs and pumping systems. In sum, liquid cooling enables AI clusters to scale safely to the next level of density and performance, while significantly reducing energy and water requirements.




