Should I Retrofit AI Racks With Rear-Door Heat Exchangers or Direct-to-Chip Cooling?

Use rear-door heat exchangers when you need a lower-disruption, liquid-assisted bridge for moderate-density racks and can add water piping behind them. Move to direct-to-chip cooling when AI racks are consistently in the 50-100+ kW range, especially near 100 kW, because it carries heat away at the components rather than asking room air to do work it cannot reliably handle.
At what rack density is air cooling no longer enough for AI servers?
Air cooling becomes a poor planning assumption as AI rack density moves beyond conventional levels, while some form of liquid cooling becomes necessary at roughly 100 kW per rack. That is the practical line drawn by the Los Alamos National Laboratory report, which contrasts conventional 10-15 kW average racks with emerging 30-50 kW racks, hyperscale deployments approaching 100 kW, and high-end HPC/AI deployments reaching 250 kW.
The important word is planning. There is no honest single number that makes every air-cooled rack fail. Rack layout, server design, room configuration, and the portion of heat captured in liquid all matter. But the density ranges provide a useful decision screen: ASHRAE characterizes legacy CPU racks as typically 5-10 kW and GPU clusters as often 40-100 kW per rack. It recommends liquid or liquid-assisted architectures to support 50-100+ kW racks.
| Rack-density signal | What the cited guidance indicates | Retrofit question to answer |
|---|---|---|
| 5-10 kW | Typical legacy CPU-rack range cited by ASHRAE. | Can the existing air system continue to serve this load? |
| 30-50 kW | Emerging workload range identified by Los Alamos. | Is a liquid-assisted zone needed before more racks arrive? |
| 50-100+ kW | ASHRAE recommends liquid or liquid-assisted architectures. | Will rear-door cooling capture enough heat, or must heat leave at the chip? |
| About 100 kW and above | Los Alamos says some form of liquid cooling becomes necessary at roughly this level. | Can the facility support a direct liquid cooling system? |
That distinction saves time. A 40 kW design does not automatically require the same retrofit as a 100 kW design. Treating them as identical is how teams buy equipment before they have answered the harder question: where will the heat go after it leaves the rack?

Can rear-door heat exchangers cool high-density GPU racks?
Rear-door heat exchangers can reduce the air-side burden substantially, but public evidence does not establish them as a universal answer for the highest-density GPU racks. They are indirect liquid cooling: air transfers server heat to a liquid-cooled door, unlike direct-to-chip cooling, where cold plates attach to heat-generating components.
A Department of Energy and Lawrence Berkeley National Laboratory case study illustrates both the appeal and the limit. Across six doors using 9 gallons per minute per door, the rear-door heat exchangers removed 31.6 kW, or about 48% of a 66 kW server waste-heat load. The six racks were operating at 10-11 kW each, and passive doors reduced server outlet air from 100-120 F to about 80 F.
That is a useful result, not a density guarantee. It shows that rear doors can capture a meaningful share of heat with no moving parts in the door, while still requiring cooling-water flow, piping, pumps, valves, sensors, pressure testing, and leak checks. The same bulletin reported equipment cost of $6,000 per rear-door device, excluding installation and infrastructure additions. Public data in these sources does not provide a comparable installed-price figure for a direct-to-chip retrofit, so a blanket claim that either option is cheaper would be guesswork.
Rear doors make the most sense when the servers remain primarily air-cooled, floor disruption must be constrained, and the goal is to extend an existing room rather than create a dedicated high-density zone. They are a bridge, not a promise. ASHRAE's guidance is more direct for purpose-built AI facilities whose densities routinely exceed 50-120 kW: use a Technology Cooling System with integrated heat transport and rejection.
Is direct-to-chip cooling practical in an existing data center?
Direct-to-chip cooling is practical in an existing data center when the team can create and operate a separated liquid system from the facility connection to the rack manifold and server cold plates. It is not a bolt-on accessory. The Open Compute Project's cooling-loop requirements describe the standard path as CDU to rack manifold to IT equipment and back to the CDU, with heat transferred to facility water through a plate-and-frame heat exchanger.
The CDU is the dividing line that makes the architecture manageable. OCP says it isolates IT-side and facility-side loops and may sit in-rack, at row level, or at facility level. It maintains pressure, flow, temperature, dew-point control, cleanliness, and leak detection. Typical in-rack and row-level Technology Cooling System pressure is 20-65 psi, while applicable IT-equipment safety testing can require testing at three times normal operating pressure.
That sounds like a long checklist because it is one. But the checklist is the decision tool. As with an infrastructure integration decision, the visible device is only one part of the operating system around it. Los Alamos also notes that direct liquid cooling remains hybrid: roughly 10-15% of rack power may still need air cooling. Replacing air cooling is not the assignment; reducing its load to a realistic job is.
Fluid selection deserves the same discipline. LBNL's liquid-cooled rack specification warns that water-based transfer-fluid chemistries, including antifreeze and corrosion or biological inhibitors, are often proprietary and should not be mixed without compatibility validation. Every wetted component, from server heat exchangers and quick connects to tubing, manifolds, and CDUs, needs compatibility with the selected fluid and with the rest of the loop.
How should a team choose between the two approaches?
Choose the architecture after mapping density, heat-rejection capacity, and operational ownership, in that order. Do not begin with a vendor category. Begin with measured and planned rack loads, then identify whether existing chilled or tower-water capacity, mechanical space, pumping, and routing can support the necessary loop.
For a staged retrofit, rear-door heat exchangers can be the sensible first move where air-cooled servers dominate and the room needs additional heat capture. For sustained GPU deployments in the 50-100+ kW range, direct-to-chip cooling provides the more deliberate path because heat is collected near the source. At roughly 100 kW and above, the Los Alamos guidance makes the direction clear: plan for liquid heat removal.
Finally, assign operating responsibility before installation. ASHRAE includes isolation valves and temperature, pressure, flow, and leak-detection sensors integrated with BMS/DCIM in its Technology Cooling System description. A design without named ownership for alarms, water treatment, fluid compatibility, inspections, and response actions is not yet a finished retrofit plan. Measure the racks, trace the water path, and choose the smallest system that can safely serve the density you are actually building for.
Frequently Asked Questions
What infrastructure does a liquid-cooling retrofit require?
A liquid-cooling retrofit requires a facility water path, piping, pumping, isolation and balancing valves, liquid interfaces, sensors, and a way to reject heat. For direct-to-chip systems, ASHRAE describes a Technology Cooling System with IT-side and facility-side loops, CDUs, variable-speed pumping, and BMS/DCIM-integrated temperature, pressure, flow, and leak sensors. Existing capacity for water, pumps, piping, and heat rejection must be checked before hardware is selected.
How do coolant distribution units and leak detection work?
A CDU separates the facility water system from the IT-side cooling loop through a plate-and-frame heat exchanger, then manages the IT loop's pressure, flow, temperature, cleanliness, and leak detection. OCP specifies that rope leak sensors can change electrical state when exposed to water-based coolant; management hardware can use that signal to alarm, notify operators, shut equipment down, or interrupt flow at a manifold or CDU.
Is immersion cooling necessary for AI racks?
No. Los Alamos identifies direct liquid cooling and immersion cooling as the two predominant liquid-cooling approaches, not as a rule that immersion is mandatory for every AI rack. Direct-to-chip cooling can address high-density deployments while retaining a hybrid air path for the roughly 10-15% of rack power that may still need air cooling.
Sources
- ASHRAE AI Data Center Energy Performance Framework
- U.S. Department of Energy and Lawrence Berkeley National Laboratory: Rear-Door Heat Exchanger Case Study
- Open Compute Project Foundation: Cold Plate Cooling Loop Requirements
- Los Alamos National Laboratory: Futureproofing Through 2035 for the AI and HPC Power Density Trend
- Lawrence Berkeley National Laboratory: Open Specification for a Liquid Cooled Server Rack