Exploring the breakthrough innovations shaping our world. From AI infrastructure and robotics to biotech, quantum computing, and spatial tech.
For decades, conventional data center thermal management relied on familiar engineering principles: move heated air away from servers, deliver cooler air back to the equipment, and maintain temperatures within the operating range specified by hardware manufacturers.
That model still works well for many workloads. It becomes harder to apply, however, when computing equipment is packed into increasingly dense AI racks.
Modern accelerator systems can place tens of kilowatts of electrical load into a single rack, with some high-density configurations approaching or exceeding 100 kilowatts. Because nearly all of the electrical power consumed by computing hardware eventually becomes heat, the cooling system has to remove a comparable thermal load continuously.
At those densities, the question is no longer simply whether a data center can cool its servers. The more useful question is which cooling architecture provides the right balance of thermal capacity, efficiency, serviceability, infrastructure cost, and future expansion.
Three approaches are particularly important: conventional air cooling, direct-to-chip liquid cooling, and immersion cooling.
Air remains the simplest cooling medium to deploy. It is readily available, does not create liquid-management concerns inside the server, and fits naturally into conventional data center maintenance practices.
Its limitation is the amount of heat it can transport through a given volume of airflow.
Water and other liquids can carry substantially more heat per unit volume than air for a comparable temperature rise. That difference is one reason liquid cooling becomes increasingly attractive as rack power density rises.
Air cooling can be pushed further through better heat sinks, higher airflow rates, containment systems, improved fan designs, and more effective heat exchangers. But those improvements come with trade-offs.
A very high-power rack may require large volumes of fast-moving air. That increases fan energy, airflow-management requirements, acoustic output, and the amount of space devoted to moving air. At some point, continuing to increase airflow becomes a less attractive way of handling additional heat.
The practical limit therefore depends on more than the rack's nominal power. Server design, inlet temperature, component-level thermal resistance, airflow path, facility climate, and cooling plant efficiency all matter.

Direct-to-chip (DTC) cooling has become one of the most important liquid-cooling approaches for high-density computing.
Instead of relying on room air to carry most of the heat away from the processor, a DTC system places a cold plate directly against major heat-producing components such as GPUs or CPUs. A purpose-designed liquid coolant flows through channels inside the cold plate, absorbs heat near its source, and carries that heat toward a heat exchanger or coolant distribution unit.
This changes the thermal path.
With conventional air cooling, heat travels from the silicon into a heat sink and then into the surrounding air. DTC cooling replaces much of that final air-based heat transfer with a liquid-based path that can move substantially more heat through a smaller physical volume.
A typical DTC installation includes cold plates, tubing, pumps, a coolant distribution unit, heat exchangers, monitoring equipment, and connections between the IT equipment and facility cooling infrastructure.
The coolant loop can be separated from the facility's primary water system. A coolant distribution unit can control flow, pressure, and temperature while transferring heat between the server-side loop and the facility-side cooling system.
The exact design varies by hardware platform and facility.
Some systems use liquid cooling for only the highest-power components while continuing to use air for other equipment. Others are designed around a much larger proportion of liquid-cooled hardware.
The biggest advantage is balance.
DTC can handle much higher component heat loads without requiring every part of the server environment to become liquid-cooled. GPUs and CPUs can receive direct thermal management while lower-power components continue to rely on airflow.
That makes DTC particularly useful for mixed environments.
It can also be incorporated into some existing facilities without completely replacing the data center's air-handling infrastructure. The degree of retrofit compatibility depends heavily on the building's available cooling capacity, floor layout, plumbing infrastructure, and the server platform being deployed.
DTC does not automatically eliminate air cooling.
Depending on the server design, components such as power supplies, storage devices, networking equipment, memory modules, or voltage regulation hardware may still release heat into the surrounding air.
This means many DTC facilities operate as hybrid systems. Liquid handles the highest thermal loads, while air removes the remaining heat.
That compromise is often a feature rather than a weakness. It allows the facility to use liquid where it delivers the greatest benefit without requiring every component to be redesigned around immersion or another more comprehensive cooling method.

Immersion takes a different approach.
Rather than attaching cold plates to selected components, immersion systems place electronic hardware in a dielectric fluid that does not conduct electricity in the same way as conventional water-based cooling loops.
The cooling medium can therefore surround a much larger portion of the hardware and remove heat from multiple components at once.
Two broad approaches are commonly discussed: single-phase and two-phase immersion.
In a single-phase system, the dielectric fluid remains liquid during normal operation.
Heat generated by the submerged hardware is transferred into the fluid, which is then circulated through an external heat exchanger or cooling loop before returning to the tank.
The design is conceptually straightforward, but the facility still needs pumps, heat exchangers, fluid management, filtration, monitoring, and appropriate containment.
Two-phase systems use a dielectric fluid selected for a controlled phase change at the intended operating temperature.
Heat from the hardware causes the fluid to vaporize. The vapor rises to a condenser, where the heat is rejected and the fluid returns to the liquid state.
The process can take advantage of latent heat transfer, allowing a large amount of thermal energy to be moved through a relatively small amount of fluid.
Two-phase systems, however, introduce their own engineering considerations, including fluid selection, containment, materials compatibility, vapor management, and environmental requirements.
There is no universal winner among air, direct-to-chip, and immersion cooling.
The best option depends on the rack density, hardware configuration, facility design, maintenance model, and expected future growth.
Air remains attractive when power densities are moderate and the existing facility already has sufficient cooling capacity.
Its strengths include familiar maintenance procedures, straightforward hardware compatibility, and relatively simple infrastructure.
Its weakness is scaling. As thermal density rises, airflow requirements and the supporting cooling infrastructure can become increasingly demanding.
DTC sits between conventional air and full immersion.
It can remove large amounts of heat directly from major components while allowing some parts of the server and facility to continue operating with air cooling.
That makes it a practical choice for many high-density AI deployments, particularly where operators want higher thermal capacity without completely changing their hardware-service model.
The trade-off is complexity. Pumps, coolant loops, quick-disconnects, cold plates, distribution units, and leak-management procedures all become part of normal operations.
Immersion provides the broadest contact between the cooling medium and the electronic hardware.
That can create substantial thermal headroom for systems where extremely high component densities make conventional airflow difficult to manage.
The price is a larger operational change. Technicians need procedures for handling fluid-filled hardware, equipment manufacturers must validate material compatibility, and facility operators need appropriate fluid management and containment systems.
Hardware warranties and maintenance practices also need to be considered before a large-scale deployment.

Cooling efficiency is only one part of the decision.
A data center can achieve excellent thermal performance and still be a poor operational fit if technicians cannot service the hardware efficiently.
Air-cooled systems have the advantage of familiarity. A failed component can usually be removed using established procedures without interacting with a liquid system.
DTC introduces additional connections and coolant-management requirements, but much of the server remains accessible through a relatively familiar rack architecture.
Immersion changes the maintenance workflow more substantially. Hardware has to be removed from the cooling environment, and technicians need procedures for fluid handling, draining, cleaning, and component replacement.
For facilities operating thousands of accelerators, those differences can affect maintenance time and operational planning.
Liquid cooling also requires a broader definition of reliability.
A conventional air-cooled server does not have a liquid loop running through its rack. A DTC installation introduces additional components such as pumps, tubing, connectors, valves, and heat exchangers.
Those components need monitoring and maintenance.
Immersion introduces another layer of considerations because the fluid must remain compatible with the materials used in the server hardware. Seals, plastics, coatings, connectors, and other components can respond differently to particular dielectric fluids over long periods.
For this reason, large-scale deployment generally requires validation across the entire hardware and cooling ecosystem rather than simply testing whether a fluid can remove heat from a processor.
It is tempting to compare cooling technologies using a single efficiency figure, but facility performance depends on much more than the cooling medium.
Power Usage Effectiveness, for example, measures the relationship between total facility energy consumption and the energy delivered to computing equipment. Cooling is one contributor, but so are UPS systems, power distribution, pumps, fans, lighting, and other infrastructure.
A liquid-cooled facility can therefore have a different PUE from another liquid-cooled facility because the two buildings use different cooling plants, climates, operating temperatures, electrical systems, or redundancy strategies.
The same principle applies to immersion.
Its thermal characteristics can reduce the amount of energy required for certain cooling functions, but the final facility efficiency depends on how the entire system is designed.

A useful way to think about the three technologies is not as a simple ranking, but as a set of engineering choices.
Air cooling remains well suited to lower-density equipment and facilities where simplicity and conventional maintenance are priorities.
Direct-to-chip liquid cooling is often attractive when accelerator density has outgrown comfortable air-cooling limits but the operator still wants a relatively familiar server architecture and maintenance model.
Immersion cooling becomes more compelling when thermal density is extremely high and the facility can accommodate the operational changes associated with fluid-filled hardware.
The decision also depends on whether the facility is new or existing.
A purpose-built greenfield data center has more freedom to design plumbing, heat rejection, rack layouts, and electrical systems around liquid cooling from the beginning. An existing building may place greater value on a cooling architecture that can coexist with its current air-handling infrastructure.
The most effective thermal design is rarely determined by the coolant alone.
Rack layout, accelerator selection, power density, inlet temperature, heat exchanger design, facility climate, redundancy requirements, maintenance procedures, and future expansion plans all influence the result.
Even workload characteristics can matter. A cluster running sustained high utilization may have very different thermal requirements from one with lower or more variable accelerator utilization.
That is why cooling should be considered together with the electrical and computing architecture rather than treated as a separate facility subsystem.

High-density AI computing is forcing data center operators to rethink a part of infrastructure that once received relatively little attention outside facilities engineering.
Air cooling remains practical for many environments, but increasing rack power densities make airflow more demanding and can reduce the room available for further scaling. Direct-to-chip liquid cooling provides a middle path, moving heat directly from major components while retaining some of the familiar characteristics of conventional servers. Immersion cooling goes further by surrounding much more of the hardware with a dielectric cooling medium, offering substantial thermal headroom at the cost of greater operational and hardware-management complexity.
The important question is therefore not which technology is universally best. It is which architecture fits the equipment, rack density, facility, maintenance model, and expansion plans.
As AI systems become denser, thermal management will increasingly be designed alongside compute, power distribution, and networking. Cooling is no longer just a way to keep hardware within temperature limits. It is part of the architecture that determines how much computing capacity a facility can realistically deploy.