English
Back
Open Account
Mega chipmaker Terafab breaks ground! Will this catalyze the AI supply chain rally?
Yee Hop Holdings
joined discussion · Aug 29 02:20

From Rack-Scale to Liquid Cooling: A Decade of Transformation in Server Design

Ten years ago, the server industry focused on rack-scale architecture. The core idea was not simply cramming more machines into a cabinet, but changing the unit of resource management. Traditional servers packaged processors, memory, storage, networking, and power supplies into a single complete device; enterprises purchased and repaired them on a per-unit basis. With the rise of cloud data centers, this model exposed issues such as low utilization rates, slow scalability, and vendor lock-in. Intel’s Rack Scale Design (RSD) therefore proposed disaggregating compute, storage, and accelerators into composable resource pools, dynamically configured by software according to workloads. The rack evolved from a mere metal cabinet for installing equipment into a schedulable system.
This shift aligned with the Open Compute Project’s promotion of open racks, shared power supplies, and modular designs. For large cloud service providers, servers no longer needed individual, complete enclosures, power supplies, and management logic for each unit; as long as the entire rack could be rapidly deployed, replaced, and scaled, total cost of ownership (TCO) could be reduced. The first decade of rack-scale architecture primarily addressed issues of standardization, resource utilization, and operational efficiency.
AI Turns the Rack into a Supercomputer
With the emergence of generative AI, the focus suddenly shifted from "how to share resources" to "how to enable massive numbers of GPUs to collaborate like a single processor." Training and inference for large models require high-bandwidth memory, GPU interconnects, and low-latency networks; if any node waits for data, it drags down the efficiency of the entire cluster. NVIDIA’s GB200 NVL72 integrates 36 Grace CPUs and 72 Blackwell GPUs into a liquid-cooled rack, forming a 72-GPU computing domain via NVLink, with total interconnect bandwidth reaching 130 TB/s. Compared to the previous generation, which primarily used eight GPUs as the high-speed interconnect unit, the design logic has shifted from "multiple servers forming a cluster" to "the entire rack is the server."
Rising performance density is pushing thermal management to its physical limits. The power requirement for a GB200 NVL72 rack is approximately 120 kW, far exceeding the low-density racks found in traditional enterprise data centers. Since air has limited heat-carrying capacity, simply increasing fan speed not only consumes more power and generates noise but also occupies chassis space, potentially forcing chips to throttle. Given liquids' higher heat capacity and thermal conductivity, direct liquid cooling has evolved from a niche solution in high-performance computing to a foundational design for AI servers.NVIDIA MGX Technical Documentationindicates that cold plates, manifolds, and quick-disconnect couplings are now integrated into the rack-level platform, rather than being retrofitted by data center operators later on.
So-called "direct-to-chip" liquid cooling involves attaching cold plates directly to primary heat sources like CPUs and GPUs. The coolant flows through in-rack manifolds to a Coolant Distribution Unit (CDU), which then transfers the heat to the facility's water loop. Memory, power supplies, and other low-heat components can remain air-cooled. This hybrid architecture improves heat capture efficiency while reducing the burden on room-level air conditioning. Dell's currentdirect liquid cooling rack solutionsalready cover ranges from 40 kW to over 80 kW, and its next-generation CDUs can handle thermal loads exceeding 220 kW, reflecting how cooling capacity is rapidly scaling up in tandem with chip power consumption.
However, liquid cooling is not as simple as swapping fans for pipes. Server motherboards, cold plates, quick-disconnect couplings, pumps, manifolds, leak detection systems, control software, and maintenance procedures must all be co-designed. Data centers must also reconsider supply water temperature, water quality, floor load-bearing capacity, and redundancy. According to ASHRAE's data center guidelines, a fully loaded liquid-cooled rack can weigh over 1,800 kg, meaning older facilities may not be able to install them directly. In other words, the competitive landscape has expanded from server manufacturers to include suppliers of power, cooling, connectors, data center engineering, and building infrastructure.
Liquid cooling is also changing the investment rhythm of data centers. While traditional server rooms allow for incremental server additions, ultra-high-density AI systems often require the simultaneous deployment of racks, power, networking, and cooling. This increases upfront capital expenditure and makes customers more reliant on vendors capable of full-rack design, factory pre-integration, and on-site maintenance. The competitive advantage in servers is shifting from merely procuring high-performance chips to the ability to reliably assemble thousands of components into a long-term operating "rack-scale computer."
The true turning point of this decade is the continuous outward expansion of server design boundaries. Engineers previously optimized individual motherboards, then entire servers, and later whole racks; now, they must view chips, networking, power, cooling, and even building infrastructure as a single system. Rack-scale computing focuses on how many resources can be scheduled per rack, while liquid cooling concerns how much effective compute can be generated per watt of power. The future winner in AI infrastructure will depend not only on who has the fastest GPU, but also on who can keep an entire rack running at full capacity stably over the long term with the lowest energy and maintenance costs.
Yee Hap Holdings Investor Relations Department
(Chips and Computing Power Series #88)
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Thumbs Up
2
Lol
1
94K Views
Report
Comments
Write a Comment...
3