Jensen Huang predicts sales will double next year, sparking a strong rebound in hardware stocks!
In recent years, the investment logic for AI infrastructure has been nearly synonymous with "hoarding GPUs." It seemed that acquiring more high-end accelerators would enable the training of larger models. However, as AI clusters expand from hundreds to tens or even hundreds of thousands of GPUs, computational efficiency is no longer determined solely by the performance of individual chips, but by the ability of GPUs to exchange data at high speeds. If network latency, congestion, or packet loss causes GPUs to wait, even the most expensive chips may end up as idle capital locked inside server racks.
Training large models does not involve evenly distributing tasks for independent completion; instead, it requires repeated parameter synchronization across different GPUs. Collective communication operations such as All-Reduce and All-to-All generate massive amounts of east-west traffic. Congestion on any single link can slow down the entire task. This differs from traditional cloud data centers: while general applications are more tolerant of traffic fluctuations, AI training demands low latency, high bandwidth, and near-lossless transmission, with the slowest node often dictating overall speed. Consequently, the network has evolved from peripheral equipment to an integral part of the AI "supercomputer."
This issue also persists during the inference phase. Next-generation inference models need to handle longer contexts and more concurrent users, leading to the gradual disaggregation of compute, memory, and storage resources. The movement of model weights, KV cache, and intermediate results between servers makes inference costs increasingly dependent on network utilization. If bandwidth is insufficient, adding more GPUs may not proportionally increase throughput, but could instead drive up the cost per token.
From the chip race to a comprehensive revaluation of networking
This explains why market focus is rapidly shifting from 400G and 800G to 1.6T networking. Arista is set to launch its next-generation 1.6T platform in June 2026, offering a single-system bandwidth of 102.4 Tbps; Broadcom's Tomahawk 6 switch ASIC also provides 102.4 Tbps capacity, reflecting that switch upgrade cycles are catching up with accelerator iterations.
Beneficiaries extend beyond switch manufacturers to include network chips, NICs and DPUs, optical modules, DSPs, connectors, optical fibers, and network management software. As speeds increase to 800G and 1.6T, the distance for stable electrical signal transmission shortens, raising the proportion of optical interconnects; however, the number of optical modules, power consumption, and failure rates also rise accordingly. This makes silicon photonics, linear pluggable optics, and Co-Packaged Optics (CPO)—which integrates optical components with switch chips—the next focal point of competition. NVIDIA’s Spectrum-X photonics switch, planned for release in the second half of 2026 with a maximum bandwidth of 409.6 Tbps, aims to break through the limitations of power consumption and signal integrity associated with electrical interconnects.
Another key theme is the rivalry between InfiniBand and Ethernet. InfiniBand holds a first-mover advantage in large-scale AI clusters due to its low latency and mature lossless transmission capabilities, while Ethernet benefits from openness, a broad ecosystem, and multi-vendor interoperability. The Ultra Ethernet Consortium has released version 1.0 of its specifications, aiming to re-optimize Ethernet—from NICs and switches to transport protocols and congestion control—to transform traditional networks into architectures suitable for AI and high-performance computing.UEC This suggests that no single technology will dominate entirely in the future; instead, we may see dedicated high-speed interconnects used within racks, while enhanced Ethernet scales connectivity between racks.
However, the narrative that 'networking is taking the baton from GPUs' should not be oversimplified to mean that all optical communication stocks will benefit in the long term. Network equipment still faces price competition, and optical modules have distinct product cycles. Factors such as the ramp-up timeline for 1.6T volumes, yield rates, thermal management, and customer certification may fall short of expectations. While CPO offers energy savings, it introduces challenges such as difficult maintenance, supply chain restructuring, and a lack of unified standards. Companies with a true moat are those that can not only produce higher-speed components but also control power consumption, enhance reliability, and create synergy among switch chips, optical hardware, and network software.
GPUs remain the core of AI data centers, but being the core does not mean they are the sole bottleneck. As chip supply gradually increases, investment returns will depend more on how effectively the entire cluster is utilized. The next phase of the AI infrastructure race is no longer just about who owns the most GPUs, but who can make tens of thousands of GPUs operate as a single computer. The capital market’s valuation focus for AI is shifting from isolated compute power to the entire network connecting that compute power.
Yee Hap Holdings Investor Relations Department
(Chips and Compute Power Series #86)
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Comments
to post a comment
