English
Back
Open Account
Yee Hop Holdings
wrote a column · Aug 10 01:39

Arista's AI Ethernet Story: A Second Path Beyond InfiniBand

In the AI race, GPUs typically steal the spotlight. However, when thousands or even hundreds of thousands of accelerators are connected into a computing cluster, the network is no longer just a supporting role for data transfer—it becomes critical to the entire system’s efficiency. If the network experiences congestion, latency, or packet loss, expensive GPUs sit idle waiting for data, and multi-billion-dollar AI computing centers cannot operate at full capacity. This is precisely the opportunity Arista Networks is aiming to seize: building high-speed Ethernet-based AI backend networks as an alternative to NVIDIA-dominated InfiniBand.
InfiniBand has long been the preferred choice for high-performance computing and large-scale AI training clusters. It offers low latency, lossless transmission, adaptive routing, and in-network computing capabilities. Combined with NVIDIA’s GPUs, ConnectX NICs, and Quantum switches, it forms a highly integrated end-to-end system. NVIDIA’s latest Quantum-X800 platform delivers 800 Gbps per port and uses SHARP technology to handle collective communications within the network—the advantage lies not just in speed, but in the co-design of hardware, software, and computational workloads.
However, this integrated model also means enterprises become more dependent on a single vendor. Large cloud service providers already use Ethernet to manage front-end, storage, and general data center networks. If they deploy separate InfiniBand architectures for AI clusters, they must maintain two sets of equipment, technical teams, and management tools. As cluster scale grows, considerations such as supply chain options, equipment costs, and operational consistency gain importance, turning Ethernet from a 'lower-performance but cheaper' option into an open standard actively championed by large customers.
Can Ethernet Turn Its Openness Advantage into Predictable Performance?
Arista’s core value proposition isn’t reinventing Ethernet, but rather adapting it for AI workloads. Traditional data center traffic tends to be dispersed, whereas AI training generates brief bursts of a small number of extremely large, synchronized data flows—prone to traffic collisions, tail latency, and localized congestion. Through its Etherlink platform, EOS network operating system, RDMA load balancing, ECN and PFC congestion management, and CloudVision monitoring tools, Arista aims to enable networks to observe AI-related traffic, identify bottlenecks, and dynamically adjust routing paths.
Its Cluster Load Balancing technology routes data flows based on RDMA queue states to balance utilization between leaf and spine switches; AI Analyzer monitors packet loss, buffer occupancy, NIC errors, and job completion times at high frequency. Etherlink currently supports 400G and 800G platforms, capable of scaling from small clusters to deployments with over 100,000 accelerators.Arista AI Networking Solutions The company has also announced its 1.6T platform, further extending Ethernet from rack-to-rack 'scale-out' networking to intra-rack 'scale-up' connectivity.
Another driving force behind the Ethernet roadmap comes from the Ultra Ethernet Consortium. The consortium's 1.0 specification covers NICs, switches, optical components, cables, and transport protocols, aiming to enhance RDMA, bandwidth, latency, congestion control, and large-scale deployment capabilities while preserving Ethernet interoperability. Once the standard matures, customers will be able to mix and match products from different vendors, avoiding vendor lock-in and enabling Broadcom, Arista, AMD, and other networking equipment providers to build a more comprehensive ecosystem.
However, openness does not equate to simplicity. InfiniBand’s advantage lies in its mature, end-to-end tuning methodology, whereas Ethernet requires precise coordination among switches, NICs, optical modules, cabling, and congestion control mechanisms. Even minor configuration deviations can lead to increased tail latency, slowing down entire training workloads. Large cloud providers possess sufficient engineering talent to optimize their networks internally and reduce costs through scale; enterprises seeking only rapid deployment of NVIDIA GPU clusters may still find pre-integrated InfiniBand solutions more reliable. Therefore, Ethernet’s total cost advantage must factor in engineering, testing, and troubleshooting expenses—not just switch pricing.
Arista’s financial performance already reflects strong market demand for high-speed networking. The company reported $9 billion in revenue for 2025, and Q1 2026 revenue rose 35.1% year-over-year to $2.709 billion. However, this does not mean Ethernet has comprehensively overtaken InfiniBand, as Arista’s revenue still includes traditional cloud, data center, campus, and routing businesses, and its AI orders remain highly dependent on capital expenditure cycles of a few hyperscale customers.
More notably, NVIDIA is also developing its own Spectrum-X Ethernet solution. This suggests the real competition may not be a binary choice between InfiniBand and Ethernet, but rather who controls the software-hardware value stack for AI networking. Cutting-edge training clusters demanding extreme performance may still favor InfiniBand, while cloud platforms prioritizing multi-vendor support, operational consistency, and cost flexibility will likely accelerate Ethernet adoption. Arista’s opportunity isn’t to eliminate InfiniBand, but to demonstrate that open networking can deliver predictable performance approaching that of proprietary architectures. If successful, AI infrastructure will no longer have only one path—NVIDIA’s.
(Chips and Compute Power Series #83)
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
88K Views
Report
Comments
Write a Comment...