Author, Source: Semiconductor Industry Hub
At GTC Taipei, Jensen Huang presented an equation: 'Compute equals revenue.' The underlying message is clear: what AI companies, cloud providers, and enterprise customers are buying isn’t just hardware—they’re investing in intelligent capacity that will generate sustainable future revenue.
NVIDIA is no longer just a chip company—it now resembles the 'general contractor' of a massive token factory. And Jensen Huang aims to be the factory’s chief architect, ensuring everything works seamlessly together. What matters in this factory isn’t the capability of any single piece of equipment, but the output of the entire production line per unit of energy consumed. Jensen Huang provided NVIDIA’s answer through the five-layer cake theory: every layer—from energy to applications—is redefined as a component of the token production system. These five layers, from bottom to top, are 'energy, chips, infrastructure, models, and applications.'

While competitors are still optimizing parameters within a single layer, NVIDIA is already optimizing and strategically positioning every layer of the cake, leveraging multiplicative effects that leave rivals far behind.
01 Energy: In AI factories, computing power starts with electricity
From AI-driven wind and solar power forecasting to high-voltage direct current (HVDC) power distribution infrastructure, and further to intelligent energy storage and stable power supply, the energy layer addresses how AI factories can operate efficiently and ensure stable token production.
The competition among future AI factorieswill first be a contest over 'how much intelligence each kilowatt-hour of electricity can produce.'. Jensen Huang introduced the concept of 'Token Factory Economics,' stating that under fixed power constraints, the core metric of competitiveness is no longer peak computing power but rather 'Tokens per Watt'—the number of tokens produced per watt of power consumed. NVIDIA's Vera Rubin high-performance computing platform has achieved a 10x improvement in performance per watt, reducing the cost per token to one-tenth of its previous level.Second is competition over energy supply capacity. In 2025, NVIDIA’s venture arm, NVentures, entered the energy sector for the first time, participating in TerraPower’s $650 million funding round and investing in Commonwealth Fusion Systems (CFS). Beyond exploring cutting-edge clean energy, NVIDIA has also invested in a range of small and large enterprises focused on optimizing data center power solutions—including chips, computing power, and grid management—such as Emerald AI and Utilidata. In August 2025, NVIDIA updated its official website to include partners for its 800V DC power architecture, with Chinese companies Innoscience and Megmeet listed among them.
When electricity becomes a hard constraint for AI factories, keeping pace with innovation at the energy layer is equivalent to securing pricing power over upstream raw materials. NVIDIA’s investments, supply chain integration, and technological enablement in the energy layer ensure the stability of supply across the entire five-layer stack.
02 Chips: The Vera Rubin Platform Completes the Computing Power Suite
As a prime example of NVIDIA’s extreme co-design philosophy, the Vera Rubin platform integrates the Vera CPU, Rubin GPU, NVLink 6 switch, ConnectX-9 SuperNIC, BlueField-4 DPU, Spectrum-6 Ethernet switch, and Groq 3 LPU into a unified system.
NVIDIA bundles these chips together because AI models are extremely sensitive to communication bandwidth and latency. NVLink 6 and CPO (Co-Packaged Optics) switches address software-level communication bottlenecks through tight physical coupling. More importantly, in the era of co-design, the cost of heterogeneous integration has risen sharply.
First, let’s look at the Vera CPU, NVIDIA’s first CPU purpose-built for agentic AI. The Vera CPU features 88 NVIDIA-designed Olympus cores, spatial multithreading technology, and an LPDDR5X memory subsystem with up to 1.2 TB/s of bandwidth. In agentic workloads, Vera completes tasks 1.8 times faster than x86 CPUs. More importantly, as part of the Vera Rubin NVL72 platform, the Vera CPU pairs with NVIDIA GPUs via NVLink-C2C interconnect technology, delivering up to 1.8 TB/s of coherent bandwidth—approximately seven times that of PCIe Gen 6—enabling high-speed data sharing between CPU and GPU.
Next, networking:NVLink 6 provides the fast, seamless GPU-to-GPU communication required by today’s large-scale MoE models. Each GPU supports 3.6 TB/s of bandwidth, and each Vera Rubin NVL72 rack delivers 260 TB/s. Spectrum-X is the world’s first mass-produced silicon photonics Ethernet switch platform. The next-generation Spectrum-X switches are built using CPO (co-packaged optics) technology, integrating silicon photonics components directly with the switch ASIC to further reduce power consumption, enhance reliability, and boost AI productivity. The ConnectX-9 SuperNIC offers up to 1.6 Tb/s of throughput, breakthrough acceleration capabilities, and optimized network performance, delivering ultra-low-latency 800 Gb/s networking that accelerates data transfer, optimizes RoCE performance, and ensures consistent, predictable network performance for demanding AI workloads.
Finally, storage:The BlueField-4 DPU is NVIDIA’s data processing unit for data centers, part of the full-stack BlueField platform and also integrated into the Rubin platform. NVIDIA has built the BlueField-4 STX storage architecture based on the BlueField-4 DPU. The first rack-scale deployment integrates the new CMX Context Memory Storage platform. CMX treats the KV Cache as a new AI-native data type, specifically designed to store and retrieve KV Cache data generated during LLM inference, making context a high-bandwidth resource shared across AI cluster systems. At GTC Taipei 2026, NVIDIA will unveil an upgraded Vera BlueField-4 STX platform that brings NVIDIA DOCA security libraries and microservices into the AI storage layer, helping enterprises protect data, agents, and context memory when deploying agentic AI in production environments.
In comparison, AMD’s Instinct MI350+EPYC disaggregated architecture follows a modular approach, but severe CPU-GPU communication bandwidth bottlenecks significantly limit data transfer efficiency between CPU and GPU—especially in agent scenarios requiring frequent data exchanges. Meanwhile, proprietary chip ecosystems like Google TPU, AWS Trainium, and Microsoft Maia are vertically integrating their own closed-loop solutions, but none can match NVIDIA’s full-stack scale. When the Vera CPU is tightly coupled with Rubin GPUs via NVLink-C2C at 1.8 TB/s bandwidth, customers will find it extremely difficult to replace Vera with x86 CPUs, as doing so would result in a substantial loss of token generation efficiency.
03 Infrastructure: From 'Buying Servers' to 'Building Factories'
Chips must operate within AI infrastructure. Jensen Huang has separately designated infrastructure as the third layer, encompassing land, power, cooling, networking, and the systems that orchestrate thousands of processors into a single machine. Essentially, this is what NVIDIA repeatedly refers to as the 'AI factory.' NVIDIA aims to shift customers from 'buying servers' to 'building AI factories.'
So how should one build an AI factory? IBM defined mainframe standards with System/360, and AWS established cloud architecture standards with its Well-Architected Framework. NVIDIA is defining the AI factory standard through its AI Infrastructure Hardware Guidelines and NVIDIA Dynamo.Leveraging its strong influence in AI computing centers, NVIDIA has released the Vera Rubin DSX AI Factory reference design—a guide for building co-designed AI infrastructure.The Vera Rubin DSX AI Factory reference design outlines how to design, build, and operate the entire AI factory infrastructure stack, covering compute, Spectrum-X Ethernet networking, and storage to deliver repeatable, scalable, and high-performing cluster performance. The documentation within the reference design also provides industry partners with best practices for designing, building, and operating power, cooling, and control systems to enable seamless hardware-software integration and scalable deployment.NVIDIA clearly aims for future AI Factories to adopt its own standards, thereby locking in long-term demand for NVIDIA’s chip products. Replacing other chips would involve more than swapping out a board—it would require overhauling the entire factory’s design logic. And the 'operating system' of this factory is Dynamo.Dynamo is an open-source, low-latency, modular inference framework designed to serve generative AI models in distributed environments. It enables seamless scaling of inference workloads across large GPU clusters, offering intelligent resource scheduling and request routing, optimized memory management, and efficient data transfer.
Section 04: Models—Accelerating Agent Deployment to Boost AI Factory Efficiency
NVIDIA not only provides hardware solutions but is also actively expanding into the model layer.Nemotron is currently NVIDIA’s most widely deployed proprietary model family.Its recently launched Nemotron 3 Ultra is a mixture-of-experts model with 550 billion total parameters and 55 billion active parameters per forward pass, capable of handling orchestration and complex reasoning tasks in autonomous workflows—such as making architectural decisions during extended coding sessions, synthesizing insights across hundreds of research sources, and validating thousands of interdependent constraints.Cosmos is NVIDIA’s world foundation model tailored for physical AI (including robotics and autonomous driving), primarily designed to generate physics-compliant synthetic data and action policies, thereby reducing industry-wide data acquisition costs.Cosmos 3 addresses a core challenge in physical AI: enabling robots, intelligent vehicles, or vision-based agents to generalize effectively in the real world despite limited training data and fragmented simulation stacks.
For vertical industries, NVIDIA also offers specialized models—for example, BioNeMo, its vertical model platform for life sciences. NVIDIA isn't aiming to become 'the next OpenAI'; instead, it seeks to demonstrate that its hardware can run these models at a lower cost per token and prove that NVIDIA’s full-stack solution delivers the best tokens-per-watt efficiency. This encourages customers—operating under the equation where 'compute power equals revenue'—to choose NVIDIA hardware for running these already-validated, highly efficient models.
05 Applications: The Real Market
At the top layer of the five-tier cake lies the domain that directly serves end users and generates tangible value—encompassing agents, enterprise automation, and physical AI applications in areas such as drug discovery, industrial robotics, and autonomous driving. This is the ultimate monetization layer of NVIDIA’s full-stack strategy and the output end of the entire 'AI factory.' By defining agent standards, providing physical AI development platforms, and certifying industry-specific skill modules, NVIDIA enables a thriving ecosystem of AI applications to flourish within its own infrastructure.
Agentic AI is NVIDIA’s current strategic focus.Following the viral success of OpenClaw, NVIDIA swiftly responded by launching NemoClaw, an enterprise-grade secure deployment software stack designed to complement OpenClaw with sandbox isolation, permission controls, and scalable operational capabilities—aiming to capitalize on OpenClaw’s momentum. To further bolster security, NVIDIA also introduced Verified Agent Skills, which embed transparency, provenance tracking, security validation, and authenticity checks directly into the agent’s capability layer, giving developers greater confidence when scaling autonomous agents.Physical AI is NVIDIA’s strategic lever for the future.Recently, NVIDIA partnered with Unitree Robotics and Sharpa to unveil an open humanoid robot reference design built on the NVIDIA Isaac GR00T platform. According to the official roadmap, Unitree Robotics will commercially launch the Isaac GR00T humanoid robot reference platform by the end of 2026.
From energy and chips to infrastructure, models, and applications, NVIDIA drives innovation across its full-stack technology. Today, all five layers of the 'cake' are firmly in place. At the Chain Expo held in Beijing on June 22, we saw localized implementation cases in China and met NVIDIA’s growing roster of Chinese partners across its full-stack ecosystem: In the energy layer, companies like Xingneng Xuanguang and Energy Singularity leverage accelerated computing and digital twin technologies to advance next-generation energy R&D, while Jinpan Technology and Chint use NVIDIA AI for full-lifecycle design and operations; in the chip layer, NVIDIA provides Chinese partners with platforms like Spectrum-X Ethernet; in infrastructure, CloudNine InfoTech has developed the iMDC liquid-cooled containerized AI computing center solution using NVIDIA products; in the model layer, SGLang achieves high-performance inference for DeepSeek-V4 on NVIDIA GPU servers; and in applications, concept humanoid robots from Zhongjian Intelligence and Lingrui P1 heavy-duty industrial quadruped robots—both built on NVIDIA solutions—were showcased simultaneously.
06 NVIDIA Wants to 'Sell the Entire Five-Layer Cake Together'
What NVIDIA truly aims to build is a complete, closed-loop system encompassing all five layers of the cake. The Vera Rubin platform integrates CPU, GPU, networking, and storage into a single package; Dynamo acts as the operating system for the AI factory, orchestrating resources across the entire stack; DSX Digital Twin validates the flow of every watt of power before construction even begins; and the Enterprise Agent Toolkit fully embeds business logic into NVIDIA’s ecosystem. NVIDIA doesn’t just supply the bricks—it also provides the architectural blueprints (DSX Blueprint), the construction crew (Dynamo orchestration), and even the skilled workforce (NemoClaw).
Jensen Huang's ambition extends far beyond being a token factory for the digital world. Today, agents like OpenClaw are beginning to serve as the 'operating systems for agent computers.' In the future, everything—from cloud-based agents to personal computers, autonomous vehicles, humanoid robots, and even satellites, base stations, and factory equipment—will run agents. And the foundational raw material powering all this is tokens. NVIDIA is redefining the metrics of AI commerce through token economics, encapsulating vertical-specific know-how into pricetag-ready agent capabilities via CUDA-X skillification, and employing full-stack co-design to ensure that no alternative solution can produce equivalent tokens under the same power budget. Customers who choose NVIDIA’s full-stack solution gain access to proven, quantifiable intelligent throughput; competitors opting for non-full-stack approaches will face endless system integration and optimization challenges.
All five layers of the cake are essential—none can be omitted. Jensen Huang’s strategy has already gone well beyond chips and data centers; he is building an integrated five-in-one technological and commercial logic.
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Comments
to post a comment
