NVIDIA's revenue doubles, beating expectations; is the AI trade narrative making a comeback?
The generative AI boom has propelled Nvidia to the throne of computing power, while simultaneously triggering a shared anxiety among global tech firms: when the most critical production tool is controlled by a single supplier, chip pricing, delivery lead times, and bargaining power become difficult to manage. Consequently, Google is expanding its TPU deployments, Amazon is advancing Trainium, Microsoft and Meta are investing in in-house chips, and AMD is directly challenging Nvidia with its Instinct series. This has given rise to market speculation about 'de-Nvidia-ization,' as if launching another accelerator alone could break Nvidia’s dominance.
Reality, however, is far more complex. Nvidia’s moat has never been just about a GPU—it’s an integrated ecosystem comprising CUDA software, development tools, model libraries, high-speed interconnects, server architectures, and engineering talent. What enterprises buy isn’t merely chip specifications but the ability to rapidly train, deploy, and continuously upgrade models. Switching to another platform—even if the per-chip cost is lower—may require rewriting code, retesting models, retraining teams, and even redesigning data center networks. While compute power may become cheaper, time-to-market could be delayed by six months, a cost potentially too steep for companies racing in the AI arena.
True competition lies not in chip price, but in cost per token
Alternatives do have opportunities. AMD’s MI300X, equipped with 192GB of HBM3 memory, is better suited for large-model inference and memory-intensive workloads; Google’s sixth-generation TPU, Trillium, delivers 4.7x higher peak compute performance per chip compared to the TPU v5e and improves energy efficiency by over 67%; AWS, meanwhile, pairs Trainium2 with its Neuron software stack to support PyTorch, vLLM, and large open-source models, significantly lowering migration barriers.
However, chip comparisons shouldn't focus solely on peak compute performance. What truly determines cost is the total cost per token, including chip purchase or rental price, memory capacity, interconnect efficiency, power consumption, cooling, cluster utilization, and software engineering effort. Even if a chip boasts exceptional theoretical performance, its actual cost may not be lower if the model can only achieve 60% of that performance or requires lengthy debugging times. This also explains why cloud giants can afford to develop their own ASICs, while most small and medium-sized enterprises continue using Nvidia: the former process trillions of tokens daily, allowing them to amortize design and software costs over massive, steady workloads; the latter lack sufficient workload volume, making mature platforms more cost-effective.
'De-Nvidiafication' will first occur in the inference market, not in cutting-edge model training. Training workloads evolve rapidly and require highly flexible general-purpose computing capabilities—Nvidia still holds an advantage in hardware-software co-optimization, cluster scale, and fault management. Inference, by contrast, is relatively standardized; once a model is finalized, companies can reduce costs through quantization, pruning, and dedicated chips. Therefore, search, recommendation, advertising, and fixed-model services are best suited for migration to TPUs, Trainium, or custom ASICs, while exploratory R&D and large-scale model training will likely remain on Nvidia platforms.
In the Chinese market, 'De-Nvidiafication' carries an additional layer of supply chain security considerations. Domestic AI chips must not only catch up in computational performance but also build out compilers, operator libraries, interconnect technologies, server ecosystems, and developer communities. While policy mandates and export controls can create demand for alternatives, they cannot automatically close the software gap. If enterprises maintain two or even three parallel technology stacks to meet self-reliance requirements, short-term costs may actually rise. The true benchmark for successful domestic substitution isn't merely whether a chip powers on, but whether it can stably run mainstream models over the long term and reduce migration and maintenance costs to acceptable levels.
Thus, the future won't see 'Nvidia being replaced,' but rather a gradual stratification of the AI compute market: Nvidia will continue dominating cutting-edge training; large cloud platforms will use in-house chips for fixed workloads; AMD will target customers seeking general-purpose GPUs while wanting to reduce vendor dependency; and domestic chips will expand within specific markets and policy environments. Nvidia’s market share may decline, but as total AI compute demand continues to grow, its revenue may not necessarily shrink.
Therefore, 'De-Nvidiafication' is more accurately described as 'reducing marginal dependence on Nvidia.' The dream is to escape reliance on a single supplier, but the reality is that the software ecosystem is hard to bypass—and ultimately, the outcome hinges on cost curves. Only when alternative platforms can consistently lower the cost per token without sacrificing model performance or development speed will De-Nvidiafication evolve from a corporate slogan into a genuine industry trend.
Yee Hap Holdings Investor Relations Department
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Comments
to post a comment
1
