Anthropic Races Toward US IPO! Can It Surpass SpaceX?
Cerebras Systems went public on Nasdaq in May 2026, marking one of the most talked-about IPOs in the AI chip market in recent years. The company priced its offering at $185 per share, selling 30 million shares to raise $5.55 billion. On its first trading day, the stock opened as high as $350—89% above the offer price—and its fully diluted valuation briefly exceeded $100 billion. However, the frenzy quickly cooled; by July 20, the share price had fallen to around $173, dipping below the IPO price. This dramatic swing reflects the market’s conflicting views on Cerebras: it may represent the most original AI computing architecture outside NVIDIA—or an expensive experiment that, while technically brilliant, remains commercially unproven.Cerebras IPO Pricing Announcement
The traditional approach with GPU clusters involves connecting hundreds or even tens of thousands of chips via high-speed networks to jointly train or run large models. The problem is that as model size grows, so does the time spent transferring parameters, synchronizing computations, and redistributing workloads among chips. Investors are familiar with NVIDIA’s GPUs but often overlook that its moat isn’t just the chips—it’s also CUDA software, NVLink interconnects, networking hardware, and the entire cluster management ecosystem.
Cerebras has taken a different path: instead of dicing wafers into numerous small chips, it fabricates an entire wafer as a single massive processor. Its third-generation Wafer-Scale Engine uses a 5-nanometer process, spans 46,225 square millimeters, integrates 4 trillion transistors and 900,000 AI compute cores, features 44GB of on-die SRAM, and delivers memory bandwidth of 21 petabytes per second. The core idea is straightforward: rather than having thousands of GPUs exchange data across a network, keep vast amounts of computation and communication within a single wafer to reduce latency and system complexity.Cerebras WSE-3 Product Specifications
The challenge isn't just a single GPU, but the entire CUDA ecosystem
The most compelling use case for wafer-scale architecture is large-model inference requiring extremely low latency. In traditional GPU clusters, model data must be moved between different GPUs during answer generation, and network communication can become a bottleneck. Cerebras, by contrast, leverages its exceptionally high on-chip bandwidth to accelerate token-by-token generation. As AI evolves from simple Q&A to multi-step reasoning, intelligent agents, and real-time code generation, users care not only about cost per million tokens but also about wait time for responses. What Cerebras truly sells is a 'time premium.'
The company has expanded beyond selling CS-3 hardware to offering cloud-based inference services. Revenue in 2025 is projected at approximately $510 million, up from $290 million in 2024; in Q1 2026, revenue grew 94% year-over-year to $193.4 million, with net losses narrowing to $14 million. Even more significantly, the company signed a multi-year partnership with OpenAI to provide 750MW of inference computing capacity, with an additional 1.25GW option. This positions Cerebras not merely as a seller of 'giant chips,' but as an aspiring operator of AI inference infrastructure.
However, order size does not equate to realized profitability. Cerebras forecasts an adjusted gross margin of only 38% to 41% for 2026—far below NVIDIA’s roughly 75%. Larger wafers increase the difficulty of manufacturing yield, thermal management, maintenance, and capacity scheduling. Moreover, its cloud business requires upfront commitments to data center leases, power procurement, and substantial capital investment. The company has even temporarily leased back its own systems from customers to meet short-term demand, indicating that production ramp-up still lags behind orders.Reuters' first post-IPO earnings report
Customer concentration is another unavoidable issue in valuation. In Q1 2026, Mohamed bin Zayed University of Artificial Intelligence (UAE) and G42 accounted for 63% and 11% of revenue, respectively; the OpenAI deal may represent a significant portion of future revenue. In other words, while Cerebras has successfully avoided dependence on NVIDIA hardware, it has concentrated its commercial risk among a few hyperscale customers, data center delivery timelines, and power supply availability.Cerebras Q1 Report
Cerebras also cannot fully replace GPUs. GPUs can be scaled incrementally based on demand and are suitable for a broad range of tasks—including training, inference, scientific computing, and graphics processing—whereas wafer-scale systems resemble highly specialized AI supercomputers. More importantly, CUDA has accumulated a vast ecosystem of development tools, skilled talent, and software libraries; the cost for enterprises to switch architectures far exceeds the theoretical speed comparison between two chips.
Therefore, Cerebras’ challenge to GPU clusters is more likely to result in partial substitution rather than full disruption. If it can demonstrate a lower total cost of ownership in low-latency inference, large-model training, and specific scientific workloads, it could carve out a meaningful share from NVIDIA’s massive market. However, if revenue growth remains heavily reliant on a few customers and continuous capital investment, its valuation will ultimately hinge on gross margins, cash flow, and capacity utilization.
(Chips and Computing Power Series #77)
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Comments
to post a comment
