"AI Bottleneck Trade" Ignites Upstream Sector—Who’s Raking in the Profits?

💡 Key Insight
– Moonshot AI has released Kimi K3, an open-source model with 2.8 trillion parameters, placing its overall capabilities among the global leaders,second only to the Claude Fable series and top-tier GPT models in performance, and on July 28 furtheropen-sourced key infrastructure (Infra) technologies。
– The market had previously worried that improved model efficiency would reduce demand for computing power, butK3 suspended new user sign-ups within 48 hours of launch due to computing capacity shortages,coupled with its exceptionally high deployment barrier (2.8 trillion parameters, requiring a 64-GPU super node),proving that demand for computing power is rising, not fallingto extendContinuously driving demand across the entire supply chain, including GPUs, HBM, and networking equipment。
– The narrowing gap between open-source and closed-source capabilities is exerting downward pressureU.S. closed-source vendors (Anthropic, OpenAI) have revised down their pricing and revenue expectations, at the same timeStrengthening a self-reliant domestic computing power ecosystem, weakening NVIDIA's pricing power in the Chinese market and its CUDA ecosystem; 'AI central bank'-style financing guarantees are also amplifying systemic risks across the industry chain.
– Looking ahead, both Chinese and U.S. large-model vendors willaccelerate R&D and price competition, with pricing shifting from technology premium tovalue for money,ultimately benefiting the entire AI infrastructure layer。
I. Event Recap
1.1 Kimi Launches K3 Model
On July 16, Moonshot AI released its new open-source foundation model, Kimi K3, with a parameter count of2.8 trillion, along with the simultaneous launch of API services and developer documentation. The model ranks among the global leaders in code generation, knowledge work, long-context retrieval, and agent tasks, with overall capabilities second only to Anthropic’s Claude Fable series and top-tier GPT models. Initial market concerns that efficiency breakthroughs would reduce GPU demand caused semiconductor stocks to dip temporarily; however, shortly after the launch, Moonshot AIWithin 48 hourswas forced to suspend new C-end user subscriptions due to insufficient computing capacity,demonstrating that compute demand has not weakened, leading NVIDIA’s stock to rebound subsequently.. Amid competition between Chinese and U.S. models,market attention has shifted to models' monetization capabilities and whether U.S. AI companies can deliver on expected returns。
On July 28, during Kimi's K3 Technology Open Day, Moonshot AI open-sourced key infrastructure technologiesMoonEP, FlashKDA, AgentEnv, prompting overseas AI infrastructure providers such as Nebius, Baseten, and Fireworks to announce Day 0 compatibility, allowing overseas cloud vendors to directly benefit from this efficiency improvement.
Chart 1: Kimi K3 Performance Summary

Source: Public information, compiled by Futu Securities
1.2 K3 Performance Specifications and Technical Architecture
Kimi K3 has a total of 2.8 trillion parameters, making itthe world's largest open-source model to date, adopting a Mixture-of-Experts (MoE) architecture, activates only 32 billion parameters per inference, resulting in computational costs equivalent to those of a smaller 32-billion-parameter model. Its core technologies include Stable Latent MoE and the KDA hybrid linear attention mechanism:
– Sparse Expert Architecture: Only 16 out of 896 experts are activated per inference, and the WideEP expert parallelism strategy distributes expert modules across multiple GPUs, expanding model capacity while controlling per-token computation.
– Hybrid Linear Attention: The KDA mechanism can reduceKV Cache bandwidth pressure by up to approximately 10x,enabling inference with ultra-long contexts of up to 1 million tokens, with overall scaling efficiency improved byapproximately 2.5x compared to the previous generation.。
The actual impact of this architecture on computing demand is not a reduction but an increase:

Figure 2: K3 Performance and Technical Architecture

Source: Compiled by Futu Securities
1.3 Open-source vs. Closed-source
The performance gap between cutting-edge Chinese and U.S. models is substantially narrowing; under Artificial Analysis benchmarks, their scores now differ by only3 points. Chinese open-source models, thanks to their exceptional cost-performance ratio, have already been integrated byover 80% of U.S.-based AI startups. Taking Kimi K3 as an example, the cost per task is approximately $0.94, roughly half that of Opus 4.8 and one-third that of Fable 5,forcing overseas closed-source vendors to lower prices.。
U.S. closed-source vendors (such as Anthropic’s Claude Fable 5, priced at $10 per million input tokens and $50 per million output tokens) rely on a"capability tiering + controlled release of safety features"model to sustain high premiums, but the foundation of their premium pricing—performance leadership and scarcity of alternatives—is being eroded by open-source models.If the performance gap continues to narrow, closed-source models will face dual pressures on both market share and pricing premiums., which is also the core rationale behind the market’s downward revision of Anthropic and OpenAI’s revenue growth trajectories.
Exhibit 3: Open-source Path vs. Closed-source Path

Source: Compiled by Futu Securities
II. AI Transmission Chain
2.1 Demand for GPUs and Other Segments
The market previously interpreted K3’s linear attention mechanism as signaling 'compute deflation,' but its extremely high deployment barrier has invalidated this logic—K3 suspended new user subscriptions within 48 hours of launch due to compute capacity shortages,Compute demand is far from peaking. In the long run, declining inference costs will drive a surge in usage volume, continuously boosting demand for GPUs, HBM, and networking equipment.
– Training and general-purpose compute chips offer the highest certainty: K3 heralds the emergence of more large-scale AI models from both China and the U.S., representing net new demand rather than substitution; with 2.8 trillion parameters, over 1.5TB of HBM memory, and a 64-GPU super-node deployment threshold, it directly drives demand for rack-scale systems such as GB200/GB300 NVL72.
– Core use cases for inference compute are shifting toward AI coding and long-horizon agent tasks: A single request consumes significantly more GPU time than a typical conversation; the more capable the model and the longer the context, the higher the computational demand for inference.
– The Jevons Paradox is now fully in effect: K3’s low-cost, high-performance offering (input priced at $0.3 per million tokens) lowers the barrier to adoption, causing a surge in total token usage. The decline in unit cost is offset by explosive growth in usage volume, leading to an overall increase—not decrease—in total resource consumption. However, if major players like OpenAI and Anthropic widely adopt KDA-like solutions, demand for storage and interconnect bandwidth in long-context inference could temporarily decline.
Exhibit 4: The Jevons Paradox

Source: Public information
Next is the memory/storage segment: K3’s 2.8 trillion parameters generate massive demand for HBM memory, and its 1-million-token context window intensifies—rather than alleviates—the pressure on DDR5 and SSDs from KV caching. Since K3 is open-source and deployable globally,deployments both in China and overseas require the same HBM and DDR5。
2.2 Domestic Computing Power
Domestically developed models are better optimized for domestic computing infrastructure, and when combined with policy mandates and supply chain security requirements,K3’s success directly boosts end-to-end demand for domestic GPUs, servers, and networking equipment; K3 demonstrates that domestic large AI models have already caught up with leading global standards,strengthening the investment narrative around domestic substitution。
Recently, the domestic computing power supply chain has been impacted by ChangXin Memory Devices' (CXMT) IPO, with effects concentrated in two areas: in the short term,a capital-siphoning effect(expected to account for approximately 25% of STAR Market trading volume, though this impact will gradually diminish); in the medium to long term, it serves asa valuation anchor—the market is concerned that memory prices may peak as CXMT ramps up production, potentially causing the memory sector’s valuation to peak and decline even before DRAM prices do. CXMT’s pricing also reflects the market’s view on the sustainability of the current cycle.
2.3 NVIDIA’s CUDA Ecosystem Challenges
NVIDIA currently faces three key challenges: domestic substitution in China, erosion of its CUDA ecosystem, and market concerns over its role as an "AI digital central bank."

The supply-demand balance for computing power has shifted from "extreme shortage" to "moderate shortage.", which is the core reason why NVIDIA's stock price, despite rebounding after K3 suspended its subscription, remains under pressure.
2.4 Investment Thesis Summary
AI stock valuations have already priced in substantial assumptions about future earnings. Sustaining these valuations requires even more optimistic assumptions—such as greater productivity gains, faster application adoption, higher profit-sharing by enterprises, or lower discount rates. The market is currently focused on two key questions: how large the total value of AI will be, and who will ultimately capture that value.
– AI Adoption Rate: Current use cases are concentrated in coding; adoption in other areas remains below expectations, raising market concerns about AI monetization rates and return on capital expenditure (Capex ROI).
– Value Distribution: K3’s pricing aligns with the standard price of Claude Sonnet 5, yet its performance is already approaching that of Fable 5 and GPT-5.6. Its cache-hit input price is only $0.30 per million tokens (one-tenth of the miss price), offering a significant cost advantage in real-world programming scenarios. This implies thatunder intensifying model competition, Chinese companies could achieve higher profit margins than their U.S. counterparts. If Chinese firms capture market share with lower pricing, U.S. AI companies and their supply chains may fail to realize their originally expected returns, potentially prompting cloud providers to slow capital expenditures, thereby impacting the entire AI investment chain.
Exhibit 5: Pricing of Major AI Models Domestically and Overseas

Source: Public information
3. Impact on Other Segments
3.1 Cloud Providers
K3 demonstrates that China's open-source approach is no less competitive than overseas models,The most certain opportunity lies in the surge in inference demand.Cloud providers' gross profit equals the AI inference fees paid by customers minus underlying compute costs. Efficiency gains enable the same hardware to handle more demand; meanwhile, under Jevons Paradox, total demand expands in tandem. The cost of long-context processing shifts from exponential to linear, while token consumption grows exponentially,further boosting profit margins.However, the U.S. Big Four cloud providers have already raised their capital expenditure forecasts. Google and Amazon’s free cash flow is expected to turn negative this year or next, and the market remains concerned in the near term about free cash flow and return on capital expenditure (Capex ROI).
3.2 Large Model Providers
Regardless of whether they pursue open-source or closed-source strategies, K3 shows that the large model industry has yet to produce a clear winner.No single vendor can sustain technological barriers or commercial premium pricing over the long term.。
– Chinese manufacturers: K3’s open-weight model and zero-cost trial directly undermine Zhipu AI’s API pricing power and its ability to command a commercial premium. Its cost advantage has also heightened market concerns about a price war, intensifying competitive pressure on MiniMax, which emphasizes cost-performance ratio. Moonshot AI plans to pursue a Hong Kong IPO within as little as six months (targeting a $30 billion valuation), which will serve as a key benchmark for Zhipu’s market cap and constrain its valuation upside.
– U.S. vendors: Prior to K3’s launch, OpenAI was already caught in a 'squeezed middle'—its high-end offerings pressured by Anthropic’s Opus series and its low-end segment eroded by Chinese open-source models. K3 delivers performance on par with or even better than GPT-5.5 at roughly one-third the price, directly compressing OpenAI’s API pricing room. Analysts forecast OpenAI’s revenue could peak as early as 2026, with annualized recurring revenue (ARR) reaching only about $11 billion by mid-2027 under a conservative scenario—essentially flat growth. Anthropic faces limited near-term impact (its ARR surged from $9 billion to $44 billion between January and May, doubling every six weeks), but signs of a potential inflection point in growth are emerging. Anthropic’s moat lies in high barriers to local deployment, product ecosystem lock-in (Claude Code, Agent, Fable), and potential regulatory barriers (Washington may restrict the use of Chinese AI models in the U.S.).
Both Chinese and U.S. vendors are accelerating R&D to counter competition, shifting pricing strategies from technology-driven premiums toward cost-performance. Ultimately, this R&D arms race benefits the entire AI infrastructure layer—GPUs, memory chips, and computing infrastructure.
Exhibit 6: Closed-Source Model ARR Comparison

Source: TickerTrends
[Investment Advisory Information]
Sun Bihan, Licensed Representative, Central Entity Reference Number: BWS708
Disclaimer
This report is prepared by Futu Securities International (Hong Kong) Limited (“Futu Securities”). Without the prior written consent of Futu Securities, this report and the information contained herein may not be (i) reproduced, copied, or stored in any form, or (ii) directly or indirectly distributed or transmitted to any other person for any purpose. The information in this report is derived from sources that Futu Securities believes to be accurate and reliable as of the date of publication. However, this report is not intended to contain all information necessary for investor decision-making and may be subject to delays, obstructions, or interceptions in transmission. Futu Securities does not expressly or impliedly warrant or represent the adequacy, accuracy, completeness, reliability, or fairness of any such information or opinions. Accordingly, Futu Securities and its affiliates (collectively, “Futu Group”) shall not be liable for any losses of any kind (including but not limited to direct, indirect, or consequential losses) arising from any actions taken by third parties relying on the content of this report. The views, recommendations, suggestions, and opinions expressed in this report do not necessarily reflect the positions of Futu Securities or its affiliates and may be changed at any time without notice. Futu Securities has no obligation to update any information or opinions contained herein. This report is provided for general informational purposes only and is intended solely for the general reading of Futu Securities’ clients, without regard to the specific investment objectives, financial situation, or particular needs of any individual recipient. Nothing in this report constitutes or should be construed as an offer, recommendation, or solicitation by any member of Futu Group to buy or sell any securities, investments, or other financial instruments. It should not be interpreted as an offer or invitation to purchase or sell securities. Any decision to purchase securities mentioned in this report should take into account publicly available information, including relevant prospectuses. The products referenced in this report may not be suitable for all investors; readers should fully consider relevant factors and seek professional advice before making any investment decisions. This report is provided to recipients on the understanding that they are capable of independently evaluating investment risks and exercising independent judgment in investment decisions. In certain jurisdictions or countries, the distribution, issuance, or use of this report may violate local laws, regulations, rules, or other registration or licensing requirements. This report is not intended for distribution to or use by any person or entity in such jurisdictions or countries. Hong Kong-based investors with questions regarding Futu Securities research reports should contact Futu Securities directly. The Central Entity Number of the Hong Kong SFC license held by the report’s author is disclosed next to the author’s name on the report’s cover page. The analyst primarily responsible for preparing this report confirms that: (i) the views expressed in this report accurately reflect his/her personal views regarding the listed corporation(s) covered; and (ii) none of the compensation he/she has received, currently receives, or will receive—directly or indirectly—is linked to any specific recommendation or view expressed in this report. The analyst further confirms that neither he/she nor any connected persons have traded in the securities of the listed corporation(s) covered in this report during the 30 calendar days prior to its publication or within three business days following its publication. Neither the analyst nor any connected persons serve as senior management of the listed corporation(s) covered in this report, nor do they hold any financial interest in such corporation(s). Futu Securities holds no financial interest amounting to 1% or more of the market value of the listed company mentioned in this report and has had no investment banking relationship with the company in the past 12 months. Employees of Futu Securities are not employees of the listed company.
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Comments (2)
to post a comment
4
