HK Stock Market Barometer | Divergence in Tech and Internet Stocks, Precious Metals Rebound! How to
[AI Key Takeaways]
Financial Performance
- Revenue in the first half of 2026 reached approximately $120 million, a year-over-year increase of 283%. H1 revenue has already reached 1.5 times the full-year revenue of last year.
- Q2 revenue increased by 82% quarter-over-quarter compared to Q1.
- Gross profit was approximately $21 million, up 465% year-over-year, with gross margin rising from 12.1% to 17.9%.
- Annual Recurring Revenue (ARR) exceeded $800 million in August, with the B2B segment contributing over 80% of the total ARR.
Business Progress
- Continuously upgraded the M2 series models and launched M3, resulting in increased inference throughput.
- Launched the H3 multimodal model, which garnered over 24 million downloads within three weeks and spawned more than 300 derivative models.
- The global base of enterprise customers and developers expanded rapidly, surpassing 2 million, which is ten times the figure at the end of last year.
- Overseas markets contribute over 60% of revenue
Next Quarter Guidance
- We expect the model release cycle to shorten significantly in the near term, accelerating the pace of intelligence enhancement
- Upcoming updates include new series such as M3.1, M3 Pro, and H3.1
- Gross margin is expected to continue improving in the second half of this year, with further room for growth next year
- The adoption rate of domestic chips is expected to rise in Q4 of this year
Opportunity
- Product Innovation: We are developing flagship models and further advancing Minimize Slive Attention 2.0, which triples computational efficiency compared to the previous generation
- Market Expansion: Multimodal capabilities and specialization in professional domains determine the depth and breadth of AI's penetration into the real world
- Operational Efficiency: Efficient scheduling has increased computing power utilization; we estimate at least threefold room for scale expansion under current capital efficiency
[AI Conference Transcript]
Operator
Hello everyone, and welcome to MiniMax's 2026 Semi-Annual Results Conference Call. I will now hand over to Olivia Sha, our Director of Investor Relations.
Olivia Sha
Good evening, dear investors and friends. Thank you for joining the MiniMax 2022 Interim Results Conference Call. Participating in today's call are members of our management team, including Yan Junjie, CEO; Yun Yeyi, COO; and Xue Zizhao, Vice President.
Before we begin, I would like to read our disclaimer. Today’s discussion may contain forward-looking statements, including but not limited to outlooks on the company’s business prospects and expected financial performance. Such statements are subject to various risks and uncertainties, and actual results may differ materially from those expressed or implied by the forward-looking statements.
This conference call will be conducted primarily in Chinese, covering both management remarks and the Q&A session, with simultaneous English interpretation provided by a third-party translation team. The English translation is for convenience only. In the event of any discrepancy between the English translation and the original Chinese remarks, the original Chinese shall prevail. I will now hand over the call to our CEO, Mr. Yan Junjie.
Yan Junjie
Good evening, investors and analysts. Thank you for joining our 2026 interim results conference call. Since the beginning of this year, we have continuously upgraded our M2 series models, launched the M3 model, and significantly increased inference throughput. On the business front, we have achieved rapid overall growth.
In July, our token consumption volume increased approximately twentyfold compared to January of this year. Revenue in the third quarter grew by 81.8% quarter-over-quarter compared to the first quarter. Furthermore, in August, our Annualized Run Rate (ARR) further increased to exceed USD 800 million.
As we continue to refine our M-series and H-series model pipelines, the company expects the model release cycle to shorten significantly in the near future, while the pace of intelligence enhancement will accelerate. Today, I would like to take this opportunity to share our positioning and strategy.
First, the intelligence enhancement driven by Large Language Models (LLMs) appears to have virtually no limits. This journey has evolved from coding and agent capabilities since the second half of last year, to more autonomous and creative execution of long-horizon tasks, and towards future end-to-end interactive outcomes. We continue to pursue higher levels of intelligence; however, computing power remains a finite resource for every company.
'Minimize inference costs and maximize intelligence for everyone' has been our consistent technological roadmap. We believe this is not merely about providing cost-effective services, but rather about building the core capability for effective scaling toward advanced intelligence. This is because human scale will play an increasingly significant role in model capabilities, and inference costs determine the scalability of such human involvement.
For models, we prioritize the level of intelligence first, followed by cost-effectiveness. The level of model intelligence is directly correlated with parameter scale. Our strategy is to continuously develop models within the largest tier in the industry, while exploring optimal architectures to balance pre-training loss and computational efficiency.
Subsequently, we leverage our computational efficiency advantages to scale up during post-training as much as possible. First, regarding pre-training and the architectural choices for models with large parameter counts, we explored two extremes. The first involves sparsity in Mixture of Experts (MoE); our expert sparsity remains within 2% while still maintaining robust convergence.
The second aspect concerns attention sparsity. Beyond the standard MaxStart attention architecture, which maintains training stability, we can significantly compress computational costs for long contexts. In our flagship model currently under development, we have further introduced version 2.0 of Minimize Slive Attention.
This further compresses the technical capacity required for activation, thereby increasing cache hit rates and batch sizes during inference, leading to higher utilization of computing power. At the maximum parameter tier, our computational efficiency has improved threefold compared to the previous generation's MMSA 1.0, with the advantage becoming more pronounced as task length increases.
Secondly, in mid-training and post-training stages, the proportion of computing power allocated is actually increasing. The computational load in these training phases increasingly stems from core data, rollout processes during reinforcement learning aggregation, and extensive evaluations supporting experiments. These three computational modes are essentially inference tasks.
Especially as the proportion of long-context tasks rises, inference efficiency under long contexts becomes key to scaling. Structures with higher inference efficiency scale faster during mid-training and post-training, securing a long-term competitive advantage. Under the same computing power scale, the multiplier effect of improved inference efficiency is basically linearly correlated with the multiplier effect of increased candidate scale.
The architectural advantages and innovations brought by our large-parameter models, along with the ability to continuously expand post-training scale, are expected to be directly reflected in our subsequent model updates. Next, let us discuss infrastructure. Building on strengthened cooperation with cloud providers to further increase our computing supply, we are one of the first two independent large-model companies in China to establish scaled, long-term stable infrastructure.
This is the key support enabling us to develop two models simultaneously. Currently, our self-built infrastructure can independently support our largest training task, with an Effective Training Ratio (ETR) now reaching 97%, which is among the highest levels known in the industry.
This full-stack capability not only enhances our training scale but also provides high flexibility in architectural design, allowing for significant improvements in AI-native training and inference. For instance, through the co-design of memory, SSDs, and networking, we can achieve a significant boost in cache hit rates, thereby halving prefetching costs.
As the proportion of agent-type tasks in post-training increases, there is a growing need for CPU-based sandboxes. Our multimodal business inference clusters contain substantial CPU resources, providing excellent space for mixed scheduling. Meanwhile, the tidal nature of online inference traffic provides abundant available computing power for small-scale validation experiments and reinforcement learning rollouts.
Based on our full-stack capabilities in inference, we estimate that there is at least threefold room for expansion in subsequent scale per unit of capital efficiency. Furthermore, improvements in inference efficiency will continue to boost our future gross margin levels. Finally, let us discuss multimodality. Unlike most current large model companies, we have consistently incorporated native considerations for dynamic information in our model design.
Regarding language models, we believe that visual understanding and generation are critical components of productivity. Dynamic generation is currently the second-largest market within AGI, aside from programming. H3, which we launched three weeks ago, represents a practical integration of language models and visual generation. It goes beyond merely using a language model as an encoder; more importantly, it leverages the language model for fine-grained context awareness, enabling comprehensive referencing and precise generation capabilities.
The breakthrough in performance of MiniMax A3, our open-source strategy, and the demonstrated cost-performance advantage have fundamentally shifted the landscape of the sector, where state-of-the-art models were previously dominated by large tech companies following closed-source approaches. Its popularity among creators and the open-source community has far exceeded our expectations. In the three-plus weeks since the open-source release of A3, downloads on platforms like ModelScope and Hugging Face have surpassed 24 million, spawning over 300 public derivative models. This is likely the most downloaded model globally this year.
Driven by strong community reputation and superior cost-performance, the usage volume of our official online services has also seen explosive growth. With deeper integration of dynamic and language models, the introduction of physical world information, and optimizations in encoder context computation and data processing, we anticipate further significant leaps in multi-modal model capabilities and editing functions.
Leveraging our unique technical reserves in both language and generative models, we expect our technological leadership, market share, and influence in this field to further increase. Looking further ahead, we believe that paradigms combining language models with design-generation models will continue to emerge in more frontier domains in the near future.
Returning to the initial question, MiniMax's goal is not merely to trade off between intelligence and cost. Only by minimizing the unit cost of intelligence can we train models to achieve maximized intelligence levels. Reducing the supply cost per unit of intelligence is the only way to enable higher-level intelligence to permeate broader production and daily life scenarios.
Model innovation determines how high we can push intelligence, while model efficiency dictates the scale of intelligence supply. Reinforcement learning and real-world feedback determine the speed of intelligence improvement in the next cycle, whereas multi-modality and specialized domain applications determine how deeply and broadly intelligence can penetrate the real world.
We are also aware that a single model does not represent the entire core; the key lies in whether the underlying infrastructure and R&D system for model development can sustain scaling. During the R&D of M3, we faced time constraints, but for us, the more critical issues are whether we can continuously define and iterate our technical roadmap, accelerate the density of intelligence improvements, and deliver higher-level intelligence to more users at lower costs. This aligns with our company's persistent vision: 'Intelligence for Everyone.' Thank you for your attention and support; we will continue to strive forward. Thank you all.
Xue Zizhao
Next, I invite Xue Zizhao, Vice President of the Company, to present the financial results. Hello everyone, I am Zizhao. I will briefly introduce the Group's financial performance for the first half of 2026. During the reporting period, we achieved revenue of approximately $120 million, representing a year-on-year growth of 283%. Revenue for the first half already reached 1.5 times the full-year revenue of last year. Specifically, Q2 revenue grew by 82% quarter-on-quarter compared to Q1.
The core driver of revenue growth lies in our unique positioning around text and multi-modal models. While providing cutting-edge model capabilities, we optimized inference efficiency to achieve attractive cost-performance ratios. Leveraging this advantage, our global base of enterprise customers and developers has expanded rapidly. The continuous growth in API call volumes and model inference demand has accelerated the conversion of model capabilities into actual product usage and commercialized revenue.
In terms of revenue composition, revenue from the open platform and other AI enterprise services amounted to approximately $74 million, representing a year-over-year increase of 703%. Its share of total revenue rose from 30% in the same period last year to 63%, driven primarily by growth in the number of paying users and enterprise clients, increased API call volume, and rapid adoption of token plans. Revenue from AI-native products was approximately $43 million, up 101% year-over-year, with overseas markets contributing more than 60% of this revenue.
Revenue growth accelerated in the third quarter. As of August, Annualized Recurring Revenue (ARR) exceeded $800 million, with the B2B segment contributing over 80% of the total ARR. On the technology and infrastructure front, we are continuously improving model computational efficiency through the implementation of architectural innovations such as MiniMax sparse attention and the ongoing enhancement of infrastructure capabilities.
Gross profit for the first half of 2026 was approximately $21 million, a year-over-year increase of 465%. The gross margin improved by 5.8 percentage points from 12.1% to 17.9%. In terms of expenses, R&D expenditures were approximately $300 million, up 139% year-over-year, mainly due to our continued investment in foundational model capabilities. The growth rate of R&D spending was significantly lower than the revenue growth rate, reflecting improved efficiency in converting R&D outcomes into business growth.
Sales and distribution expenses were approximately $30 million, down 18% year-over-year, primarily due to the continued promotion of organic user growth strategies, which reduced marketing spend. Administrative expenses were approximately $30 million, up 104% year-over-year, but their proportion of revenue decreased from 49% to 26%, gradually demonstrating operating leverage. [Note: The final sentence fragment regarding the narrowing of a figure by 11% to ~$360 million appears disconnected in the source; translated as is:] The figure narrowed by 11% from approximately $400 million in the same period last year to approximately $360 million.
After adding back share-based compensation, losses from the fair value of financial liabilities, and listing expenses, the adjusted net loss for the first half of 2026 was approximately $290 million, compared to approximately $140 million in the same period last year, an increase of 111%. This primarily reflects our continued investment in foundational model capabilities and infrastructure. To further strengthen the company's cash reserves and financial flexibility, the company completed a placement issuance in July 2026, raising a total of HK$16 billion.
The company currently holds cash reserves exceeding $3 billion, providing a solid foundation for our continued investment in cutting-edge model R&D, computing infrastructure construction, and attracting top-tier talent. Overall, the company made significant progress in revenue scale, business structure, and gross margin in the first half, with expense metrics continuing to improve. We will continue to enhance R&D, operational, and commercialization efficiency while maintaining long-term strategic investments. Thank you all.
Operator
This concludes the company's remarks. We will now proceed to the Q&A session. Thank you to management for the sharing. We now enter the interactive exchange session. Investors are welcome to raise their hands for voice interaction or send text questions. Hello everyone. For participants joining via phone, please press the '*' key on your handset followed by '1' to request to speak. For online participants, you may type questions in the interaction area of the live stream or click the 'raise hand' button to request voice interaction. Thank you.
Yang Zeyuan
Next, we invite Yang Zeyuan from CITIC Securities to ask questions. Please go ahead. Okay, thank you. Thanks to MiniMax and Mr. Wen Teng for giving me this opportunity to ask questions. I also congratulate the company on making substantial and rapid progress in the first half. I have two questions, which I will ask together. First, the market has recently been closely watching the intensifying competition in coding assistance, as well as changes in the commercialization growth rates of certain overseas models.
During the presentation just now, we were listening carefully. I noticed you mentioned that you believe there is still significant room for improvement in global model intelligence. What is the basis for this judgment? Furthermore, beyond coding capabilities, which specific capability breakthroughs or application scenarios are most likely to become the next source of scaled demand? This is my first question, regarding the outlook post-coding.
My second question is as follows: The company just disclosed that token consumption in July reached a level approximately twenty times that of January. Furthermore, August's Annualized Revenue (AR) is expected to have grown rapidly, surpassing $800 million. Could management please break down the drivers of this growth? Specifically, is it primarily driven by new customer acquisition, expansion among existing customers, improvements in model capabilities (such as dynamic or text-based features), or changes in product structure?
Additionally, regarding our current business structure, particularly the geographic or regional breakdown, could leadership please help us dissect the growth data mentioned earlier? Thank you. Those are my two questions.
Yan Junjie
Thank you for the question. First, regarding coding-related capabilities, I believe there is significant room for improvement in intelligence, driven by both technological advancements and market dynamics. Technologically speaking, a key metric for evaluating model intelligence over the past few months has been the ability to handle complex tasks.
By 'complex tasks,' we refer to workflows that may require hundreds of iterations, such as tool calls, coding, and structural validation. Based on external progress and our extensive internal experiments over the past few months, we have observed a clear trend: if a capability can be precisely evaluated and constructed into a reinforcement learning environment, the model can essentially learn it.
I consider this a highly certain outcome. The core competitive advantage or strategic focus lies in how we abstract domain-specific problems into high-quality reinforcement learning environments that allow for verification and continuous model learning, thereby enhancing these capabilities.
Building on a sufficiently large pre-training foundation, the next steps involve Supervised Fine-Tuning (SFT) and reinforcement learning. SFT typically involves lower data efficiency, requiring larger volumes of data, whereas reinforcement learning utilizes the highest quality data. As mentioned earlier, I believe the scale and quality of reinforcement learning environments will differentiate the pace of model improvement across companies in the coming period. We have already invested significant effort into optimizing this area.
From a market perspective, we view coding as currently the most core capability. However, even in this area, there is substantial room for improvement. Currently, the use of AI in coding is largely limited to code generation and auxiliary development. We believe the market potential for more end-to-end delivery tasks remains very significant.
That addresses the first point. Secondly, we are seeing an increasing number of generalized coding-like tasks. For example, cybersecurity has been a typical case recently. Essentially, when a model's scientific capabilities are strong enough, combined with extensive security-related environments and iterations, it demonstrates remarkable capabilities in cybersecurity, potentially transforming the industry.
We are now observing that wherever a data closed-loop can be established and verifiable reinforcement learning environments can be constructed, similar trends emerge across many industries, such as chip design and related fields. We anticipate that such breakthroughs may occur several more times in the coming months.
Secondly, regarding commercialization, we have observed that once model capabilities突破 a certain threshold, growth—both for the industry and our company—is not continuous or linear. Instead, the release of value from a superior model tends to occur in step-wise increments rather than following a simple linear trajectory.
In July this year, our token consumption increased significantly compared to January. This growth was primarily driven by two factors: first, a substantial increase in agent-related tasks during the Spring Festival period; and second, following the release of M3, we introduced native dynamic capability expansion, which also enhanced performance. Additionally, faster computation speeds further drove an increase in general usage effects.
Moving on to the second point, let's discuss the breakdown of revenue and tax rates. The core driver behind August's AR exceeding $800 million was the improvement in model capabilities. We have consistently pursued ultimate intelligence at the best price-performance ratio. With M3, while keeping prices unchanged compared to M2, we first delivered higher levels of intelligence. Secondly, we attracted many new customers; the number of enterprise and other clients has exceeded two million, which is approximately ten times the figure from the end of last year.
On the other hand, existing users and legacy customers have unlocked more scenarios due to improved model capabilities. For example, beyond coding and agents, many clients are increasingly using office productivity scenarios. We have observed that model consumption is shifting from human-AI interactions to multi-turn interactions between agents and AI.
Human queries are expanded into multi-turn model requests, tool calls, and agent tasks via AI. Consequently, the growth in inference demand driven by agents is significantly faster than the growth in the number of human users and messages. These changes are also reflected in our model usage; structurally, the general consumption per single user is growing rapidly.
Regarding the structure of AR, from a customer perspective, the B2B segment accounted for approximately 80% in August, compared to around 30% last year. We can see that the adoption rate of models by enterprises and developers is accelerating significantly. We believe that in the AI era, the model itself is the product. Therefore, whether B2B or B2C, customers are essentially paying for model capabilities.
As a large language model company, our core objective is to continuously enhance model intelligence and refine superior models for broader adoption. From a regional perspective, internationalization remains a key characteristic of our company. In the first half of this year, overseas revenue accounted for over 60%. From a modality perspective, we have observed significant growth in dynamic capabilities, particularly after the release of S3.
However, language models continue to grow, primarily because we offer excellent price-performance ratios. We have observed that an increasing number of customers are simultaneously utilizing language, vision, and audio capabilities. Therefore, from June to August, we believe that both language models and dynamic models jointly drove AR growth.
Looking ahead, we believe the most important source of growth in the next stage will be changes in model capabilities. Stronger models will undoubtedly attract more users and tasks. Furthermore, higher inference efficiency will allow customers to expand their usage scope at more acceptable costs. We look forward to upcoming updates in our new series, including M3.1, M3 Pro, and H3.1.
Ronald Kong
Thank you. We look forward to more innovations from the company. Thanks to management for the answers. Next, I invite Ronald Kong from Goldman Sachs to speak. Thank you, Mr. Yuan, for addressing Olivia's two questions. Some competitors are continuously raising prices, while others are lowering them. Could the company further explain why reducing inference costs is not just a commercial price-performance strategy but also directly impacts maximizing intelligence? Could you share the current R&D and release roadmap? What verifiable progress in coding agents and long-context tasks should the market focus on? Thank you.
Yan Junjie
First, regarding this initial question: I believe inference is not merely a service providing intelligence via network capabilities; it is also a crucial internal process for generating next-generation intelligence. Originally, we could clearly see that during the pre-training phase, scaling continued—from models like the 3T parameter model to even larger ones. The scale of pre-training, in terms of parameter count and corresponding token volume (maintaining a roughly 1:20 ratio), remains capable of continuous scaling.
However, on the other hand, the proportion of post-training is becoming increasingly significant. Post-training relies heavily on various factors, such as data synthesis, the construction and sampling of agent-based environments, and extensive rollouts during reinforcement learning. Most importantly, due to the high level of uncertainty in post-training, a large number of experiments are required.
These experiments necessitate highly precise evaluations. In this context, it becomes evident that the majority of computational load actually occurs during post-training, specifically within inference. For this reason, I believe inference efficiency determines the scale at which we can effectively scan or iterate during post-training. Essentially, under models of the same scale, higher inference efficiency allows us to reap greater benefits from candidate models.
The core of inference efficiency is not just about reducing the cost per single token; it refers to how many effective inferences can be completed per unit of time with the same computing power. More inferences mean more training trajectories, more experiments, and faster evaluations, thereby accelerating iteration. This is why we place such high importance on inference efficiency—it is critical both for controlling external service costs and for accelerating internal model R&D and scaling.
We consider this to be a core competitive advantage for our next-generation models and the foundation for achieving performance leadership in future iterations. Regarding the pace of upcoming model updates, I believe model release is only one aspect. More practically, it is about whether we can establish solid, accumulative infrastructure during the R&D process—spanning from data and compute to pre-training frameworks, reinforcement learning frameworks, evaluation systems, and inference infrastructure.
Furthermore, there are deeper layers, including large-scale cluster construction and software optimization. This entire process requires close collaboration among infrastructure, R&D, algorithm, and optimization teams. It is a comprehensive workflow. Previously, during the M3 phase, we faced some challenging moments, which reflected deficiencies in our evaluation infrastructure. This is precisely why we dedicated significant effort in the second half of the year to essentially rebuild our infrastructure.
As this infrastructure has become increasingly efficient, we have seen preliminary validation starting with the H3 model. Moving forward, models such as M3.1, M3 Pro, and H3.1 are nearing completion. Specifically, M3.1 is positioned to fully validate our entire infrastructure pipeline, aiming to deliver a widely adoptable model that excels in stability, quality, inference efficiency, and agent generalization.
If M3.1 meets our expectations, it will largely validate our subsequent pipeline, leading into M3 Pro. M3 Pro represents a comprehensive upgrade in model scale, parameter count, and capabilities. It is a model with parameters approaching 3 trillion (3T). From pre-training data volume and training efficiency to reinforcement learning capabilities and long-task training on candidate models, M3 Pro pushes the limits at every stage. Beyond reaching the scale limit of four tokens (context/sequence), we have implemented significant innovations, including major architectural upgrades and enhancements in inference efficiency and speed. We are currently observing progress in its performance ceiling relative to candidate models.
What we can confirm now is that its inference speed and cost should offer a significant advantage within the 3T-class model tier. Additionally, regarding H3.1: following the release of H3, which saw extensive adoption by both corporate clients and the research community, H3.1 delivers a substantial capability upgrade on top of that foundation. These are the positioning strategies for our upcoming models.
Alright, thank you. Let me add a brief point: our core strategy is to solidify every stage of the process and strengthen our infrastructure to ensure high-quality execution. We are also working hard to enable users to adopt these new models as quickly as possible. Thank you.
Gary
Thank you for the clarification, management. Next, we invite Gary from Morgan Stanley. Thank you, iotimo and nollivia, for your insights. My question concerns computing power. Currently, computing capacity is arguably the most critical resource for large model companies, and the state is actively stockpiling AI chips. What is our current reserve of computing power? Is the capacity required for M3Pro and subsequent models already in place? How are resources allocated between training and inference? Finally, what is the status of adaptation for domestic chips? Thank you.
Yan Junjie
Yes, that is indeed a very key question. Overall, I believe that our existing computing power, as well as the capacity currently in planning, is sufficient to support the subsequent iterations of the models we see today and the next phase of development. Leveraging our experience in self-built infrastructure, we are also planning reserves for larger-scale model training. Currently, for text models at the 3T parameter scale and our video models, we have prepared ample training compute and a portion of inference compute. By combining our self-controlled core clusters with partnerships with cloud providers and general-purpose factories, we have established a comprehensive computing supply network. Through our large-scale autonomous computing scheduling, we integrate these three sources—our mechanical clusters, cloud provider partnerships, and factory collaborations—to achieve flexible scalability.
On our self-controlled core clusters, our primary goal is to support pre-training, post-training, and reinforcement learning tasks for our flagship models. Large-scale training clusters require long-duration continuous operation, demanding high network stability and efficient software coordination. To support such clusters, we need to independently build surrounding infrastructure, including high-speed networks, dedicated lines, servers, and storage. While this may not be immediately reflected in short-term financial reports, I believe we achieved substantial progress in this area during the first half of this year. This forms the core technical foundation for our future larger-scale operations. Secondly, our deepened cooperation with cloud providers offers stronger guarantees and significant elasticity for both model training and online inference.
At the inference level, many of our online businesses rely on partnerships with cloud providers. We allocate resources based on our model tasks, experiments, and cost considerations, which now offer more options. Thirdly, general-purpose factories serve as an excellent supplement to our self-built clusters and scientific collaborations for inference compute. Based on actual head-end demand, we have increased supply through these factories to support business growth and handle periodic peak loads. This combination not only ensures our service quality but also helps our partner factories expand their service scale.
Internally, since we are developing different models, internal computing allocation and resource scheduling present a unique challenge we must address. Different models undergo dynamic changes at various stages according to needs. For instance, at the current stage, text R&D is critical, and its training volume is significantly larger than that of video. On the other hand, improvements in large language models provide a foundation for understanding our dynamic models and task planning.
Also, regarding your earlier mention, we have already started using a portion of domestic chips. We expect the proportion of domestic chip usage to increase further in Q4 of this year, applying to both foundational and dynamic models. We observe that domestic chips can improve our overall utilization rate and even reduce unit computing costs. More importantly, they provide a reliable option for ensuring supply security.
Yu Zhonghai
Thank you for the answer, management. Next, we invite Yu Zhonghai from CICC to ask questions. Welcome, Mr. Yan and distinguished leaders. Thank you for giving me this opportunity to ask a question. A highlight for the company this year has been the significant improvement in gross margin during the first half. My question is: as model prices continue to decline, how does the company maintain and improve gross margins by reducing unit intelligence costs? Will the trend of gross margin improvement in the second half mainly stem from inference optimization and improved cluster utilization, or from advantages in supply chain management? Thank you.
Yan Junjie
Yes, I believe this is a metric we attach great importance to. Its improvement is achieved through a combination of several factors. We believe there is still significant room for enhancement in the near term. First and foremost, the most critical factor is our ability to optimize computing power. For instance, over the past two months, the throughput per unit of computing power for our text models has nearly tripled. Admittedly, rising costs for computing power have somewhat offset these gains, but overall, we have seen marked improvements. This is the first point.
Secondly, regarding our computing infrastructure, which supports both training and inference, we achieve high utilization rates through highly efficient scheduling. For example, during nighttime hours, when traffic for text inference decreases compared to daytime business hours, these computing resources are automatically reallocated for tasks such as evaluations and algorithmic validation experiments. This allows our overall utilization rate to remain high and continue improving. Since we have built some of this infrastructure in-house, we enjoy considerable flexibility, enabling us to implement these adjustments rapidly.
Thirdly, as you mentioned earlier regarding procurement and supply chains, as the scale of our self-built infrastructure expands, we have gained a much more direct understanding of the supply chain compared to simply purchasing from cloud providers. This presents significant room for improvement. Following cost reductions, another key point—reiterated earlier—is our strong emphasis on inference costs in model architecture design. This focus is central to our technical strategy and unique advantage, facilitating better scaling. Our strategy is that after achieving cost reductions, we will pass on part of these savings to customers by lowering prices and reducing usage barriers, thereby attracting more users and expanding investment scale. Simultaneously, we will retain a portion of these savings to improve our gross margin. We do not view price reductions and gross margins as contradictory; as long as unit economics remain positive and gross margins continue to improve, scale will expand. Increased scale, in turn, provides further room for higher gross margins.
From a business structure perspective, while our multimodal generation and voice businesses have historically had lower gross margins than our text business, the revenue share of the text business is rising rapidly. We believe its gross margin still has substantial room for improvement. As we expand our optimization scale and further improve cluster utilization, we expect the text business to become a significant driver for gross margin improvement. Looking ahead, we anticipate our gross margins will continue to improve in the second half of this year, with further upside potential next year.
Michael
Thank you for your response, management. Next, I’d like to invite Michael Alley. Thank you, Mr. Yan, and the management team for the opportunity to ask questions. Congratulations on the company’s strong performance in August. My question has two parts. First, regarding the release and open-sourcing of H3: What new validations does management see for the development of multimodal models and Artificial General Intelligence (AGI)? Will the decision to open-source H3 weaken the company’s future commercialization and pricing capabilities? Second, following up on the resource allocation issue mentioned earlier: As the company simultaneously advances both language models and multimodal models, how does it plan to balance resource investment? Thank you.
Yan Junjie
Indeed, for quite some time, we have received many inquiries from investors asking why we are committed to developing generative models. I believe the impact achieved by H3 has alleviated some of these concerns. Let me explain our rationale for open-sourcing. First, we view this field as a productivity tool, which aligns closely with the broader goal of AGI to enhance productivity. Second, when we released version 3, our primary intention was to change the industry. This sector has long been dominated by closed ecosystems with high pricing, leading to near-monopoly conditions. We believe that, similar to language models, a healthy ecosystem that generates greater social impact and commercial growth requires openness. Therefore, we chose to open-source it. Furthermore, the technology is still in its early stages. While attention to its design has far exceeded our expectations and it can currently produce good content, it remains far from our ideal of higher-level intelligence characterized by greater stability, richer content expression, and real-time speed.
Our approach is threefold: First, use open-source strategies to transform the industry. Second, drive advancements in intelligence within this field through continuous scaling. Third, enhance our commercialization capabilities by continuously delivering tangible value to users, developers, and creators with each improvement in capability. We consider U3 a very good start. Regarding resource allocation, we do not view language models and generative models as competing entities. Instead, we see them as mutually reinforcing in terms of model capabilities, data, and infrastructure. Our goal is to foster synergistic development between the two paths to jointly advance intelligence and deliver better productivity tools to our users.
Finally, I would like to briefly address the relationship between language models and visual models. As I mentioned, we see many frontier areas in the future, such as chip design and pharmaceuticals, where integrating different modalities and information sources is key. We have already established a pathway in combining language and generative models: understanding better to generate better, linking data across domains, conducting collaborative training, and transforming industries. As our language models continue to improve, I believe that in Q4 of this year and Q1 of next year, regardless of whether it is us or our peers, in China or the US, there will be increasing opportunities for cross-domain integration. Scientific development often arises from the intersection of different fields.
Jeffrey
Understood. Thank you, management, for the clarification. Next, I’d like to invite Jeffrey to speak. Good evening, Jeffrey. Thank you to the management team for taking my questions. Against the backdrop of China’s AI models rapidly catching up with global peers, domestic capacity awaiting production, and intensifying competition in the AI pure-play sector: First, how does the company view the competitive landscape in the next phase? Second, what does MiniMax consider to be the core capabilities that determine long-term competitiveness? Thank you.
Yan Junjie
Thank you for the question. We believe this industry is not a simple zero-sum game. At any given point, we are really just getting started, as the potential for intelligence enhancement seems almost endless. Looking three to six months ahead, current AI capabilities are still in their early stages. For this reason, the most practical metric is the speed of progress. Since the second half of last year, we have seen coding agents begin to deliver tangible utility. Over the past few months and likely in the near future, models will certainly become more autonomous and creative, capable of handling longer-horizon tasks, ultimately moving toward end-to-end delivery with verifiable results.
In this process, I believe that as long as any company’s technological advancements push the boundaries forward, it will ultimately expand the value of the entire industry and benefit more people. Therefore, the core of competition in the next phase may not be solely about computing power scale or user base size—in fact, it is certainly not just about user scale. More importantly, it is about whether a company can uniquely define the right problems and effectively choose the right paths to solve them. There are too many choices here, ranging from parameter scale and architecture to training compute allocation, reinforcement learning, inference, and evaluation. This leaves ample room for exploration. Consequently, each company’s definition of its model and its chosen technical roadmap will vary significantly. Thus, this field will undoubtedly exhibit a state of diverse development and innovation.
For MiniMax, our roadmap is very clear: while continuously raising the upper bound of intelligence, we aim to maximize the intelligence generated per unit of computing power, achieving extreme cost-performance ratio. For us, this is not merely a pricing strategy but also a technical roadmap that allows our subsequent chain to scale better. This is the technical path we are most committed to and which distinguishes us. Ultimately, since the essence of AI is productivity enhancement, the value of a model should be judged by how many real-world problems it can solve. Whether in coding or content generation, it has been proven that users are willing to pay for higher task completion quality, lower usage costs, and better production efficiency. Therefore, we should not simply view today’s market as competition between tech giants and independent large companies. Instead, we should focus on more fundamental aspects: who can continuously enhance intelligence, who can reduce costs, and who can achieve higher intelligence conversion efficiency. This is our perspective on the matter. Thank you.
Operator
Thank you. Due to time constraints, we will not address further questions. Thank you all for participating in this conference call. The content of this meeting is based on publicly available information as of today. Relevant data is for reference only. If there are discrepancies with the company’s officially disclosed information, investors should refer to the company’s official financial disclosures. That concludes our session. We wish you all a pleasant life. Goodbye.
More details:MINIMAX-W IR
Disclaimer: The above content is generated by an AI language model based on public data and third-party automatic subtitles. The above content does not represent any position of Futu and does not constitute any investment advice. Futu Group makes no express or implied warranties or representations regarding the accuracy, timeliness, or completeness of the above content.
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Comments (506)
to post a comment
240
12
