On July 6, Tencent officially launched HunYuan Hy3. Compared to the preview version, Hy3 demonstrates significantly stronger performance than models of similar size and matches the intelligence level of flagship models with 2–5 times more parameters, while further reducing pricing and substantially improving overall stability and cost-effectiveness. Hy3 has already been integrated into multiple business applications including WorkBuddy/CodeBuddy, Yuanbao, Marvis, and IMA. Its API is now available on Tencent Cloud TokenHub, with several international API platforms set to onboard it soon.

Model Intelligence Leaps Forward Across the Board; Agent Practicality Achieves a Qualitative Breakthrough
Hy3 is a model that integrates fast and slow thinking, built on a Mixture-of-Experts (MoE) architecture with 295 billion total parameters and 21 billion activated parameters, supporting a context length of up to 256K tokens. The Hy3 preview released on April 23 marks the first version after the complete overhaul of HunYuan, delivering qualitative improvements over Hy2 in complex reasoning, instruction following, in-context learning, code generation, and agent capabilities. Hy3 continues along a steep and clear trajectory of capability growth, achieving another significant leap across various tasks compared to the Hy3 preview by further scaling up post-training compute resources and enhancing data quality and diversity—now matching the performance of large flagship models both domestically and internationally despite its relatively compact size.


Rapid advancement in HunYuan model capabilities
Hy3 shows particularly notable progress in productivity tasks such as software development, office productivity, financial modeling, front-end design, and game development, making it a cost-effective and reliable choice. In an internal blind evaluation conducted by 270 domain experts based on real-world work scenarios, Hy3 achieved an average score of 2.67 out of 4, outperforming GLM-5.1 (2.51/4), with especially significant advantages in front-end development, data and storage, and CI/CD categories.
Refined through massive real-world user scenarios, significantly enhancing user experience
Hy3 has been rigorously tested through global developer adoption and Tencent’s extensive real-world business applications. Since the preview launch, its average daily token consumption has increased 20-fold, reflecting growing market recognition of its positioning as a high-value, practical model.
As China’s most popular AI-powered office assistant, WorkBuddy presents genuinely complex and demanding use cases—such as automated script generation and workflow orchestration—that provide high-value direction for HunYuan model iteration. Since its release, the number of users who voluntarily selected Hy3 preview on WorkBuddy has grown sixfold. In internal evaluations of office scenarios on WorkBuddy, Hy3 raised task success rates from 72% to 90% and reduced average task completion time by 34%, enabling more efficient and worry-free problem resolution.
Yuanbao’s conversational interaction scenarios also provide highly valuable feedback for model improvement. To address hallucination issues in long-document understanding and AI search, Hy3 employs fine-grained data cleaning and training constraints, enabling the model to produce reliable outputs even under complex evidentiary conditions. In evaluations based on real user logs, Hy3 reduced factual error rates by half and hallucination rates by more than half compared to the preview version. Leveraging Hy3’s substantially enhanced agent capabilities, Yuanbao has simultaneously launched its Agent feature. In internal assessments, Hy3 now surpasses numerous leading domestic models—including GLM-5.1—in both comprehensive office and lifestyle service scenarios, demonstrating sufficient stability to support real-world business workflows. Users can simply state their needs in everyday conversation, and Yuanbao will directly execute complex tasks and deliver files in formats such as PPT, Word, Excel, PDF, and HTML—all free of charge.
ima evaluated Hy3 across two core scenarios: knowledge base question answering and agent tasks. In agent tasks, Hy3 delivered outstanding overall performance with system stability reaching 95.1%. Its tool orchestration capability stood out in particular, significantly reducing ineffective operations such as blind retries or failure to terminate when appropriate—enabling more accurate, one-step planning for complex office tasks. Performance in knowledge base question answering also improved markedly, with a nearly 19% net gain in reasoning quality. The model now demonstrates more systematic thinking, broader information coverage, and notably enhanced structural integrity and usability in long-form writing and proposal generation.
In evaluations of Marvis Agent’s core scenarios—including document editing/generation, file management, computer diagnostics, and operations—Hy3 achieved a task completion rate of 93.7%, a 12.7% improvement over the preview version. When coordinating six agents simultaneously, task assignment accuracy reached 92%, up 13.5% from the preview. Additionally, Hy3 demonstrated significant upgrades in executing complex tasks, managing multi-step agent workflows, enabling office automation, and reducing hallucinations—providing stable support for real business pipelines and delivering tangible improvements in execution reliability, low-latency user experience, and cost reduction for Marvis Agent.
HunYuan continues to actively support various WeChat and gaming businesses. In specialized evaluations of AI avatars and customer service for WeChat Official Accounts, Hy3 improved intent recognition accuracy to 98.94%, capable of reasonably inferring user intent from incomplete expressions by leveraging account context—avoiding both excessive speculation and rigid template-based responses. In WeRead, Hy3 increased tag annotation accuracy by 14.1% and boosted overall classification efficiency by 8.4% compared to the Hy3 preview. In gaming, the recently launched Path of Exile: Descend AI Game Assistant on WeGame integrated Hy3, raising multi-turn reasoning and tool scheduling success rates to 92% and cutting hallucination rates from 4.5% to 2.8%, significantly improving output accuracy and enhancing player experience.
Tencent's highly diversified product portfolio provides extensive and real-world feedback for optimizing the HunYuan model, while advancements in the HunYuan model’s capabilities are in turn fed back into all products, creating a virtuous cycle of transferable and generalizable improvements. For example, optimizations in dialogue scenarios—such as intent understanding and output style—can also empower various agent applications, while enhanced tool-calling capabilities of intelligent agents significantly benefit businesses like conversational search.
Committed to an open strategy centered on high cost-performance, making advanced AI accessible to more users and industries
Hy3 continues the practical and inclusive positioning of its predecessors, offering services at a highly competitive price: RMB 1 per million input tokens, RMB 4 per million output tokens, and RMB 0.25 per million tokens for cache-hit inputs.
On the open-source front, Hy3 is released under the commercially friendly Apache 2.0 license, allowing developers worldwide to download and use it freely for commercial purposes. To further facilitate global developer adoption, Hy3 will be progressively launched on multiple international platforms, including OpenRouter, Hermes, Kilo, Cline, OpenClaw, OpenCode, and CherryStudio, and will be integrated into open-source model communities Hugging Face and ModelScope from day one.
From the infrastructure overhaul completed by the end of January 2026, to the release of the Hy3 preview in April, and today’s official launch of Hy3, HunYuan has completed an entire model development pipeline—from foundational reconstruction to product-level value creation—in less than six months. Going forward, Tencent HunYuan will continue accelerating its technology iterations, constantly pushing the boundaries of model intelligence, with a strong focus on enabling real-world application deployments and translating cutting-edge large-model capabilities into broadly accessible productivity across countless industries.
Relevant links:
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Comments (6)
to post a comment
40
75
