English
Back
Open Account
ME News
wrote a column · Jul 14 00:50

Six new trends to understand what’s happening in the AI race

Article author, source: World Model Factory
Over the past few days, the AI community has seen another wave of explosive developments:
OpenAI merged Codex and ChatGPT, lifted Codex’s five-hour usage limit, and reset users’ quotas;
Anthropic extended the promotional period for Fable 5 once again, and the offer granting a 50% weekly quota boost for Claude Code remains active;
Grok 4.5 suddenly gained strong praise, with many users saying it outperforms Opus 4.8—Musk may finally be breaking into the top tier;
Tencent’s HunYuan Hy3, activating only 21 billion parameters, is already challenging much larger flagship models, and WorkBuddy’s user base continues to grow rapidly;
Zhipu AI has launched its "TouchHigh Ambition Plan," committing significant resources over the next two years to long-horizon tasks, autonomous agents, fully self-supervised training, and safety governance.
There’s so much news it’s dizzying—but when viewed together, clear trends emerge.
The stronger the model, the shorter the window Article author, source: World Model Factory Over the past few days, the AI community has seen another wave of explosive developments: OpenAI merged Codex and ChatGPT, lifted Codex’s five-hour usage limit, and reset users’ quotas; Anthropic extended the promotional period for Fable 5 once again, and the offer granting a 50% weekly quota boost for Claude Code remains active; Grok 4.5 suddenly gained strong praise, with many users saying it outperforms Opus 4.8—Musk may finally be breaking into the top tier; Tencent’s HunYuan Hy3, activating only 21 billion parameters, is already challenging much larger flagship models, and WorkBuddy’s user base continues to grow rapidly; Zhipu AI has launched its "TouchHigh Ambition Plan," committing significant resources over the next two years to long-horizon tasks, autonomous agents, fully self-supervised training, and safety governance. There’s so much news it’s dizzying—but when viewed together, clear trends emerge.  Trend 1: The lead window for cutting-edge models is shrinking rapidly Not long ago, everyone was still benchmarking Fable 5 and GPT-5.6, but now many are already saying Grok 4.5 is 'more usable than Opus 4.8'; Tencent’s HunYuan Hy3, with only 21 billion activated parameters, has matched flagship-level performance on certain agent and office tasks. In most people’s perception, foundational models like Grok and HunYuan weren’t initially strong and even went through issues with their...
Not long ago, everyone was still benchmarking Fable 5 and GPT-5.6, but now many are already saying Grok 4.5 is 'more usable than Opus 4.8';
Tencent’s HunYuan Hy3, with only 21 billion activated parameters, has matched flagship-level performance on certain agent and office tasks.
In most people’s perception, foundational models like Grok and HunYuan weren’t strong to begin with—even undergoing model overhauls—and previously lagged far behind the state-of-the-art. Yet they’ve caught up in just a few months. Why?
This is no coincidence.
Core technical know-how for models has now diffused widely, post-training barriers have lowered, and coding and agent tasks benefit from inherent automated scoring mechanisms—enabling newcomers to find faster paths to catch up.
As a result, leading models are now scoring increasingly closer on traditional knowledge-based Q&A and simple coding tasks, making it hard to discern real-world experience differences.
More importantly, users don’t interact with raw models—they engage with a complete system combining "model + prompts + tools + memory + retry mechanisms."
A slightly weaker model, when paired with a superior framework, can easily outperform a stronger raw model in actual user experience.
Thus, while a new model used to maintain a lead of several months to half a year, today that advantage may last only a few weeks—or even days.
As this lead window continues to shrink, users will no longer focus on a single globally dominant model, but rather on capability tiers—for example, some excel at coding, others at long-horizon tasks, and others at office productivity deliverables.
Users commonly adopt multi-model routing, and enterprises won’t put all their eggs in one basket.
Ordinary users can now clearly notice that previously high-performing models have suddenly become much more generous.
For example, Anthropic extended the usage period of Claude Fable 5 until July 19 and removed its separate 50% quota limit;
OpenAI temporarily lifted Codex’s five-hour usage limit outright and reset user quotas; Codex now has 6 million active users.
This resembles the early internet era of burning cash to capture traffic—except now companies are burning GPUs and tokens, treating free quotas as customer acquisition costs.
Model providers are willing to sacrifice gross margins at this stage to cultivate user habits because the switching costs for their current flagship Agent products are far higher than those for chatbots.
Once this round of subsidy wars ends, only customers truly willing to pay for long-term value will remain.
As benchmark scores across various models continue converging, the commercial gap could widen even more rapidly in the future.
User base, agent task data, enterprise entry points, inference costs, developer ecosystems, cash flow—these factors will accelerate the data flywheel for market leaders.
Next, the real threat to model companies isn't stronger models—it's competitors rapidly replicating capabilities and capturing users through pricing and access points.
Tencent’s Hy3 is a typical example:
With a total of 295 billion parameters but only 21 billion activated, its Mixture-of-Experts (MoE) architecture emphasizes practical performance and low cost in agent and office productivity applications—and it’s integrated into WeChat’s ecosystem.
Models like Hy3, which combine 'highly efficient architecture + super-app distribution,' are already more than sufficient for the average office worker, delivering exceptional value for money.
This is why many say Tencent has quietly played a game-changing move—offering adequate capability, massive user access, and controllable costs.
This shows that parameter scale is no longer a bragging point for models; users and enterprises now care more about the total cost of completing a task, its success rate, and how smoothly it works in practice.
Grok 4.5 and Cursor are trained together on real-world software engineering data; Hy3 iterates based on feedback from over 50 Tencent products; Codex and Claude Code learn from vast amounts of agent interaction trajectories, including failure cases...
These real user interaction data are becoming the scarcest fuel for next-generation models.
The more the product is used, the faster the data flywheel spins, and the smarter the model becomes—creating a virtuous cycle.
A truly defensible moat emerges only when a model company has massive real-world users, deploys its models to execute tasks daily, and feeds the results back into training.
Tencent represents the path of drilling down into products and real-world scenarios.
By emphasizing activated parameters, agent success rates, and actual products like WorkBuddy and CodeBuddy—and using user feedback to refine training—it exemplifies an engineering-efficiency-plus-super-app-portal strategy.
Zhipu AI represents the path of pushing the upper limits of intelligence.
In an internal open letter titled 'The Great Wave Is Here,' Zhipu founder Tang Jie stated that over the next two years, the company will forgo chasing short-term monetization and instead focus on long-horizon tasks, autonomous agents, fully self-supervised training, and safety governance—and plans to allocate billions in resources toward mechanistic interpretability.
Although their approaches appear opposite, both companies share highly aligned ultimate goals—and will eventually converge on agents.
The difference lies only in direction: Tencent iterates upward from products, while Zhipu first elevates the technological ceiling before opening it up downward—one prioritizes widespread adoption first, the other pursues the performance frontier first. Both paths merit close observation.
While this generation of models is still competing for users, the next-generation race has already begun.
OpenAI is investing in new pretraining architectures, synthetic reinforcement learning, long-horizon agents, robotics, and consumer hardware;
Anthropic is focusing on long-horizon tasks, AI-assisted AI research and development, and mechanistic interpretability;
Google continues to advance world models, virtual environments, and robotics.
Given that latecomers in the model space can now catch up to the leading cohort within just a few months, will future models become increasingly homogeneous?
In fact, the next-generation model competition could still create significant performance gaps.
The true drivers of generational differentiation are likely to be long-horizon agents, automated AI R&D, world models, and robotics.
The gap may be especially pronounced in world models and robotics.
Real-world data cannot be scraped directly from the internet—it requires robots, sensors, simulation environments, and long-term deployment to accumulate.
Whoever first establishes a closed loop of 'devices-data-world model-action model' will gain an advantage even harder to replicate than that of language models.
However, progress in this direction is also slower. Hardware costs, security concerns, and real-world deployment cycles all mean it won't explode overnight.
Therefore, the next 'GPT-4 moment' is likely to arrive—but in a different form.
Yet as technology spreads faster and faster, the window of model leadership could still be very brief.
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Thumbs Up
1
39K Views
Report
Comments
Write a Comment...
1
1