From July 17 to 20, crowds at the Shanghai Expo Exhibition Hall likely reached a nine-year high. At this year’s World Artificial Intelligence Conference (WAIC) 2026, the total exhibition area surpassed 100,000 square meters for the first time, with over 1,100 companies showcasing more than 3,000 exhibits. Embodied intelligence emerged as a core track on par with intelligent computing, and over 200 robotics companies competed side by side.
The most visible change on the show floor was that cloud-based large models—once dominating center stage—were now surrounded by AI devices equipped with physical bodies: ones that play piano, film sports games, or tighten screws. Exhibitors’ pitches also shifted—from 'Look how cool I am' to 'I can do this task and save you money.' Additionally, Nubia unveiled the world’s first AI agent smartphone, Honor launched Agentic OS, and Jieyue Xingchen’s lightweight on-device Step Edge model has already been installed on 42 million devices. This wave of product launches by device makers made on-device AI deployment one of the hottest topics at WAIC 2026.
The surge in on-device AI had clear precursors. Just two days before the conference opened, China’s Cyberspace Administration released the first batch of seven registered on-device large models for smartphones. Major brands including Apple, Huawei, Xiaomi, and Samsung were all included, along with Mianbi Intelligence—the only large model company listed specifically for its on-device capabilities. On-device AI has officially moved past the era of 'driving without a license.'
Policy approval, compute power decentralization, device manufacturers entering en masse—these three forces are shifting AI from 'centralized cloud generation' to 'distributed on-device execution.'。
But beneath the surface buzz, the real questions are just beginning to emerge: As AI develops a 'body,' which capabilities must reside directly on the device? What hurdles lie between crowdfunding success and mass production? And once the boundary between software and hardware blurs, how will companies defend their respective ecological niches?
On the opening night of the exhibition’s first day, during an event co-hosted by TMTPost and the WAIC Organizing Committee,at the 'WAIC UP! AI Three Poles Night Talk,',“the roundtable discussion titled 'The Awakening of On-Device Intelligence'featured four founders who each represented a distinct facet of on-device intelligence: Jia Shuo, Vice President of Quwan Tech, gave AI the fingertip dexterity to 'pluck'; Zhang Yizhen, Co-founder and COO of XbotGo, equipped AI with eyes that can 'capture'; Zhu Haohua, CTO and Co-founder of Lingqiao Intelligence, endowed AI with palms capable of 'grasping'; and Yin Jilei, Founder and CEO of Weefan Intelligence, is building the 'thinking' brain for these AI 'bodies.' Together, they offered an honest, on-the-spot exploration of the three questions above.

July 17 – WAIC UP! AI Three Poles Night Talk – Event Photos
Once AI enters physical environments, the first critical question becomes how to allocate tasks between edge devices and the cloud: Which capabilities belong in the cloud, and which must be embedded directly into the device? Four different scenarios yielded four answers—but all shared a surprisingly consistent underlying logic:Work backward from user needs and draw boundaries based on real-time performance requirements.。
Quwan Tech’s journey progressed from models all the way to hardware: starting with the Tianpu Music large model, followed by Tunee, a music agent, and culminating in the world’s first generative AI-powered guitar—which was brought to the event that evening. The choice of guitar as a physical medium stems from Jia Shuo’s assessment of music consumption patterns: AI-generated music output has already exploded—according to his citation, one leading platform alone produces over 7 million AI-composed songs per day—but listening through headphones remains a passive experience, whereas the continued popularity of karaoke and smart musical instruments demonstratesThe demand among ordinary people for 'actively participating in music-making' represents a rapidly growing incremental market.Regarding the division of labor between on-device and cloud processing, Jia Shuo draws a clear line: 'In terms of stability, quality, efficiency, cost, and other factors, music generation has only just reached the threshold of acceptability for average users on the cloud.' Therefore, generative tasks remain in the cloud, while prompt responses, guidance, and feedback—requiring high real-time sensitivity during strong user interaction—are handled on-device.The cloud handles creation; the device handles interaction.When software and hardware conflict, how does he decide? He follows one principle: 'Music is an aesthetic form of content consumption—it must be of sufficiently high quality before anything else matters.' Quality takes priority; everything else can be compromised as needed.
XbotGo has undergone a complete product evolution from 'cloud' to 'device.' The first two generations of XbotGo’s AI sports camera relied primarily on smartphone computing power for shooting and motion detection. By the third-generation product, 'Falcon,' it integrated an independent AI processor delivering 6 TOPS of performance, moving its core capabilities entirely onto the device itself. This shift wasn’t driven by technological showmanship but by users’ 'zero tolerance': 'Users simply won’t accept waiting several seconds after recording for cloud processing while the camera still hasn’t tracked the action—the ball’s already in the net.' Thus, the division became simple and clear: millisecond-level tasks like player recognition, ball trajectory tracking, and decisions between wide shots or zoom-ins must happen on-device; post-processing, highlight editing, and data analytics—value-added services users can tolerate with slight latency—are delegated to the cloud.
If the division of labor between device and cloud for consumer hardware is a single line, then for embodied intelligence it’s a multi-layered network. Zhu Haohua refers to dexterous robotic hands as the 'second brain of the human body'—they integrate the richest tactile and force-sensing capabilities along with the most precise motor outputs. His overarching principle is:On-device autonomy executes tasks; the cloud drives evolution.but even within the on-device side, further layering is required based on timescales:Perception-layer reflexes operate at kilohertz speeds, while decision-layer planning functions at hundreds of milliseconds.He cites a counterintuitive example: your hand releases a scalding-hot cup before your brain even registers that it’s hot.Action precedes perception—this is precisely the value of millisecond-level instinctive responses on the device side.Haptic feedback poses an even more unique challenge: its signal volume is enormous and highly latency-sensitive, requiring smaller, dedicated models deployed at the haptic layer to extract pulse signals and perform pattern recognition within a closed loop on the device. Architecturally, the path toward dexterous intelligence involves using a single foundational large model that, through intermediate and post-training, adapts to different timescales—hierarchical adaptation, rather than merely enabling simple 'dialogue' among multiple separate models.
From a chip perspective, Yin Jilei further abstracts this division-of-labor map into a compute spectrum: smart hardware represents the 'small end,' smartphones and vehicles occupy progressively larger positions, and robots sit at the 'large end' with the greatest demand. In his view,the essence of device-side hierarchy isn’t about the physical size of devices, but rather the alignment between algorithms and compute capacity—AI algorithms evolve far faster than chips. Thus, massive models that rely on brute-force scaling must remain in the cloud, with edge devices collaborating via the network. Embodied intelligence dramatically intensifies this tension: traditional dexterous hands, powered by cerebellum-like control combined with reinforcement learning, require only a few TOPS; yet once VLA (Vision-Language-Action) models and world models are distilled onto the edge, their parameter count balloons from a few billion to tens of billions, driving compute requirements from a few TOPS to hundreds or even thousands of TOPS.
These four answers together form a continuous spectrum of edge-cloud collaboration: the higher the real-time requirement, the greater the weight on the edge; the more a model depends on brute-force scaling, the heavier the reliance on the cloud. The boundary isn’t drawn by technology—it’s defined by users’ intolerance for waiting.
Identifying a use case is merely securing an entry ticket; what truly determines commercial success is the ability to move from 'building one product' to 'achieving scalable mass production.' While many industry players have dubbed 2026 the 'inaugural year of embodied intelligence commercialization,' this roundtable revealed a stark reality:mass production isn’t achieved through a breakthrough at a single technical node, but through systemic breakthroughs under the quadruple constraints of use case viability, cost, supply chain, and engineering execution。
The first gate is scalability in both use case and user base.The ceiling for edge-side hardware is largely determined by the breadth of user personas. Jia Shuo validated this with real-world data: Tianpu AI Guitar launched spot sales at the end of last year and 'surpassed RMB 100 million in GMV in a very short time.' The growth driver wasn’t traditional instrument enthusiasts, but rather ordinary people who 'enjoy music and dabble casually in instruments'—a far larger incremental market than professional users. He even challenged the popular industry phrase 'lowering the learning barrier' at the product-definition level: 'Aside from those studying for school entrance exams or civil service tests, who actually wants to pay money to suffer?' Users never pay for an 'easier learning process'; they pay to skip the pain and reach a sense of achievement directly. This insight aligns with broader industry trends: device manufacturers are collectively embedding AI capabilities as seamless, default experiences. Before WAIC, Jieyue Stars released its full suite of Step Edge on-device AI models targeting smartphones and automobiles, while Mosaic Intelligence’s on-device agent SuperMate is expected to be deployed in over 300,000 mass-produced vehicles by the end of 2026. Most large-scale adoption of on-device AI begins precisely in scenarios where 'users don’t need to learn.'
B2B scenario selection follows a different algorithm. Zhu Haohua proposed a '1+3+N' framework for dexterous hands: one foundational technology platform, anchored in three key scenarios—industrial, power, and scientific research/education—and then extended to N additional use cases.Three criteria are used to screen scenarios: the degree of environmental structuring, technological maturity coupled with safety and ethics considerations, and the clarity and scalability of ROI.Industrial scenarios have the clearest SOPs; high-level decisions are already mapped out by humans, so edge devices only need to handle tasks akin to 'segment-level' operations in autonomous driving, resulting in low replication costs. In contrast, home environments are highly unstructured and require substantial cloud computing support as a safety net—which is why B2B industrial applications lead the way, while C2C home applications follow later. Yin Jilei added nuance regarding timing differences: certain C2C sub-segments like robotic vacuum cleaners have been validated for over a decade and are now expanding overseas into categories such as lawn mowing, pool cleaning, and snow shoveling. Meanwhile, B2B applications concentrate on essential needs where 'humans can’t go' (e.g., hazardous biochemical environments) or 'don’t want to go' (e.g., dirty, repetitive logistics sorting tasks).
Regardless of whether it’s B2B or C2C, the first principle for achieving scale is that the scenario must address genuine industry pain points—not fabricated demands.
The second gate is cost.The cost ceiling for mass-market consumer electronics is often not dictated by the technical solution itself, but by 'how much users are willing to pay.' Jia Shuo has felt this acutely: edge-side computing and storage costs continue rising due to supply-demand dynamics—even Apple has had to raise prices. 'Product developers are essentially dancing in chains,' he noted. Pursuing the ultimate experience without regard for cost may be idealistic, but commercial mass production must find the optimal solution within the cost structure defined by user willingness to pay. Yin Jilei further explained the cost transmission mechanism from the chip supply chain upstream: a single high-end edge AI chip can cost several thousand to tens of thousands of yuan. When combined with peripheral components, the total system cost remains prohibitively high—humanoid robots often carry price tags of hundreds of thousands of yuan, 'sometimes even more than a car,' placing them far beyond most consumers’ reach. This creates a vicious cycle: 'Prices will naturally drop once volume scales up, but without lower prices, it’s hard to achieve scale in the first place.' His conclusion was candid: 'This isn’t something our company alone can solve—it’s an industry-wide challenge.'
The third gate is the supply chain.This is the biggest hurdle in transitioning from a 'crowdfunding hit' to 'tens of thousands sold annually.' Zhang Yizhen experienced this most directly: last year, XbotGo’s third-generation product, Falcon, raised USD 2.5 million on Kickstarter, but 'crowdfunding is merely a pre-sale validation—the real battle begins with mass production.' Delivery delays, across-the-board price hikes for components like memory, and production ramp-up challenges each inflicted significant pain. Even more difficult was the shift in user expectations: crowdfunding backers are forgiving and accept imperfect products, but after mass production, regular consumers 'demand extremely high reliability and expect every promised feature to work flawlessly.' The transition from 'high tolerance' to 'zero tolerance' drastically compresses the window available for bug fixes and software iteration.
The fourth gate is engineering implementation.Mass production delivery isn’t about lab prototypes—it’s about industrial-grade consistency. Amid an industry-wide rush to declare this the ‘Year One of Mass Production,’ Lingqiao Intelligence has deliberately chosen not to blindly scale output. The prototype of their new generation tendon-driven dexterous hand, unveiled at WAIC, was actually first introduced as early as May 2025. Yet the team spent a full 14 months afterward resolving engineering challenges related to lifespan, reliability, thermal rise, and more. ‘Many manufacturers probably wouldn’t choose to invest this much time in such issues,’ said Zhu Haohua. But in his view, customers aren’t buying demos—they’re buying production tools capable of running continuously for thousands of hours on factory lines. Those 14 months were a necessary cost under their ‘scenario-first’ strategy.
Yin Jilei ultimately highlighted, from a chip perspective, the deeper dual bottlenecks underlying mass production. On the algorithm side, cross-scenario generalization of VLA and world models ‘hasn’t even reached convergence yet—it’s still a period of diverse exploration.’ Without stable algorithms, chip architectures lack a fixed definition anchor. On the data side, ‘even after algorithms stabilize, we still face the “last-mile” data challenge’: synthetic simulation data can reduce some costs, but real-world deployment in specific scenarios still requires real-robot data, which remains expensive. Moreover, multimodal sensing—such as tactile feedback—‘has only been incorporated in the past year or two,’ so data accumulation is still in its infancy. This means that mass production of robot ‘brain’ chips has never been about ‘proactively scaling output’; rather, it’s driven by upstream and downstream forces: algorithm convergence defines the chip specification window, robot body scaling determines order volume, and supply chain cost reductions shape pricing room. Weifan Intelligence’s targeted Q4 2027 mass production milestone is less a single company’s achievement and more an optimistic expectation of synchronized convergence across four variables: algorithms, data, robot bodies, and the supply chain.

Only when these four gates align does the complete picture of mass production become clear: scenarios define the ceiling, cost determines survival, the supply chain tests organizational capability, and engineering excellence builds competitive moats. Breaking the vicious cycle of ‘no volume means no cost reduction, and high prices prevent scaling’ requires not just one company’s sprint, but the gradual establishment of a virtuous industry-wide loop.
As software companies begin building hardware and hardware firms dive into software, once-clear industry boundaries are blurring—ByteDance launched AI-powered recording hardware, OpenAI ventured into hardware, and DJI, though hardware-native, perfected its flight control software. ‘Everyone is crossing over and going full-stack,’ summarized Zhang Yizhen. Post-crossover, the real test is no longer ‘can you do it?’ but ‘what anchors your ecological niche?’ The four panelists offered four distinct answers.
Quwan’s answer is ‘open APIs, closed-loop hardware’. Quwan opens its model and agent layers to partners: ‘Any capability that can be easily delivered via API, we keep open.’ Yet the company emphasizes, ‘many solutions and user experiences simply can’t be fully delivered through cloud endpoints alone.’ On-device hardware thus serves both as an ecosystem gateway and a vehicle for a self-contained commercial loop. This approach has already shown early validation: Quwan’s pure application layer has achieved operational breakeven, while deeper integration between hardware and agent applications ‘will proceed at a measured pace.’ Openness is the posture; closed-loop capability is the foundation.
XbotGo is betting on the value loop created by the integration of hardware and software, using ByteDance’s AI-powered voice recording hardware as an example:Strong software capabilities combined with a hardware platform can funnel more users from hardware into the software ecosystem.XbotGo is following the same path, evolving from 'a single camera' to 'vertical multi-platform': beyond shooting, it offers subscription-based software services tailored to different user roles, such as highlight editing, team management, and coaching.Hardware is the entry point; software is the second growth curve.—Although she admits the monetization model is 'still being validated,' she adds, 'it’s clearly something that can happen.'
Lingqiao Intelligence is placing its bet on the hardest-to-collect type of data: physical interaction data. Zhu Haohua believes that 'data essentially determines the upper limit of any model. Nearly all easily accessible data today is video-based, but real-world operations inevitably involve physical contact with objects—and that contact data is extremely difficult to collect.' Even harder are endogenous force data like muscle tension. When asked whether the entry point for embodied intelligence will lie in tactile data or models, his answer was cautious yet firm: 'It’s hard to say where the operating system’s entry point will emerge, but to some extent, I agree that tactile sensing is a critical data gateway.' In today’s world, where video data is becoming increasingly commoditized, whoever can effectively capture and represent physical data—such as tactile feedback and force control—and bridge sensors with models will hold the next competitive moat. Dextrous hands, in essence, are the shovels for mining this gold mine.
Yin Jilei’s vision of ecosystem building carries the boldest 'challenger' spirit. Faced with NVIDIA’s dominance—holding 70–80%, or even over 90%, market share in certain segments—his answer is to trade openness for ecosystem growth: just as Android challenged Apple and RISC-V challenged ARM, 'we should open-source our foundational IP as much as possible, enabling broader feedback and adoption, so the ecosystem naturally flourishes.' He even made a controversial claim: AGI is a 'so-called pseudo-problem'—different scenarios demand different algorithms, just as humans have different professions; yet all these professions share the same brain. Whether it’s robotic hands or computing chips, a unified platform may emerge, but it must be sufficiently open. The loosening grip of CUDA already signals a trend: underlying hardware is becoming more通用 (general-purpose), and tools like OpenAI’s open-source Triton are driving software stacks toward greater diversity. Still, he issued a pragmatic warning: 'If it’s not user-friendly, nobody will use it.'Open-source is merely an amplifier; usability is the prerequisite.。”
These four approaches differ in strategy but converge on the same ultimate conclusion:The endpoint of ecosystems isn’t closure—it’s differentiation within openness. The degree of openness may vary, but core capabilities must remain firmly in one’s own hands.. As for who will win, Jia Shuo’s summary may come closest to the answer: 'It doesn’t matter where you start—any path could potentially create a formidable advantage.'Ultimately, the real deciding factor is who can first capture application-side scenarios and identify genuine, mass-producible, and scalable demand.。”
Looking back at this discussion, the four founders were essentially answering the same question together: what does the 'awakening' of on-device AI truly mean?
It’s not simply about stuffing cloud-based models into devices. Rather, it means AI must confront, for the first time, all the constraints of the physical world—millisecond-level real-time performance, razor-thin BOM costs, every supply chain price hike, and thousands of hours of continuous operation on production lines. Cloud AI can forever live inside demo videos, but on-device AI must actually 'work' in the real world. As Zhu Haohua put it at the outset:In the exam hall of on-device awakening, 'it’s not about having an all-capable body, but about having a pair of hands that can get things done on-site.'。
A decade ago, the mobile internet boom wasn’t triggered by faster smartphone chips, but because touchscreens, the App Store, and 4G networks converged within the same window of opportunity.Today’s on-device intelligence stands on the eve of a similar inflection point: regulatory filings have opened compliance pathways, on-device compute has crossed the usability threshold, and model developers, device makers, and chipmakers are now seated at the same table. What’s missing isn’t any single technology—it’s the precise moment when algorithms, data, hardware, and cost structures all converge simultaneously.。
Awakening happens in an instant, but growing up demands patience from the entire industry. When the noisy exhibition lights dim, the real race begins—in the milliseconds users can’t afford to wait, on the balance sheets of supply chains, and through thousands of operational hours on production lines. Whoever endures this rite of passage first will define how AI interacts with the physical world for the next decade. (This article is based on live notes from the 'WAIC UP! AI Three Poles Night Talk' event.)
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Comments
to post a comment
1
