AI is undergoing a pivotal transformation: shifting from generating a piece of content or providing a single response to completing long-horizon tasks and delivering system-level outcomes.
Code-related products have already pioneered this path: evolving from Copilot assisting professional developers, to Vibe Coding expressing intent through natural language, and now to Agentic Coding—planning, executing, verifying, and delivering around defined objectives. Today, multimodal content creation is undergoing the same evolutionary shift.
On July 18, at an event hosted by the Artificial Intelligence Committee of the All-China Federation of Industry and Commerce, the Artificial Intelligence Special Committee of the Shanghai Federation of Industry and Commerce, and the World Artificial Intelligence Conference Organizing Committee, and organized by SenseTime,"Foundation Large Model Architecture Innovation and Ecosystem Collaboration Forum"SenseTime officially launchedits delivery-grade, native multimodal agent foundation model for long-horizon tasks—SenseNova U1 Pro.The industry is now formally moving beyond traditional 'point-in-time content generation' and fully entering a new era of 'system-level content delivery,' capable of closing the loop on complex, long-horizon tasks.

Dr. Xu Li, Chairman and CEO of SenseTime,stated in his speech: 'Multimodal AI is following a clear evolutionary trajectory toward long-horizon capabilities: from initially solving single-point creative tasks focused on 'generating beauty,' to interactive and controllable generation, and now, decisively advancing toward system-level delivery. When understanding, planning, generation, verification, and correction form a complete long-horizon closed loop, the longstanding industry bottleneck—where controllability does not guarantee deliverability—will finally be broken.'
The long-horizon evolution of SenseNova U1 Pro:
From 'Gacha' to 'Delivery'
From SenseNova U1’s native unification of understanding and generation to SenseNova-Vision’s native integration of classical vision capabilities, SenseTime has continuously refined and iterated its multimodal technology roadmap, officially unveiling a flagship addition to its large model family—SenseNova U1 Pro(hereinafter referred to as U1 Pro).
For more information, visit the SenseNova official website (click 'Read More' to redirect):
Currently, most multimodal AI remains stuck in the 'vibe creation' phase—capable of generating and modifying content, but minor adjustments often disrupt the overall structure. Professionally oriented outputs fail scrutiny at the detail level and fall short of real-world delivery standards.
Real-world visual tasks extend far beyond 'generating an image.' A complete visual task must simultaneously satisfy multiple requirements: content comprehension, information structuring, text-image alignment, layout hierarchy, textual accuracy, fine-grained detail rendering, and stylistic consistency.Click to view enlarged image↓↓ (Image quality compressed due to upload specifications)





Centered on 'a complex objective, continuously understood, planned, executed, reviewed, and refined,' SenseNova U1 Pro delivers four core production-ready capabilities:
Peak design aesthetics—free from the 'AI look': Image generation has advanced from 'photorealistic detail' to a new stage of 'professional design aesthetics,' producing works with exceptional composition, color harmony, and layout—delivering production-ready quality directly.
Native 8K Ultra-High-Definition Output: Supports native 8K resolution output at maximum settings, effortlessly handling ultra-long and ultra-large image creation with high information density. Text, lines, icons, and modular relationships remain sharp and stable even when zoomed in, meeting the rigorous detail standards required for printing and exhibitions.
Precise Control Over Text-Image Integration and Details: Significantly enhanced capabilities in analyzing and integrating textual and visual information. Even under extremely high information density, U1 Pro maintains coherent overall layout and content expression, with exceptionally low text rendering error rates.
Long-Horizon Agentic Closed-Loop Reasoning:Inherently supports extended reasoning that interleaves text and visuals, enabling dozens of rounds of Agentic Generation Loops around complex objectives. In the latest version, it further achieves synchronized, precise control over both global style and local text editing, ensuring deliverables can be iteratively refined and reused.
This powerful visual creation capability and production-grade quality unlock numerous commercial applications—from enterprise services, brand marketing, and educational publishing to creative entertainment. SenseNova U1 Pro excels across diverse content types including infographics, presentation slides, promotional posters, explanatory illustrations, film storyboards, and digital art.
Following the roadmap of 'evolving from controllable generation toward production-ready output,' SenseTime has rapidly iterated U1 three times over the past two months, significantly improving complex layout stability, small-text rendering accuracy, and background consistency.


July 15 Version: Further enhanced editing capabilities now support simultaneous adjustments to global style and partial text, along with precise local modifications via selection tools—transitioning from stable generation to fully controllable editing.
Real user data shows that in June this year, the average daily image generation per U1 user tripled to 107 images. On GitHub, the repositories for U1 and Skills have seen continuous growth in stars, surpassing a combined total of 8,500. A large number of users have already integrated these tools into their real-world workflows to understand, generate, and iteratively refine complex content, completing a wide range of content creation tasks.

Recently open-sourced SenseNova-Vision breaks away from traditional approaches that stitch together multiple expert models, natively embedding classic vision tasks—such as detection, segmentation, and depth estimation—into a large model. This marks its evolution from a mere 'execution tool' into a 'world-understanding model,' establishing a foundational perception layer for multimodal deep understanding of the physical world.
Building upon U1’s unified understanding-and-generation capability and SenseNova-Vision’s native visual competence, the new SenseNova U1 Pro serves as the foundational native multimodal agent platform designed for long-horizon, deliverable-grade tasks. The WAIC Ninth Anniversary Scroll: Proof of a system-level delivery
WAIC Ninth Anniversary Scroll
Proof of a system-level delivery
To illustrate what 'deliverable-grade for long-horizon tasks' means, Xu Li demonstrated on-site a WAIC ninth anniversary scroll generated by SenseNova U1 Pro.
Four years ago, Professor Qiu Zhijie from the Central Academy of Fine Arts created 'The Map of Intelligent Convergence' for WAIC’s fifth anniversary. The intricate map integrated the conference’s journey, developments in artificial intelligence, Shanghai’s industrial ecosystem, and numerous key figures and events into a single visual system.
Today, SenseNova U1 Pro attempted an even more complex ninth-anniversary task on-site: generating a panoramic ink-wash scroll in Eastern aesthetic style that weaves together key milestones across nine years of AI industry development through a cohesive visual narrative.

Panoramic ink-wash scroll commemorating WAIC’s ninth anniversary (2018–2026). This is a system-level visual delivery task characterized by high information density, an exceptionally wide aspect ratio, and a long temporal span. It must encompass the entire nine-year history of the conference from 2018 to 2026 while accurately reflecting each year’s unique thematic focus, landmark events, and industrial inflection points. Maintaining a 4:1 ultra-wide format, the composition seamlessly integrates mountains, rivers, cities, roads, and mist to form a stylistically unified, continuous narrative—free of visible seams or modern UI-style hard-coded text overlays. From afar, it appears as a complete landscape-and-city scroll; up close, it reveals abundant legible details in both Chinese and English. (Image quality has been compressed due to upload specifications.)
Ultimately, the full scroll generated by SenseNova U1 Pro demonstrates coherent chronological storytelling, clear information hierarchy, and consistent Eastern aesthetic style. Even when zoomed in, the details remain precise—marking a leap from mere information stacking to a fully realized visual narrative.


WAIC 2018: The inaugural World Artificial Intelligence Conference (WAIC) launched with visuals highlighting WAIC’s starting point, AI PARK, Shanghai’s iconic landmarks, and the gateway through which AI entered the city—demonstrating how AI, for the first time in 2018, systematically integrated into urban life, industry, and public spheres through a dedicated conference. The seven 'AI+' themes displayed below vividly illustrate that WAIC focuses not only on technical concepts but also on transforming AI into tangible, experiential real-world scenarios.
The capabilities of SenseNova U1 Pro extend beyond static images—it can perform comprehensive visual planning at the front end of video creation, including world-building, character design, shot planning, and storyboard consistency, providing a stable visual blueprint for subsequent video generation.

Worldview Setting & 22-Panel Storyboard Overview for 'Blades of Sand and Shadow' (image quality compressed due to upload specifications)
Taking the desert ambush mission from 'Blades of Sand and Shadow,' showcased onsite, as an example: SenseNova U1 Pro first generated a complete worldview setting and 22 consecutive storyboard panels, ensuring consistent character appearance across shots, coherent action and plot progression, and unified storytelling through both wide-angle and close-up shots.
With a complete visual blueprint in place, video generation is then carried out via Seko—transforming video creation from a game of chance into a more controllable content delivery process where characters, narrative, and visual language are all precisely managed.

As more modalities—such as images, videos, audio, layout, and spatial elements—enter the same long-horizon closed loop, the scope of tasks expands significantly. AI will gain the ability to tackle far more complex real-world tasks and unlock a vast array of new applications.
It doesn’t just understand instructions and write code—it can also interpret visual content, spatial relationships, sequential shots, and real-world contexts, continuously verifying and refining its outputs during task execution to ultimately achieve 'multimodal system-level delivery.'
One more thing!
A month ago, SenseTime’s Little Raccoon identified Cape Verde as a dark horse with potential to advance to the Round of 32, becoming one of the few AIs to successfully predict this outcome.

For the highly anticipated final match, our Little Raccoon analyzed the teams’ advancement paths, key players, tactical matchups, historical head-to-head records, and critical statistics on-site, then used SenseNova U1 Pro to compile these insights into a comprehensive visual report. Calling all football fans to join us in examining the details~

(Due to upload specifications, video quality has been compressed)
Let’s wait and see the final outcome—together, we’ll test the limits of AI capabilities!
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Comments
to post a comment
4
