AI's Evolving World Models: A New Era Beyond Content Generation
AI's Evolving World Models: A New Era Beyond Content Generation
In a groundbreaking declaration made during the World Artificial Intelligence Conference (WAIC) on July 19, 2026, Fang Han, Chairman and CEO of Kunlun Tech, proclaimed that the year marks the first year of world models. This announcement indicates a significant shift in the artificial intelligence landscape, moving from the realm of mere content generation to a profound understanding and interaction with the physical world. For the past two years, the focus within the AI industry has predominantly revolved around generating various forms of content, including text, images, videos, and music. However, this new phase aims to serve a more significant purpose: comprehending and interacting with reality in ways that were previously unattainable.
I. Transitioning from Generation to Understanding
The transition from generating content to understanding the physical world can be likened to the evolution of photography as explained by Jiao Juan, Chief Analyst for Internet and Media at Founder Securities. When photography was invented, it didn’t merely create a new form of visual representation; it dismantled the economic underpinnings of realist painting and birthed cinema. This analogy effectively articulates a pattern where genuine revolutions pave the way by first dismantling existing supply structures before creating new forms.
II. Breakthrough in Long-Term Memory
Skywork AI's recent technical advancements underscore a remarkable transformation in world models, particularly concerning long-term memory. Traditionally, these models “remembered” information by capturing entire frames, leading to significant challenges when the camera moved, rendering previously recognized objects unidentifiable. The evaluation data indicated that mainstream world models could not achieve an object reappearance score above 0.6. However, with the introduction of Matrix-Game 3.5, an innovative interactive world model, Skywork AI has shifted the paradigm. Instead of merely storing individual frames, this model now segments each frame into numerous small patches, each with assigned three-dimensional coordinates. This novel approach, termed “Patch Memory,” allows the model to retain spatial awareness rather than simply visual representation, addressing the long-standing issue of disappearing objects within the industry.
The advancements in positional encoding further enhance this capability; transforming the previous conceptualizations of time and space into a comprehensive geometric coordinate system. This architectural shift results in a model capable of real-time video generation at a resolution of 720P and around 20 frames per second with one-minute memory retention. As articulated in the technical report, the model's interactive elements elevate it from being a video generator to producing a “living world” that responds instantaneously to user input.
III. Data-Driven Insights
An essential foundation for AI's understanding of the physical world is robust data. The Skywork team realized that training intelligent models necessitates not just visual data but videos enriched with physical information. Their innovative approach includes a multi-agent system that continuously assesses thousands of games, generating over 10,000 hours of rich, structured training data with integrated physical attributes, such as camera pose and depth maps.
IV. Architecture Over Parameters
Interestingly, the design team made a deliberate choice not to overload their model with new parameters, opting instead for a modular and transferable architecture. Their strategy involved developing a lightweight “plugin system” that retains the core interactive capabilities without compromising the foundational model's open-ended content generation ability. Such an approach highlights a distinct shift toward architectural finesse over mere parameter accumulation—an essential consideration in competitive AI development.
V. Academic Contributions
The academic community also enriches this field with theoretical advancements. Zhou Zhihua from the Chinese Academy of Sciences outlined three layers of world models—perception, understanding, and decision-making. Noteworthy breakthroughs in these areas promise to align theoretical models with practical applications, advancing the potential for AI systems to evolve into more adaptable and inclusive frameworks.
VI. Industry Impact and Innovations
A turning point in AI's evolution was pointed out by Huang Xiaoming, Vice Chairman of the China Film Association, who described witnessing the quality of AI-generated cinema evolve to a level indistinguishable from traditional cinema. This marks a critical inflection point where AI has matured enough to compete directly with conventional production techniques.
VII. Conclusion
Fang Han’s assertion regarding this year as the first for world models is not a conclusion but rather an invitation for exploration and innovation. As the gaming sector gears up for transformation and the boundaries of content creation expand, the pathways for AI's integration into everyday experiences are becoming more dynamic. This signals that the continued evolution of AI's role in media and interaction is only beginning, with the potential to reshape entire industries in the process.