ShengShu Technology Launches Vidu S2 for Enhanced Real-Time Interaction and Video Editing
On September 15, 2026, ShengShu Technology took a significant leap into the world of AI video generation by unveiling their latest creation, Vidu S2. This innovative product consists of two distinct models:
Vidu S2-Avatar and
Vidu S2-Editing. Together, these models aim to revolutionize how users interact with digital video content by introducing unprecedented real-time interaction and editing capabilities.
Vidu S2-Avatar: The Future of Interactive Characters
Vidu S2-Avatar focuses on improving continuous interaction with digital characters. This model significantly enhances the interactivity features first explored in Vidu S1, allowing users to upload reference images while a character is talking. This feature enables the character to dynamically interact with the provided content in real-time. One of the most notable advancements in S2-Avatar is the increase in output resolution, now capable of delivering videos in
720p, a step up from its predecessor's 540p. Additionally, users can update reference images at any moment during video generation, offering greater freedom in character customization.
The model can recognize and respond to a variety of commands, including more complex movements like dancing. For instance, if a user instructs the character to “pick up a cup from the image and then smile,” the character executes the actions in succession while maintaining the cup. This feature is made possible through advanced visual feedback mechanisms and state preservation technologies, ensuring seamless interaction.
Vidu S2-Editing: Intelligent Video Editing in Real Time
On the editing front, Vidu S2-Editing presents a powerful tool for anyone dealing with incoming video streams. This model edits continuous video feeds based on user-defined instructions and can work with optional reference images. It is capable of performing four primary editing tasks:
1.
Real-Time Style Transfer: Adjusts the visual aesthetics of a video based on a reference image while maintaining key elements intact.
2.
Real-Time Outfit Change: Enables characters to switch outfits by applying new textures and styles from reference images seamlessly.
3.
Real-Time Subject Replacement: Allows the facial features and overall appearance of subjects in the video to be altered based on user input.
4.
Real-Time Background Replacement: Swaps out the environment in a video while accurately maintaining the arrangement of subjects within that environment.
These editing capabilities are crucial for streamers and content creators looking to enhance their videos in real-time without the need for laborious post-production techniques.
Exploring New Dimensions with Spatial Video
ShengShu Technology is not stopping at conventional uses; they are pioneering the integration of spatial video, allowing users to perceive depth and movement in a three-dimensional space. This feature can enhance virtual reality experiences by providing synchronized left and right eye views during video streaming, lending an immersive quality to interactions with characters in the virtual realm.
Technical Innovations Driving Vidu S2
The functioning of Vidu S2 is backed by robust technical improvements. To ensure coherent action, the video models utilize
Self-Replay Forcing (SRF)—a training method that corrects errors across video segments. Additionally, a sophisticated
Vision-Language Model Agent (VLM Agent) evaluates user commands, adjusting outputs based on the feedback loops established through interaction.
Furthermore, the architecture supporting
S2-Avatar’s high-resolution outputs has been carefully optimized to ensure efficiency in processing while maintaining visual quality. These techniques are translated into real-time performance, distinguishing ShengShu's models in an increasingly competitive field.
Evaluating Performance and Next Steps
Initial performance evaluations show promising results for both models, surpassing expectations in user preference tests and maintaining high-quality output across several metrics. As they move forward, ShengShu aims to refine spatial video technologies, addressing challenges related to latency in VR environments and exploring panoramic spatial video capabilities.
With Vidu S2, ShengShu Technology solidifies its position at the forefront of AI video creation and editing, paving the way for rich and immersive experiences that both engage and empower users. This breakthrough can redefine how we conceive interactivity in digital media, making it an exciting time for creators and consumers alike.
For further exploration of Vidu S2 and its capabilities, visit
Vidu's official platform and delve into the latest in AI-driven video content.