Vidu S2: Real-time digital person, real-time video editing.

CN
1 hour ago

Just now, the latest Real-time Video Model S2 has been launched, output improved to 720p, maintaining 25—42 FPS, including the following two:

  • S2-Avatar, Real-time Digital Human You can talk to it, make it do actions, like change clothes, pick up things, change scenes, supporting adding reference images at any time

  • S2-Editing, Real-time Video Editing Input a segment of ongoing video, the model alters the art style and content according to instructions, while retaining the original video's actions and temporal relationships

For example... the left side is style transfer, and the right side is it changed into another outfit

This can already be played, experience it here: https://vidu.com/vidu-stream

In addition, an interesting product was launched this time called Real-Time Spatial Video: Generating real-time spatial video based on S2 for use with headsets, the principle is to turn the generated image into left and right eye views, serving as a VR input source

This can be seen in conjunction with the news below, where everyone is exploring the " Path yet to be completed by Zuck"

S2-Avatar: Real-time Digital Human

Avatar continues the route of S1, generating an interactive digital human with just one image. But S2 has made a significant improvement called "Dynamic Reference," allowing you to change character settings, scenes, clothing, and props at any time during the dialogue

For example, during the conversation, you can give it a picture of a cup so it can pick it up, or a photo of a coat so it can put it on

Additionally, I checked the tech report and found an interesting detail about how Vidu ensures the consistency of character actions: They created a layer of VLM Agent to generate subsequent prompts For example, when instructing the character to "pick up the cup and then smile," this Agent retains the hint "continue holding the cup," only changing the expression, avoiding new instructions from overwriting the previous state

The video model is responsible for generation, the Agent is responsible for maintaining stateThe video model is responsible for generation, the Agent is responsible for maintaining state

S2 has achieved a significant improvement in delivery effects. During the training process, additional solo dance and 2D, 3D animation data were added, and the instructions were expanded to include more complete body movements. A low-resolution backbone combined with a single-step Refiner was used for rendering, with the backbone emphasizing motion and timing, and the Refiner preserving appearance details. The output improved from 540p to 720p, with a reported speed of 25—42 FPS

To ensure stability in long-term generation, they employed their proposed Self-Replay Forcing, abbreviated as SRF, which means the model first generates a long trajectory according to the actual reasoning method, then re-adds noise to the generated segments, conducting causal playback training to improve error accumulation in continuous generation

S2-Editing: Real-time Video Editing

Editing directly receives video streams, changing art styles and content in real-time based on text instructions and optional reference images. Unlike Avatar, here the actions are provided by the original video, that is, used for "P video"

For instance, if you're doing a gesture dance in front of the camera and you upload a picture of a cow, it can let the cow jump in the output video; it currently supports four types of tasks: style transfer, virtual try-on, character replacement, background replacement

The following cases come from the tech report

Style transfer: watercolor and fine brush styleStyle transfer: watercolor and fine brush style

Virtual try-on: white shirt and denim jacketVirtual try-on: white shirt and denim jacket

Character replacement: real person and cartoon characterCharacter replacement: real person and cartoon character

Background replacement: Paris street scene and bedroomBackground replacement: Paris street scene and bedroom

Swipe left and right to view · Total of 4 images

This technology has a challenge: how to ensure the modified images do not distort, such as when pulling a shirt, if I change the shirt to another outfit, it also means the patterns, textures, etc., on the clothing should change as well, and the model must understand the clothing features in the reference image while retaining the interaction between the person and clothes in the original video

To address this, the Vidu team used the Frame-Aligned Attention solution, where the target frame only exchanges information with the source frame at the same time when reading the source video; reference images can be read by each target frame, continuously providing new appearance conditions, avoiding mixing source images from different time points together. Once in the streaming generation, the model continues to output by combining effective historical segments.

In training, Editing also reused the previous causal streaming solution and SRF to improve consistency in continuous editing and performed special memory optimization to reduce the overall process delay

Real-Time Spatial Video: Delivered to Headsets

Building on the previous two models, Vidu has taken another step forward, providing slightly different images for the left and right eyes, achieving "Real-Time Spatial Video"

In this process, ordinary video is generated first, and then depth information is used to create left and right eye views, continuously sent to headsets; Monocular image, transformed into left and right eye viewsMonocular image, transformed into left and right eye views

If it is real-time edited video, it divides into two scenarios:

  • If the original video is monocular, it will first modify the ordinary video, then convert it to binocular

  • If the original video is binocular, the left and right eye images are stitched together to form a wide video for editing, then split back to binocular to display with two images based on the same principle

Swipe left and right to view · Total of 2 images

BenchMark

According to official statements, based on BenchMark evaluation, the Vidu S2 model has achieved state-of-the-art in the field

The full chain R&D of Vidu S2 is led by Professor Zhu Jun's doctoral student, Zhang Jintao, who is also the head of Vidu Technology's streaming video generation and reasoningThe full chain R&D of Vidu S2 is led by Zhang Jintao, a doctoral student of Professor Zhu Jun, who is also the head of Vidu Technology's streaming video generation and reasoning

The BenchMark here includes the corresponding StreamAV-Bench for Avatar, as well as Sparkle-Bench, OpenVE, RefVIE, and ViViD virtual try-on test sets for Editing, summarized as follows

StreamAV-BenchStreamAV-Bench

Sparkle-BenchSparkle-Bench

OpenVE and RefVIEOpenVE and RefVIE

ViViD virtual try-onViViD virtual try-on

Swipe left and right to view · Total of 4 images

Finally

The product has been released today through Vidu's official website, available as an online demo and API, you can try one related to the technical report here, those interested can read the technical report: https://arxiv.org/pdf/2609.11638

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink