律动BlockBeats
律动BlockBeats|Jul 31, 2026 02:35
[MiniMax Releases All-Modal Generation Model H3: A Single Model Handles Images, Videos, and Audio] According to monitoring by Beating, MiniMax has released the all-modal generation model H3. It can simultaneously understand text, images, videos, and audio, and then generate or edit videos based on a single natural language command. Users can instruct H3 to mimic camera movements from one video, use a character from another image, and reference audio from a third clip. Previously, actions like motion transfer, character referencing, audio referencing, and video editing had to be invoked separately, but now they can be integrated into a single task. H3 can generate videos up to 15 seconds long, supporting 2K resolution and native stereo sound. The official API pricing for 2K resolution is $0.13 per second, meaning a 15-second video costs approximately $1.95; for 768P resolution, the cost is $0.09 per second. A single task can input up to 9 images, 3 videos, and 3 audio clips, with a total limit of 12 files. MiniMax plans to release the model weights in the coming days. The official evaluation results comparing H3 with models like Seedance and Veo have not yet been disclosed. [Original Link]
+6
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads