r/LocalLLaMA · · 1 min read

MiniMax H3: A New Open-Weight Video Model, Live in ComfyUI

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

MiniMax H3: A New Open-Weight Video Model, Live in ComfyUI

MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio.
In ComfyUI, you can use H3 for text-to-video, image-to-video, first- and last-frame generation, and reference-driven creation. H3 jointly generates the visuals and synchronized stereo audio, including dialogue, sound effects, ambience, and music, rather than adding audio afterward.
The open-weight H3 checkpoints support clips up to 15 seconds at 768p. MiniMax’s hosted H3 model also supports generation at up to 2K resolution.
During the stream, we’ll test the model live and discuss how H3 brings multiple generation tasks into one architecture, how its high-compression video representation improves efficiency, and what developers should know when setting it up locally through ComfyUI.

submitted by /u/pmttyji
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA