Open-weight video gen that actually delivers. Five days with MiniMax H3 on local hardware.
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
H3 weights went live on HuggingFace August 3rd and I started pulling them immediately. An omni-modal video model with native stereo audio in the same forward pass, where audio can actually drive the video generation? On open weights? I had to try it.
Five days in, the quality is legit. 2K at 24fps, 5 to 15 second clips, text and image and video and audio all in one shared context. You can throw 9 reference images, 3 videos, 3 audio clips at it per generation. The audio-driven mode where you feed it a track and it generates video that moves to the sound is genuinely unlike anything I've run from open weights before. But here's the catch: running this locally on a consumer GPU means your iteration speed is painful. Every failed generation costs you real time and power, and you will need retakes. Motion consistency is much better than I expected, but complex camera moves and longer clips still need multiple attempts to get right.
What actually saved my workflow is that APOB AI is self-hosting MiniMax H3 unlimited and free right now since the weights are open. I stopped grinding my local card for every test run. I prototype on their hosted H3, figure out which prompts and references produce clean results, then switch back to my own hardware when I want full control or need to run something custom.
For comparison, Seedance 2.5 from ByteDance also launched July 31st. Completely different approach: up to 30 seconds in a single pass at 4K with native audio, 50 multimodal references, and region-level editing to fix part of a shot without regenerating the whole clip. But API-only through BytePlus ModelArk, no published weights. If you actually want to run the model yourself, H3 is the one. Having open-weight video gen at this quality level sitting on HuggingFace is a real milestone for local inference.
[link] [comments]
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.