Skip to content
HN On Hacker News ↗

MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

▲ 334 points 94 comments by vblanco 3w ago HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

62 %

AI likelihood · overall

Mixed
45% human-written 55% AI-generated
SEGMENTS · HUMAN 1 of 2
SEGMENTS · AI 1 of 2
WORD COUNT 446
PEAK AI % 95% · §1
Analyzed
Aug 3
backend: pangram/v3.3
Segments scanned
2 windows
avg 223 words each
Distribution
45 / 55%
human / AI fraction
Verdict
Mixed
Pangram v3.3

Article text · 446 words · 2 segments analyzed

Human AI-generated
§1 AI · 95%

MiniMax H3 dropped today with open weights, and it’s natively supported in ComfyUI as of this morning. Day zero.This is a next-generation open-weights video model. Feed it text, images, video, or audio and it generates video with real stereo sound, up to 2K, up to 15 seconds a clip. It is MiniMax’s third-generation video model, following Hailuo 01 and Hailuo 02, and the first the company has released with open weights.Try on Comfy CloudText-to-video — prompt only.Image-to-video — bring an image to life.First-and-last-frame — control the opening frame, the closing frame, or both, and let the model fill in the rest.Reference-to-video — supply reference images, video, or audio and carry a subject, a motion, or a voice through the clip.Output runs to 2K and up to 15 seconds. Audio is generated with the video in the same pass, in stereo, not bolted on afterward.This is the capability MiniMax leads with, and it’s what collapses five separate tasks into one model. Real work rarely draws on one modality. H3 takes images, audio, and video together and resolves them against a prompt that explains how they relate. Describe the relationship between your inputs and the shot you want, and the model handles the cross-modal work itself.Audio is a property of the model, not a post-process. Every audio output is native stereo.Motion transfer is the one that matters most for graph work. A reference video can supply movement — a camera move, a performance, a cutting rhythm — while the subject and style come from elsewhere.

§2 Human · 22%

Combined with in-place editing, that means iterating on a shot.Getting H3 to run well on consumer hardware took significant machine learning engineering. We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality.On top of that, the weights ship with an accurate and efficient int8 convrot quantization, and custom kernels reduce the peak VRAM use during inference.The result gives a total memory footprint reduced by 66%, from 123.6 GB in full precision to 42.5 GB with the smallest models variants. Combining this with our dynamic VRAM offloading enables a next-generation 2K video model to run locally on a GPU like the RTX 3060.Update ComfyUI to the latest version 0.30.0 or go to Comfy CloudDownload the workflows below, or find them in the template library.Download MiniMax H3 I2V WorkflowDownload MiniMax H3 R2V WorkflowDownload MiniMax H3 T2V WorkflowFollow the note in the workflow to download the models and save them in the correct model directory.Write your prompt, connect any frame or reference inputs, and run.Model weights: 🤗 Comfy-Org/MiniMax-H3As always, enjoy creating!No posts