MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

316 points · 89 comments on HN · read original →

Points and comments are a snapshot, not live.

MiniMax H3's open-weights video model runs locally on a 3060 after a 66% memory reduction.

ComfyUI offers day-0 support for MiniMax H3, an open-weights omni-modal video model accepting text, images, video, or audio. It generates up to 15-second clips at 2K resolution with native stereo audio in a single pass. Features include text-to-video, image-to-video, first-and-last-frame control, and reference-to-video for subject or motion transfer. ComfyUI optimized the model by pruning modulation weights and applying int8 quantization, reducing memory footprint from 123.6 GB to 42.5 GB, enabling local inference on a GPU like the RTX 3060 with dynamic VRAM offloading. Workflows are available for immediate use.

What commenters are saying

Commenters are impressed by the output quality, calling it a leap forward and a win for open-source video generation. Performance benchmarks show 10 minutes for a 10-second 480p clip on a 4070 Ti Super and 3 minutes on a 5080; optimizations like SageAttention and EasyCache can reduce times further, though quality trade-offs exist.

Some debate whether H3 surpasses Seedance 2.0/2.5, with one commenter arguing it is still behind. Others note licensing restrictions in the US, EU, and UK. Several commenters describe the outputs as 'slop' and express concern about AI-generated content displacing human creative work.