MiniMax H3 packs multi‑modal video and native stereo audio generation into a single model pass; this repository makes those assets accessible inside ComfyUI by providing converted model files, LoRA adapters, example presets and demo outputs. The practical effect is a lower-friction path for ComfyUI users to experiment with MiniMax‑H3-style T2VA/FL2VA/Ref2VA workflows without hand-building node stacks from scratch.
What Sets It Apart
- ComfyUI-ready conversions: packaged weights and processor/tokenizer files formatted for ComfyUI nodes so users can drop them into existing Comfy workflows.
- LoRA adapters and recommended strengths: includes distilled LoRAs (example strengths noted) to tweak style or resource usage without retraining full weights.
- Demo content and presets: example videos and preset node configurations help reproduce expected outputs and iterate quickly.
- Focused on inference and integration: aimed at using MiniMax‑H3 within ComfyUI rather than providing new training pipelines or novel model research.
Who It's For and Tradeoffs
Great fit if you use ComfyUI and want to prototype short multimodal video+audio generations with ready-made MiniMax‑H3 conversions and LoRAs. Look elsewhere if you need official upstream training code, licensed original checkpoints bundled here, or lightweight CPU-only execution — running full generative video models typically requires a capable GPU and/or INT8 optimizations and may rely on obtaining original MiniMax weights or partner runtimes separately.