Provides GGUF-quantized, ComfyUI-ready MiniMax‑H3 model files (FL2VA/REF2VA, text encoder, audio/video VAEs) to enable local ComfyUI inference for short video + stereo audio generation; requires the official VAEs and sufficient VRAM.
Provides ComfyUI-ready INT8 MiniMax‑H3 checkpoints (conditioning encoder plus optional generation tail) for a Heretic-edited Qwen3‑VL‑32B source; preserves the vision tower in BF16 and uses row-wise ConvRot INT8 quantization to reduce VRAM needs for ~32GB GPUs. Not a full Transformers generation repository.
ComfyUI-ready H3 conditioning encoder builds for Qwen3-VL-32B: a BF16 full-precision checkpoint, an INT8 ConvRot quantized checkpoint, and an optional generation tail (layers 50–63). Retains vision tower in BF16 and targets H3 workflows and lower-VRAM systems.