Provides GGUF-quantized, ComfyUI-ready MiniMax‑H3 model files (FL2VA/REF2VA, text encoder, audio/video VAEs) to enable local ComfyUI inference for short video + stereo audio generation; requires the official VAEs and sufficient VRAM.
A continuous-latent diffusion language model that preserves a high-capacity, decodable text latent and directly models its distribution via a block-causal diffusion transformer and query-based encoder–decoder; achieves top results on OpenWebText and XSum while scaling to 1B parameters.
Installation-oriented dataset that packages ComfyUI-ready files and instructions for running MiniMax H3 locally — includes pruned/INT8/BF16 checkpoints, matching Qwen3-VL text encoders, video/audio VAEs, and official ComfyUI workflow templates for joint audio+video generation.
Experimental MiniMax H3 variant that injects learned stylistic and motion 'character' from LTX 2.3, Wan 2.2 and Krea 2 into H3 by surgically grafting attention and MLP components; preserves H3 modality routing while shifting t2v/i2v aesthetics, with limited audio impact and community-license constraints.
A LoRA adapter for MiniMax-H3 that enables joint video + synchronized stereo audio generation in as few as 4 sampler steps, cutting sampling time roughly ~5×; early prototype under-trained, so 6–8 steps or newer checkpoints give better sharpness.
Provides ComfyUI-compatible pruned/curve-form LoRA conversions of the MiniMax‑H3 Turbo 4-step audio‑video generation preview, including further-trained ckpt500 EMA and non‑EMA variants and an example ComfyUI workflow for low-step experiments.
Packaged diffusers checkpoint of MiniMax H3 for image/text-to-short-video generation with native stereo audio; provided for direct use in diffusers image-to-video pipelines and aimed at easy integration into prototyping and production workflows.
Provides ComfyUI-compatible conversions and LoRA adapters of the MiniMax‑H3 video+audio generative model, with example presets and demo videos to run short stereo audio+video inference inside ComfyUI workflows.
Generates complete songs (up to five minutes) from lyrics and a music description, producing 32 kHz stereo WAV with expressive vocals and long-range musical structure. Uses hierarchical LLMs fused with flow-matching/Flow-VAE synthesis for coherent arrangement and timbre; requires CUDA and integrates with Diffusers and SGLang-Omni.
Provides 1,000 five-second video clips generated by MiniMax H3 for lightweight evaluation of multimodal generation and understanding. Clips are roughly 768p base resolution with diverse aspect ratios and themes, produced with a pruned int8 minimax_h3_fl2va checkpoint at 30 steps.
A 2.9B-parameter text-to-image model fine-tuned from CircleStone Labs' Anima for anime and illustration; trained on an additional 1.7M samples with a July 2026 knowledge cutoff. Designed for non-commercial creative image generation and ComfyUI integration; weights released under the CircleStone Labs Non-Commercial (derivative) license.
Contains ~2 million human pairwise preference judgments comparing images generated from text prompts; each example pairs two images with a preferred/tie label and is formatted for preference learning, reward-model training, and evaluation.