Provides ComfyUI-ready INT8 MiniMax‑H3 checkpoints (conditioning encoder plus optional generation tail) for a Heretic-edited Qwen3‑VL‑32B source; preserves the vision tower in BF16 and uses row-wise ConvRot INT8 quantization to reduce VRAM needs for ~32GB GPUs. Not a full Transformers generation repository.
Generates complete songs (up to five minutes) from lyrics and a music description, producing 32 kHz stereo WAV with expressive vocals and long-range musical structure. Uses hierarchical LLMs fused with flow-matching/Flow-VAE synthesis for coherent arrangement and timbre; requires CUDA and integrates with Diffusers and SGLang-Omni.
Provides 2-bit quantized weights of Qwen3.8-27B (~10.15 GB) for local deployment, enabling the full 27B parameter model to run on a single 24 GB GPU with long-context support. Delivered as safetensors plus a companion SGLang runtime; measured to match FP8 reference on common benchmarks with small or no quality loss.
Generates low-latency, instruction-driven English and Chinese speech for voice cloning, voice design, and directed performances; supports real-time streaming, reference-free voice creation, and reference-guided cloning. Open-weight PyTorch model released under a research/non-commercial license with GPU recommendations.
Turns past discovery traces into replayable simulators so alternative exploration policies can be evaluated offline ('dreaming'), enabling fast, low-cost meta-level policy improvement for agent-driven discovery across coding, optimization, and GPU-kernel tasks.
Compresses a 27B-class multimodal model into end-to-end ternary weights to run 27B reasoning on-device: 5.9–8.6 GB deployed footprint, 262K-token context, ~98.2% of FP16 benchmark performance; ships MLX and GGUF packs and runs on Apple MLX and CUDA.
A 27B-class language model packaged in GGUF with end-to-end ternary weights for on-device or single-GPU llama.cpp inference; reduces FP16 footprint to ~5.9–7.2 GB while retaining ~98% of baseline performance and supporting up to 262K tokens.
Reduces self-attention complexity to O(N log N) by using a coarse-to-fine (pyramid) Top-K block selection with LogSumExp scoring, implemented with hardware-aware Triton kernels for fused routing and scoring—aimed at long-context LMs and retrieval tasks.
Converts Turkish text to 48 kHz speech offline using a compact 5.6M DiT acoustic model plus a 3M vocoder (8.6M parameters, ~34 MB). Streams audio with very low latency (~3.86 ms first audio on an RTX 4090), supports fast batching and runs under an Apache‑2.0 license.