Generates real-time, infinite-length portrait video from one reference image on a 12GB GPU. Combines implicit facial signals and 3D keypoints with step-distilled diffusion and autoregressive micro-chunk streaming for low-latency live use.
Converts images (and other conditions) into high-fidelity, fully textured 3D assets using a 4B-parameter generative model and a field‑free sparse voxel format (O‑Voxel). Handles arbitrary topology, PBR materials, and near real-time mesh/voxel conversions; requires Linux and an NVIDIA GPU with >=24GB memory.
Provides a DiT-based audio–video foundation model plus an official Python inference and LoRA trainer. Ships multiple production-ready pipelines (text/image/audio→video), checkpoints, and performance optimizations (FP8, distilled pipelines) for high-fidelity synchronized audio–video generation.
Real‑time full‑duplex speech‑to‑speech system that controls conversational role via text prompts and voice timbre via audio-conditioned embeddings. Built on Moshi; optimized for low-latency, persona-consistent spoken interactions.
Generates low-latency, streaming text-to-speech entirely on CPUs (no GPU or cloud API required), using an ~100M-parameter model with voice cloning and multilingual support. Optimized for low resource use (2 CPU cores, ~200ms to first audio chunk) — suited for local, privacy-sensitive, or embedded TTS.
Provides a conditional memory module that performs O(1) N‑gram lookups and fuses static embeddings into transformer hidden states — enables offloading large embedding tables to host memory with minimal inference overhead.
Generates controllable multilingual speech from text with nine predefined timbres and custom-voice control; supports voice design, quick voice cloning and low-latency streaming (first audio packet after a single character), suitable for real-time TTS and voice-design workflows.
Local-first voice cloning studio that runs on your machine to clone voices, generate speech in 23 languages, apply audio effects, and compose multi-voice projects. Includes five switchable TTS engines, a REST API, and native GPU/MLX support for privacy-sensitive offline workflows.
Enables research-grade character animation with neural networks in a single NumPy/PyTorch environment — train models, run inference, and visualize results without leaving Python. Includes ECS-style architecture, mocap import (GLB/FBX/BVH), built-in renderer, and headless/standalone modes for rapid prototyping.
Generates high‑fidelity, expressive speech and environmental sounds from text. The MOSS‑TTS Family provides specialized models for long‑form TTS, multi‑speaker dialogue, voice design and realtime streaming, plus torch‑free inference paths (llama.cpp / ONNX) and Hugging Face releases.
A challenge repository for training the best language model that fits inside a 16,000,000‑byte (16MB) submission artifact; provides baseline training code, FineWeb bpb evaluation, a public leaderboard, and compute-grant instructions for short 8×H100 runs.
Provides a one-command CLI to fine-tune and post-train LLMs, with layer streaming that lets an 8B model be fine-tuned on a 4 GB laptop GPU. Auto-configures quantization, LoRA adapters, batching and evaluation gates, and supports export and serving workflows.