Fine-tunes and deploys 600+ LLMs and 400+ multimodal models in one framework, with SFT, pretraining, RLHF (DPO, PPO, GRPO), and lightweight methods like LoRA and QLoRA. Adds Megatron parallelism, vLLM/SGLang/LMDeploy inference, and a training web UI.
Runs Stable Diffusion XL behind a Midjourney-style interface, hiding samplers, model swaps, and LoRA weights. A built-in GPT2 expander rewrites prompts into richer styling, and it works fully offline on as little as 4GB of Nvidia VRAM.
Applies deep learning workflows to geospatial data, covering imagery search, dataset preparation, model training, inference, visualization, and QGIS integration for remote sensing.
Converts microphone or streamed audio to text with sub-second latency, pairing WebRTC/Silero voice-activity detection and wake-word activation with swappable local backends — faster-whisper by default, plus whisper.cpp, Moonshine, and sherpa-onnx.
Generates expressive multilingual speech from text, with sub-word control over prosody and emotion via inline tags like [whisper] or [angry]. Handles multi-speaker, multi-turn dialogue; the weights ship under a research-only license.
Terminal CLI for on-device Whisper ASR using Hugging Face Transformers + Optimum, with optional Flash Attention 2, batching, and diarization support — focused on high-throughput transcription on NVIDIA GPUs and Apple Silicon (mps).
Provides a scalable physics-and-rendering simulation interface for robotics and embodied-AI research — unified multi-physics solvers, the Nyx renderer, and the Quadrants compiler. Runs from laptop to datacenter GPUs; suited for sensor-rich data generation and RL/robotics prototyping.
GPU-native physics engine unifying rigid-body, fluid, cloth, and deformable solvers in one Python framework for robotics and embodied-AI research. Built by a 20+ lab collaboration, now backed by Genesis AI, with generative tools to author 4D scenes.
Multilingual automatic speech recognition and speech-translation model that transcribes and translates audio. Trained on a mix of weakly labeled and pseudo-labeled data (1M + 4M hours), uses 128 Mel bins and adds a Cantonese token, and supports timestamps and long-form chunking for offline ASR and translation.
Performs speaker diarization (who spoke when) with pyannote-audio: combines voice-activity detection, speaker-change and overlapped-speech detection to produce time-stamped speaker segments; compatible with Hugging Face Endpoints and ASR pipelines.
Provides a NumPy-like array framework for building and training ML on Apple silicon, with Python, C/C++, and Swift APIs plus PyTorch-style higher-level modules. Features lazy evaluation, composable AD/vectorization, and a unified-memory multi-device model so arrays can be used on CPU and GPU without explicit copies.
A selective State Space Model architecture and PyTorch implementation for linear-time sequence modeling. Hardware-aware, designed for information-dense tasks (e.g. language modeling), with pretrained weights on Hugging Face; requires CUDA-enabled PyTorch.