Generates controllable multilingual speech from text with nine predefined timbres and custom-voice control; supports voice design, quick voice cloning and low-latency streaming (first audio packet after a single character), suitable for real-time TTS and voice-design workflows.
Local-first voice cloning studio that runs on your machine to clone voices, generate speech in 23 languages, apply audio effects, and compose multi-voice projects. Includes five switchable TTS engines, a REST API, and native GPU/MLX support for privacy-sensitive offline workflows.
Fetches multi-source content (webpages, YouTube, PDFs, WeChat, paywalled articles, podcasts), uploads it to Google NotebookLM, and generates outputs such as podcasts, PPTs, mind maps, or quizzes. Differentiators: automatic paywall-bypass pipeline, Claude Code Skill integration, and CLI + MCP components for WeChat and document scraping.
Provides a Spotify-like local UI for running ACE-Step 1.5 to generate full songs (including vocals), batch variations, and manage a local music library, with reference-audio styling and built-in editing/stem tools for users running the model locally.
Generates high‑fidelity, expressive speech and environmental sounds from text. The MOSS‑TTS Family provides specialized models for long‑form TTS, multi‑speaker dialogue, voice design and realtime streaming, plus torch‑free inference paths (llama.cpp / ONNX) and Hugging Face releases.
Provides multi-task long-speech evaluation data for eight speech-understanding tasks (ASR, summarization, QA, translation, emotion, speaker counting, content separation, language detection). Includes 101,822 long audio files and ~204,881 annotated examples with JSONL task splits for easy loading.
Extracts derived keys from running WeChat 4.x processes to decrypt SQLCipher 4 databases and .dat media files, and provides a real-time message monitor with a Web UI. Cross-platform (Windows/Linux/macOS) but requires process-memory or local-data access and is intended for decrypting your own WeChat data only.
An instruction‑tuned Gemma 4 E4B multimodal model on Hugging Face that accepts text, images and audio and generates text; notable for 128K long context support, built-in thinking mode, and an on‑device‑friendly E4B architecture under an Apache‑2.0 license.
Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.
Provides a unified 615k-hour English speech corpus for TTS training, aggregating 11 public datasets and web-sourced recordings into 16 kHz Opus WebDataset shards. Includes a quality-filtered core subset (510.1k hours), metadata splits, and mixed licenses across sources.
Provides 3,000+ hours (≈611K utterances) of transcribed 16 kHz multi-dialect Arabic speech across 13 dialects for ASR and spoken-dialect identification. Transcripts preserve dialectal orthography (partial diacritics); the train split is ~337 GB in parquet, so streaming is recommended.
Converts text to natural-sounding speech across 600+ languages in a zero-shot way, with short-reference voice cloning and fine-grained voice-design controls; uses a diffusion language-model-style architecture to balance quality and very low inference latency.