Converts text to natural-sounding speech across 600+ languages in a zero-shot way, with short-reference voice cloning and fine-grained voice-design controls; uses a diffusion language-model-style architecture to balance quality and very low inference latency.
Runs local speech-to-text inference for a wide range of ASR model families using GGUF models on the ggml runtime. Supports Metal, Vulkan, and CUDA GPU backends plus a tinyBLAS-accelerated CPU path, prebuilt GGUFs on Hugging Face, and a quantization tool.
Desktop app for local voice cloning, real-time dictation, and end-to-end video dubbing using zero-shot TTS across 600+ languages; features multi-engine TTS/ASR, speaker diarization, vocal isolation, batch pipelines, and invisible audio watermarking — all run fully offline.
Generates expressive, prompt-driven text-to-speech audio with optional 10-second voice cloning; prompts control speaker identity, emotion, pauses and nonverbal sounds. An IC‑LoRA fine-tune of LTX‑2.3 that applies an imperceptible Resemble Perth watermark.
Benchmarks ASR on long-form English call-center conversations with wide accent coverage; 128.6 hours across 14 accent groups and 16 service domains, designed for segmentation-sensitive evaluation and intended for evaluation/analysis (CC BY‑SA 4.0).
Provides 100 English–Khasi parallel sentence pairs with aligned studio-quality WAV recordings for ASR, TTS and translation evaluation; curated by Medharvix as a restricted public sample—full corpus available by request.
Multilingual 2B speech–language model for ASR and bidirectional speech translation (EN, FR, DE, ES, PT, JA), providing punctuation/truecasing, keyword biasing, and a dual-head CTC encoder to boost transcription accuracy.
Converts text into natural-sounding speech locally using compact ONNX TTS assets. Optimized for CPU/edge inference (~99M params) with support for 31 languages, expression tags (e.g., <laugh>), and improved stability versus Supertonic 2 — suitable for on-device multilingual TTS.
Generates high-quality Japanese speech from text with zero-shot voice cloning and emoji-based style controls; uses a flow-matching diffusion transformer over DACVAE continuous latents, includes a duration predictor and integrated SilentCipher watermarking. Japanese-only.
Multilingual streaming ASR that transcribes 40 language-locales using a cache-aware FastConformer‑RNNT architecture. Supports language-ID prompting (or auto-detect), punctuation/capitalization, and configurable chunk sizes to trade latency vs. accuracy for production transcription and streaming voice agents.
Generates music, sound effects, and general audio from text prompts using a medium-size Stable Audio 3 diffusion model — a balance of generation quality and inference cost suitable for prototyping, demo assets, and creative sound design workflows.
Provides a large-scale ASR corpus organized by normalized acoustic subsets for robustness training and evaluation. About 645,925 examples across 54 acoustic conditions (noise, echo, far-field, recording distortions) with many distortion/dropout/noise Parquet splits. Distributed as split Parquet files; license not specified on the dataset page.