Converts text into natural-sounding speech locally using compact ONNX TTS assets. Optimized for CPU/edge inference (~99M params) with support for 31 languages, expression tags (e.g., <laugh>), and improved stability versus Supertonic 2 — suitable for on-device multilingual TTS.
Generates high-quality Japanese speech from text with zero-shot voice cloning and emoji-based style controls; uses a flow-matching diffusion transformer over DACVAE continuous latents, includes a duration predictor and integrated SilentCipher watermarking. Japanese-only.
Multilingual streaming ASR that transcribes 40 language-locales using a cache-aware FastConformer‑RNNT architecture. Supports language-ID prompting (or auto-detect), punctuation/capitalization, and configurable chunk sizes to trade latency vs. accuracy for production transcription and streaming voice agents.
Generates music, sound effects, and general audio from text prompts using a medium-size Stable Audio 3 diffusion model — a balance of generation quality and inference cost suitable for prototyping, demo assets, and creative sound design workflows.
Provides a large-scale ASR corpus organized by normalized acoustic subsets for robustness training and evaluation. About 645,925 examples across 54 acoustic conditions (noise, echo, far-field, recording distortions) with many distortion/dropout/noise Parquet splits. Distributed as split Parquet files; license not specified on the dataset page.
Converts long-form multi-speaker audio/video into a compact, speaker-aware transcript with timestamps and anonymous speaker labels in one pass. Combines ASR and diarization in a single model, supports custom prompts/hotwords, and targets meetings, podcasts, interviews and long recordings.
Generates conversational speech and voice continuation from text and optional audio context, outputting Mimi audio codes. Built on a Sesame-style CSM with an 8B Llama-like backbone plus a smaller autoregressive audio decoder. Suited for local TTS inference and voice-cloning workflows.
A 12B unified, encoder-free multimodal model that directly ingests text, images and audio and returns text; supports very long contexts (up to 256K tokens), native function-calling/thinking modes, and small-model deployment for local or on-device use.
Large streaming-audio dataset for training and evaluating audio-LLMs and audio agents. About 2.28M clips grouped into multi-turn “streams” across six task subsets (ASR, speech translation, audio understanding, voice chat, proactive response, environment-aware); audio shipped as tar shards.
Generates multilingual text-to-speech with zero-shot voice cloning, token-level duration control, and inline pause markers. v1.5 improves multilingual fidelity (with language tags), cloning stability, and long-reference handling—suitable for research and production TTS pipelines.
Provides ~100 hours of expert-annotated, multi-channel Chinese conversational speech with per-segment timestamps, speaker IDs and paralinguistic labels for turn-taking, overlap/interruption detection and full‑duplex dialogue research. Licensed for academic/non-commercial use (CC BY‑NC 4.0).
Zero-shot TTS for expressive long-form monologue and multi-speaker dialogue, designed to preserve acoustic consistency, conversational coherence, and affective continuity. Trained on SwanData-Speech and using a 25 Hz VAE, pause-aware text conditioning, and a flow-matching DiT with DiffusionNFT fine-tuning.