Turns web reading into an in-context language-learning experience by injecting context-aware translations, explanations, subtitle translation, and TTS directly into the browser. Supports selection translation, batch requests and configurable AI providers to balance cost and quality.
Provides a 10,000-hour Sichuanese (Chuan-Yu) speech corpus with rich annotations (timestamps, speaker age/gender/emotion, SNR, DNSMOS) and unified metadata for ASR and TTS research; includes metadata.jsonl, evaluation benchmarks, and an LLM-assisted transcription pipeline.
Build and self-host production voice agents with a drag-and-drop workflow builder, real-time telephony integration, and pluggable LLM/STT/TTS backends. Docker-first with an optional managed cloud offering for teams that want faster onboarding.
Delivers multilingual, on-device text-to-speech via ONNX Runtime with prebuilt ONNX assets and cross-platform SDKs (Python, Node, mobile); targets low-latency, privacy-preserving TTS with ready demos and 31-language support in v3.
Generates low-latency, streaming text-to-speech entirely on CPUs (no GPU or cloud API required), using an ~100M-parameter model with voice cloning and multilingual support. Optimized for low resource use (2 CPU cores, ~200ms to first audio chunk) — suited for local, privacy-sensitive, or embedded TTS.
Orchestrates low-latency, multi-stage pipelines for omni and multimodal models by running each stage with its own scheduler and using zero-copy shared memory for tensor transfer. Emphasizes per-stage bottleneck tuning and OpenAI-compatible streaming endpoints, suitable for TTS and multimodal serving.
Provides open ASR and TTS speech data for 24 Sub‑Saharan African languages to train and evaluate speech models. Includes ~1,250 hours of transcribed ASR and ~235 hours of single‑speaker TTS with train/validation/test/unlabeled splits and mixed CC-BY licenses.
Extracts derived keys from running WeChat 4.x processes to decrypt SQLCipher 4 databases and .dat media files, and provides a real-time message monitor with a Web UI. Cross-platform (Windows/Linux/macOS) but requires process-memory or local-data access and is intended for decrypting your own WeChat data only.
An instruction‑tuned Gemma 4 E4B multimodal model on Hugging Face that accepts text, images and audio and generates text; notable for 128K long context support, built-in thinking mode, and an on‑device‑friendly E4B architecture under an Apache‑2.0 license.
Provides a unified 615k-hour English speech corpus for TTS training, aggregating 11 public datasets and web-sourced recordings into 16 kHz Opus WebDataset shards. Includes a quality-filtered core subset (510.1k hours), metadata splits, and mixed licenses across sources.
Provides 3,000+ hours (≈611K utterances) of transcribed 16 kHz multi-dialect Arabic speech across 13 dialects for ASR and spoken-dialect identification. Transcripts preserve dialectal orthography (partial diacritics); the train split is ~337 GB in parquet, so streaming is recommended.
Converts text to natural-sounding speech across 600+ languages in a zero-shot way, with short-reference voice cloning and fine-grained voice-design controls; uses a diffusion language-model-style architecture to balance quality and very low inference latency.