Small speech models usually trade off accuracy for size; this one narrows that gap by matching much of its full‑precision teacher's accuracy from a download 15× smaller. The core insight is that aggressive learned quantization and a five‑level encoder allow storing weights at about 2.1 bits each while preserving leaderboard-grade accuracy and producing punctuated, capitalized transcripts suitable for dictation and offline workflows.
What Sets It Apart
- Teacher-level accuracy at tiny size: achieves a 5.21% average WER across seven public English test sets while downloading as just 164 MB — comparable accuracy to a 2.5 GB full‑precision teacher on several surfaces. So what: you can run near‑state performance on edge devices and modest servers.
- Extremely compact quantization: encoder weights are represented as one of five learned levels (~2.1 bits per weight). So what: storage, memory, and I/O costs drop substantially, enabling fast cold starts and smaller container images.
- Broad runtime support and throughput: runs 174× realtime on an M5 MacBook Air (MLX), ~143× on eight Zen5 cores, and scales to very high throughput on modern NVIDIA GPUs. So what: feasible for local dictation, on‑prem transcription, and large batched cloud workloads.
- Practical outputs and tooling: emits punctuation, capitalization and word timestamps (with --json), shipped with command‑line tooling and Docker images for CPU and CUDA. So what: ready for integration into dictation apps, local pipelines, and batch transcription jobs.
Who it's for — Fit and tradeoffs
Great fit if: you need high‑quality English ASR with minimal download/installation footprint for on‑device or low‑resource server deployment; you want deterministic, punctuated transcripts out of the box; or you need a fast, local alternative to large cloud models.
Look elsewhere if: you need multilingual coverage beyond English, absolute lowest WER regardless of model size (very large full‑precision models still lead in some sets), or you require a specific commercial licence different from CC‑BY‑4.0.
Where it fits
Positioned between tiny on‑device models and large cloud ASR: it’s a practical choice when you want most of the accuracy of a half‑gig+ teacher but must ship under a few hundred megabytes and run reliably on CPUs or Apple silicon.
How it works (brief)
Built as a quantized derivative of NVIDIA's Parakeet TDT 0.6B v3, the model uses learned five‑value quantization in the encoder and greedy decoding to produce punctuated text. Weights are released under CC‑BY‑4.0; command‑line tooling and runtime code are provided under permissive licenses for integration and deployment.