Why this matters
Mobile, embedded and privacy-sensitive applications increasingly need accurate ASR without cloud round-trips. Whistle compresses transcription, word-level alignment and per-frame speech embeddings into a single 16.9 MB deployable file that runs on the same Needle CPU runtime — trading model size and latency against the full accuracy of large CPU/GPU models.
Key Capabilities
- On-device transcription: 16 kHz mono input, up to 30 seconds per call, automatic language detection across English, German, French, Spanish, Italian, Dutch and Polish. Silence and steady noise return an empty transcript rather than invented text — useful for trigger/voice-activity use cases.
- Word timestamps and confidences: every word includes start, end and probability aligned from the decoder attention, so apps can highlight, seek or cut on words without external forced-alignment.
- Speech embeddings: encoder outputs one row per ~80 ms frame for retrieval, speaker/match tasks or downstream classification without decoding the transcript.
- Deployment and performance: a single 16.9 MB .cact file, laddered decoder (selectable depth at load time), 2–4 bit quantization and SIMD kernels via the Needle engine. Designed for low latency (time-to-first-token in single-digit ms on modern mobile CPUs) and tiny-device memory/CPU budgets.
Who it fits — and tradeoffs
Great fit if you need offline, low-latency ASR and embeddings on phones, wearables, robots or air-gapped devices and want a tiny, single-file runtime integration that reuses Needle. It’s also suitable when word-level timestamps and keyword biasing are required locally.
Look elsewhere if you need continuous/long-form transcription without chunking (clips >30s), the absolute top accuracy on every academic benchmark across dozens of languages, or GPU-backed heavy-duty models for large-vocabulary transcription and multilingual coverage beyond the seven supported languages.
Where it sits in the stack
Whistle is positioned as a tiny on-device ASR: significantly smaller and quicker-to-first-token than Whisper base while accepting a tradeoff in model capacity. It’s most useful as an edge/inference component paired with server-side or larger models for post-processing, re-ranking or cross-checking when higher recall is required.