AIAny
AI Audio2026
Icon for item

Whistle

On-device speech-to-text for short clips (up to 30s) in seven languages, yielding transcripts, word-level timestamps and per-frame speech embeddings in a single 16.9 MB model. Runs on the Needle CPU engine with 2–4 bit quantization, supports keyword biasing and returns empty transcripts for silence.

Introduction

Why this matters

Mobile, embedded and privacy-sensitive applications increasingly need accurate ASR without cloud round-trips. Whistle compresses transcription, word-level alignment and per-frame speech embeddings into a single 16.9 MB deployable file that runs on the same Needle CPU runtime — trading model size and latency against the full accuracy of large CPU/GPU models.

Key Capabilities
  • On-device transcription: 16 kHz mono input, up to 30 seconds per call, automatic language detection across English, German, French, Spanish, Italian, Dutch and Polish. Silence and steady noise return an empty transcript rather than invented text — useful for trigger/voice-activity use cases.
  • Word timestamps and confidences: every word includes start, end and probability aligned from the decoder attention, so apps can highlight, seek or cut on words without external forced-alignment.
  • Speech embeddings: encoder outputs one row per ~80 ms frame for retrieval, speaker/match tasks or downstream classification without decoding the transcript.
  • Deployment and performance: a single 16.9 MB .cact file, laddered decoder (selectable depth at load time), 2–4 bit quantization and SIMD kernels via the Needle engine. Designed for low latency (time-to-first-token in single-digit ms on modern mobile CPUs) and tiny-device memory/CPU budgets.
Who it fits — and tradeoffs

Great fit if you need offline, low-latency ASR and embeddings on phones, wearables, robots or air-gapped devices and want a tiny, single-file runtime integration that reuses Needle. It’s also suitable when word-level timestamps and keyword biasing are required locally.

Look elsewhere if you need continuous/long-form transcription without chunking (clips >30s), the absolute top accuracy on every academic benchmark across dozens of languages, or GPU-backed heavy-duty models for large-vocabulary transcription and multilingual coverage beyond the seven supported languages.

Where it sits in the stack

Whistle is positioned as a tiny on-device ASR: significantly smaller and quicker-to-first-token than Whisper base while accepting a tradeoff in model capacity. It’s most useful as an edge/inference component paired with server-side or larger models for post-processing, re-ranking or cross-checking when higher recall is required.

Information

  • Websitehuggingface.co
  • OrganizationsCactus Compute, Inc.
  • AuthorsJakub Mroz, Henry Ndubuaku, Karen Mosoyan, Noah Cylich, Satyajit Kumar, Parkirat Sandhu, Roman Shemet, Justin H. Lee
  • Published date2026/09/30

Categories

More Items

Hugging Face
AI Audio2026

Transcribes English speech into punctuated, capitalized text — a 164 MB quantized ASR model that averages 5.21% WER across seven Open ASR Leaderboard sets. Optimized for on-device and CPU/GPU inference, with fast runtimes on Apple M5 and Docker/GPU support.

Hugging Face

Provides 3,451 hours (2,051,810 clips) of AI‑generated 48 kHz Turkish speech with transcripts, spoken forms and per‑clip voice descriptions for TTS and ASR development. Includes 2,752 designed voices and is licensed CC BY 4.0 / CC BY‑SA 4.0 (attribution to PatientDesk AI required).

Hugging Face
AI Audio2026

Performs speaker diarization (who spoke when) for live and recorded audio using an open-weight, 100M-parameter streaming-capable model that supports up to eight anonymous speaker channels, overlapping speech, chunked processing, and configurable latency for ASR integration.