AIAny
AI Audio2026
Icon for item

Phonon-2

Transcribes English speech into punctuated, capitalized text — a 164 MB quantized ASR model that averages 5.21% WER across seven Open ASR Leaderboard sets. Optimized for on-device and CPU/GPU inference, with fast runtimes on Apple M5 and Docker/GPU support.

Introduction

Small speech models usually trade off accuracy for size; this one narrows that gap by matching much of its full‑precision teacher's accuracy from a download 15× smaller. The core insight is that aggressive learned quantization and a five‑level encoder allow storing weights at about 2.1 bits each while preserving leaderboard-grade accuracy and producing punctuated, capitalized transcripts suitable for dictation and offline workflows.

What Sets It Apart
  • Teacher-level accuracy at tiny size: achieves a 5.21% average WER across seven public English test sets while downloading as just 164 MB — comparable accuracy to a 2.5 GB full‑precision teacher on several surfaces. So what: you can run near‑state performance on edge devices and modest servers.
  • Extremely compact quantization: encoder weights are represented as one of five learned levels (~2.1 bits per weight). So what: storage, memory, and I/O costs drop substantially, enabling fast cold starts and smaller container images.
  • Broad runtime support and throughput: runs 174× realtime on an M5 MacBook Air (MLX), ~143× on eight Zen5 cores, and scales to very high throughput on modern NVIDIA GPUs. So what: feasible for local dictation, on‑prem transcription, and large batched cloud workloads.
  • Practical outputs and tooling: emits punctuation, capitalization and word timestamps (with --json), shipped with command‑line tooling and Docker images for CPU and CUDA. So what: ready for integration into dictation apps, local pipelines, and batch transcription jobs.
Who it's for — Fit and tradeoffs

Great fit if: you need high‑quality English ASR with minimal download/installation footprint for on‑device or low‑resource server deployment; you want deterministic, punctuated transcripts out of the box; or you need a fast, local alternative to large cloud models.

Look elsewhere if: you need multilingual coverage beyond English, absolute lowest WER regardless of model size (very large full‑precision models still lead in some sets), or you require a specific commercial licence different from CC‑BY‑4.0.

Where it fits

Positioned between tiny on‑device models and large cloud ASR: it’s a practical choice when you want most of the accuracy of a half‑gig+ teacher but must ship under a few hundred megabytes and run reliably on CPUs or Apple silicon.

How it works (brief)

Built as a quantized derivative of NVIDIA's Parakeet TDT 0.6B v3, the model uses learned five‑value quantization in the encoder and greedy decoding to produce punctuated text. Weights are released under CC‑BY‑4.0; command‑line tooling and runtime code are provided under permissive licenses for integration and deployment.

Information

  • Websitehuggingface.co
  • OrganizationsFermionResearch, NVIDIA
  • Published date2026/09/28

More Items

Hugging Face
AI Model2026

Open-weight 309B Mixture-of-Experts causal LLM with 15.5B active parameters and a native 1M-token context for coding and AI R&D. Combines Sliding-Window Attention and DeepSeek Sparse Attention (no full-attention layers), supports FP8 inference; weights under MIT license.

Hugging Face
AI Video2026

Turns a single photo into a geometry-consistent, frozen-time 360° camera orbit that returns to the exact start frame. Implemented as a LoRA for MiniMax‑H3 FL2VA — use identical first+last keyframes to produce seamless orbit clips; trained on a small human-centric square orbit dataset, so results are domain-limited.

AI Model2026

Explains how Jev turns input state into typed decisions and probabilities without generating text. Introduces parallel sampling and RLCD training, with workflow evaluations and caveats for software automation.