AIAny
AI Audio2025
Icon for item

OpenSuperWhisper

Provides real-time, local audio recording and transcription on macOS using Whisper and Parakeet engines, with global hotkeys and hold-to-record behavior. Includes model download, microphone selection, drag-and-drop file transcription, multilingual auto-detection and Asian-language autocorrect; Apple Silicon only.

Introduction

Why this matters Most desktop speech tools either rely on cloud APIs or provide limited, single-engine local transcribe. OpenSuperWhisper targets macOS users who want local, immediate transcription control: record from any mic, get near-real-time transcripts, and paste or save results without switching apps.

What Sets It Apart
  • Local Whisper + Parakeet engines: lets you run transcription locally with downloadable models, so you can avoid cloud dependency and reduce latency. This means better privacy and offline use for sensitive audio.
  • Global hotkeys & hold-to-record: record from any application using a single modifier or key combination, and release to stop — so capturing quick voice notes or dictation becomes frictionless.
  • Flexible I/O and mic management: supports built-in, external, Bluetooth and iPhone mics, plus drag-and-drop file queueing and improved handling of common audio containers; useful for both live dictation and batch file transcription.
  • Multilingual support with Asian-language autocorrect: auto-detects languages and applies autocorrect for Japanese/Chinese/Korean, improving transcription quality for those languages.
Who it's for and tradeoffs

Great fit if you need local, privacy-minded speech-to-text on Apple Silicon Macs, want quick global-hotkey dictation, or need mic selection and file-queue transcription without cloud uploads. Look elsewhere if you require Intel macOS support, enterprise deployment features, cloud-model accuracy tradeoffs, or built-in streaming transcription and advanced keyword boosting (these are noted as TODOs). The app is open-source (MIT) and expects users to manage model downloads and storage locally.

Information

  • Websitegithub.com
  • AuthorsStarmel
  • Published date2025/02/06

Categories

More Items

Hugging Face
AI Audio2026

Transcribes English speech into punctuated, capitalized text — a 164 MB quantized ASR model that averages 5.21% WER across seven Open ASR Leaderboard sets. Optimized for on-device and CPU/GPU inference, with fast runtimes on Apple M5 and Docker/GPU support.

Hugging Face

Provides 3,451 hours (2,051,810 clips) of AI‑generated 48 kHz Turkish speech with transcripts, spoken forms and per‑clip voice descriptions for TTS and ASR development. Includes 2,752 designed voices and is licensed CC BY 4.0 / CC BY‑SA 4.0 (attribution to PatientDesk AI required).

Hugging Face
AI Audio2026

Performs speaker diarization (who spoke when) for live and recorded audio using an open-weight, 100M-parameter streaming-capable model that supports up to eight anonymous speaker channels, overlapping speech, chunked processing, and configurable latency for ASR integration.