Most open Amharic speech resources are tiny or noisy; this release supplies a carefully screened 22.7‑hour corpus of read Amharic designed for reliable ASR development and evaluation. The data were crowdsourced, peer-validated, acoustically screened, and split so no speaker or prompt appears in more than one partition, reducing inflated evaluation scores.
What Sets It Apart
- Quantified, curated scale: 7,405 clips (22.706 hours) from 320 volunteer contributors with 7,145 distinct prompts — large enough for fine-tuning and robust validation but still compact compared with major languages. So what? You get a middle‑scale, high‑signal Amharic resource that fits typical ASR fine-tuning and benchmark workflows.
- Community-driven collection and peer validation: recordings were submitted via a Telegram bot and accepted only after unanimous peer approval, then acoustically screened. So what? The pipeline emphasises real-user contributions with community quality checks rather than single-annotator transcriptions.
- Evaluation-friendly splits and metadata: speaker- and prompt-disjoint train/validation/test splits, per-clip speech timing, LUFS loudness, gender/age/region with k-anonymity protections, and salted pseudonymous IDs. So what? You can run honest generalisation evaluations and reproduce pre-processing choices deterministically.
- Permissive audio licence with constrained text reuse: audio is CC BY 4.0; prompt text is reproduced as transcripts but may carry third-party rights, so redistributing text separately requires separate review. So what? Models trained on the audio can be shared under CC BY, but check prompt-text licensing for separate text redistribution.
Who It's For and Tradeoffs
Great fit if you need a well‑screened, read-speech Amharic corpus for ASR training, fine-tuning, or benchmarking, and you value speaker-disjoint evaluation and reproducible per-clip metadata. Also useful for demographic analysis within provided privacy constraints.
Look elsewhere if you require spontaneous conversational speech, broad demographic representativeness, or studio-quality recordings: contributors skew young and urban (majority 18–24, ~47.5% Addis Ababa), recordings originate as Telegram Opus messages (codec/device artefacts), and prompts bias formal vocabulary. Also note the dataset is not loudness-normalised — use the provided LUFS values if you need deterministic normalization.