AIAny
Icon for item

Dataset.ET Amharic Speech

Provides 22.7 hours of read Amharic speech (7,405 clips, 320 speakers) for ASR, collected via a crowdsourced Telegram bot and peer-validated; speaker- and prompt-disjoint train/validation/test splits, 16 kHz audio under CC BY 4.0.

Introduction

Most open Amharic speech resources are tiny or noisy; this release supplies a carefully screened 22.7‑hour corpus of read Amharic designed for reliable ASR development and evaluation. The data were crowdsourced, peer-validated, acoustically screened, and split so no speaker or prompt appears in more than one partition, reducing inflated evaluation scores.

What Sets It Apart
  • Quantified, curated scale: 7,405 clips (22.706 hours) from 320 volunteer contributors with 7,145 distinct prompts — large enough for fine-tuning and robust validation but still compact compared with major languages. So what? You get a middle‑scale, high‑signal Amharic resource that fits typical ASR fine-tuning and benchmark workflows.
  • Community-driven collection and peer validation: recordings were submitted via a Telegram bot and accepted only after unanimous peer approval, then acoustically screened. So what? The pipeline emphasises real-user contributions with community quality checks rather than single-annotator transcriptions.
  • Evaluation-friendly splits and metadata: speaker- and prompt-disjoint train/validation/test splits, per-clip speech timing, LUFS loudness, gender/age/region with k-anonymity protections, and salted pseudonymous IDs. So what? You can run honest generalisation evaluations and reproduce pre-processing choices deterministically.
  • Permissive audio licence with constrained text reuse: audio is CC BY 4.0; prompt text is reproduced as transcripts but may carry third-party rights, so redistributing text separately requires separate review. So what? Models trained on the audio can be shared under CC BY, but check prompt-text licensing for separate text redistribution.
Who It's For and Tradeoffs

Great fit if you need a well‑screened, read-speech Amharic corpus for ASR training, fine-tuning, or benchmarking, and you value speaker-disjoint evaluation and reproducible per-clip metadata. Also useful for demographic analysis within provided privacy constraints.

Look elsewhere if you require spontaneous conversational speech, broad demographic representativeness, or studio-quality recordings: contributors skew young and urban (majority 18–24, ~47.5% Addis Ababa), recordings originate as Telegram Opus messages (codec/device artefacts), and prompts bias formal vocabulary. Also note the dataset is not loudness-normalised — use the provided LUFS values if you need deterministic normalization.

Information

  • Websitehuggingface.co
  • OrganizationsSnapwre Technologies PLC, Dataset.ET
  • Published date2026/08/25

Categories

More Items

Hugging Face

Evaluation dataset for comparing eight text-to-image models using 8,000 generated images with source prompts and per-image scores for aesthetic quality, emotional resonance, and content integrity. Includes model labels, shared prompts, GPT-5.6 Sol automated scores, embedded images in Parquet, and an Apache-2.0 license.

Hugging Face

Provides a 1 trillion-token multimodal interleaved dataset (HTML subset updated as data_v1_1 with 742B HTML tokens) and 3.4B images drawn from HTML/PDF/ArXiv sources for multimodal pretraining; released under CC-BY-4.0 with safety and deduplication guidance.

Hugging Face

Provides 369 Harbor sandbox tasks ported from OpenAI's openai/math: each task is a Lean theorem with missing `sorry` proofs that an agent must complete, graded by a strict Comparator exact-match verifier. Includes task definitions, generator, and manifest for RL/code-agent evaluation.