AIAny
Icon for item

Dataset.ET Amharic Speech

Provides 22.7 hours of read Amharic speech (7,405 clips, 320 speakers) for ASR, collected via a crowdsourced Telegram bot and peer-validated; speaker- and prompt-disjoint train/validation/test splits, 16 kHz audio under CC BY 4.0.

Introduction

Most open Amharic speech resources are tiny or noisy; this release supplies a carefully screened 22.7‑hour corpus of read Amharic designed for reliable ASR development and evaluation. The data were crowdsourced, peer-validated, acoustically screened, and split so no speaker or prompt appears in more than one partition, reducing inflated evaluation scores.

What Sets It Apart
  • Quantified, curated scale: 7,405 clips (22.706 hours) from 320 volunteer contributors with 7,145 distinct prompts — large enough for fine-tuning and robust validation but still compact compared with major languages. So what? You get a middle‑scale, high‑signal Amharic resource that fits typical ASR fine-tuning and benchmark workflows.
  • Community-driven collection and peer validation: recordings were submitted via a Telegram bot and accepted only after unanimous peer approval, then acoustically screened. So what? The pipeline emphasises real-user contributions with community quality checks rather than single-annotator transcriptions.
  • Evaluation-friendly splits and metadata: speaker- and prompt-disjoint train/validation/test splits, per-clip speech timing, LUFS loudness, gender/age/region with k-anonymity protections, and salted pseudonymous IDs. So what? You can run honest generalisation evaluations and reproduce pre-processing choices deterministically.
  • Permissive audio licence with constrained text reuse: audio is CC BY 4.0; prompt text is reproduced as transcripts but may carry third-party rights, so redistributing text separately requires separate review. So what? Models trained on the audio can be shared under CC BY, but check prompt-text licensing for separate text redistribution.
Who It's For and Tradeoffs

Great fit if you need a well‑screened, read-speech Amharic corpus for ASR training, fine-tuning, or benchmarking, and you value speaker-disjoint evaluation and reproducible per-clip metadata. Also useful for demographic analysis within provided privacy constraints.

Look elsewhere if you require spontaneous conversational speech, broad demographic representativeness, or studio-quality recordings: contributors skew young and urban (majority 18–24, ~47.5% Addis Ababa), recordings originate as Telegram Opus messages (codec/device artefacts), and prompts bias formal vocabulary. Also note the dataset is not loudness-normalised — use the provided LUFS values if you need deterministic normalization.

Information

  • Websitehuggingface.co
  • OrganizationsSnapwre Technologies PLC, Dataset.ET
  • Published date2026/08/25

Categories

More Items

Evaluates AI agents' ability to complete end-to-end scientific workflows by releasing and assessing 97 tasks from a 300-task FrontierChallenge suite across chemistry, materials, life science, and electrochemistry. Finds that top agent configurations achieved only a 20.6% pass rate despite high partial scores, revealing a gap between partial progress/confident completion claims and actual complete scientific deliverables.

Hugging Face

Evaluates whether tool-using LLM agents reliably complete stateful business workflows via 507 executable agent–tool–user tasks across retail, travel, auto insurance, neobank, and IT/HR consulting. Provides browsable Parquet tables for tasks, scenarios, and agent instructions; v1.0 is intended for evaluation-only.

Hugging Face

A human‑curated corpus of AI‑generated music with MP3s, cover art, exact generation prompts and a 32‑column metadata schema; uses a 70/30 quality vs. mainstream split and a three‑level taxonomy to support fine‑grained audio‑ML, prompt‑fidelity and recommendation research.