AIAny
Icon for item

Waxal NLP Datasets

Provides open ASR and TTS speech data for 24 Sub‑Saharan African languages to train and evaluate speech models. Includes ~1,250 hours of transcribed ASR and ~235 hours of single‑speaker TTS with train/validation/test/unlabeled splits and mixed CC-BY licenses.

Introduction

WAXAL addresses a major gap in speech resources for Sub‑Saharan African languages by releasing large, curated ASR and TTS collections suitable for model training and evaluation. The release bundles both natural, image‑prompted ASR recordings and studio‑quality single‑speaker TTS scripts, enabling a range of speech tasks from recognition to synthesis while foregrounding local partnerships and ethical considerations.

What Sets It Apart
  • Scale and breadth: roughly 1,250 hours of transcribed, natural ASR audio and ~235 hours of high‑quality TTS across 24 languages, with metadata on speaker age, gender and recording environment. This makes it one of the largest open multilingual African speech resources.
  • Dual modalities: includes both ASR (diverse, spontaneous speech; 10% of audio transcribed) and TTS (phonetically balanced scripts, single‑speaker studio recordings) designed for different downstream needs.
  • Open licensing and curation: datasets are released under CC‑BY / CC‑BY‑SA variants, collected with local partners and paid annotators; quality control and PII removal were applied during curation.
Who It's For and Tradeoffs

Great fit if you need training or benchmarking data for ASR/TTS models in low‑resource African languages, multilingual transfer experiments, or linguistic analysis of speech patterns. Look elsewhere if you require fully transcribed corpora for every sample (only ~10% of collected ASR audio is transcribed) or if you need exhaustive dialectal coverage—dialectal and socio‑linguistic variation may be underrepresented. Also note mixed licensing across providers; check per‑language license before commercial use.

Information

  • Websitehuggingface.co
  • OrganizationsGoogle Research, Makerere University, University of Ghana, Digital Umuganda, Media Trust, Loud and Clear, AIMS Senegal, Bill & Melinda Gates Foundation
  • Published date2026/01/19

More Items

Hugging Face

Aggregated, screened corpus of 55,050 normalized Indian public-information text bodies and 65,209 source records for retrieval and question-answering. Exports include deduplicated CSV/Parquet with provenance, topic labels, extraction quality flags and a private SQLite backup.

Hugging Face

Provides 3,451 hours (2,051,810 clips) of AI‑generated 48 kHz Turkish speech with transcripts, spoken forms and per‑clip voice descriptions for TTS and ASR development. Includes 2,752 designed voices and is licensed CC BY 4.0 / CC BY‑SA 4.0 (attribution to PatientDesk AI required).

Hugging Face

Provides imagined interaction segments generated by world models for RoboTwin2.0 tasks, stored as fixed-length HDF5 chunks (21 observation frames, 20 actions, rewards and episode flags). Useful for training and evaluating world-model-based policies; currently limited to the RoboTwin2.0 subdataset.