AIAny
Icon for item

Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus

100-hour, single-narrator Egyptian Arabic speech corpus with 15,653 aligned clips at 24 kHz for TTS and ASR fine-tuning; studio-consistent audio, machine-generated undiacritized transcripts, CC BY-NC 4.0 (research/non-commercial use).

Introduction

Why this matters

Large, studio-quality single-speaker collections in dialectal Arabic are rare. This corpus supplies ~100 hours of Egyptian Arabic narration from one consistent voice with native 24 kHz audio and aligned transcripts—enough data to fine-tune production TTS models rather than only perform few-shot cloning, and valuable for ASR adaptation to Egyptian (Masri) dialect.

What Sets It Apart
  • Single-speaker scale: ~100 hours (15,653 clips) recorded across 247 source episodes, concentrating long-form narrated utterances that support long-context prosody learning rather than conversational fragments.
  • Production consistency: studio narration on one microphone chain (consistent loudness, low noise), delivered as 24 kHz, 16-bit PCM WAV—no resampling required for modern neural vocoders.
  • Dialectal transcripts: machine-generated, normalized Egyptian Arabic (undiacritized, no punctuation), preserving colloquial orthography (e.g. عايز, بقى, ازاي) that differs from MSA corpora.
  • Rich metadata: per-clip ASR confidence, source video id, start/end offsets and a promo flag to filter sponsor segments; split assignment is disjoint by source video for robust evaluation.
  • Licensing & provenance: manifests/transcripts/segmentation released under CC BY-NC 4.0; underlying source recordings remain with original rights holders—dataset intended for research/non-commercial use.
Who It's For and Trade-offs

Great fit if you need to fine-tune TTS or adapt ASR to Egyptian Arabic using a single, consistent voice (audiobooks, IVR, conversational agents), or to study dialectal phonology and prosody with studio-quality narration. Use the provided confidence/promo fields and timestamps to build stricter training subsets.

Look elsewhere if you need multi-speaker, spontaneous conversational, noisy/telephony, or commercially licensed voice assets: transcripts are machine-generated (expect errors), text is undiacritized (models must learn vowelization), most clips are long-form (~18–19s), and the CC BY-NC 4.0 license prohibits commercial use without further permission. If a guaranteed single-speaker acoustic guarantee is required, run speaker-embedding verification and filter out outliers.

Information

  • Websitehuggingface.co
  • OrganizationsHugging Face
  • AuthorsEhab Negm
  • Published date2026/08/08

Categories

More Items

Hugging Face

Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.

Hugging Face

Provides a reproducible, deduplicated corpus of text extracted from PDFs for LLM pretraining—about 3 trillion tokens from ~475 million documents in 1733 language-script pairs. Includes OCR and text extraction pipelines, per-page language IDs, MinHash deduplication, and is released under ODC‑By 1.0.

Hugging Face

Provides OCR full text for 11.55M public-domain German newspaper pages (1638–1964) with per-page IIIF scans and ALTO XML coordinates; suited for historical NLP, language-model training, and OCR research. Pages carry explicit per-page public-domain licenses.