AIAny
Icon for item

Dialectal Arabic Lahgtna v2

Provides 3,000+ hours (≈611K utterances) of transcribed 16 kHz multi-dialect Arabic speech across 13 dialects for ASR and spoken-dialect identification. Transcripts preserve dialectal orthography (partial diacritics); the train split is ~337 GB in parquet, so streaming is recommended.

Introduction

Most Arabic ASR quality problems stem from dialectal variation and lack of large, labeled dialectal corpora. Lahgtna v2 addresses that by providing a single, large-scale collection of per‑utterance dialect labels and authentic dialectal transcripts so models can be fine-tuned and evaluated on real spoken varieties rather than MSA-normalized text.

What Sets It Apart
  • Scale and dialect coverage: ~3,000+ hours (≈611K utterances) across 13 Arabic dialects, enabling cross-dialect pretraining and fine‑tuning at a scale many public Arabic corpora do not reach. This makes it possible to both improve ASR robustness and train spoken-dialect identification heads from the same data.
  • Authentic dialect orthography: transcripts keep dialectal spellings and often include partial diacritics to reflect pronunciation, rather than converting to Modern Standard Arabic—so models learn real surface forms used by speakers.
  • Practical data layout: 16 kHz mono audio with a small test split held out for evaluation; train split is distributed as large parquet files (~337 GB), so streaming loading is recommended unless you need a full local copy.
  • Task-ready: packaged for ASR, spoken dialect identification, and downstream TTS/data-pipeline use; commonly used to fine-tune CTC/encoder–decoder ASR variants such as Whisper-like and OmniASR models.
Who It's For and Tradeoffs

Great fit if you need large amounts of real, labeled dialectal speech to fine-tune ASR or to build a dialect classifier; it reduces the gap between lab evaluations and in-the-wild dialectal performance. Look elsewhere if you need a small, balanced dataset per-dialect (Lahgtna v2’s dialect distribution is not uniform) or strictly standardized orthography—minor orthographic inconsistency exists and intra-dialect regional sublabels are not provided. Also plan for storage and I/O: the train split is large, so use streaming workflows or cluster storage for experiments.

Where It Fits

Use Lahgtna v2 when scale and authentic dialectal transcriptions matter (e.g., adapting large ASR models or training dialect-aware decoders). For low-resource per-dialect benchmarking you might prefer smaller, balanced datasets; for non-dialectal Arabic ASR research, MSA-focused corpora remain complementary.

Information

  • Websitehuggingface.co
  • Organizationsoddadmix, Hugging Face
  • Published date2026/03/29

Categories

More Items

Transforms source code into verifiable, reusable agent skills by extracting atomic operations, workflows, and recurring patterns and validating them via source-body-blind reconstruction. Produces CodeSkillBank (1,006,822 accepted records from 19,769 GitHub repos) and yields ~11.7% average downstream improvement.

Hugging Face

Synthetic English-language dataset of 100k+ B2B SaaS sales conversations with turn-by-turn conversion outcomes, engagement and sales-effectiveness metrics, and 3072-dim embeddings for training conversion-prediction and RL-based conversation models.

Hugging Face

Snapshot delivery of arXiv metadata, submission files and rendered documents in multiple Parquet configs (metadata, paper_text, latex, source, pdf, ps). Includes ~3.15M papers, full-text TeX assemblies and indexes to fetch large assets for training, retrieval and analysis.