Why this matters now Web-scale speech research increasingly needs high-fidelity and spatial audio rather than downsampled mono clips. YODAS v3 supplies a uniquely large public collection of 48kHz multi-channel web audio with weak supervision (captions and translations) so practitioners can train or evaluate models that rely on true stereo/high-bandwidth signals at scale.
What Sets It Apart
- Scale + fidelity: roughly 1.1 million hours of audio saved in original 48kHz multi-channel OPUS format, enabling experiments that require real high-frequency content and spatial cues. Over 92% of recordings have effective sampling rates ≥32kHz and ~67% reach ~44kHz. 99.8% of containers are two-channel and >70% contain two distinct channels.
- Language balance and coverage: metadata spans 100+ languages (147 languages reported in publication-level summaries), with dozens of languages having thousands of hours each (22 languages >10k hours; 73 >5k hours). This balance helps medium- and low-resource language work that standard crawls miss.
- Weak supervision at scale: roughly 75% of items have transcripts and many non-English items include English translations with sentence-level timestamps; ~95% of transcripts include word-level timestamps where available. Metadata also reports estimated bandwidth, channel counts, and audio bitrates to help filter by quality.
- Web-ready layout: distributed as sharded WebDataset audio tarballs plus matched parquet metadata so large-scale streaming pipelines (webdataset, parquet, Dask, polars) can pair audio and labels efficiently.
Who should use it — and when to look elsewhere
Great fit if you need large-scale training or evaluation data for ASR/TTS, self-supervised audio representation learning, or research that benefits from genuine multi-channel/high-bandwidth web audio (e.g., spatial audio, codec research, realistic noise/mixture modeling). The dataset’s size and metadata let you subsample by language, effective sampling rate, channel distinctness, or transcript availability. Look elsewhere if you need fully curated, human-verified transcripts across the whole corpus, guaranteed speaker-level annotations, or a small, curated benchmark set with strict licensing constraints — YODAS v3 is web-crawled and weakly labeled, so per-sample quality varies and transcripts are often auto-generated. Also plan storage and bandwidth carefully — the full dataset is very large and sharded for streaming rather than single-download convenience.
Quick practical notes
- Licensing: released under CC BY 3.0; follow citation and redistribution guidance in the repository.
- Structure: language-organized directories with audio tar shards and corresponding parquet metadata; each audio file has a unique hex id and timestamped segment/word annotations when available.
- Tradeoffs: the dataset’s strength is scale and fidelity; its weakness is heterogeneity in transcript quality and web-origin noise. Expect to filter and validate subsets for high-quality supervised training.