AIAny
Icon for item

YODAS v3

Provides over 1.1M hours of high-bandwidth, multichannel multilingual speech with segment- and word-level timestamps, English translations, and per-file metadata for ASR, TTS and audio-representation research. Preserves original 48kHz multichannel OPUS audio and is released under CC BY 3.0.

Introduction

Why this matters now Web-scale speech research increasingly needs high-fidelity and spatial audio rather than downsampled mono clips. YODAS v3 supplies a uniquely large public collection of 48kHz multi-channel web audio with weak supervision (captions and translations) so practitioners can train or evaluate models that rely on true stereo/high-bandwidth signals at scale.

What Sets It Apart
  • Scale + fidelity: roughly 1.1 million hours of audio saved in original 48kHz multi-channel OPUS format, enabling experiments that require real high-frequency content and spatial cues. Over 92% of recordings have effective sampling rates ≥32kHz and ~67% reach ~44kHz. 99.8% of containers are two-channel and >70% contain two distinct channels.
  • Language balance and coverage: metadata spans 100+ languages (147 languages reported in publication-level summaries), with dozens of languages having thousands of hours each (22 languages >10k hours; 73 >5k hours). This balance helps medium- and low-resource language work that standard crawls miss.
  • Weak supervision at scale: roughly 75% of items have transcripts and many non-English items include English translations with sentence-level timestamps; ~95% of transcripts include word-level timestamps where available. Metadata also reports estimated bandwidth, channel counts, and audio bitrates to help filter by quality.
  • Web-ready layout: distributed as sharded WebDataset audio tarballs plus matched parquet metadata so large-scale streaming pipelines (webdataset, parquet, Dask, polars) can pair audio and labels efficiently.
Who should use it — and when to look elsewhere

Great fit if you need large-scale training or evaluation data for ASR/TTS, self-supervised audio representation learning, or research that benefits from genuine multi-channel/high-bandwidth web audio (e.g., spatial audio, codec research, realistic noise/mixture modeling). The dataset’s size and metadata let you subsample by language, effective sampling rate, channel distinctness, or transcript availability. Look elsewhere if you need fully curated, human-verified transcripts across the whole corpus, guaranteed speaker-level annotations, or a small, curated benchmark set with strict licensing constraints — YODAS v3 is web-crawled and weakly labeled, so per-sample quality varies and transcripts are often auto-generated. Also plan storage and bandwidth carefully — the full dataset is very large and sharded for streaming rather than single-download convenience.

Quick practical notes
  • Licensing: released under CC BY 3.0; follow citation and redistribution guidance in the repository.
  • Structure: language-organized directories with audio tar shards and corresponding parquet metadata; each audio file has a unique hex id and timestamped segment/word annotations when available.
  • Tradeoffs: the dataset’s strength is scale and fidelity; its weakness is heterogeneity in transcript quality and web-origin noise. Expect to filter and validate subsets for high-quality supervised training.

Information

  • Websitehuggingface.co
  • AuthorsWilliam Chen, Shinnosuke Takamichi, Sayaka Shiota, Satoru Fukayama, Samuele Cornell, Shinji Watanabe
  • Published date2026/06/19

Categories

More Items

Hugging Face

Synthetic, clinician-verified ChatML dataset of 2,194 doctor–patient encounters covering 2,194 unique human diseases; each JSONL record includes 20 structured fields, verified PubMed references, realistic vitals/labs, and is intended for RAG and model fine-tuning (not medical advice).

Hugging Face

Provides 2,000 synthetic multiple-choice items designed for continuation log-likelihood scoring to evaluate small language models' Theory of Mind (social-cognitive) abilities; 40 constructs, balanced answer positions, and easy/medium difficulty.

Hugging Face

Provides 5.5K+ self-contained data-analysis RL tasks: each row bundles a real tabular dataset, a question, and a deterministically-gradable gold answer. Verified from jupyter-agent notebooks; splits for training, held-out testing, and quick eval; intended for prompting, fine-tuning, and agent RL.