Why this matters Suno V5.5 made it practical to produce coherent, production‑style tracks across many niche subgenres; this dataset packages those outputs with unusually detailed prompts and dense metadata so researchers can evaluate prompt‑to‑audio fidelity, fine‑grained genre classifiers, and cross‑genre transfer without scraping noisy community uploads.
What Sets It Apart
- Curatorial 70/30 split: 70% focuses on musically sophisticated territory (neo‑soul, progressive psy, contemporary jazz) while 30% targets high‑traffic mainstream styles with a deliberate “sophistication overlay” — so what: enables transfer and robustness studies between high‑quality and commercial material.
- Prompt‑paired audio: every track ships with the exact descriptive prompt (30–80 tokens typical) plus BPM, key, mood and a taxonomic label — so what: lets you re‑render prompts in other generators or quantify prompt‑fidelity systematically.
- Three‑level taxonomy and rich metadata (100+ sub‑sub‑genres, 32 columns): so what: supports hierarchical classification, auto‑tagging benchmarks and listener‑context research using the explicit
tax_radio_destinofield. - Small, curated collection instead of a broad scrape: so what: lower noise and clearer genre boundaries for supervised experiments, at the cost of scale.
Who it's for
Great fit if you need a compact, well‑labelled dataset to benchmark text‑to‑music alignment, train fine‑grained audio classifiers, or study generator capability gaps across nuanced subgenres. The CC‑BY‑4.0 license permits derivative training with attribution. Look elsewhere if you require very large scale (tens of thousands to millions of samples), stems/multi‑track separation, faithful odd‑meter or ethnomusicologically rigorous performances, or long‑form compositions — the collection skews short (≈2 min), favors 4/4 and electronic timbres, and provides only mixed MP3s and auto‑generated cover art.