AIAny
Icon for item

Suno AI Music Dataset — Multi-Genre Curated

A human‑curated corpus of AI‑generated music with MP3s, cover art, exact generation prompts and a 32‑column metadata schema; uses a 70/30 quality vs. mainstream split and a three‑level taxonomy to support fine‑grained audio‑ML, prompt‑fidelity and recommendation research.

Introduction

Why this matters Suno V5.5 made it practical to produce coherent, production‑style tracks across many niche subgenres; this dataset packages those outputs with unusually detailed prompts and dense metadata so researchers can evaluate prompt‑to‑audio fidelity, fine‑grained genre classifiers, and cross‑genre transfer without scraping noisy community uploads.

What Sets It Apart
  • Curatorial 70/30 split: 70% focuses on musically sophisticated territory (neo‑soul, progressive psy, contemporary jazz) while 30% targets high‑traffic mainstream styles with a deliberate “sophistication overlay” — so what: enables transfer and robustness studies between high‑quality and commercial material.
  • Prompt‑paired audio: every track ships with the exact descriptive prompt (30–80 tokens typical) plus BPM, key, mood and a taxonomic label — so what: lets you re‑render prompts in other generators or quantify prompt‑fidelity systematically.
  • Three‑level taxonomy and rich metadata (100+ sub‑sub‑genres, 32 columns): so what: supports hierarchical classification, auto‑tagging benchmarks and listener‑context research using the explicit tax_radio_destino field.
  • Small, curated collection instead of a broad scrape: so what: lower noise and clearer genre boundaries for supervised experiments, at the cost of scale.
Who it's for

Great fit if you need a compact, well‑labelled dataset to benchmark text‑to‑music alignment, train fine‑grained audio classifiers, or study generator capability gaps across nuanced subgenres. The CC‑BY‑4.0 license permits derivative training with attribution. Look elsewhere if you require very large scale (tens of thousands to millions of samples), stems/multi‑track separation, faithful odd‑meter or ethnomusicologically rigorous performances, or long‑form compositions — the collection skews short (≈2 min), favors 4/4 and electronic timbres, and provides only mixed MP3s and auto‑generated cover art.

Information

  • Websitehuggingface.co
  • OrganizationsHugging Face, Suno
  • AuthorsKukito
  • Published date2026/05/24

Categories

More Items

Hugging Face

Evaluation dataset for comparing eight text-to-image models using 8,000 generated images with source prompts and per-image scores for aesthetic quality, emotional resonance, and content integrity. Includes model labels, shared prompts, GPT-5.6 Sol automated scores, embedded images in Parquet, and an Apache-2.0 license.

Hugging Face

Provides a 1 trillion-token multimodal interleaved dataset (HTML subset updated as data_v1_1 with 742B HTML tokens) and 3.4B images drawn from HTML/PDF/ArXiv sources for multimodal pretraining; released under CC-BY-4.0 with safety and deduplication guidance.

Hugging Face

Provides 369 Harbor sandbox tasks ported from OpenAI's openai/math: each task is a Lean theorem with missing `sorry` proofs that an agent must complete, graded by a strict Comparator exact-match verifier. Includes task definitions, generator, and manifest for RL/code-agent evaluation.