AIAny
Icon for item

Nemotron-Personas-Vietnam

Provides 600,000 synthetic Vietnamese persona texts (100,000 records, 6 personas per record) aligned to Vietnam's 2024 census and surveys for training and evaluating NLP / text-generation models; includes 21 demographic and persona fields, CC BY 4.0, single train split.

Introduction

Why this matters

Large-scale, regionally grounded persona data for Vietnamese is scarce; this release fills that gap by producing synthetic persona narratives aligned to Vietnam's official 2024 demographic sources. The dataset emphasizes realistic demographic distributions (age, sex, education, occupation, province) across six major provinces and supplies multiple persona types per record to increase conversational diversity in downstream models.

What Sets It Apart
  • Census-grounded synthesis: persona attributes are generated to match distributions from Vietnam's Population & Housing Census 2024 and VHLSS 2024, so demographic coverage reflects recent official statistics rather than generic web crawls — useful when you need regionally representative behavior priors.
  • Multi-persona per record: each record contains six persona variants (professional, sports, arts, travel, culinary, and a concise persona), enabling augmentation strategies that preserve contextual demographic fields while varying persona voice and intent.
  • Auditability & reproducibility: produced with a NeMo Data Designer pipeline and a probabilistic graphical model augmented by the SaoLa4-Small component, with an explicit schema (21 fields) and a single train split (100k records). This makes it straightforward to sample, filter, or integrate into training pipelines.
  • Licensing & scope clarity: CC BY 4.0 license and explicit exclusion of enterprise-only fields (e.g., names/personality trait details) make reuse for research and commercial model training straightforward while highlighting limitations.
Who It's For & Trade-offs

Great fit if you need synthetic, demographically grounded Vietnamese personas to augment training data, reduce sampling bias, or test model behavior across population slices (age, education, occupation, urban/rural, province). It is also useful for benchmarking Vietnamese text-generation and persona-conditioned response diversity.

Look elsewhere if you require: fine-grained real personal identifiers (the dataset omits real names and sensitive enterprise fields), child personas (only ages 18+), or fully public-source provenance for every seed record (the release mixes public statistics with proprietary Data Designer workflows). Also avoid using synthetic personas as direct substitutes for audited, consented human subject data in high-stakes domains without additional review.

Information

  • Websitehuggingface.co
  • AuthorsNVIDIA Corporation, FPT Smart Cloud, Quantum AI & Cyber Security Institute (FPT Corporation)
  • Published date2026/06/04

Categories

More Items

Hugging Face

Provides 315,000 pairwise human-preference votes comparing 15 English TTS models over 300 operational prompts, with 4,500 high‑quality audio renders and structured vote/pair/prompt records for training or evaluating preference/reward models. Metadata under CC-BY-4.0; audio use governed by model providers' terms.

Hugging Face

A 10‑billion‑document retrieval benchmark with per‑document 768‑dim unit‑norm dense embeddings and mGTE sparse embeddings, FineWeb text/metadata, and exact top‑1000 MS MARCO ground truth for ~120k queries. Built for large‑scale evaluation of dense/sparse/hybrid retrieval, filtered search, indexing, ANNS algorithms, and embedding compression.

Hugging Face

Provides 4.5 billion TikTok video records with captions, timestamps, music IDs and engagement counts for research; split across 27 zstd-compressed Parquet files (~289 GB) and sampled via TikTok's mobile API; released for research-use only with privacy and ToS caveats.