AIAny
Icon for item

Nemotron-Personas-Vietnam

Provides 600,000 synthetic Vietnamese persona texts (100,000 records, 6 personas per record) aligned to Vietnam's 2024 census and surveys for training and evaluating NLP / text-generation models; includes 21 demographic and persona fields, CC BY 4.0, single train split.

Introduction

Why this matters

Large-scale, regionally grounded persona data for Vietnamese is scarce; this release fills that gap by producing synthetic persona narratives aligned to Vietnam's official 2024 demographic sources. The dataset emphasizes realistic demographic distributions (age, sex, education, occupation, province) across six major provinces and supplies multiple persona types per record to increase conversational diversity in downstream models.

What Sets It Apart
  • Census-grounded synthesis: persona attributes are generated to match distributions from Vietnam's Population & Housing Census 2024 and VHLSS 2024, so demographic coverage reflects recent official statistics rather than generic web crawls — useful when you need regionally representative behavior priors.
  • Multi-persona per record: each record contains six persona variants (professional, sports, arts, travel, culinary, and a concise persona), enabling augmentation strategies that preserve contextual demographic fields while varying persona voice and intent.
  • Auditability & reproducibility: produced with a NeMo Data Designer pipeline and a probabilistic graphical model augmented by the SaoLa4-Small component, with an explicit schema (21 fields) and a single train split (100k records). This makes it straightforward to sample, filter, or integrate into training pipelines.
  • Licensing & scope clarity: CC BY 4.0 license and explicit exclusion of enterprise-only fields (e.g., names/personality trait details) make reuse for research and commercial model training straightforward while highlighting limitations.
Who It's For & Trade-offs

Great fit if you need synthetic, demographically grounded Vietnamese personas to augment training data, reduce sampling bias, or test model behavior across population slices (age, education, occupation, urban/rural, province). It is also useful for benchmarking Vietnamese text-generation and persona-conditioned response diversity.

Look elsewhere if you require: fine-grained real personal identifiers (the dataset omits real names and sensitive enterprise fields), child personas (only ages 18+), or fully public-source provenance for every seed record (the release mixes public statistics with proprietary Data Designer workflows). Also avoid using synthetic personas as direct substitutes for audited, consented human subject data in high-stakes domains without additional review.

Information

  • Websitehuggingface.co
  • AuthorsNVIDIA Corporation, FPT Smart Cloud, Quantum AI & Cyber Security Institute (FPT Corporation)
  • Published date2026/06/04

Categories

More Items

Evaluates whether video models reason according to physical laws by treating generated videos as visible reasoning traces and using a three-stage Perception–Formulation–Deduction protocol. Includes Orchard (400 mechanics videos), chain-of-frames prompting on annotated first frames, and a hybrid MLLM-plus-objective scoring suite for stage-resolved diagnostics.

Hugging Face

Provides intermediate pretraining checkpoints for the Aether-7B-5Attn base model to enable reproducible training-dynamics research. Includes three raw checkpoints (110k, 115k, 162k steps) packaged with model.safetensors, config, and tokenizer; uses a custom aether_v2_7way architecture requiring the aether_pkg loader.

Hugging Face

Provides 2,056 penetration-free cloth simulation trajectories (240 frames each, 493,440 frames, ~33 GB) across human garments, robotic manipulation, and object-collision scenarios. Includes per-vertex positions, per-frame displacements, mesh topology and collision fields under CC BY 4.0 — useful for training and evaluating learning-based cloth simulators.