AIAny
Icon for item

Meddies Persona VIE

Provides 150,000 synthetic Vietnamese patient personas to condition clinical text generation. Each persona bundles demographics, socioeconomic context, health and behavior fields, and prompt-ready narratives; intended for research and simulation, not clinical decision-making.

Introduction

Most clinical-generation pipelines fail upstream because the patient "behind" the text is underspecified. This release supplies 150,000 machine-generated Vietnamese personas designed to be the persona-first conditioning layer for downstream note/dialogue generation and triage simulations—rich demographics and narrative context to reduce brittle, context-light outputs.

What Sets It Apart
  • Persona-first schema: dense coverage for demographics, healthcare behavior, and LLM-facing narrative fields so prompts get realistic social and cultural anchors rather than isolated symptoms.
  • Scale and variety: 150k rows spanning the full life course, regional dialect clusters, long-tail symptom distributions, and varied socioeconomic contexts to stress-test generation across realistic edge cases.
  • Prompt-ready fields and metadata: includes chief complaint, HPI-style narrative snippets, social-barrier cues, and generation metadata (seeds, model ids) to support reproducible scenario pipelines.
  • Purposeful limits: medication and deep medical-history fields are intentionally lighter—these personas are anchoring context for synthetic generation, not substitutes for clinical charts.
Who It's For and Trade-offs

Great fit if you build synthetic doctor–patient consultations, intake/HPI note generators, triage simulations, or prompt-stress QA pipelines for Vietnamese clinical workflows. Look elsewhere if you need real-world prevalence estimates, clinical-grade patient records, or a dataset cleared for unrestricted commercial healthcare use—the release is machine-generated, can contain implausible or biased combinations, and is licensed CC-BY-NC-4.0, so apply QA and domain review before downstream use.

Information

Categories

More Items

Hugging Face

Provides a machine-readable catalog of 117 AI/AX safety and deployment-readiness diagnostic criteria for assessing model intrinsic and serving/infrastructure risks. Includes MODEL-SCAN and AX-SCAN axes, bilingual source fields, per-item evidence guidance, severity/assurance metadata, and a CC BY-NC 4.0 release-candidate.

Hugging Face

Evaluates schema-guided structured extraction from documents: given a document and a JSON schema, systems must return a schema-valid JSON with page-and-box grounding. Covers 370 documents (4,869 pages) across 8 business domains and 67 document types; scores value accuracy, word/page grounding, and long-list completeness.

Hugging Face

Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.