AIAny
Icon for item

Doctor-Patient Conversations — All Human Diseases (Opus 5.5)

Synthetic, clinician-verified ChatML dataset of 2,194 doctor–patient encounters covering 2,194 unique human diseases; each JSONL record includes 20 structured fields, verified PubMed references, realistic vitals/labs, and is intended for RAG and model fine-tuning (not medical advice).

Introduction

Opus 5.5 is a purpose-built synthetic clinical-dialogue dataset designed to be a grounded, auditable source of training and retrieval context for medical AI workflows. Rather than short examples, each record simulates a full clinical encounter from presentation through workup, treatment and disposition, with numeric consistency checks and live verification of cited PubMed IDs.

What Sets It Apart
  • One record per disease (2,194 unique diseases) in a single JSONL file; each line is a validated JSON object with 20 keys (personas, scenario, conversation, structured clinical fields, pubmed_refs, executive summary, etc.).
  • Long, realistic conversations: average ~33.4 turns and ~5,700 tokens (~22.8k characters) per conversation, designed for ChatML workflows and SFT conversion.
  • Rigorous provenance and verification: generated with Claude Opus 5.5 under automated gates (>=15 turns, >=8,000 characters, strict role alternation, numeric/lab consistency checks) and every PMID cited in-text was live-checked against NCBI; pubmed_refs rebuilt from the actual citations.
  • Practical engineering choices: compact retrieval keys (name+aliases+description+executive_summary) recommended for embeddings, full conversation stored as payload for RAG to budget large-context models.
Key uses
  • Retrieval-augmented generation (RAG) grounding for clinical Q&A and assistants — long conversations act as payloads while compact keys power retrieval.
  • Supervised fine-tuning / SFT for dialog models using the included ChatML-ready conversation lists.
  • Safety, pedagogy and numeric-consistency benchmarks: labs, vitals, dosing and interactions were enforced to match internal chemistry checks (Henderson–Hasselbalch, anion gap, dose-vs-weight, known drug interactions).
Who it's for and trade-offs
  • Great fit if you need high-fidelity, long-form clinical dialogues with structured metadata for RAG or model training and can responsibly handle medical content. Use with retrieval strategies that avoid embedding full ~5.7k-token conversations to fit model windows.
  • Not appropriate as medical advice or as a replacement for primary literature — PMIDs were verified to exist and titles matched, but claims should be validated against source papers before clinical use. Some records intentionally include a single patient obfuscation as a pedagogic feature. Licensing: MIT (dataset card lists Apache-2.0 and MIT notes; confirm on the dataset page before redistribution).

Information

Categories

More Items

Hugging Face

Provides over 1.1M hours of high-bandwidth, multichannel multilingual speech with segment- and word-level timestamps, English translations, and per-file metadata for ASR, TTS and audio-representation research. Preserves original 48kHz multichannel OPUS audio and is released under CC BY 3.0.

Hugging Face

Provides 2,000 synthetic multiple-choice items designed for continuation log-likelihood scoring to evaluate small language models' Theory of Mind (social-cognitive) abilities; 40 constructs, balanced answer positions, and easy/medium difficulty.

Hugging Face

Provides 5.5K+ self-contained data-analysis RL tasks: each row bundles a real tabular dataset, a question, and a deterministically-gradable gold answer. Verified from jupyter-agent notebooks; splits for training, held-out testing, and quick eval; intended for prompting, fine-tuning, and agent RL.