Cancer screening and treatment-response research depend on longitudinal, multimodal clinical data rather than single-timepoint images. CancerVerse addresses this gap by providing a large, de-identified corpus of longitudinal CT scans paired with the original radiology reports and expert voxel-level tumor masks, enabling models to learn temporal trajectories and grounded image–text signals rather than isolated snapshots.
What Sets It Apart
- Longitudinal scale: tens of thousands of scans (≈24.4k) from >14k patients with up to 14.4 years and up to 26 timepoints per patient — suitable for modeling progression and treatment response.
- Multimodal alignment: every scan is paired with the radiologist's free-text report, plus pathology and clinical metadata when available — enables vision–language and report-generation tasks grounded in real clinical language.
- Cancer-centric, voxel-level supervision: expert-drawn tumor segmentations across 13 cancer types (totaling >12k annotated lesions) so you can benchmark detection and segmentation at clinically relevant false-positive rates.
- Real-world heterogeneity & verification: multi-scanner data, four contrast phases, and a verified healthy cohort for realistic specificity estimation; responsibly de-identified for research use.
Who It's For and Trade-offs
Great fit if you are developing or evaluating AI for multicancer screening, longitudinal disease modeling, image–text pretraining, tumor detection/segmentation, or clinical-report generation. It is less suitable if you need a permissive commercial license (the public release is CC BY‑NC‑ND 4.0), if you require small/disk-light datasets (downloaded CT data require multiple terabytes), or if you lack the domain expertise and tooling to process 3D medical volumes and DICOM/NIfTI imaging formats.
Where It Fits
CancerVerse complements organ- or task-specific public sets (e.g., KiTS, LiTS) by offering large-scale longitudinal and multimodal coverage across many abdominal/pelvic/chest cancers, making it a natural choice for researchers focused on multi-organ screening and temporal modeling rather than single-organ segmentation benchmarks.