AIAny
Icon for item

CancerVerse

De-identified longitudinal multimodal CT dataset for multicancer screening that pairs ~24k CT volumes with radiology reports and voxel-wise tumor annotations across 13 cancer types. Designed for longitudinal disease modeling, detection/segmentation and vision–language research; CC BY‑NC‑ND 4.0 for non-commercial use.

Introduction

Cancer screening and treatment-response research depend on longitudinal, multimodal clinical data rather than single-timepoint images. CancerVerse addresses this gap by providing a large, de-identified corpus of longitudinal CT scans paired with the original radiology reports and expert voxel-level tumor masks, enabling models to learn temporal trajectories and grounded image–text signals rather than isolated snapshots.

What Sets It Apart
  • Longitudinal scale: tens of thousands of scans (≈24.4k) from >14k patients with up to 14.4 years and up to 26 timepoints per patient — suitable for modeling progression and treatment response.
  • Multimodal alignment: every scan is paired with the radiologist's free-text report, plus pathology and clinical metadata when available — enables vision–language and report-generation tasks grounded in real clinical language.
  • Cancer-centric, voxel-level supervision: expert-drawn tumor segmentations across 13 cancer types (totaling >12k annotated lesions) so you can benchmark detection and segmentation at clinically relevant false-positive rates.
  • Real-world heterogeneity & verification: multi-scanner data, four contrast phases, and a verified healthy cohort for realistic specificity estimation; responsibly de-identified for research use.
Who It's For and Trade-offs

Great fit if you are developing or evaluating AI for multicancer screening, longitudinal disease modeling, image–text pretraining, tumor detection/segmentation, or clinical-report generation. It is less suitable if you need a permissive commercial license (the public release is CC BY‑NC‑ND 4.0), if you require small/disk-light datasets (downloaded CT data require multiple terabytes), or if you lack the domain expertise and tooling to process 3D medical volumes and DICOM/NIfTI imaging formats.

Where It Fits

CancerVerse complements organ- or task-specific public sets (e.g., KiTS, LiTS) by offering large-scale longitudinal and multimodal coverage across many abdominal/pelvic/chest cancers, making it a natural choice for researchers focused on multi-organ screening and temporal modeling rather than single-organ segmentation benchmarks.

Information

Categories

More Items

Hugging Face

Provides 98,877 historical newspaper page images (1700s–1940s) paired with ALTO OCR, per-word confidences and line/word bounding boxes in image pixels — ready for line-level OCR training, OCR quality estimation, re‑OCR comparisons and layout analysis. OCR is library-produced (silver); image resolutions and OCR quality vary.

Hugging Face

Provides 1.21M densely annotated desktop screenshots and 159.7M element instances for training and evaluating GUI grounding and screen-parsing models. Includes per-element accessibility-derived annotations, 917K recorded click transitions, multi-application scenes across seven appearance presets and resolutions; distributed as WebDataset shards with Parquet indexes.

Hugging Face

Generates humanoid robot motion references that preserve object contact locations/timing by solving windowed trajectory optimizations against contact targets in the object frame. Releases retargeted trajectories for two Unitree robots across 75 objects (≈13.9k robot–motion pairs); CC BY‑NC‑SA 4.0.