AIAny
Icon for item

SenseNova Vision Corpus 50M

Provides ~50M multimodal annotations organized for unified training across structured visual understanding, segmentation, dense geometric prediction, and multi-view reconstruction — released as task-specific JSONL files that reference original image assets rather than redistributing raw images.

Introduction

Most public vision corpora are task-siloed and inconsistently annotated, which complicates training unified multimodal models. SenseNova Vision Corpus 50M reorganizes open-source visual data into task-specialized, training-ready supervision across structured perception, segmentation, dense geometry, and multi-view geometry, making it practical to train or evaluate models that need consistent, multi-task visual signals.

What Sets It Apart
  • Unified task families: consolidates 73 dataset-task entries across 10 task types (structured understanding, segmentation, dense geometric prediction, multi-view geometry), so you can assemble multi-task training mixes without manual reformatting.
  • Task-aware curation pipelines: uses pipelines (Rex-Omni adaptations, MoGe-2 densification, LingBot-Depth) to turn sparse or inconsistent annotations into more training-compatible labels — this reduces per-dataset preprocessing overhead.
  • Asset referencing strategy: JSONL annotations keep relative file paths instead of embedding RGB assets, avoiding redistribution issues but requiring users to align a local root with original image sources.
  • Scale and balance: provides tens of millions of frames distributed across complementary task families (e.g., ~18.9M structured, ~17.3M dense-geometry frames), enabling both dense-prediction and high-level multimodal supervision at scale.
Who It's For and Tradeoffs

Great fit if you are training or evaluating multimodal vision models that need consistent supervision across geometry, segmentation, and structured tasks, or if you want to build multi-task curricula without stitching dozens of ad-hoc dataset formats. Look elsewhere if you cannot obtain or reconcile the original image assets (annotations reference file paths, not raw images), if you require a permissive commercial license (dataset uses CC BY-NC 4.0), or if you need datasets exclusively containing proprietary or private image collections. The corpus reduces annotation heterogeneity but shifts effort to dataset rooting and asset alignment.

Information

Categories

More Items

Hugging Face

De-identified longitudinal multimodal CT dataset for multicancer screening that pairs ~24k CT volumes with radiology reports and voxel-wise tumor annotations across 13 cancer types. Designed for longitudinal disease modeling, detection/segmentation and vision–language research; CC BY‑NC‑ND 4.0 for non-commercial use.

Hugging Face

Provides 98,877 historical newspaper page images (1700s–1940s) paired with ALTO OCR, per-word confidences and line/word bounding boxes in image pixels — ready for line-level OCR training, OCR quality estimation, re‑OCR comparisons and layout analysis. OCR is library-produced (silver); image resolutions and OCR quality vary.

Hugging Face

Provides 1.21M densely annotated desktop screenshots and 159.7M element instances for training and evaluating GUI grounding and screen-parsing models. Includes per-element accessibility-derived annotations, 917K recorded click transitions, multi-application scenes across seven appearance presets and resolutions; distributed as WebDataset shards with Parquet indexes.