AIAny
Icon for item

Gen-HumanEgo

Provides 1,800+ hours of synchronized egocentric multi-view recordings with 3D hand reconstructions, wide‑FOV depth, and hierarchical task/subtask annotations for embodied AI and robot learning. Includes six fisheye views, hand meshes, and per-episode temporal labels across 44k+ episodes.

Introduction

Gen-HumanEgo matters because modern embodied AI and robot learning need large-scale, realistic human manipulation data that includes both visual observations and structured supervision. By combining synchronized egocentric multi-view video with dense 3D hand geometry, depth, and hierarchical task labels, this dataset bridges the gap between raw human demonstrations and training-ready signals for policies and perception.

What Sets It Apart
  • Scale and diversity: ~1,847.7 hours across 44,632 episodes and 10,257 unique tasks spanning home, business, industry, and agriculture—suitable for broad generalization. This volume supports both pretraining and downstream fine-tuning.
  • Rich multimodal supervision: six synchronized fisheye RGB streams at 1600×1300@30FPS, Ego-Depth (large field of view), per-hand 3D keypoints, MANO parameters and full hand meshes, plus SLAM/pose metadata—so models can learn appearance, geometry, and dynamics concurrently.
  • Structured semantics: hierarchical annotations (video/episode descriptions, task segments, fine-grained subtasks with temporal boundaries and success flags) provide dense, temporally localized supervision for behavior cloning, imitation learning, and task understanding.
  • Robot-ready packaging: data stored in MCAP with standardized stream paths and an official toolkit for parsing, enabling direct use in embodied AI pipelines and sim-to-real workflows.
Who It's For and Tradeoffs

Great fit if you are training or evaluating embodied-perception models, hand/object pose estimators, imitation or behavior policies, or multi-modal representation learners that require synchronized vision + geometry + semantics at scale. The dataset is particularly valuable for zero-shot or few-shot human-to-robot transfer research that needs dense hand-object interaction signals.

Look elsewhere if you need calibrated external motion-capture ground truth for entire body full-scene metrics, proprietary-label legal clearances for commercial sensitive domains, or extremely low-storage footprints—the dataset’s scale (multi-terabyte class) and MCAP format imply significant storage and processing demands.

More Items

Provides WROP: a 1.5M-sample synthetic video corpus and a 300-question exam for training and evaluating object permanence and solidity in video world models. Includes 150 Blender task generators, a human Elo benchmark across 14 models, and a fine-tuned 16B continuation model (PWM-WROP).

Hugging Face

Provides 10,000 agentic multi-turn coding and reasoning traces from Fable 5.1 with step-by-step chain-of-thought and tool-use, heavily deduplicated and filtered; totals ~500M tokens (1.97 GB). Suited for SFT, distillation, and training long-horizon reasoning and code-generation models.

Hugging Face

Benchmark for typed probabilistic decisions: given one shared state, a model answers multiple typed questions at once and returns full probability distributions (noul/choice/score). Contains four workflows with train/test splits and soft gold labels from teacher samples, designed to evaluate accuracy, calibration, and latency trade-offs.