Most embodied-AI and robot-learning gaps come from a lack of large-scale, real-world streams that jointly capture vision, motion, audio and spatial geometry. This dataset converts everyday first‑person interactions into a unified 4D training corpus so models can learn perception, dynamics, and interaction from the same synchronized experience stream.
What Sets It Apart
- Truly multimodal and large-scale: 10M interaction episodes with synchronized vision (four fisheye streams), audio, stereo/monocular depth, IMU, hand and full-body mocap, and camera poses — totals on the order of 2.88B RGB frames, 720M depth frames, and ~1PB of data. This scale supports pretraining across temporal, spatial and kinematic modalities.
- Structured 3D/4D annotations: per-episode annotation files include calibration, geometry, trajectories, hierarchical natural-language captions, object instances and dense pose/mocap — enabling cross-modal grounding (language↔action↔3D) and long-horizon trajectory learning.
- Egocentric, in-the-wild focus: first-person captures of human interactions emphasize human-object interaction, manipulation, and embodied behavior rather than lab-constrained, staged scenes — useful for imitation learning, real-to-sim transfer, and world models that must reason about agent-centric observations.
- Designed for downstream embodied tasks: the data supports SLAM/pose estimation, action recognition/localization, multimodal pretraining (vision+language+motion), sim-to-real pipelines, and robotics imitation learning with dense kinematic supervision.
Who it's for — and tradeoffs
Great fit if you need large-scale, synchronized multimodal experience traces for training embodied or spatially-aware models (e.g., multimodal foundation models, robot policies, or 3D reconstruction systems). The dataset's scale and annotation density make it especially valuable for pretraining or for supervised tasks requiring motion/pose labels aligned with video and language.
Look elsewhere if you need fully open commercial use or very low-bandwidth samples: access is research-only and gated, the dataset is enormous (~1PB) so storage and compute costs are substantial, and many users will rely on provided samples rather than the full corpus. Also, because captures are egocentric, tasks requiring third‑person multi-view cinematic footage are not the primary fit.
Practical notes
- Language annotations are in English and vocabulary/annotation schemas are hierarchical for multi-granularity supervision.
- The dataset is released for non-commercial research use and may require an access agreement.
- Typical uses: embodied model pretraining, action-language grounding, 3D/4D reconstruction and tracking, imitation learning and sensor-fusion research.