Embodied agents fail when perception and action streams are assembled from disparate sources; ACE-Data-0 tackles this by recording whole interactions on one clock and in one world frame so every visual, kinematic, audio and contact signal aligns exactly with physical object and body states. That measured alignment exposes perceptual failure modes (occlusion, egomotion, long horizons) that image-only benchmarks miss and makes the dataset useful both for evaluation and for training cross-modal, action-capable systems.
What Sets It Apart
- Unified, measured supervision rather than pieced-together estimates: body, articulated hands, object 6-DoF poses, camera intrinsics/poses, and tactile pressure are tracked or sensed in the capture frame, so annotations remain valid under occlusion and extreme viewpoints — this reduces label noise for downstream imitation, policy learning, and world-model training.
- Large, behaviorally rich takes with multi-scale capture: the planned release lists 150+ hours, ~17M video frames, ~75k interaction episodes across 200+ task categories, spanning short atomic manipulations to long-horizon chains (up to 20–30 minutes) — so models can be benchmarked on temporal consistency and goal-driven behavior, not just per-frame accuracy.
- Paired egocentric and dense exocentric views with tactile and audio: 4 headset fisheye cameras plus 8+ exocentric views synchronized with mocap and tactile gloves enable cross-view reconstruction, tactile inference from vision, and analysis of egomotion vs. finger articulation errors.
- Gated, research-only release with measured calibration: access is granted to named individuals for non-commercial academic research under a binding license; captures include per-take calibration and millisecond-level sync, enabling exact projection of any tracked 3D point into any camera frame.
Who It's For and Trade-offs
Great fit if you develop or evaluate embodied perception-action systems (imitation learning, world models, vision-language-action), cross-modal prediction (vision→tactile, video→motion), or robust pose/interaction recovery under occlusion and long horizons. The dataset’s strengths are synchronized multisensory fidelity and long, goal-driven behavior captures.
Look elsewhere if you need broad geographic or cultural diversity (data from two sites only), unlabelled large-scale web video, or an immediate unrestricted public download — ACE-Data-0 is gated, contains identifiable participants, disallows commercial use and redistribution, and the mocap rigs and gloves are visible in recordings (which may introduce dataset-specific visual cues).