Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.
Provides 1,800+ hours of synchronized egocentric multi-view recordings with 3D hand reconstructions, wide‑FOV depth, and hierarchical task/subtask annotations for embodied AI and robot learning. Includes six fisheye views, hand meshes, and per-episode temporal labels across 44k+ episodes.
10,000-hour head-and-wrist egocentric dataset pairing synchronized head and wrist video with left/right 3D hand pose and optional full-body pose; provided in LeRobot/MCAP formats with episode-level semantic annotations and automated de-identification.
A small public sample of egocentric human demonstration video with synchronized 3D hand and body pose annotations for imitation learning and embodied-AI research. Delivered in Parquet and common multimodal packages (LeRobot, MCAP) for schema inspection before requesting gated access to larger EgoSuite releases.
Provides 90,000 hours of head-mounted egocentric video paired with synchronized 3D hand pose and an optional 3D full‑body pose add-on, with event-level semantic labels available as a complimentary layer — designed for embodied AI and robotics training at scale.
Provides a large-scale benchmark and a human-aligned metric for humanoid whole-body motion tracking — about 153 hours of optical mocap from professional performers plus HumanScore trained on 12K human-labeled preference pairs to reveal contact, timing, and stability failures.
Provides 617.5 hours of high-precision optical motion-capture with synchronized object trajectories and standardized 55-joint BVH for whole-body and human–object interaction research. Frame‑LU indexed and paired with natural-language descriptions; designed for humanoid learning, motion priors, and interaction-aware benchmarks.
Turns an uncalibrated monocular actor video into multiview-consistent novel-view videos and lifts them into 4D Gaussian Splatting assets. Introduces Reference Context Packing to keep reference conditioning fixed-size and Target Context Routing to exchange context across target groups, improving large-view reconstruction consistency.
Generates synchronized spoken dialogue and explicit full-body co-speech motion (facial expressions, hands, upper- and lower-body) end-to-end from the same hidden states, replacing the speech-then-motion cascade. Trains with a scalable pseudo-labeling pipeline (422,856 ranked pairs) and supports real-time inference (RTF 0.78) while matching teacher motion metrics within ~2%.
Provides 1,274 hours of head-mounted egocentric video paired with seven-point IMU arm tracking (24 Hz orientation; raw accel/gyro/mag on a subset), packaged for embodied-AI and egocentric-vision research. Key features: torso-relative pose via chest reference, separate Parquet IMU repo for efficient joins, CC-BY-4.0 license; heavy class skew and limited contributor diversity are important constraints.