Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.
Transforms articulated 3D asset creation into a programmatic, LLM-driven code-generation workflow that produces objects with semantic parts, robust geometry, and physical joints. Includes CLI generation, a local viewer, and pipelines for large-scale dataset contribution.
Large-scale in-the-wild robot manipulation dataset with ~76K teleoperated trajectories (~350 hours) that provides synchronized multi-view video, depth, camera calibration, robot state/action traces, and natural-language task instructions to train and evaluate manipulation policies and dynamics models. Collected across 564 scenes, 86 tasks, 52 buildings, on a uniform Franka Panda hardware stack and released in LeRobotDataset v3.0 format (≈707 GB, OpenMDW1.1).
Reconstructs camera poses and dense 3D point clouds from video streams using a feed‑forward foundation model. Combines a Geometric Context Transformer (anchor + local window + trajectory memory) with paged KV‑cache attention to enable stable, long‑sequence streaming inference (~20 FPS at 518×378).
Provides a large-scale multimodal embodied dataset (vision, depth, hand/arm kinematics, tactile) captured with an exoskeleton glove and egocentric sensors; organized as clip-level Zarr volumes for manipulation, imitation learning, and vision–action research. Includes both high-precision glove measurements and natural bare-hand clips; sizable storage required.
A library of reusable agent skills that generate, inspect, and hand off CAD, robot-description, and fabrication artifacts. Exports STEP/STL/3MF, writes URDF/SDF/SRDF, slices meshes to G-code, previews files in-browser, and includes off-the-shelf STEP part lookup and benchmarks.
Provides aligned urban driving sensor streams (camera frames, LiDAR, radar and HD‑map / lanelet2 annotations) for multimodal perception, tracking and mapping research. Expert-generated labels under CC BY‑NC‑4.0 and hosted on Hugging Face.
Large-scale synthetic video dataset of physically simulated multi-object interaction scenes for training and evaluating models on physical reasoning, depth and optical-flow estimation, instance segmentation, and physics-grounded captioning. Provides RGB + lossless depth, per-frame instance masks, per-object physics annotations (NPZ), VLM-grounded captions, and USD scene files — useful for world-model and simulation-to-real work; commercial use permitted.
Open egocentric multimodal dataset for embodied AI and robot learning captured on commodity iPhone Pro: ~200 hours and ~10M RGB frames with LiDAR depth, ARKit 6‑DoF poses, IMU, two‑hand MANO mocap, room meshes, and hierarchical action captions.
Provides 10,000 articulated 3D objects in URDF for robotics and embodied-AI research. Generated by the Articraft agent and released under CC-BY-4.0, the dataset targets simulation, manipulation, kinematics evaluation, and training of embodied agents.
Performs fast, high-quality vision–language grounding: given an image plus a natural-language prompt it returns bounding boxes or points for referred objects. Uses Parallel Box Decoding for parallel coordinate prediction (higher throughput) and targets research/non-commercial use.
Generates and reasons about multimodal physical-world content—text, images, video, audio, and robot/action trajectories—conditioned on combinations of text, image, video and action inputs. The 64B “Super” variant targets Physical AI use cases and supports vLLM‑Omni, Diffusers, and action prediction.