AIAny
Icon for item

EgoPro

10,000-hour head-and-wrist egocentric dataset pairing synchronized head and wrist video with left/right 3D hand pose and optional full-body pose; provided in LeRobot/MCAP formats with episode-level semantic annotations and automated de-identification.

Introduction

Why this matters

Egocentric manipulation research needs dense, close-range views of hands interacting with objects plus temporal context. This release provides 10,000 hours of synchronized head- and wrist-mounted video with per-hand 3D pose (and a 2,000‑hour body variant), enabling learning of fine-grained contact, grasping, and whole-body coordination at scale.

What Sets It Apart
  • Scale and viewpoint: 10,000 total hours split into EgoProStandard (8,000 h: head+wrist + left/right hand pose) and EgoProStandard-body (2,000 h: adds full-body pose). This Pro line complements a larger 100k-hour family focused on head-only capture.
  • Multimodal packaging: Distributed in LeRobot v3.0-compatible and MCAP packages with Parquet metadata, synchronized video streams, and verified sidecar files for easy ingestion by multimodal training pipelines.
  • Task and scene diversity: Episodes span 15,000+ tasks and distinct collection scenes (collection-level coverage shared across the wider release), designed for transfer beyond lab demonstrations to real environments.
  • Responsible release: Recordings underwent automated de-identification with human verification (faces, plates blurred) and participant consent; semantic event-level annotations are included as a complimentary add-on.
  • Practical constraints surfaced: the full-body pose is only present in the body SKU; clients should rely on declared feature/topic names when wrist 6DoF pose is included.
Who It's For and Trade-offs

Great fit if you build perception or imitation systems that require close-up hand views, hand-object contact labels, or cross-view synchronization (e.g., manipulation perception, robotic grasping, action segmentation, self-supervised representation learning). It’s also suitable for benchmarking multimodal pipelines that consume LeRobot/MCAP and Parquet metadata.

Look elsewhere if you need small-scale, label-light benchmarks (this release is very large—multi-terabyte storage and substantial compute required) or if every episode must include full-body annotations (only the body SKU provides that). Licensing is repository-governed (non-standard license tag), so verify permitted uses for downstream applications.

Where It Fits

Positioned as the Pro/wrist-view arm of the broader EgoSuite family, this dataset complements head-only collections by adding wrist-camera close-ups for high-fidelity hand-object interaction, making it a natural choice when comparing head-only models to multi-view embodied agents.

Information

  • Websitehuggingface.co
  • OrganizationsLightwheelAI
  • Published date2026/08/07

Categories

More Items

Hugging Face

A small public sample of egocentric human demonstration video with synchronized 3D hand and body pose annotations for imitation learning and embodied-AI research. Delivered in Parquet and common multimodal packages (LeRobot, MCAP) for schema inspection before requesting gated access to larger EgoSuite releases.

Hugging Face

Provides 90,000 hours of head-mounted egocentric video paired with synchronized 3D hand pose and an optional 3D full‑body pose add-on, with event-level semantic labels available as a complimentary layer — designed for embodied AI and robotics training at scale.

Hugging Face

Provides 1,021.64 hours across 597 CAD/BIM workflows with synchronized screen recordings and interaction logs; each workflow includes video, timestamped input events, task specs, source files, final outputs, and evaluation rubrics for training or evaluating desktop CAD agents.