AIAny
Icon for item

ACE-Data-0

Captures synchronized multimodal embodied-human data in real homes — egocentric and multi-view video, metric body/hand/object motion, audio, and tactile signals. Released under a gated non-commercial research license with identifiable participants and strict non-redistribution/privacy constraints.

Introduction

Embodied agents fail when perception and action streams are assembled from disparate sources; ACE-Data-0 tackles this by recording whole interactions on one clock and in one world frame so every visual, kinematic, audio and contact signal aligns exactly with physical object and body states. That measured alignment exposes perceptual failure modes (occlusion, egomotion, long horizons) that image-only benchmarks miss and makes the dataset useful both for evaluation and for training cross-modal, action-capable systems.

What Sets It Apart
  • Unified, measured supervision rather than pieced-together estimates: body, articulated hands, object 6-DoF poses, camera intrinsics/poses, and tactile pressure are tracked or sensed in the capture frame, so annotations remain valid under occlusion and extreme viewpoints — this reduces label noise for downstream imitation, policy learning, and world-model training.
  • Large, behaviorally rich takes with multi-scale capture: the planned release lists 150+ hours, ~17M video frames, ~75k interaction episodes across 200+ task categories, spanning short atomic manipulations to long-horizon chains (up to 20–30 minutes) — so models can be benchmarked on temporal consistency and goal-driven behavior, not just per-frame accuracy.
  • Paired egocentric and dense exocentric views with tactile and audio: 4 headset fisheye cameras plus 8+ exocentric views synchronized with mocap and tactile gloves enable cross-view reconstruction, tactile inference from vision, and analysis of egomotion vs. finger articulation errors.
  • Gated, research-only release with measured calibration: access is granted to named individuals for non-commercial academic research under a binding license; captures include per-take calibration and millisecond-level sync, enabling exact projection of any tracked 3D point into any camera frame.
Who It's For and Trade-offs

Great fit if you develop or evaluate embodied perception-action systems (imitation learning, world models, vision-language-action), cross-modal prediction (vision→tactile, video→motion), or robust pose/interaction recovery under occlusion and long horizons. The dataset’s strengths are synchronized multisensory fidelity and long, goal-driven behavior captures.

Look elsewhere if you need broad geographic or cultural diversity (data from two sites only), unlabelled large-scale web video, or an immediate unrestricted public download — ACE-Data-0 is gated, contains identifiable participants, disallows commercial use and redistribution, and the mocap rigs and gloves are visible in recordings (which may introduce dataset-specific visual cues).

Information

  • Websitehuggingface.co
  • OrganizationsS-Lab, Nanyang Technological University, Singapore, ACE Robotics
  • AuthorsYukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian
  • Published date2026/07/28

More Items

Hugging Face

Provides 999,847 persona records—599,847 grounded from real sources and 400,000 synthetic—each encoded as 1,290 categorical attributes packed into 645-byte Parquet blobs. Includes a codebook, postings index, and calibration/audit artifacts; decode with pyarrow and persona_codes.schema.json.

Enables tactile-aware robot manipulation by pretraining a vision–tactile–language–action foundation model and improving offline policies with ALTER. Combines large-scale NeoData visuo-tactile pretraining, a latent tactile pathway for predictive touch signals, and advantage‑conditioned offline RL for contact-rich tasks.

Hugging Face

A compact evaluation dataset and harness for testing agentic AI on safety-critical robotics tasks. Includes multimodal episodes in parquet format, task-specific eval scripts (gauge reading, human safety monitoring, VLA estimators), and TFDS/Hugging Face integration for reproducible safety evaluations.