AIAny
Icon for item

Xperience-10M

Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.

Introduction

Most embodied-AI and robot-learning gaps come from a lack of large-scale, real-world streams that jointly capture vision, motion, audio and spatial geometry. This dataset converts everyday first‑person interactions into a unified 4D training corpus so models can learn perception, dynamics, and interaction from the same synchronized experience stream.

What Sets It Apart
  • Truly multimodal and large-scale: 10M interaction episodes with synchronized vision (four fisheye streams), audio, stereo/monocular depth, IMU, hand and full-body mocap, and camera poses — totals on the order of 2.88B RGB frames, 720M depth frames, and ~1PB of data. This scale supports pretraining across temporal, spatial and kinematic modalities.
  • Structured 3D/4D annotations: per-episode annotation files include calibration, geometry, trajectories, hierarchical natural-language captions, object instances and dense pose/mocap — enabling cross-modal grounding (language↔action↔3D) and long-horizon trajectory learning.
  • Egocentric, in-the-wild focus: first-person captures of human interactions emphasize human-object interaction, manipulation, and embodied behavior rather than lab-constrained, staged scenes — useful for imitation learning, real-to-sim transfer, and world models that must reason about agent-centric observations.
  • Designed for downstream embodied tasks: the data supports SLAM/pose estimation, action recognition/localization, multimodal pretraining (vision+language+motion), sim-to-real pipelines, and robotics imitation learning with dense kinematic supervision.
Who it's for — and tradeoffs

Great fit if you need large-scale, synchronized multimodal experience traces for training embodied or spatially-aware models (e.g., multimodal foundation models, robot policies, or 3D reconstruction systems). The dataset's scale and annotation density make it especially valuable for pretraining or for supervised tasks requiring motion/pose labels aligned with video and language.

Look elsewhere if you need fully open commercial use or very low-bandwidth samples: access is research-only and gated, the dataset is enormous (~1PB) so storage and compute costs are substantial, and many users will rely on provided samples rather than the full corpus. Also, because captures are egocentric, tasks requiring third‑person multi-view cinematic footage are not the primary fit.

Practical notes
  • Language annotations are in English and vocabulary/annotation schemas are hierarchical for multi-granularity supervision.
  • The dataset is released for non-commercial research use and may require an access agreement.
  • Typical uses: embodied model pretraining, action-language grounding, 3D/4D reconstruction and tracking, imitation learning and sensor-fusion research.

Information

  • Websitehuggingface.co
  • OrganizationsRopedia, Hugging Face
  • Published date2026/03/11

More Items

Hugging Face

Provides 2,000 synthetic multiple-choice items designed for continuation log-likelihood scoring to evaluate small language models' Theory of Mind (social-cognitive) abilities; 40 constructs, balanced answer positions, and easy/medium difficulty.

Hugging Face

Provides 5.5K+ self-contained data-analysis RL tasks: each row bundles a real tabular dataset, a question, and a deterministically-gradable gold answer. Verified from jupyter-agent notebooks; splits for training, held-out testing, and quick eval; intended for prompting, fine-tuning, and agent RL.

Hugging Face

Provides mixed-domain, verifiable RL training environments for LLM agents (code, cyber, knowledge work, web dev, music) as Parquet datasets, with domain-specific verifiers, Docker artifacts and links to training code for reproducible agentic RL experiments.