AIAny
Icon for item

Xperience-10M

Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.

Introduction

Most embodied-AI and robot-learning gaps come from a lack of large-scale, real-world streams that jointly capture vision, motion, audio and spatial geometry. This dataset converts everyday first‑person interactions into a unified 4D training corpus so models can learn perception, dynamics, and interaction from the same synchronized experience stream.

What Sets It Apart
  • Truly multimodal and large-scale: 10M interaction episodes with synchronized vision (four fisheye streams), audio, stereo/monocular depth, IMU, hand and full-body mocap, and camera poses — totals on the order of 2.88B RGB frames, 720M depth frames, and ~1PB of data. This scale supports pretraining across temporal, spatial and kinematic modalities.
  • Structured 3D/4D annotations: per-episode annotation files include calibration, geometry, trajectories, hierarchical natural-language captions, object instances and dense pose/mocap — enabling cross-modal grounding (language↔action↔3D) and long-horizon trajectory learning.
  • Egocentric, in-the-wild focus: first-person captures of human interactions emphasize human-object interaction, manipulation, and embodied behavior rather than lab-constrained, staged scenes — useful for imitation learning, real-to-sim transfer, and world models that must reason about agent-centric observations.
  • Designed for downstream embodied tasks: the data supports SLAM/pose estimation, action recognition/localization, multimodal pretraining (vision+language+motion), sim-to-real pipelines, and robotics imitation learning with dense kinematic supervision.
Who it's for — and tradeoffs

Great fit if you need large-scale, synchronized multimodal experience traces for training embodied or spatially-aware models (e.g., multimodal foundation models, robot policies, or 3D reconstruction systems). The dataset's scale and annotation density make it especially valuable for pretraining or for supervised tasks requiring motion/pose labels aligned with video and language.

Look elsewhere if you need fully open commercial use or very low-bandwidth samples: access is research-only and gated, the dataset is enormous (~1PB) so storage and compute costs are substantial, and many users will rely on provided samples rather than the full corpus. Also, because captures are egocentric, tasks requiring third‑person multi-view cinematic footage are not the primary fit.

Practical notes
  • Language annotations are in English and vocabulary/annotation schemas are hierarchical for multi-granularity supervision.
  • The dataset is released for non-commercial research use and may require an access agreement.
  • Typical uses: embodied model pretraining, action-language grounding, 3D/4D reconstruction and tracking, imitation learning and sensor-fusion research.

Information

  • Websitehuggingface.co
  • OrganizationsRopedia, Hugging Face
  • Published date2026/03/11

More Items

Hugging Face

Provides a reproducible, deduplicated corpus of text extracted from PDFs for LLM pretraining—about 3 trillion tokens from ~475 million documents in 1733 language-script pairs. Includes OCR and text extraction pipelines, per-page language IDs, MinHash deduplication, and is released under ODC‑By 1.0.

Hugging Face

Provides OCR full text for 11.55M public-domain German newspaper pages (1638–1964) with per-page IIIF scans and ALTO XML coordinates; suited for historical NLP, language-model training, and OCR research. Pages carry explicit per-page public-domain licenses.

Hugging Face

Provides 545,431 math problems with model-generated solution traces (chain-of-thought and Python tool-integrated reasoning) verified against reference answers for training and evaluating LLM mathematical reasoning. Parquet-format dataset; DeepSeek‑V4‑Pro generated traces and mixed CC BY / CC BY‑SA licensing.