Most embodied-learning pipelines rely on expensive real rollouts or simulator resets; imagined rollouts let models acquire multi-step visual and control trajectories at scale without additional real execution. This dataset curates HDF5 chunks of imagined interaction segments produced by a world model for RoboTwin2.0 tasks, exposing multi-view images, joint/gripper states, end-effector poses, actions, and sparse rewards in a chunked temporal format that is ready for model training and offline evaluation.
What Sets It Apart
- Chunked imagined rollouts with temporal alignment: each HDF5 file encodes one fixed-length chunk (21 frames, 20 actions) where obs[t] -- action[t] --> obs[t+1], making it straightforward to train transition models or sequence predictors without reassembling episodes.
- Multi-view perceptual and control state: synchronized RGB from head, left- and right-wrist cameras plus joint positions, gripper targets, and 7-DoF end-effector poses enable joint perception–control research (e.g., world models conditioned on proprioception and multisensor vision).
- Designed for world-model workflows: generated frames are the model’s predictions (not replayed ground-truth), and sparse binary rewards reflect an automated reward model, which mirrors common imagined-data training setups and supports contrastive evaluation between imagined and real trajectories.
- Task-organized and reproducible indexing: data are grouped by task (50 RoboTwin2.0 tasks) and include provenance via a root attribute (branch_id, chunk_index, transition_id) to reconnect chunks into longer rollouts when available.
Who It's For & Trade-offs
Great fit if you develop or evaluate world models, model-based RL, or multimodal transition predictors that need large numbers of imagined visual-control rollouts without running physical robots. It is especially useful for research comparing imagined vs. real rollouts, training video-conditioned planners, or probing multi-view consistency in predictions.
Look elsewhere if you need human-annotated frame-level rewards, complete real-world episodes, or a broad set of embodied domains beyond RoboTwin; the current release contains only the RoboTwin2.0 subdataset and the generated observations should not be treated as recorded sensor feedback. The dataset's license is not specified on the card, so verify licensing before redistribution or commercial use.