AIAny
Icon for item

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Provides a portable, robot-free UMI capture pipeline and shows that policies post-trained only on this high-fidelity data deploy directly on real robots matching teleoperation baselines. Capture achieves ~3 mm end-effector accuracy, microsecond sync, ultra-wide FOV, and releases 2,000h HiFi-UMI-2K.

Introduction

Raising the fidelity of robot-free UMI demonstrations — rather than merely increasing the small fraction of real-robot anchors — can enable zero-robot post-training: policies trained or post-trained only on a sufficiently accurate UMI corpus can transfer to real hardware with minimal real-robot data.

Key Findings
  • Zero-robot post-training works in practice: a policy post-trained solely on the HiFi-UMI corpus deploys on a real robot and matches in-domain teleoperation baselines across three backbone families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points.
  • High-fidelity capture matters: the HiFi-UMI system attains ~3 mm workspace-local end-effector accuracy without external trackers by combining head-mounted stereo-inertial SLAM, native inter-gripper relative pose, microsecond GPIO synchronization, and two wide-angle cameras per hand (~200° coverage).
  • Large-scale pretraining amplifies gains: pretraining on 4,000 hours from the same corpus reduced action error on ten unseen tasks by 41% and improved real-robot success on StarVLA-QwenPI by 18.1 percentage points.
  • Practical resource release: the authors open-source HiFi-UMI-2K, a 2,000-hour, microsecond-synchronized ultra-wide-FoV demonstration corpus with automatic reconstruction and simulation-validated replay.
Who it's for and trade-offs

Great fit if you develop robotic manipulation policies and need scalable, low-cost demonstration data that can transfer to real robots without large teleoperation budgets. It benefits teams focused on vision-language-action and world-action-model architectures seeking to reduce real-robot collection. Look elsewhere if your domain demands contact-rich dynamics measured by force/torque sensors, sub-millimeter positional accuracy, or specialized hardware not covered by head-mounted capture; HiFi-UMI emphasizes trajectory/pose fidelity and visual coverage rather than instrumented force sensing.

Where it fits

HiFi-UMI sits between small-scale teleoperation datasets (high fidelity but costly) and large, low-fidelity synthetic or proxy corpora: it aims to deliver near-teleop realism at scale, trading expensive robot-hours for a calibrated, synchronized human-mounted capture rig and large annotation-free hours of demonstrations.

Information

  • Websitearxiv.org
  • AuthorsYuteng Wei, Jinming Ma, Jiawei Wang, Weitao Zhou, Yushen Zuo, Ke Rui, Minglei Li, Jinhao Zhang, Zhikang Pan, Xiang Wang
  • Published date2026/07/28

More Items

Trains LVLMs to improve 3D spatial reasoning using a synthetic block-stacking dataset of 15,000 problems that cover 3D→2D projection, viewpoint transforms, and structural combination; includes color-anchored cues and provides code and data.

An 8B native unified multimodal model that understands, reasons about, and generates high-fidelity visual content without a separate visual encoder or VAE—focused on spatially coherent patch reconstruction, native 4K generation, multi-expert distillation, and improved text rendering and image editing.

Hugging Face
AI Model2026

Integrates a pretrained vision–language model with a BEV perception head and a Planning Expert to provide 3D perception, driving VQA and motion planning for autonomous driving while keeping the base VLM architecture unchanged.