AIAny
Icon for item

HOI-Retarget — Humanoid Human–Object Interaction Motion

Generates humanoid robot motion references that preserve object contact locations/timing by solving windowed trajectory optimizations against contact targets in the object frame. Releases retargeted trajectories for two Unitree robots across 75 objects (≈13.9k robot–motion pairs); CC BY‑NC‑SA 4.0.

Introduction

Why contact matters

Robotic retargeting that only matches poses often drops or penetrates the object once physics are enabled — contact geometry and timing are the missing link. HOI‑Retarget makes contact the primary constraint: every labeled contact is expressed in the object's frame and used as a target in a windowed trajectory optimizer so the resulting humanoid trajectories actually recover the intended grasps and supports under the robot's kinematic limits.

What Sets It Apart
  • Contact‑centric retargeting: contacts live in object coordinates and are treated as explicit targets during optimization, enabling the same demonstration to re-solve correctly when the object is resized or when different robot embodiments are used.
  • Large curated corpus: 6,952 distinct motions retargeted to two robots (Unitree G1 and H2) producing 13,904 robot–motion rows, covering 75 distinct object meshes and ≈13.8 hours of motion data; most rows include per-link contact flags and rendered preview videos.
  • Practical solver design: a windowed trajectory optimization balances contact satisfaction, body tracking, foot support and smoothness under joint limits (and outputs QC flags that mark wrist pinning, excessive trunk folding, or other failure modes).
  • Multi-source and extendable: assembled from five HOI capture datasets (OMOMO, ParaHome, NeuralDome, CoRoleHOI, IMHD²), supports collaborative two‑robot clips, and the pipeline can accept contact reconstructions from monocular video or new object meshes.
  • Reproducible assets and licensing: code and robot models are BSD-3-Clause; trajectories and derived assets are released under CC BY-NC-SA 4.0 (non-commercial, share-alike), with OMOMO object meshes redistributed under MIT.
Who it's for — and trade-offs

Great fit if you need realistic, contact-aware humanoid interaction references for embodied policy training, imitation learning, or simulation-based validation of whole‑body manipulation and loco‑manipulation. The dataset is especially useful when contact timing/locations must survive scaling across object sizes or robot morphologies.

Look elsewhere if you require permissive commercial licensing (the corpus is CC BY‑NC‑SA), need fine-finger grasp annotations (the robots use palm/contact links rather than per‑finger grasp detail), or require raw human mocap only (these are retargeted, robot‑centric trajectories and thus derivative of source captures).

Practical notes
  • Columns include per-frame joint angles, floating base pose, object 6‑DoF poses, per‑link contact flags, and a small preview video per row; the repo provides schema docs and a Hugging Face 3D viewer for clip playback.
  • QC metrics are provided per (motion, robot) pair; failing rows are retained but flagged so users can filter for the curated subset.
  • Because contacts are encoded in the object frame, the same clip can be re‑solved at different object scales or retargeted to additional robots with the same pipeline.

Information

  • Websitehuggingface.co
  • OrganizationsRobotic Systems Lab, ETH Zürich
  • AuthorsJihwan Shin, Adrià López Escoriza, Junzhe He, Matthias Heyrman, Marco Hutter
  • Published date2026/09/18

More Items

Hugging Face

Collection of 1.44M unique Turkish voice‑assistant sentences (≈1,819 hours estimated), normalized for TTS and organized by service domains (appointments, banking, e‑commerce). Designed for training and evaluating TTS and text-generation models; licensed CC BY 4.0 with required attribution.

Hugging Face

Provides a monthly Parquet snapshot of ~5.6 billion public TikTok videos (2014–Oct 2026), including captions, hashtags, sounds, engagement metrics and TikTok Shop links. Designed for large-scale querying (DuckDB/Pandas/Polars); licensed CC BY-NC 4.0 for research and personal use.

Lets general-purpose vision-language models directly command robots via a compact mid-level action interface and asynchronous monitoring, enabling zero-shot manipulation without task-specific policy training; demonstrates strong sim benchmarks and real xArm6 transfer.