Omnimodal world model that jointly processes and generates text, images, video, audio, and action trajectories for physical AI. Uses a mixture-of-transformers to combine autoregressive reasoning and diffusion-based multimodal generation; released open-source with checkpoints, datasets and benchmarks for robotics and simulation.
Evaluates multimodal LLMs on streaming egocentric video for spatial intelligence using 1,680 human-annotated questions across 348 videos; organizes tasks into four hierarchical levels (perception → tracking → simulation → allocentric mapping) and highlights allocentric mapping as the main bottleneck.
Trains a GPT-style causal Transformer on a 2-billion-frame retargeted motion corpus to enable zero-shot whole-body motion tracking and control. By scaling both data and model capacity, it tracks highly dynamic behaviors while generalizing to unseen motions; accepted to CVPR 2026.
Simulates egocentric, embodied human–world interactions and enables customizable, self-evolving local scenes by defining anchor views and text-driven evolution. Uses exogenous viewpoints and full-body motion supervision to improve spatial grounding and interaction consistency.
A benchmark that evaluates interactive spatial reasoning for multimodal agents in realistic tasks. It unifies eight heterogeneous simulators under a simulator-agnostic protocol, provides 760 human-annotated tasks with vision-only partial observability, and uses text-based actions plus terminal-state verification to measure task success.
Synthesizes scalable, photoreal 3D Earth tiles from georeferenced satellite imagery using a generative 3D Gaussian Splatting representation; trained on urban reconstructions, it generates novel scenes at under 10 minutes/km² with hierarchical LOD for real-time web map visualization and Embodied AI use cases.
Provides 500+ hours of human whole-body teleoperation demonstrations for humanoid robot learning in real homes, with synchronized video, joint states, action traces and language annotations. Includes 23K+ episodes, fine-grained subtask labels, and raw ROS/MCAP plus compressed LeRobot formats.
Provides 500+ hours of human whole-body teleoperation recordings of a Unitree G1 in real homes, packaged in LeRobot v3.0 for robot learning. Contains 23K+ episodes, ~40M frames, multi-view 480p@30 video, 29-DoF states, actions and language annotations; CC BY 4.0 and large download size.
Provides 1000+ hours of high-precision optical motion-capture for humanoid robotics and embodied AI, including full-body skeleton, 20+DoF hands, object 6D, and multi-view video at 120 Hz. Sub-mm spatial accuracy, BVH/CSV/NPZ outputs and Unitree G1 retargets; ideal for imitation learning and sim-to-real, with some raw captures gated by license.
Learns, maintains, and runs unified world models for Physical AI using a cross-embodiment pretraining curriculum and a hybrid linear temporal-attention architecture. Emphasizes long-horizon state persistence, theoretical bounds on error accumulation, and deployment-aware low-latency inference for real-world embodied agents.
Language-conditioned robot policy that reuses a pretrained geometric foundation model and inserts a causal future predictor at an intermediate layer so the same backbone produces future 3D-aware features and action outputs, enabling geometry-aware temporal prediction with minimal architectural change.
Converts large-scale egocentric human videos into robot-format pseudo-action trajectories and introduces ACE-EGO-0, a VLA pretraining framework that unifies camera-space actions, morphology conditioning, and reliability-aware weighting to jointly learn from noisy human and high-quality robot data for improved robotic manipulation transfer.