AIAny
Icon for item

OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs

Converts static 3D Gaussian Splatting scenes into endlessly looping 3D cinemagraphs by inferring plausible dynamics with a vision-language model, synthesizing a reference video, lifting it to multi-view videos, and fitting a Fourier-parameterized Periodic Deformation Field with a Grounded Drift Field—mask-free capture of deformation, object motion, and illumination changes.

Introduction

Most photorealistic 3D world models produce static, time-frozen scenes; adding believable, view-consistent motion without seams is hard because single-view videos are only approximately periodic and multi-view video syntheses are mutually inconsistent. OuroWorld's core insight is to use a generated, approximately looping reference video as weak supervision and distill it into a 4D Gaussian Splatting representation that is periodic by construction while tolerating cross-view inconsistencies during training.

Key Findings
  • Loop-by-construction deformation: representing temporal change as a Fourier-series Periodic Deformation Field guarantees seamless looping at inference, avoiding visible seam artifacts common when fitting non-periodic video outputs.
  • Robustness to inconsistent supervision: a Grounded Drift Field absorbs cross-view inconsistencies from multi-view generated videos during training and is discarded at inference, preserving visual sharpness and coherent motion.
  • Mask-free and general dynamics: unlike Eulerian flow methods constrained to fluid-like motion and requiring masks, the method models general deformation, rigid object motion, and illumination changes directly in the 4DGS representation.
  • Practical pipeline: a vision-language model proposes plausible dynamics and prompts a video generator for a reference clip; a 3D foundation model lifts that clip to a dynamic point cloud; inpainting completes multi-view videos; the 4DGS fitting stage consolidates everything into an endlessly looping 3D cinemagraph.
Who it's for and tradeoffs

Great fit if you need photorealistic, explorable 3D scenes that exhibit continuous, repeatable dynamics (e.g., virtual worlds, architectural visualizations, interactive cinematics) and you can afford the compute and data pipeline required for multi-stage video lifting and 4DGS fitting. Look elsewhere if you need real-time on-device synthesis, strict frame-by-frame physical accuracy, or if only low-resource, single-view motion augmentation is required.

Where it fits

Positioned between video diffusion-based motion synthesis and 3D scene representations: it leverages modern video generators and 3D foundation models to add temporality to high-quality static world models, bridging generative video research and 4D scene representations.

Methodology (brief)

Training supervision is created by (1) prompting a vision-language model to infer plausible cyclic dynamics for a chosen reference view, (2) generating an approximately looping reference video with a controllable video model, (3) lifting the reference video to a per-frame dynamic point cloud with a 3D foundation model and completing nearby viewpoints via a video inpainting model, and (4) fitting a canonical 3D Gaussian Splatting scene deformed by a Fourier-parameterized Periodic Deformation Field plus a temporary Grounded Drift Field to handle cross-view inconsistency. The Grounded Drift Field is used only during training and removed at inference so the final scene is inherently periodic and view-consistent.

Information

  • Websitearxiv.org
  • OrganizationsNational Yang Ming Chiao Tung University, Alaya Lab
  • AuthorsYou-Zhe Xie, Ting-Wei Chou, Yu-Hsuan Li, Kaipeng Zhang, Zhixiang Wang, Yu-Lun Liu
  • Published date2026/10/08

More Items

Predicts identity-preserving dense pixel correspondences between image pairs that violate spatio-temporal priors (e.g., edits and reference-guided generation). Fuses generative (FLUX2) and semantic (DINOv3) foundation representations with heterogeneous supervision and teacher-guided iterative refinement to generalize beyond classical optical-flow assumptions.

Generates synchronized egocentric video streams for multiple interacting agents in a shared environment, enforcing cross-view action consistency, shared environment memory, and consistent propagation of interaction-induced state changes — aimed at embodied AI, VR, and multi-agent vision research.

Lets a pretrained multimodal LLM interpret navigation requests and orchestrate motion via tool calls for generalist robot navigation across unfamiliar scenes. Key features: an agent harness with Navigation Skills, a unified visual-point interface, task-progress tracking, and tool-based motion execution without navigation-specific fine-tuning.