AIAny
Icon for item

AlayaWorld: Long-Horizon and Playable Video World Generation

Autoregressively synthesizes long-horizon, playable video worlds conditioned on current state and user actions for real-time interaction. Ships as an open-source, full-stack framework covering data preparation, model architectures, training, inference acceleration, and deployment for interactive generative worlds.

Introduction

Generative world models aim to replace laborious manual world authoring by synthesizing future observations conditioned on state and user actions. AlayaWorld demonstrates how a research prototype becomes a practical, reproducible toolkit: it packages model design, data pipelines, training recipes, inference optimizations, and evaluation into a single open-source stack for creating playable, long-horizon video worlds.

Key Findings
  • Trains on both gameplay recordings and real-world video to capture diverse visual styles and dynamics, so what: models generalize across stylized and photoreal scenes and support richer interactions than navigation-only systems.
  • Supports open-ended real-time interaction including navigation, combat-like actions, and event-driven effects, so what: users can perform mid-rollout actions that meaningfully change subsequent frames rather than only following prerecorded trajectories.
  • Provides reproducible pipelines and evaluation tools, so what: researchers can compare models on standardized inputs and reproduce training/inference steps instead of relying on bespoke demos.
  • Emphasizes deployment: inference acceleration and modular deployment components are included, so what: the project targets usable latency budgets for interactive applications, not just offline benchmarks.
Who It's For and Trade-offs

Great fit if you want a reproducible starting point for research or prototypes in interactive generative worlds, need integrated data-to-deploy tooling, or want to study action-conditioned video generation at scale. Look elsewhere if you require provably accurate physics, production-grade large-scale multiplayer backend systems, or very low-VRAM inference on commodity devices—these areas still demand substantial engineering and compute. The code and models lower the barrier to experimentation but expect significant compute for training and careful dataset curation to avoid biases.

Information

  • Websitearxiv.org
  • AuthorsAlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Ruicong Liu …
  • Published date2026/07/07

More Items

Introduces LoHi, a training-free, single-pass method that mixes dense low-resolution video streams with sparse high-resolution frames to improve long-video vision-language model accuracy under strict token budgets while cutting front-end decoding latency.

Generates 5-second text- or image-conditioned videos with synchronized 44 kHz audio (including lip-sync) and built-in super-resolution to 1920×1080; available in Lite (3B) and Pro (29B) variants with code and checkpoints released under an MIT license.

Hugging Face

Generates humanoid robot motion references that preserve object contact locations/timing by solving windowed trajectory optimizations against contact targets in the object frame. Releases retargeted trajectories for two Unitree robots across 75 objects (≈13.9k robot–motion pairs); CC BY‑NC‑SA 4.0.