AIAny
Icon for item

Long-WAM: Scaling the Context of World-Action Models

Scales visual history for real-time robot control by combining autoregressive video pretraining with a streaming, asynchronous predict-then-act pipeline; shows longer context improves long-horizon manipulation and runs full inference in 107.4 ms per action chunk on an RTX 5090.

Introduction

Real-time robot control faces a fundamental tradeoff: more visual history can disambiguate motion and task progress, but processing extra frames often increases latency and degrades closed-loop responsiveness. This paper's core insight is counterintuitive yet actionable: longer visual context yields substantial control gains only when the video foundation is autoregressively pretrained to preserve causal history→future structure, and those gains require co-designed runtime optimizations to keep inference real-time.

Key Findings
  • Autoregressive (AR) video pretraining makes history usable, not just accessible: on RoboCasa GR-1, extending history from 0.0 to 19.2s raises success from 63.3% to 78.7%; bidirectional pretraining shows no net gain. This demonstrates that pretraining temporal factorization matters for control.
  • Robot-domain AR pretraining (LongLive2.0-Robot) trained on ≈10k window-equivalent hours further improves peak performance on benchmarks including LIBERO-Long and GR-1.
  • System co-design preserves real-time behavior: streaming VAE observation encoding, asynchronous execution, NVFP4 quantization, KV reuse and device-specific kernel tuning reduce end-to-end latency so the full predict-then-act pipeline runs in 107.4 ms per action chunk on an RTX 5090 (including future-video latent prediction).
  • Practical deployment: achieves state-of-the-art results on LIBERO-Long, RoboTwin 2.0 and DOMINO; real-robot demos include 95% success on dynamic cup stacking and sustained multi-second tasks on Unitree G1 and YAM.
Who it's for and tradeoffs

Great fit if you build robotic manipulation systems that need long-horizon, memory-informed execution and you can invest in matching pretraining data plus hardware optimizations. It prioritizes causal video pretraining and low-latency inference; if you cannot perform AR-style pretraining on relevant robot/egocentric data or cannot deploy device-specific acceleration, the context scaling benefits will be limited. The approach focuses on latent prediction and action coupling rather than pixel decoding, so visual fidelity and some pixel-level diagnostics are deprioritized.

Where it fits

Long-WAM sits between pure model-free controllers and high-level planners: it acts as a memory-aware executor that benefits downstream planning (example: pairing with an LLM planner increased composite-task success). Methodologically, it emphasizes combining large-scale AR video pretraining with systems engineering (quantization, streaming encoders, async execution) to convert model capacity into usable control improvements.

Information

  • Websitearxiv.org
  • OrganizationsNVIDIA, MIT, HKU, UCSD
  • AuthorsWei Huang, Bohan Zhang, Chenzhi Liu, Isabella Liu, Shuai Yang, Weian Mao, Luozhou Wang, Yicheng Xiao, Weifeng Lin, Qixin Hu …
  • Published date2026/10/07

More Items

Combines joint distribution distillation from a video teacher with marginal (frame-level) distillation from an image teacher to improve few-step video generation. Introduces LatentBridge to align incompatible latents and Latent Variation Sampling to distribute frame supervision, boosting per-frame visual quality and semantic alignment while largely preserving motion.

Introduces LoHi, a training-free, single-pass method that mixes dense low-resolution video streams with sparse high-resolution frames to improve long-video vision-language model accuracy under strict token budgets while cutting front-end decoding latency.

Generates 5-second text- or image-conditioned videos with synchronized 44 kHz audio (including lip-sync) and built-in super-resolution to 1920×1080; available in Lite (3B) and Pro (29B) variants with code and checkpoints released under an MIT license.