Real-time robot control faces a fundamental tradeoff: more visual history can disambiguate motion and task progress, but processing extra frames often increases latency and degrades closed-loop responsiveness. This paper's core insight is counterintuitive yet actionable: longer visual context yields substantial control gains only when the video foundation is autoregressively pretrained to preserve causal history→future structure, and those gains require co-designed runtime optimizations to keep inference real-time.
Key Findings
- Autoregressive (AR) video pretraining makes history usable, not just accessible: on RoboCasa GR-1, extending history from 0.0 to 19.2s raises success from 63.3% to 78.7%; bidirectional pretraining shows no net gain. This demonstrates that pretraining temporal factorization matters for control.
- Robot-domain AR pretraining (LongLive2.0-Robot) trained on ≈10k window-equivalent hours further improves peak performance on benchmarks including LIBERO-Long and GR-1.
- System co-design preserves real-time behavior: streaming VAE observation encoding, asynchronous execution, NVFP4 quantization, KV reuse and device-specific kernel tuning reduce end-to-end latency so the full predict-then-act pipeline runs in 107.4 ms per action chunk on an RTX 5090 (including future-video latent prediction).
- Practical deployment: achieves state-of-the-art results on LIBERO-Long, RoboTwin 2.0 and DOMINO; real-robot demos include 95% success on dynamic cup stacking and sustained multi-second tasks on Unitree G1 and YAM.
Who it's for and tradeoffs
Great fit if you build robotic manipulation systems that need long-horizon, memory-informed execution and you can invest in matching pretraining data plus hardware optimizations. It prioritizes causal video pretraining and low-latency inference; if you cannot perform AR-style pretraining on relevant robot/egocentric data or cannot deploy device-specific acceleration, the context scaling benefits will be limited. The approach focuses on latent prediction and action coupling rather than pixel decoding, so visual fidelity and some pixel-level diagnostics are deprioritized.
Where it fits
Long-WAM sits between pure model-free controllers and high-level planners: it acts as a memory-aware executor that benefits downstream planning (example: pairing with an LLM planner increased composite-task success). Methodologically, it emphasizes combining large-scale AR video pretraining with systems engineering (quantization, streaming encoders, async execution) to convert model capacity into usable control improvements.