Evaluates whether generative world models maintain consistent, controllable, and physically plausible simulated environments under exploration, interaction, and intervention. Introduces a six-level W1–W6 capability taxonomy across three tracks (video, spatial, embodied) with human A/B Arena and automated metrics to measure behavioral correctness.
Surveys memory mechanisms for autoregressive video generation, framing memory as persistent historical information that influences future generation. Organizes work by Forms, Functions, Operations, Learning, and Evaluation, and synthesizes challenges for long-horizon consistency and memory-aware learning.