Compresses a geometry-foundation model’s multi-level features into a compact latent that decodes jointly to RGB, depth, cameras and point maps — enabling a conditional flow to generate 3D-consistent video and novel views with measurably improved coherence.
Surveys memory mechanisms for autoregressive video generation, framing memory as persistent historical information that influences future generation. Organizes work by Forms, Functions, Operations, Learning, and Evaluation, and synthesizes challenges for long-horizon consistency and memory-aware learning.