Compresses a geometry-foundation model’s multi-level features into a compact latent that decodes jointly to RGB, depth, cameras and point maps — enabling a conditional flow to generate 3D-consistent video and novel views with measurably improved coherence.