Uses 360° panoramic observations to improve vision-and-language navigation by predicting longer action sequences, confidence-guided execution, and combined semantic–geometric panorama features; yields large SR gains on R2R-CE and RxR-CE Val-Unseen.
Groups visually grounded appearances of the same physical instance into persistent, retrievable “biographies” so agents can follow objects across hours or days for long-video question answering. Links identity-aware observations to episodic context and visual evidence; improves EgoLifeQA to 72.0% and increases evidence-window reach from 37.6% to 58.9%.