AIAny
Icon for item

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Groups visually grounded appearances of the same physical instance into persistent, retrievable “biographies” so agents can follow objects across hours or days for long-video question answering. Links identity-aware observations to episodic context and visual evidence; improves EgoLifeQA to 72.0% and increases evidence-window reach from 37.6% to 58.9%.

Introduction

Long-video QA often fails not because moments are missing, but because the same physical object or person is described differently across time and thus treated as distinct. The core insight of this paper is that organizing memory around grounded physical instances—rather than solely around chronologies or text captions—lets retrieval follow an entity’s life across clips and surface the exact encounters that matter for a question.

Key Findings
  • Grounded Entity Biographies (GEB): a memory representation that groups visually grounded observations of a single physical instance into a temporally ordered biography while preserving each encounter's episodic context and source frames. This preserves who interacted with the object and where each observation came from.
  • Retrieval mechanism: a controller can enter a retrieved encounter, follow same-instance identity edges to other encounters, and return biography excerpts plus episodic evidence to the answer model, enabling follow-up search targeted at the same inferred instance.
  • Empirical gains: across four long-video benchmarks (day- and week-long recordings), GEB improves multiple-choice and open-ended QA. On EgoLifeQA it achieves 72.0% accuracy (vs. 67.6% from the prior best), and the fraction of questions whose evidence window reaches the answering context rises from 37.6% to 58.9%.
  • Ablations: gains rely on grounded identity association and biography reading; adding extra textual descriptions without instance linking recovers only part of the benefit, and removing same-instance edges or episodic provenance reduces accuracy.
Who it's for & Trade-offs

Great fit if you build multimodal agents, wearable/egocentric assistants, or retrieval-augmented systems that must resolve persistent object/person identity across long horizons. GEB is particularly useful when identity mismatches between captions hamper retrieval or when evidence must be traced back to visual frames. Look elsewhere if your data is already perfectly labeled with persistent IDs (e.g., strong 3D reconstruction or tracking across all sessions) or if storage/compute for maintaining identity graphs and source-frame indices is prohibitive for your deployment; GEB adds indexing and association complexity compared to plain clip-level caption memories.

Where it fits

GEB sits between clip-centric caption memories and full 4D scene reconstructions: it avoids heavy geometric alignment while delivering object-centric persistence that improves retrieval and reasoning in long-horizon video QA settings.

Information

  • Websitearxiv.org
  • OrganizationsUniversity of Illinois Urbana-Champaign, Amazon.com, Inc.
  • AuthorsHui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia, Ying Chen, Alexander Schwing, Gang Hua
  • Published date2026/09/29

More Items

Uses 360° panoramic observations to improve vision-and-language navigation by predicting longer action sequences, confidence-guided execution, and combined semantic–geometric panorama features; yields large SR gains on R2R-CE and RxR-CE Val-Unseen.

Trains a single unified multimodal model with reinforcement learning to perform end-to-end self-reflection and iterative image repair — jointly learning the diagnostic (textual) reflection and the flow-based image revisions so credit propagates across rounds without an external verifier.

Uses diffusion-model forking moments as a proxy for perceptual distance to automatically generate pointwise reference-grounded labels, enabling annotation-free training of reference-based image quality assessment metrics.