Long-video QA often fails not because moments are missing, but because the same physical object or person is described differently across time and thus treated as distinct. The core insight of this paper is that organizing memory around grounded physical instances—rather than solely around chronologies or text captions—lets retrieval follow an entity’s life across clips and surface the exact encounters that matter for a question.
Key Findings
- Grounded Entity Biographies (GEB): a memory representation that groups visually grounded observations of a single physical instance into a temporally ordered biography while preserving each encounter's episodic context and source frames. This preserves who interacted with the object and where each observation came from.
- Retrieval mechanism: a controller can enter a retrieved encounter, follow same-instance identity edges to other encounters, and return biography excerpts plus episodic evidence to the answer model, enabling follow-up search targeted at the same inferred instance.
- Empirical gains: across four long-video benchmarks (day- and week-long recordings), GEB improves multiple-choice and open-ended QA. On EgoLifeQA it achieves 72.0% accuracy (vs. 67.6% from the prior best), and the fraction of questions whose evidence window reaches the answering context rises from 37.6% to 58.9%.
- Ablations: gains rely on grounded identity association and biography reading; adding extra textual descriptions without instance linking recovers only part of the benefit, and removing same-instance edges or episodic provenance reduces accuracy.
Who it's for & Trade-offs
Great fit if you build multimodal agents, wearable/egocentric assistants, or retrieval-augmented systems that must resolve persistent object/person identity across long horizons. GEB is particularly useful when identity mismatches between captions hamper retrieval or when evidence must be traced back to visual frames. Look elsewhere if your data is already perfectly labeled with persistent IDs (e.g., strong 3D reconstruction or tracking across all sessions) or if storage/compute for maintaining identity graphs and source-frame indices is prohibitive for your deployment; GEB adds indexing and association complexity compared to plain clip-level caption memories.
Where it fits
GEB sits between clip-centric caption memories and full 4D scene reconstructions: it avoids heavy geometric alignment while delivering object-centric persistence that improves retrieval and reasoning in long-horizon video QA settings.