Most streaming video systems either keep long visual histories (expensive) or only use a short recent window (losing past facts). The core insight of this paper is that a streaming video LLM can proactively turn observed evidence into compact, time-grounded textual records that serve as reusable factual memory — supporting later queries without revisiting raw frames.
Key Findings
- Proactive Hierarchical Caption Memory (PHCM): generates dense, timestamped local-detail captions for recent visual content and sparser summaries of completed events, so the model accumulates textual facts that extend the effective context beyond a short visual window.
- Proactive State Transition Learning (PSTL): reduces dominance of repeated “waiting” states by preserving supervision at anchored outputs and selectively supervising representative state-change and persistence tokens, improving timely response learning while supervising only a fraction of state tokens.
- Streaming data synthesis and OneStreamer-1M: converts offline video annotations into evidence-aligned streaming caption and QA targets; OneStreamer-1M contains over one million streaming records used to train the model.
- Empirical performance: a 4B OneStreamer model outperforms several compared baselines on eight streaming video understanding benchmarks and shows that retaining generated captions improves historical QA without hurting realtime perception.
Who it's for and trade-offs
Great fit if you need an LLM-based streaming video agent that must answer questions about past events without storing full frame histories, or if you want a single causal generation interface that unifies perception, memory formation, and response timing. Look elsewhere if your application requires lossless visual replay (e.g., fine-grained pixel-level forensic analysis) or if you cannot tolerate any risk from model-generated memory inaccuracies — the approach trades compact textual memory for potentially imperfect summarization of past frames.
How it works (brief)
The system processes incoming clips with a vision encoder and projects tokens into the LLM embedding space; a Recent-NN FIFO retains only the latest visual tokens and interleaves them with accumulated textual records. During training, causal streaming caption and QA targets are aligned to evidence-available times so the model learns both what to record and when to output user-visible responses within a single causal sequence.