AIAny
Icon for item

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

Turns live video into reusable textual memory and timely responses by training a streaming video LLM to proactively generate time-grounded captions and event summaries. Key components: Proactive Hierarchical Caption Memory (PHCM) for multi-scale records and Proactive State Transition Learning (PSTL) to balance response timing; trained on the OneStreamer-1M streaming dataset.

Introduction

Most streaming video systems either keep long visual histories (expensive) or only use a short recent window (losing past facts). The core insight of this paper is that a streaming video LLM can proactively turn observed evidence into compact, time-grounded textual records that serve as reusable factual memory — supporting later queries without revisiting raw frames.

Key Findings
  • Proactive Hierarchical Caption Memory (PHCM): generates dense, timestamped local-detail captions for recent visual content and sparser summaries of completed events, so the model accumulates textual facts that extend the effective context beyond a short visual window.
  • Proactive State Transition Learning (PSTL): reduces dominance of repeated “waiting” states by preserving supervision at anchored outputs and selectively supervising representative state-change and persistence tokens, improving timely response learning while supervising only a fraction of state tokens.
  • Streaming data synthesis and OneStreamer-1M: converts offline video annotations into evidence-aligned streaming caption and QA targets; OneStreamer-1M contains over one million streaming records used to train the model.
  • Empirical performance: a 4B OneStreamer model outperforms several compared baselines on eight streaming video understanding benchmarks and shows that retaining generated captions improves historical QA without hurting realtime perception.
Who it's for and trade-offs

Great fit if you need an LLM-based streaming video agent that must answer questions about past events without storing full frame histories, or if you want a single causal generation interface that unifies perception, memory formation, and response timing. Look elsewhere if your application requires lossless visual replay (e.g., fine-grained pixel-level forensic analysis) or if you cannot tolerate any risk from model-generated memory inaccuracies — the approach trades compact textual memory for potentially imperfect summarization of past frames.

How it works (brief)

The system processes incoming clips with a vision encoder and projects tokens into the LLM embedding space; a Recent-NN FIFO retains only the latest visual tokens and interleaves them with accumulated textual records. During training, causal streaming caption and QA targets are aligned to evidence-available times so the model learns both what to record and when to output user-visible responses within a single causal sequence.

Information

  • Websitearxiv.org
  • AuthorsXiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li, Qingyi Si, Dingyu Yao, Changlian Ma, Haoran Chen, Xinyu Chen …
  • Published date2026/10/01

More Items

Proposes a forward-process RL method that dynamically localizes gradient updates and adapts multi-reward coordination for joint audio–video diffusion models. Uses bidirectional cross-attention responses for token/layer routing and preference-preserving reweighting to balance competing objectives, improving modality quality, alignment, and synchronization.

Uses a multimodal model's own critiques as privileged context and applies on-policy self-distillation over diffusion sampling trajectories to internalize corrective guidance, improving text-to-image generation without an external teacher; shows measurable gains on GenEval and GenEval2.

Groups visually grounded appearances of the same physical instance into persistent, retrievable “biographies” so agents can follow objects across hours or days for long-video question answering. Links identity-aware observations to episodic context and visual evidence; improves EgoLifeQA to 72.0% and increases evidence-window reach from 37.6% to 58.9%.