AIAny
Icon for item

FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation

Predicts compact 'prospective tokens' that summarize upcoming information needs and uses them to select a small set of past frames for conditioning long-horizon video generation, improving long-range consistency, visual quality, and action alignment while remaining plug-and-play across diverse generators.

Introduction

Most long-horizon video generators struggle because carrying the entire generated history becomes costly and noisy; the right historical evidence is not the one most relevant to the present but the one that will matter for the near future. FrameMorrow's core insight is to anticipate "what will matter next" with a compact representation and use that anticipation to pick explicit historical frames to carry forward — instead of heuristics based on present content or preserving all past states.

Key Findings
  • A small causal transformer predicts M=4 prospective tokens that summarize future information needs from the eligible history, recent context, and rollout condition; these tokens act as attention queries to rank candidate historical frames and select top-K explicit frames for conditioning.
  • Training uses future-grounded ranking distillation: a frozen DINOv2 teacher ranks historical frames by similarity to realized continuation frames, and the selector matches those rankings with listwise and margin-based pairwise losses. At inference the selector runs without teacher supervision, regenerating tokens from available inputs.
  • Across five benchmarks and 11 generators (including closed-source models), FrameMorrow consistently improves long-range consistency, visual quality, and action alignment while adding only ~2% inference overhead. Ablations show future-grounded supervision provides most gains: removing it drops consistency improvement from 4.35 to 1.25 points, while an oracle reaches 5.65 points (leaving a 1.30-point gap to the oracle). On action-conditioned world models, action-alignment improved by roughly 0.019, 0.023, and 0.029 across three backbones.
How it works
  • Prospective tokens: a small causal Transformer F_ψ autoregressively produces M=4 tokens Q_t that summarize upcoming needs.
  • Retrieval-by-attention: tokens query a pool of candidate frame embeddings (frozen encoder), logits are aggregated by a smooth maximum and the top-K frames are selected explicitly.
  • Future-grounded distillation: during training a frozen DINOv2 encoder scores candidates against H realized continuation frames; the selector learns to match the teacher's ranking via listwise softmax cross-entropy plus a pairwise softplus margin loss.
  • Plug-and-play conditioning: FrameMorrow returns explicit frames (not internal model states), so selected frames can be fed into diverse generators through their native conditioning interfaces, including closed-source backbones.
Who it's for and tradeoffs

Great fit if you build or evaluate long-horizon video generation, interactive multi-shot generation, or action-conditioned world models and need a lightweight, model-agnostic way to maintain relevant historical evidence for long rollouts. It is especially useful when generator internals are inaccessible or when retaining the full history is impractical. Look elsewhere if your generator already maintains an efficient, learned internal memory that outperforms explicit-frame conditioning in your specific domain, if you cannot afford any extra inference cost, or if the task requires predicting highly detailed future frames (FrameMorrow targets compact future cues rather than full future synthesis).

Information

  • Websitearxiv.org
  • OrganizationsNational University of Singapore, Harbin Institute of Technology (Shenzhen)
  • AuthorsBo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan
  • Published date2026/09/30

More Items

Lets general-purpose vision-language models directly command robots via a compact mid-level action interface and asynchronous monitoring, enabling zero-shot manipulation without task-specific policy training; demonstrates strong sim benchmarks and real xArm6 transfer.

Turns live video into reusable textual memory and timely responses by training a streaming video LLM to proactively generate time-grounded captions and event summaries. Key components: Proactive Hierarchical Caption Memory (PHCM) for multi-scale records and Proactive State Transition Learning (PSTL) to balance response timing; trained on the OneStreamer-1M streaming dataset.

Proposes a forward-process RL method that dynamically localizes gradient updates and adapts multi-reward coordination for joint audio–video diffusion models. Uses bidirectional cross-attention responses for token/layer routing and preference-preserving reweighting to balance competing objectives, improving modality quality, alignment, and synchronization.