Most long-horizon video generators struggle because carrying the entire generated history becomes costly and noisy; the right historical evidence is not the one most relevant to the present but the one that will matter for the near future. FrameMorrow's core insight is to anticipate "what will matter next" with a compact representation and use that anticipation to pick explicit historical frames to carry forward — instead of heuristics based on present content or preserving all past states.
Key Findings
- A small causal transformer predicts M=4 prospective tokens that summarize future information needs from the eligible history, recent context, and rollout condition; these tokens act as attention queries to rank candidate historical frames and select top-K explicit frames for conditioning.
- Training uses future-grounded ranking distillation: a frozen DINOv2 teacher ranks historical frames by similarity to realized continuation frames, and the selector matches those rankings with listwise and margin-based pairwise losses. At inference the selector runs without teacher supervision, regenerating tokens from available inputs.
- Across five benchmarks and 11 generators (including closed-source models), FrameMorrow consistently improves long-range consistency, visual quality, and action alignment while adding only ~2% inference overhead. Ablations show future-grounded supervision provides most gains: removing it drops consistency improvement from 4.35 to 1.25 points, while an oracle reaches 5.65 points (leaving a 1.30-point gap to the oracle). On action-conditioned world models, action-alignment improved by roughly 0.019, 0.023, and 0.029 across three backbones.
How it works
- Prospective tokens: a small causal Transformer F_ψ autoregressively produces M=4 tokens Q_t that summarize upcoming needs.
- Retrieval-by-attention: tokens query a pool of candidate frame embeddings (frozen encoder), logits are aggregated by a smooth maximum and the top-K frames are selected explicitly.
- Future-grounded distillation: during training a frozen DINOv2 encoder scores candidates against H realized continuation frames; the selector learns to match the teacher's ranking via listwise softmax cross-entropy plus a pairwise softplus margin loss.
- Plug-and-play conditioning: FrameMorrow returns explicit frames (not internal model states), so selected frames can be fed into diverse generators through their native conditioning interfaces, including closed-source backbones.
Who it's for and tradeoffs
Great fit if you build or evaluate long-horizon video generation, interactive multi-shot generation, or action-conditioned world models and need a lightweight, model-agnostic way to maintain relevant historical evidence for long rollouts. It is especially useful when generator internals are inaccessible or when retaining the full history is impractical. Look elsewhere if your generator already maintains an efficient, learned internal memory that outperforms explicit-frame conditioning in your specific domain, if you cannot afford any extra inference cost, or if the task requires predicting highly detailed future frames (FrameMorrow targets compact future cues rather than full future synthesis).