AIAny
Icon for item

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

Introduces LoHi, a training-free, single-pass method that mixes dense low-resolution video streams with sparse high-resolution frames to improve long-video vision-language model accuracy under strict token budgets while cutting front-end decoding latency.

Introduction

Long videos (minutes to hours) break common VLM pipelines because visual-token budgets explode and front-end decoding costs dominate wall-clock time. The paper's surprising starting point is that, under a strict token budget, trading per-frame resolution for denser temporal coverage often yields better end-task accuracy — while front-end decoding latency scales with the candidate pool size, not the final token budget.

Key Findings
  • Dense low-resolution sampling beats sparse native-resolution sampling when visual-token budgets are matched, improving average accuracy by large margins. This means systems should consider temporal coverage over always using native per-frame resolution.
  • Resolution-sensitive queries still benefit from a small set of selected high-resolution frames. LoHi mixes a dense low-res stream with sparse high-res images to capture both temporal context and fine-grained detail.
  • Front-end decoding (building the candidate pool) dominates end-to-end runtime for hour-long videos; reducing candidate-pool size or using low-res streams can reduce decoding latency by up to 7x in their experiments.
  • Empirical results: LoHi improves average accuracy by about 10.6 percentage points over a native-resolution baseline at matched token budgets and by ~5.2 points versus the strongest prior efficiency method across three long-video benchmarks.
What Sets It Apart
  • Training-free, single-pass design: integrates directly into an existing VLM's image and video pathways without retraining; practical for deploying on top of off-the-shelf backbones.
  • Two concrete selection strategies: LoHi-Anchor uses codec I-frame metadata to pick high-res frames cheaply; LoHi-SemDiv selects frames by query relevance and visual diversity using CLIP features, balancing informativeness and redundancy.
  • Cost-aware insight: distinguishes selector-side candidate-pool costs from answerer-side token budgets and demonstrates that jointly choosing pool size and per-frame resolution yields better tradeoffs than fixing one dimension.
Who It's For — Fit & Tradeoffs

Great fit if you build or evaluate vision-language systems for long-form video QA, summarization, or retrieval where processing every high-res frame is infeasible. LoHi is practical for researchers and engineers seeking immediate gains without model retraining and for systems where decoding latency is a key bottleneck. Look elsewhere if your task requires consistent per-frame high-fidelity detail across the whole video (e.g., tiny-object tracking, per-frame super-resolution), or if reliable I-frame metadata and low-res encodings are unavailable; mixing resolutions also introduces pipeline complexity for downstream fusion and may require tuning selection thresholds per dataset.

Information

  • Websitearxiv.org
  • OrganizationsUniversity of Central Florida, Meta Reality Labs, Axon
  • AuthorsSixun Dong, Wei Li, Andong Deng, Qi Qian, Victor Zhu, Zhengping Ji, Chen Chen
  • Published date2026/10/03

More Items

Generates 5-second text- or image-conditioned videos with synchronized 44 kHz audio (including lip-sync) and built-in super-resolution to 1920×1080; available in Lite (3B) and Pro (29B) variants with code and checkpoints released under an MIT license.

Predicts compact 'prospective tokens' that summarize upcoming information needs and uses them to select a small set of past frames for conditioning long-horizon video generation, improving long-range consistency, visual quality, and action alignment while remaining plug-and-play across diverse generators.

Lets general-purpose vision-language models directly command robots via a compact mid-level action interface and asynchronous monitoring, enabling zero-shot manipulation without task-specific policy training; demonstrates strong sim benchmarks and real xArm6 transfer.