Long videos (minutes to hours) break common VLM pipelines because visual-token budgets explode and front-end decoding costs dominate wall-clock time. The paper's surprising starting point is that, under a strict token budget, trading per-frame resolution for denser temporal coverage often yields better end-task accuracy — while front-end decoding latency scales with the candidate pool size, not the final token budget.
Key Findings
- Dense low-resolution sampling beats sparse native-resolution sampling when visual-token budgets are matched, improving average accuracy by large margins. This means systems should consider temporal coverage over always using native per-frame resolution.
- Resolution-sensitive queries still benefit from a small set of selected high-resolution frames. LoHi mixes a dense low-res stream with sparse high-res images to capture both temporal context and fine-grained detail.
- Front-end decoding (building the candidate pool) dominates end-to-end runtime for hour-long videos; reducing candidate-pool size or using low-res streams can reduce decoding latency by up to 7x in their experiments.
- Empirical results: LoHi improves average accuracy by about 10.6 percentage points over a native-resolution baseline at matched token budgets and by ~5.2 points versus the strongest prior efficiency method across three long-video benchmarks.
What Sets It Apart
- Training-free, single-pass design: integrates directly into an existing VLM's image and video pathways without retraining; practical for deploying on top of off-the-shelf backbones.
- Two concrete selection strategies: LoHi-Anchor uses codec I-frame metadata to pick high-res frames cheaply; LoHi-SemDiv selects frames by query relevance and visual diversity using CLIP features, balancing informativeness and redundancy.
- Cost-aware insight: distinguishes selector-side candidate-pool costs from answerer-side token budgets and demonstrates that jointly choosing pool size and per-frame resolution yields better tradeoffs than fixing one dimension.
Who It's For — Fit & Tradeoffs
Great fit if you build or evaluate vision-language systems for long-form video QA, summarization, or retrieval where processing every high-res frame is infeasible. LoHi is practical for researchers and engineers seeking immediate gains without model retraining and for systems where decoding latency is a key bottleneck. Look elsewhere if your task requires consistent per-frame high-fidelity detail across the whole video (e.g., tiny-object tracking, per-frame super-resolution), or if reliable I-frame metadata and low-res encodings are unavailable; mixing resolutions also introduces pipeline complexity for downstream fusion and may require tuning selection thresholds per dataset.