Lune

CVPR2025Top-tier venue

Flexible Frame Selection for Efficient Video Reasoning

Shyamal Buch, Arsha Nagrani, Anurag Arnab, Cordelia Schmid

2025Year
14Top-tier citations

Abstract

Video-language models have shown promise for addressing a range of multimodal tasks for video reasoning, such as video question-answering. However, the inherent computational challenges of processing long video data and increasing model sizes have led to standard approaches that are limited by the number of frames they can process. In this work, we propose the Flexible Frame Selector (FFS), a learnable policy model with a new flexible selection operation, that helps alleviate input context restrictions by enabling video-language models to focus on the most informative frames for the downstream multimodal task, without adding undue processing cost. Our method differentiates from prior work due to its learnability, efficiency, and flexibility. We verify the efficacy of our method on standard video-question answering and reasoning benchmarks, and observe our model can maintain or improve base videolanguage model accuracy while significantly reducing the number of downstream processed frames.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext beaceeb4-2142-48aa-8f9b-d9769d880cbe

Cited by top-tier papers14

Ask how each one uses it

Builds on47

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines