FrameOracle: Learning What to See and How Much to See in Videos
Chaoyu Li, Tianzhi Li, Fei Tao, ZHENYU ZHAO, Ziqian Wu, Maozheng Zhao, Juntong Song, Cheng Niu, Pooyan Fazli
Abstract
Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of frames. Existing frame sampling strategies, such as uniform or fixed-budget selection, fail to adapt to variations in content density or task complexity. To address this, we present FrameOracle, a lightweight, plug-and-play module that predicts both (1) which frames are most relevant to a given query and (2) how many frames are needed. FrameOracle is trained via a curriculum that progresses from weak proxy signals, such as crossmodal similarity, to stronger supervision with FrameOracle-41K, the first large-scale VideoQA dataset with validated keyframe annotations specifying minimal sufficient frames per question. Extensive experiments across five VLMs and six benchmarks show that FrameOracle reduces 16frame inputs to an average of 10.4 frames without accuracy loss. When starting from 64-frame candidates, it reduces inputs to 13.9 frames on average while improving accuracy by 1.5%, achieving state-of-the-art efficiency-accuracy trade-offs for scalable video understanding. Our project is available here: https://people-robots.github. io/frameoracle .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 281 citations
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 173 citations
Related papers
- Frame-Voyager: Learning to Query Frames for Video Large Language ModelsSicheng Yu, Chengkai Jin, Huanyu Wang, Zhenghao Chen et al.ICLR 2025
- QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video ComprehensionYongdong Luo, Wang Chen, Weizhong Huang, Shukang Yin et al.AAAI 2026
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question AnsweringYuanhao Zou, Shengji Jin, Andong Deng, Youpeng Zhao et al.ICLR 2026
- Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal AdaptationYachun Mi, Yu Li, Yanting Li, Chen Hui et al.ICML 2026
- M-LLM Based Video Frame Selection for Efficient Video UnderstandingKai Hu, Feng Gao, Xiaohan Nie, Peng Zhou et al.CVPR 2025
