Adaptive Keyframe Sampling for Long Video Understanding
Xi Tang, Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, Qixiang Ye
摘要
Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. However, when the visual input changes from a single image to a long video, the above paradigm encounters difficulty because the vast amount of video tokens has significantly exceeded the maximal capacity of MLLMs. Therefore, existing video-based MLLMs are mostly established upon sampling a small portion of tokens from input data, which can cause key information to be lost and thus produce incorrect answers. This paper presents a simple yet effective algorithm named Adaptive Keyframe Sampling (AKS). It inserts a plug-and-play module known as keyframe selection, which aims to maximize the useful information with a fixed number of video tokens. We formulate keyframe selection as an optimization involving (1) the relevance between the keyframes and the prompt, and ( 2 ) the coverage of the keyframes over the video, and present an adaptive algorithm to approximate the best solution. Experiments on two long video understanding benchmarks validate that AKS improves video QA accuracy (beyond strong baselines) upon selecting informative keyframes. Our study reveals the importance of information pre-filtering in video-based MLLMs. Our codes are available at https://github.com/ncTimTang/AKS * Equal contribution. A: The panda is eating food. A: The panda is rolling down. Q: What does the panda do in the video? Uniform Sampling Adaptive Keyframe Sampling MLLM MLLM Q: After entering the museum, where did the short-haired woman in a black coat, black long skirt, and black mask visit first? Q: A white-lettered title says 'How much data do we need?'. The word 'dog' appears on the far right, When the image of a shaking dog appears, what changes occur in the black scene? Q: Two men wearing straw hats and grey clothes stand in a grass field holding long knives. What does the house behind them look like?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper56
- WorldMM: Dynamic Multimodal Memory Agent for Long Video ReasoningWoongyeong Yeo, Kangsan Kim, Jaehong Yoon, Sung Ju HwangCVPR 2026 · 被引用 53 次
- Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal EvidenceJiahao Meng, Xiangtai Li, Haochen Wang, Tan Yue 等ICML 2026 · 被引用 43 次
- VideoITG: Multimodal Video Understanding with Instructed Temporal GroundingShihao Wang, Guo Chen, De-An Huang, Zhiqi Li 等CVPR 2026 · 被引用 35 次
- FOCUS: Efficient Keyframe Selection for Long Video UnderstandingZirui Zhu, Hailun Xu, Yang Luo, Yong Liu 等ICLR 2026 · 被引用 32 次
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang 等NeurIPS 2025 · 被引用 30 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- Threading Keyframe with Narratives: MLLMs as Strong Long Video ComprehendersBo Fang, Yuxin Song, Haoyuan Sun, Qiangqiang Wu 等ICLR 2026 · 被引用 13 次
- Efficient Frame Selection for Long Video Understanding via Reinforcement LearningYaxuan Qin, Hefei Li, Wenqi Mu, Yancheng HeCVPR 2026 · 被引用 6 次
- MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video UnderstandingWenhui Tan, Xiaoyi Yu, Jiaze Li, Yijing Chen 等CVPR 2026 · 被引用 6 次
- Q-Frame: Query-Aware Frame Selection and Multi-Resolution Adaptation for Video-LLMsShaojie Zhang, Jiahui Yang, Jianqin Yin, Zhenbo Luo 等ICCV 2025 · 被引用 15 次
- B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal TokensZhuqiang Lu, Zhenfei Yin, Mengwei He, Zhihui Wang 等ICCV 2025 · 被引用 3 次
