Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
Tianyuan Qu, Longxiang Tang, Bohao Peng, Senqiao Yang, Bei Yu, Jiaya Jia
摘要
The rise of Large Vision-Language Models (LVLMs) has significantly advanced video understanding. However, efficiently processing long videos remains a challenge due to the ``Sampling Dilemma'': low-density sampling risks missing critical information, while high-density sampling introduces redundancy. To address this issue, we introduce LSDBench, the first benchmark designed to evaluate LVLMs on long-video tasks by constructing high Necessary Sampling Density (NSD) questions, where NSD represents the minimum sampling density required to accurately answer a given question. LSDBench focuses on dense, short-duration actions to rigorously assess the sampling strategies employed by LVLMs. To tackle the challenges posed by high-NSD questions, we propose a novel Reasoning-Driven Hierarchical Sampling (RHS) framework, which combines global localization of question-relevant cues with local dense sampling for precise inference. Additionally, we develop a lightweight Semantic-Guided Frame Selector to prioritize informative frames, enabling RHS to achieve comparable or superior performance with significantly fewer sampled frames. Together, our LSDBench and RHS framework address the unique challenges of high-NSD long-video tasks, setting a new standard for evaluating and improving LVLMs in this domain. Our benchmark and evaluation codes has been released at: https://github.com/dvlab-research/LSDBench
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- VisionThink: Smart and Efficient Vision Language Model via Reinforcement LearningSenqiao Yang, Junyi Li, Xin Lai, Jinming Wu 等NeurIPS 2025 · 被引用 43 次
- FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame SpotlightingZefeng He, Xiaoye Qu, Yafu Li, Siyuan Huang 等ICLR 2026 · 被引用 34 次
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video ReasoningYang Ding, Xin Lai, Yizhen Zhang, Wei Li 等ICLR 2026 · 被引用 26 次
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand 等ICLR 2026 · 被引用 11 次
- Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video UnderstandingJialuo Li, Bin Li, Jiahao Li, Yan LuCVPR 2026 · 被引用 11 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
相关 Paper
- Select Less, Reason More: Prioritizing Evidence Purity for Video ReasoningXuchen Li, Xuzhao Li, Shiyu Hu, Kaiqi HuangCVPR 2026 · 被引用 5 次
- One Token per Highly Selective Frame: Towards Extreme Compression for Long Video UnderstandingZheyu Zhang, Ziqi Pang, Shixing Chen, Xiang Hao 等NeurIPS 2025 · 被引用 5 次
- LensWalk: Agentic Video Understanding by Planning How You See in VideosKeliang Li, Yansong Li, Hongze Shen, Mengdi Liu 等CVPR 2026 · 被引用 13 次
- VideoBrain: Learning Adaptive Frame Sampling for Long Video UnderstandingJunbo Zou, Ziheng Huang, Shengjie Zhang, Liwen Zhang 等ICML 2026 · 被引用 4 次
- Seeing Is Believing: Grounding Long-Video Understanding in Spatio-Temporal Visual EvidenceZhaoyang Wei, Guoliang Wang, Guohua Gao, Yanchao Hao 等AAAI 2026
