Re-thinking Temporal Search for Long-Form Video Understanding
Jinhui Ye, Zihan Wang, Haosen Sun, Keshigeyan Chandrasegaran, Zane Durante, Cristóbal Eyzaguirre, Yonatan Bisk, Juan Carlos Niebles, Ehsan Adeli, Li Fei-Fei, Jiajun Wu, Manling Li
摘要
Efficiently understanding long-form videos remains a significant challenge in computer vision. In this work, we revisit temporal search paradigms for long-form video understanding and address a fundamental issue pertaining to all state-of-the-art (SOTA) long-context vision-language models (VLMs). Our contributions are twofold: First, we frame temporal search as a Long Video Haystack problem -finding a minimal set of relevant frames (e.g., one to five) from tens of thousands based on specific queries. Upon this formulation, we introduce LV-HAYSTACK, the first dataset with 480 hours of videos, 15,092 human-annotated instances for both training and evaluation aiming to improve temporal search quality and efficiency. Results on LV-HAYSTACK highlight a significant research gap in temporal search capabilities, with current SOTA search methods only achieving 2.1% temporal F 1 score on the LONGVIDEOBENCH subset.
Next, inspired by visual search in images, we propose a lightweight temporal search framework, T* that reframes costly temporal search as spatial search. T* leverages powerful visual localization techniques commonly used in images and introduces an adaptive zooming-in mechanism that operates across both temporal and spatial dimensions. Extensive experiments show that integrating T* with existing methods significantly improves SOTA long-form video understanding. Under an inference budget of 32 frames, T* improves GPT-4o's performance from 50.5% to 53.1% and LLaVA-OneVision-OV-72B's performance from 56.5% to 62.4% on the LONGVIDEOBENCH XL subset. Our code, benchmark, and models are provided in the Supplementary material. memory based large language models. arXiv preprint
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language ModelsGuo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang 等NeurIPS 2025 · 被引用 69 次
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video UnderstandingYiyang Zhou, Yangfan He, Yaofeng Su, Siwei Han 等NeurIPS 2025 · 被引用 55 次
- Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video UnderstandingWeiyu Guo, Ziyang Chen, Shaoguang Wang, JianXiang He 等NeurIPS 2025 · 被引用 35 次
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang 等NeurIPS 2025 · 被引用 30 次
- VideoChat-A1: Thinking with Long Videos by Chain-of-Shot ReasoningZikang Wang, Boyu Chen, Zhengrong Yue, Yi Wang 等AAAI 2026 · 被引用 25 次
它引用的顶会 Paper42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
相关 Paper
- Temporal Chain of Thought: Long-Video Understanding by Thinking in FramesAnurag Arnab, Ahmet Iscen, Mathilde Caron, Alireza Fathi 等NeurIPS 2025 · 被引用 31 次
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun 等CVPR 2024 · 被引用 83 次
- Visual Context Window Extension: A New Perspective for Long Video UnderstandingHongchen Wei, Zhenzhong ChenACM MM 2025
- One Token per Highly Selective Frame: Towards Extreme Compression for Long Video UnderstandingZheyu Zhang, Ziqi Pang, Shixing Chen, Xiang Hao 等NeurIPS 2025 · 被引用 5 次
- Sharper and Faster mean Better: Towards More Efficient Vision-Language Model for Hour-scale Long Video UnderstandingDaoze Zhang, Yuze Zhao, Jintao Huang, Yingda ChenACL 2025 · 被引用 3 次
