ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
Tanveer Hannan, Md Mohaiminul Islam, Jindong Gu, Thomas Seidl, Gedas Bertasius
摘要
Large language models (LLMs) excel at retrieving information from lengthy text, but their vision-language counterparts (VLMs) face difficulties with hour-long videos, especially for temporal grounding. Specifically, these VLMs are constrained by frame limitations, often losing essential temporal details needed for accurate event localization in extended video content. We propose ReVisionLLM, a recursive vision-language model designed to locate events in hour-long videos. Inspired by human search strategies, our model initially targets broad segments of interest, progressively revising its focus to pinpoint exact temporal boundaries. Our model can seamlessly handle videos of vastly different lengths-from minutes to hours. We also introduce a hierarchical training strategy that starts with short clips to capture distinct events and progressively extends to longer videos. To our knowledge, ReVision-LLM is the first VLM capable of temporal grounding in hour-long videos, outperforming previous state-of-the-art methods across multiple datasets by a significant margin (e.g., +2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingYe Wang, Ziheng Wang, Boshen Xu, Yang Du 等NeurIPS 2025 · 被引用 143 次
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video ReasoningHaoji Zhang, Xin Gu, Jiawen Li, Chixiang Ma 等CVPR 2026 · 被引用 92 次
- 4DP-QA: Scalable QA for 4D Perception in Vision Language ModelsSeokju Cho, Abhishek Badki, Hang Su, Jindong Jiang 等CVPR 2026 · 被引用 1 次
- CACR: Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal ReasoningMuge Qi, Rong Fu, Pengbin Feng, Xianda Li 等ICML 2026
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
相关 Paper
- VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal GroundingYongxin Guo, Jingyu Liu, Mingda Li, Dingxin Cheng 等AAAI 2025 · 被引用 27 次
- HiTeA: Hierarchical Temporal Alignment for Training-Free Long-Video Temporal GroundingXinyi Xu, Hongsong Wang, Guo-Sen Xie, Caifeng Shan 等ICLR 2026
- VTimeLLM: Empower LLM to Grasp Video MomentsBin Huang, Xin Wang, Hong Chen, Zihan Song 等CVPR 2024
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang 等NeurIPS 2025 · 被引用 30 次
- Enrich and Detect: Video Temporal Grounding With Multimodal LlmsShraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa 等ICCV 2025 · 被引用 4 次
