VTimeLLM: Empower LLM to Grasp Video Moments
Bin Huang, Xin Wang, Hong Chen, Zihan Song, Wenwu Zhu
Abstract
Large language models (LLMs) have shown remarkable text understanding capabilities, which have been extended as Video LLMs to handle video data for comprehending visual details. However, existing Video LLMs can only provide a coarse description of the entire video, failing to capture the precise start and end time boundary of specific events. In this paper, we solve this issue via proposing VTimeLLM, a novel Video LLM designed for fine-grained video moment understanding and reasoning with respect to time boundary. Specifically, our VTimeLLM adopts a boundary-aware three-stage training strategy, which respectively utilizes image-text pairs for feature alignment, multiple-event videos to increase temporalboundary awareness, and high-quality video-instruction tuning to further improve temporal understanding ability as well as align with human intents. Extensive experiments demonstrate that in fine-grained time-related comprehension tasks for videos such as Temporal Video Grounding and Dense Video Captioning, VTimeLLM significantly outperforms existing Video LLMs. Besides, benefits from the fine-grained temporal understanding of the videos further enable VTimeLLM to beat existing Video LLMs in video dialogue benchmark, showing its superior cross-modal understanding and reasoning abilities. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01011ddc-d55d-4dd6-b63d-af02e0bd33d6Cited by top-tier papers102
- Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingYe Wang, Ziheng Wang, Boshen Xu, Yang Du et al.NeurIPS 2025 · 143 citations
- StreamForest: Efficient Online Video Understanding with Persistent Event MemoryXiangyu Zeng, Kefan Qiu, Qingyu Zhang, Xinhao Li et al.NeurIPS 2025 · 79 citations
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming AssistantHaibo Wang, Bo Feng, Zhengfeng Lai, Mingze Xu et al.NeurIPS 2025 · 63 citations
- V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction TuningHang Hua, Yunlong Tang, Chenliang Xu, Jiebo LuoAAAI 2025 · 61 citations
- OneThinker: All-in-one Reasoning Model for Image and VideoKaituo Feng, Manyuan Zhang, Hongyu Li, Kaixuan Fan et al.CVPR 2026 · 55 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
Related papers
- VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video CaptioningJi Soo Lee, Jongha Kim, Jeehye Na, Jinyoung Park et al.AAAI 2025 · 11 citations
- Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningLong Qian, Juncheng Li, Yu Wu, Yaobo Ye et al.ICML 2024 · 121 citations
- Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive DecodingDaiqing Qi, Dongliang Guo, Hanzhang Yuan, Handong Zhao et al.NeurIPS 2025 · 5 citations
- DisTime: Distribution-Based Time Representation for Video Large Language ModelsYingsen Zeng, Zepeng Huang, Yujie Zhong, Chengjian Feng et al.ICCV 2025 · 2 citations
- MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment GroundingFuwen Luo, Shengfeng Lou, Chi Chen, Ziyue Wang et al.ACL 2026 · 10 citations
