VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
Ye Liu, Kevin Qinghong Lin, Chang Wen Chen, Mike Zheng Shou
Abstract
Videos, with their unique temporal dimension, demand precise grounded understanding, where answers are directly linked to visual, interpretable evidence. Despite significant breakthroughs in text-based reasoning with large language models, multi-modal reasoning - especially for videos - remains limited. In this work, we fill this gap by introducing VideoMind, a novel video-language agent for temporal-grounded video reasoning. Our method involves two key innovations: (1) We identify four essential capabilities for grounded video reasoning and propose a role-based agentic workflow, comprising a planner to coordinate roles, a grounder for temporal event localization, a verifier to assess event candidates, and an answerer for question answering. (2) To efficiently integrate these roles during inference, we propose a novel Chain-of-LoRA mechanism, where a unified base model with multiple LoRA adapters is leveraged to enable seamless role switching, balancing efficiency and flexibility. Extensive experiments on 15 benchmarks across Grounded VideoQA, Video Temporal Grounding, and General VideoQA tasks demonstrate the effectiveness of the proposed scheme in advancing video agent, test-time scaling, and long-form video reasoning. Code, models, datasets, and demos are available at https://videomind.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75f86506-4527-4557-b4dd-50d8f5dc35ddCited by top-tier papers14
- Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingYe Wang, Ziheng Wang, Boshen Xu, Yang Du et al.NeurIPS 2025 · 143 citations
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMsJun Zhang, Teng Wang, Yuying Ge, Yixiao Ge et al.CVPR 2026 · 48 citations
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual ReasoningYe Liu, Zongyang Ma, Junfu Pu, Zhongang Qi et al.NeurIPS 2025 · 39 citations
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang et al.NeurIPS 2025 · 30 citations
- Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-ThoughtChao Huang, Benfeng Wang, Wei Wang, Jie Wen et al.NeurIPS 2025 · 30 citations
Builds on46
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video ReasoningHaoji Zhang, Xin Gu, Jiawen Li, Chixiang Ma et al.CVPR 2026 · 92 citations
- VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video CaptioningJi Soo Lee, Jongha Kim, Jeehye Na, Jinyoung Park et al.AAAI 2025 · 11 citations
- Agentic Spatio-Temporal Grounding via Collaborative ReasoningHeng Zhao, Yew-Soon Ong, Joey Tianyi ZhouSIGIR 2026 · 1 citation
- LongVT: Incentivizing "Thinking with Long Videos" via Native Tool CallingZuhao Yang, Sudong Wang, Kaichen Zhang, Keming Wu et al.CVPR 2026 · 63 citations
- WorldMM: Dynamic Multimodal Memory Agent for Long Video ReasoningWoongyeong Yeo, Kangsan Kim, Jaehong Yoon, Sung Ju HwangCVPR 2026 · 53 citations
