Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
Yuan Xie, Tianshui Chen, Zheng Ge, Lionel Ni
摘要
Long-form video understanding remains a formidable challenge due to the complexity of modeling long-range temporal dependencies and multi-event narratives. Existing methods often rely on static reasoning or external Visual-Language Models (VLMs), resulting in high computational complexity and sub-optimal performance. In this paper, we propose Video-MTR, a reinforced multi-turn reasoning framework that operates solely through data-efficient, pure RL post-training. Video-MTR reformulates video understanding as a dynamic decision-making process, where the agent iteratively selects key segments conditioned on the evolving context of previously processed frames and the query. To ensure effective intermediate reasoning and training stability, we introduce a novel gated bi-level reward system, which synergizes trajectory-level rewards (answer correctness) with turn-level rewards (frame-query relevance). This mechanism eliminates the need for data-intensive supervised fine-tuning, thereby substantially reducing reliance on large-scale datasets. Remarkably, Video-MTR achieves competitive or superior performance using only 8K training samples, compared to existing approaches that require 257K to 4.4M examples. Extensive experiments on benchmarks including VideoMME, MLVU, LongVideoBench, LVBench, and EgoSchema demonstrate that Video-MTR surpasses state-of-the-art methods in both accuracy and efficiency. Code is available at https://github.com/Xyuan13/Video-MTR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- WorldMM: Dynamic Multimodal Memory Agent for Long Video ReasoningWoongyeong Yeo, Kangsan Kim, Jaehong Yoon, Sung Ju HwangCVPR 2026 · 被引用 53 次
- Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal EvidenceJiahao Meng, Xiangtai Li, Haochen Wang, Tan Yue 等ICML 2026 · 被引用 43 次
- VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering TwiceShuming Liu, Mingchen Zhuge, Changsheng Zhao, Jun Chen 等CVPR 2026 · 被引用 18 次
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual EvidenceKun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai 等CVPR 2026 · 被引用 17 次
- REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video UnderstandingJiaze Li, Hao Yin, Wenhui Tan, Jingyang Chen 等CVPR 2026 · 被引用 14 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
相关 Paper
- Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingYe Wang, Ziheng Wang, Boshen Xu, Yang Du 等NeurIPS 2025 · 被引用 143 次
- Scaling RL to Long VideosYukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu 等NeurIPS 2025 · 被引用 91 次
- LongVideo-R1: Smart Navigation for Low-cost Long Video UnderstandingJihao Qiu, Lingxi Xie, Xinyue Huo, Qi Tian 等CVPR 2026 · 被引用 7 次
- Select Less, Reason More: Prioritizing Evidence Purity for Video ReasoningXuchen Li, Xuzhao Li, Shiyu Hu, Kaiqi HuangCVPR 2026 · 被引用 5 次
- VAST: Video Ability-Stratified Taxonomy for Data-Efficient Video ReasoningZhongan Wang, Xiaoyu Wen, Lingxiao Du, Kun Li 等CVPR 2026
