TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
Junwen Pan, Qizhe Zhang, Rui Zhang, Ming Lu, Xin Wan, Yuan Zhang, Chang Liu, Qi She
摘要
Temporal search aims to identify a minimal set of relevant frames from tens of thousands based on a given query, serving as a foundation for accurate long-form video understanding. Many existing works attempt to progressively narrow the search space. However, these approaches typically rely on a hand-crafted search process, lacking end-to-end optimization for learning optimal search strategies. In this paper, we propose TimeSearch-R, which reformulates temporal search as interleaved text–video thinking, seamlessly integrating searching video clips into the reasoning process through reinforcement learning (RL). However, applying RL training methods, such as Group Relative Policy Optimization (GRPO), to video reasoning can result in unsupervised intermediate search decisions. This leads to insufficient exploration of the video content and inconsistent logical reasoning. To address these issues, we introduce GRPO with Completeness Self-Verification (GRPO-CSV), which gathers searched video frames from the interleaved reasoning process and utilizes the same policy model to verify the adequacy of searched frames, thereby improving the completeness of video reasoning. Additionally, we construct datasets specifically designed for the SFT cold-start and RL training of GRPO-CSV, filtering out samples with weak temporal dependencies to enhance task difficulty and improve temporal search capabilities. Extensive experiments demonstrate that TimeSearch-R achieves substantial improvements on temporal search benchmarks such as Haystack-LVBench and Haystack-Ego4D, long-form video understanding benchmarks like VideoMME, MLVU, and LongVideoBench, as well as video reasoning benchmarks such as Video-Holmes, consistently and significantly outperforming other existing temporal search approaches and text-only reasoning models. Our code is available at https://github.com/Time-Search/TimeSearch-R.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth 等NeurIPS 2024 · 被引用 373 次
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao 等ICLR 2026 · 被引用 321 次
相关 Paper
- VideoTrace-R1: Long Video-based Retrieval-Augmented Generation via Reinforcement LearningZongsheng Cao, Anran Liu, Jun Xie, Feng Chen 等ICML 2026
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement LearningTao Wu, Li Yang, Gen Zhan, Yabin ZHANG 等CVPR 2026 · 被引用 7 次
- Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal GroundingJin-Seop Lee, Sungjoon Lee, SeongJun Jung, Boyang Li 等CVPR 2026 · 被引用 2 次
- Reinforcing Structured Chain-of-Thought for Video UnderstandingPeiyao Wang, Haotian Xu, Noranart Vesdapunt, Rui Hou 等CVPR 2026 · 被引用 1 次
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video ReasoningHaoji Zhang, Xin Gu, Jiawen Li, Chixiang Ma 等CVPR 2026 · 被引用 92 次
