Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Video Temporal Grounding
Jian Hu, Zixu Cheng, Shaogang Gong, Isabel Guan, Jianye Hao, Jun Wang, Kun Shao
Abstract
Video Temporal Grounding (TG) aims to temporally locate video segments matching a natural language description (a query) in a long video. While Vision-Language Models (VLMs) are effective at holistic semantic matching, they often struggle with fine-grained temporal localisation. Recently, Group Relative Policy Optimisation (GRPO) reformulates the inference process as a reinforcement learning task, enabling fine-grained grounding and achieving strong in-domain performance. However, GRPO relies on labelled data, making it unsuitable in unlabelled domains. Moreover, because videos are large and expensive to store and process, performing full-scale adaptation introduces prohibitive latency and computational overhead, making it impractical for real-time deployment. To overcome both problems, we introduce a Data-Efficient Unlabelled Cross-domain Temporal Grounding method, from which a model is first trained on a labelled source domain, then adapted to a target domain using only a small number of unlabelled videos from the target domain. This approach eliminates the need for target annotation and keeps both computational and storage overhead low enough to run in real time. Specifically, we introduce Uncertainty-quantified Rollout Policy Adaptation (URPA) for crossdomain knowledge transfer in learning video temporal grounding without target labels. URPA generates multiple candidate predictions using GRPO rollouts, averages them to form a pseudo label, and estimates confidence from the variance across these rollouts. This confidence then weights the training rewards, guiding the model to focus on reliable supervision. Experiments on three datasets across six cross-domain settings show that URPA generalises well using only a few unlabelled target videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 028fd3ca-5c0f-45b8-be85-03eab95ada6dBuilds on30
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Efficient Test-Time Model Adaptation without ForgettingShuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen et al.ICML 2022 · 579 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 425 citations
Related papers
- Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal GroundingJin-Seop Lee, Sungjoon Lee, SeongJun Jung, Boyang Li et al.CVPR 2026 · 2 citations
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMsYunheng Li, Jing Cheng, Shaoyong Jia, Hangyi Kuang et al.NeurIPS 2025 · 18 citations
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement LearningTao Wu, Li Yang, Gen Zhan, Yabin ZHANG et al.CVPR 2026 · 7 citations
- CACR: Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal ReasoningMuge Qi, Rong Fu, Pengbin Feng, Xianda Li et al.ICML 2026
- Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed VideosJie Wu, Guanbin Li, Xiaoguang Han, Liang LinACM MM 2020 · 70 citations
