Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos
Jie Wu, Guanbin Li, Xiaoguang Han, Liang Lin
摘要
Temporal grounding of natural language in untrimmed videos is a fundamental yet challenging multimedia task facilitating cross-media visual content retrieval. We focus on the weakly supervised setting of this task that merely accesses to coarse video-level language description annotation without temporal boundary, which is more consistent with reality as such weak labels are more readily available in practice. In this paper, we propose a Boundary Adaptive Refinement (BAR) framework that resorts to reinforcement learning (RL) to guide the process of progressively refining the temporal boundary. To the best of our knowledge, we offer the first attempt to extend RL to temporal localization task with weak supervision. As it is non-trivial to obtain a straightforward reward function in the absence of pairwise granular boundary-query annotations, a cross-modal alignment evaluator is crafted to measure the alignment degree of segment-query pair to provide tailor-designed rewards. This refinement scheme completely abandons traditional sliding window based solution pattern and contributes to acquiring more efficient, boundary-flexible and content-aware grounding results. Extensive experiments on two public benchmarks Charades-STA and ActivityNet demonstrate that BAR outperforms the state-of-the-art weakly-supervised method and even beats some competitive fully-supervised ones.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Activation Modulation and Recalibration Scheme for Weakly Supervised Semantic SegmentationJie Qin, Jie Wu, Xuefeng Xiao, Lujun Li 等AAAI 2022 · 被引用 137 次
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng 等CVPR 2022 · 被引用 108 次
- Video Corpus Moment Retrieval with Contrastive LearningHao Zhang, Aixin Sun, Wei Jing, Guoshun Nan 等SIGIR 2021 · 被引用 88 次
- Cross-Sentence Temporal and Semantic Relations in Video Activity LocalisationJiabo Huang, Yang Liu, Shaogang Gong, Hailin JinICCV 2021 · 被引用 77 次
- Video Moment Retrieval from Text Queries via Single Frame AnnotationRan Cui, Tianwen Qian, Pai Peng, Elena Daskalaki 等SIGIR 2022 · 被引用 42 次
它引用的顶会 Paper4
- Dynamic Graph Attention for Referring Expression ComprehensionSibei Yang, Guanbin Li, Yizhou YuICCV 2019 · 被引用 251 次
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang 等AAAI 2020 · 被引用 170 次
- Tree-Structured Policy Based Progressive Reinforcement Learning for Temporally Language Grounding in VideoJie Wu, Guanbin Li, Si Liu, Liang LinAAAI 2020 · 被引用 117 次
- Graph-Structured Referring Expression Reasoning in the WildSibei Yang, Guanbin Li, Yizhou YuCVPR 2020
相关 Paper
- Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence GroundingJiaming Chen, Weixin Luo, Wei Zhang, Lin MaAAAI 2022 · 被引用 33 次
- HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in VideosTingting Han, Xinsong Tao, Yufei Yin, Min Tan 等CVPR 2026
- Weakly Supervised Temporal Sentence Grounding with Uncertainty-Guided Self-trainingYifei Huang, Lijin Yang, Yoichi SatoCVPR 2023
- DeCo: Decomposition and Reconstruction for Compositional Temporal Grounding via Coarse-to-Fine Contrastive RankingLijin Yang, Quan Kong, Hsuan-Kung Yang, Wadim Kehl 等CVPR 2023
- SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video GroundingHong Gao, Xiangkai Xu, Bin Zhong, Junjie Yin 等CVPR 2026
