Adversarial Video Moment Retrieval by Jointly Modeling Ranking and Localization
Da Cao, Yawen Zeng, Xiaochi Wei, Liqiang Nie, Richang Hong, Zheng Qin
摘要
Retrieving video moments from an untrimmed video given a natural language as the query is a challenging task in both academia and industry. Although much effort has been made to address this issue, traditional video moment ranking methods are unable to generate reasonable video moment candidates and video moment localization approaches are not applicable to large-scale retrieval scenario. How to combine ranking and localization into a unified framework to overcome their drawbacks and reinforce each other is rarely considered. Toward this end, we contribute a novel solution to thoroughly investigate the video moment retrieval issue under the adversarial learning paradigm. The key of our solution is to formulate the video moment retrieval task as an adversarial learning problem with two tightly connected components. Specifically, a reinforcement learning is employed as a generator to produce a set of possible video moments. Meanwhile, a pairwise ranking model is utilized as a discriminator to rank the generated video moments and the ground truth. Finally, the generator and the discriminator are mutually reinforced in the adversarial learning framework, which is able to jointly optimize the performance of both video moment ranking and video moment localization. Extensive experiments on two well-known datasets have well verified the effectiveness and rationality of our proposed solution.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Fast Video Moment RetrievalJunyu Gao, Changsheng XuICCV 2021 · 被引用 132 次
- Multi-Modal Knowledge Hypergraph for Diverse Image RetrievalYawen Zeng, Qin Jin, Tengfei Bao, Wenfeng LiAAAI 2023 · 被引用 40 次
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li 等ACM MM 2023 · 被引用 31 次
- Keyword-Based Diverse Image Retrieval by Semantics-aware Contrastive Learning and TransformerMinyi Zhao, Jinpeng Wang, Dongliang Liao, Yiru Wang 等SIGIR 2023 · 被引用 3 次
- Multi-Modal Relational Graph for Cross-Modal Video Moment RetrievalYawen Zeng, Da Cao, Xiaochi Wei, Meng Liu 等CVPR 2021
相关 Paper
- STRONG: Spatio-Temporal Reinforcement Learning for Cross-Modal Video Moment LocalizationDa Cao, Yawen Zeng, Meng Liu, Xiangnan He 等ACM MM 2020 · 被引用 47 次
- MS-DETR: Natural Language Video Localization with Sampling Moment-Moment InteractionJing Wang, Aixin Sun, Hao Zhang, Xiaoli LiACL 2023 · 被引用 13 次
- Natural Language Video Localization with Learnable Moment ProposalsShaoning Xiao, Long Chen, Jian Shao, Yueting Zhuang 等EMNLP 2021 · 被引用 44 次
- CONQUER: Contextual Query-aware Ranking for Video Corpus Moment RetrievalZhijian Hou, Chong-Wah Ngo, Wing Kwong ChanACM MM 2021 · 被引用 45 次
- Counterfactual Cross-modality Reasoning for Weakly Supervised Video Moment LocalizationZezhong Lv, Bing Su, Ji-Rong WenACM MM 2023 · 被引用 23 次
