Progressive Semantic Matching for Video-Text Retrieval
Hongying Liu, Ruyi Luo, Fanhua Shang, Mantang Niu, Yuanyuan Liu
摘要
Cross-modal retrieval between texts and videos is important yet challenging. Until recently, previous works in this domain typically rely on learning a common space to match the text and video, but it is difficult to match due to the semantic gap between videos and texts. Although some methods employ coarse-to-fine or multi-expert networks to encode one or more common spaces for easier matching, they almost directly optimize one matching space, which is challenging, because of the huge semantic gap between different modalities. To address this issue, we aim at narrowing semantic gap by a progressive learning process with a coarse-to-fine architecture, and propose a novel Progressive Semantic Matching (PSM) method. We first construct a multilevel encoding network for videos and texts, and design some auxiliary common spaces, which are mapped by the outputs of encoders in different levels. Then all the common spaces are jointly trained end to end. In this way, the model can effectively encode videos and texts into a fusion common space by a progressive paradigm. Experimental results on three video-text datasets (i.e., MSR-VTT, TIGF and MSVD) demonstrate the advantages of our PSM, which achieves significant performance improvement compared with state-of-the-art approaches.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang 等ACM MM 2022 · 被引用 65 次
- Progressive Spatio-Temporal Prototype Matching for Text-Video RetrievalPandeng Li, Chen-Wei Xie, Liming Zhao, Hongtao Xie 等ICCV 2023 · 被引用 62 次
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust LearningYabing Wang, Jianfeng Dong, Tianxiang Liang, Minsong Zhang 等ACM MM 2022 · 被引用 26 次
- Gravitation-Driven Semantic Alignment for Text Video RetrievalYi Yang, Zheng Wang, Xing Xu, Jingkuan Song 等CVPR 2026
- Multilateral Semantic Relations Modeling for Image Text RetrievalZheng Wang, Zhenwei Gao, Kangshuai Guo, Yang Yang 等CVPR 2023
相关 Paper
- Fine-grained Cross-modal Alignment Network for Text-Video RetrievalNing Han, Jingjing Chen, Guangyi Xiao, Hao Zhang 等ACM MM 2021 · 被引用 47 次
- Learning Semantics-Grounded Vocabulary Representation for Video-Text RetrievalYaya Shi, Haowei Liu, Haiyang Xu, Zongyang Ma 等ACM MM 2023 · 被引用 3 次
- GHAN: Graph-Based Hierarchical Aggregation Network for Text-Video RetrievalYahan Yu, Bojie Hu, Yu LiEMNLP 2022 · 被引用 7 次
- Concept Propagation via Attentional Knowledge Graph Reasoning for Video-Text RetrievalSheng Fang, Shuhui Wang, Junbao Zhuo, Qingming Huang 等ACM MM 2022 · 被引用 10 次
- Dual Alignment Unsupervised Domain Adaptation for Video-Text RetrievalXiaoshuai Hao, Wanqian Zhang, Dayan Wu, Fei Zhu 等CVPR 2023
