Dual Alignment Unsupervised Domain Adaptation for Video-Text Retrieval
Xiaoshuai Hao, Wanqian Zhang, Dayan Wu, Fei Zhu, Bo Li
Abstract
Video-text retrieval is an emerging stream in both computer vision and natural language processing communities, which aims to find relevant videos given text queries. In this paper, we study the notoriously challenging task, i.e., Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), wherein training and testing data come from different distributions. Previous works merely alleviate the domain shift, which however overlook the pairwise misalignment issue in target domain, i.e., there exist no semantic relationships between target videos and texts. To tackle this, we propose a novel method named Dual Alignment Domain Adaptation (DADA). Specifically, we first introduce the cross-modal semantic embedding to generate discriminative source features in a joint embedding space. Besides, we utilize the video and text domain adaptations to smoothly balance the minimization of the domain shifts. To tackle the pairwise misalignment in target domain, we propose the Dual Alignment Consistency (DAC) to fully exploit the semantic information of both modalities in target domain. The proposed DAC adaptively aligns the videotext pairs which are more likely to be relevant in target domain, enabling that positive pairs are increasing progressively and the noisy ones will potentially be aligned in the later stages. To that end, our method can generate more truly aligned target pairs and ensure the discriminability of target features. Compared with the state-of-the-art methods, DADA achieves 20.18% and 18.61% relative improvements on R@1 under the setting of TGIF→MSR-VTT and TGIF→MSVD respectively, demonstrating the superiority of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Uncertainty-Aware Alignment Network for Cross-Domain Video-Text RetrievalXiaoshuai Hao, Wanqian ZhangNeurIPS 2023 · 26 citations
- A Pedestrian is Worth One Prompt: Towards Language Guidance Person Re- IdentificationZexian Yang, Dayan Wu, Chenming Wu, Zheng Lin et al.CVPR 2024 · 26 citations
- Unsupervised Domain Adaptative Temporal Sentence Localization with Mutual Information MaximizationDaizong Liu, Xiang Fang, Xiaoye Qu, Jianfeng Dong et al.AAAI 2024 · 9 citations
- Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person RetrievalBingjun Luo, Jinpeng Wang, Zewen Wang, Junjie Zhu et al.AAAI 2025 · 8 citations
- FTF-ER: Feature-Topology Fusion-Based Experience Replay Method for Continual Graph LearningJinhui Pang, Changqing Lin, Xiaoshuai Hao, Rong Yin et al.ACM MM 2024 · 5 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Domain Adaptation for Structured Output via Discriminative Patch RepresentationsYi-Hsuan Tsai, Kihyuk Sohn, Samuel Schulter, Manmohan ChandrakerICCV 2019 · 333 citations
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan et al.ACM MM 2022 · 314 citations
- Temporal Attentive Alignment for Large-Scale Video Domain AdaptationMin-Hung Chen, Zsolt Kira, Ghassan Alregib, Jaekwon Yoo et al.ICCV 2019 · 205 citations
Related papers
- Mind-the-Gap! Unsupervised Domain Adaptation for Text-Video RetrievalQingchao Chen, Yang Liu, Samuel AlbanieAAAI 2021 · 28 citations
- Progressive Semantic Matching for Video-Text RetrievalHongying Liu, Ruyi Luo, Fanhua Shang, Mantang Niu et al.ACM MM 2021 · 20 citations
- Visual Co-Occurrence Alignment Learning for Weakly-Supervised Video Moment RetrievalZheng Wang, Jingjing Chen, Yu-Gang JiangACM MM 2021 · 74 citations
- UATVR: Uncertainty-Adaptive Text-Video RetrievalBo Fang, Wenhao Wu, Chang Liu, Yu Zhou et al.ICCV 2023 · 98 citations
- T2VLAD: Global-Local Sequence Alignment for Text-Video RetrievalXiaohan Wang, Linchao Zhu, Yi YangCVPR 2021
