Uncertainty-Aware Alignment Network for Cross-Domain Video-Text Retrieval
Xiaoshuai Hao, Wanqian Zhang
Abstract
Video-text retrieval is an important but challenging research task in the multimedia community. In this paper, we address the challenge task of Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), assuming that training (source) data and testing (target) data are from different domains. Previous approaches are mostly derived from classification based domain adaptation methods, which are neither multi-modal nor suitable for retrieval task. In addition, as to the pairwise misalignment issue in target domain, i.e., no pairwise annotations between target videos and texts, the existing method assumes that a video corresponds to a text. Yet we empirically find that in the real scene, one text usually corresponds to multiple videos and vice versa. To tackle this one-to-many issue, we propose a novel method named Uncertainty-aware Alignment Network (UAN). Specifically, we first introduce the multimodal mutual information module to balance the minimization of domain shift in a smooth manner. To tackle the multimodal uncertainties pairwise misalignment in target domain, we propose the Uncertainty-aware Alignment Mechanism (UAM) to fully exploit the semantic information of both modalities in target domain. Extensive experiments in the context of domain-adaptive video-text retrieval demonstrate that our proposed method consistently outperforms multiple baselines, showing a superior generalization ability for target data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17f653da-61e7-4d5d-95f7-cb1a471e176cCited by top-tier papers7
- Mind the Discriminability Trap in Source-Free Cross-domain Few-shot LearningZhenyu Zhang, Yixiong Zou, Yuhua Li, Ruixuan Li et al.CVPR 2026 · 6 citations
- FTF-ER: Feature-Topology Fusion-Based Experience Replay Method for Continual Graph LearningJinhui Pang, Changqing Lin, Xiaoshuai Hao, Rong Yin et al.ACM MM 2024 · 5 citations
- Learning Source-Free Domain Adaptation for Visible-Infrared Person Re-IdentificationYongxiang Li, Yanglin Feng, Yuan Sun, Dezhong Peng et al.NeurIPS 2025 · 4 citations
- Synergistic Prompting for Robust Visual Recognition with Missing ModalitiesZhihui Zhang, Luanyuan Dai, Qika Lin, Yunfeng Diao et al.ICCV 2025 · 2 citations
- Question-Adaptive Graph Learning for Multi-hop Retrieval Augmented GenerationYuchen Yan, Peiyan Zhang, Zhihua Liu, Hao Wang et al.SIGIR 2026
Builds on21
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- Domain Adaptation for Structured Output via Discriminative Patch RepresentationsYi-Hsuan Tsai, Kihyuk Sohn, Samuel Schulter, Manmohan ChandrakerICCV 2019 · 333 citations
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan et al.ACM MM 2022 · 314 citations
- Egocentric Video-Language PretrainingKevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray et al.NeurIPS 2022 · 306 citations
- X-Pool: Cross-Modal Language-Video Attention for Text-Video RetrievalSatya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan et al.CVPR 2022 · 190 citations
Related papers
- Dual Alignment Unsupervised Domain Adaptation for Video-Text RetrievalXiaoshuai Hao, Wanqian Zhang, Dayan Wu, Fei Zhu et al.CVPR 2023
- Mind-the-Gap! Unsupervised Domain Adaptation for Text-Video RetrievalQingchao Chen, Yang Liu, Samuel AlbanieAAAI 2021 · 28 citations
- Unsupervised Domain Adaptative Temporal Sentence Localization with Mutual Information MaximizationDaizong Liu, Xiang Fang, Xiaoye Qu, Jianfeng Dong et al.AAAI 2024 · 9 citations
- Text-Adaptive Multiple Visual Prototype Matching for Video-Text RetrievalChengzhi Lin, Ancong Wu, Junwei Liang, Jun Zhang et al.NeurIPS 2022 · 52 citations
- Relative Alignment Network for Source-Free Multimodal Video Domain AdaptationYi Huang, Xiaoshan Yang, Ji Zhang, Changsheng XuACM MM 2022 · 18 citations
