Mind-the-Gap! Unsupervised Domain Adaptation for Text-Video Retrieval
Qingchao Chen, Yang Liu, Samuel Albanie
Abstract
When can we expect a text-video retrieval system to work effectively on datasets that differ from its training domain? In this work, we investigate this question through the lens of unsupervised domain adaptation in which the objective is to match natural language queries and video content in the presence of domain shift at query-time. Such systems have significant practical applications since they are capable generalising to new data sources without requiring corresponding text annotations. We make the following contributions:
(1) We propose the UDAVR (Unsupervised Domain Adaptation for Video Retrieval) benchmark and employ it to study the performance of text-video retrieval in the presence of domain shift. (2) We propose Concept-Aware-Pseudo-Query (CAPQ), a method for learning discriminative and transferable features that bridge these cross-domain discrepancies to enable effective target domain retrieval using source domain supervision. (3) We show that CAPQ outperforms alternative domain adaptation strategies on UDAVR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0f4b49b-1cc8-4cbb-a8cf-fbe015b6a027Cited by top-tier papers10
- Tagging before Alignment: Integrating Multi-Modal Tags for Video-Text RetrievalYizhen Chen, Jie Wang, Lijian Lin, Zhongang Qi et al.AAAI 2023 · 39 citations
- Uncertainty-Aware Alignment Network for Cross-Domain Video-Text RetrievalXiaoshuai Hao, Wanqian ZhangNeurIPS 2023 · 26 citations
- Token Mixing: Parameter-Efficient Transfer Learning from Image-Language to Video-LanguageYuqi Liu, Luhui Xu, Pengfei Xiong, Qin JinAAAI 2023 · 10 citations
- Crossing the Gap: Domain Generalization for Image CaptioningYuchen Ren, Zhendong Mao, Shancheng Fang, Yan Lu et al.CVPR 2023
- Test-time Adaptation for Cross-modal Retrieval with Query ShiftHaobin Li, Peng Hu, Qianjun Zhang, Xi Peng et al.ICLR 2025
Builds on6
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Temporal Attentive Alignment for Large-Scale Video Domain AdaptationMin-Hung Chen, Zsolt Kira, Ghassan Alregib, Jaekwon Yoo et al.ICCV 2019 · 205 citations
- Fine-Grained Action Retrieval Through Multiple Parts-of-Speech EmbeddingsMichael Wray, Gabriela Csurka, Diane Larlus, Dima DamenICCV 2019 · 185 citations
- Structure-Aware Feature Fusion for Unsupervised Domain AdaptationQingchao Chen, Yang LiuAAAI 2020 · 26 citations
- Multi-Modal Domain Adaptation for Fine-Grained Action RecognitionJonathan Munro, Dima DamenCVPR 2020
Related papers
- Dual Alignment Unsupervised Domain Adaptation for Video-Text RetrievalXiaoshuai Hao, Wanqian Zhang, Dayan Wu, Fei Zhu et al.CVPR 2023
- Hierarchical Debiasing and Noisy Correction for Cross-domain Video Tube RetrievalJingqiao Xiu, Mengze Li, Wei Ji, Jingyuan Chen et al.ACM MM 2024 · 5 citations
- Unsupervised Domain Adaptative Temporal Sentence Localization with Mutual Information MaximizationDaizong Liu, Xiang Fang, Xiaoye Qu, Jianfeng Dong et al.AAAI 2024 · 9 citations
- Contrast and Mix: Temporal Contrastive Video Domain Adaptation with Background MixingAadarsh Sahoo, Rutav Shah, Rameswar Panda, Kate Saenko et al.NeurIPS 2021 · 89 citations
- Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video RetrievalJianfeng Dong, Minsong Zhang, Zheng Zhang, Xianke Chen et al.ICCV 2023 · 35 citations
