Adaptive Cross-Modal Prototypes for Cross-Domain Visual-Language Retrieval
Yang Liu, Qingchao Chen, Samuel Albanie
摘要
In this paper, we study the task of visual-text retrieval in the highly practical setting in which labelled visual data with paired text descriptions are available in one domain (the "source"), but only unlabelled visual data (without text descriptions) are available in the domain of interest (the "target"). We propose the ADAPTIVE CROSS-MODAL PROTOTYPES framework which seeks to enable target domain retrieval by learning cross-modal visual-text representations while minimising both uni-modal and cross-modal distribution shift across the source and target domains. Our approach is built upon two key ideas: first, we encode the inductive bias that the learned cross-modal representations should be compositional with respect to concepts in each modality-this is achieved through clustering pretrained uni-modal features across each domain and designing a careful regularisation scheme to preserve the resulting structure. Second, we employ mutual information maximisation between cross-modal representations in the source and target domains during learning-this provides a mechanism that preserves commonalities between the domains while discarding signal in each that cannot be inferred from the other. We showcase our approach for the task of cross-domain visual-text retrieval, outperforming existing approaches for both images and videos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Cross Modal Retrieval with Querybank NormalisationSimion-Vlad Bogolin, Ioana Croitoru, Hailin Jin, Yang Liu 等CVPR 2022 · 被引用 84 次
- Uncertainty-Aware Alignment Network for Cross-Domain Video-Text RetrievalXiaoshuai Hao, Wanqian ZhangNeurIPS 2023 · 被引用 26 次
- Noise Is Also Useful: Negative Correlation-Steered Latent Contrastive LearningJiexi Yan, Lei Luo, Chenghao Xu, Cheng Deng 等CVPR 2022 · 被引用 20 次
- Diffusion-Inspired Truncated Sampler for Text-Video RetrievalJiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan 等NeurIPS 2024 · 被引用 16 次
- Token Transformation Matters: Towards Faithful Post-Hoc Explanation for Vision TransformerJunyi Wu, Bin Duan, Weitai Kang, Hao Tang 等CVPR 2024 · 被引用 8 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- Support-set bottlenecks for video-text representation learningMandela Patrick, Po-Yao Huang, Yuki Markus Asano, Florian Metze 等ICLR 2021 · 被引用 269 次
- Fine-Grained Action Retrieval Through Multiple Parts-of-Speech EmbeddingsMichael Wray, Gabriela Csurka, Diane Larlus, Dima DamenICCV 2019 · 被引用 185 次
- Mind-the-Gap! Unsupervised Domain Adaptation for Text-Video RetrievalQingchao Chen, Yang Liu, Samuel AlbanieAAAI 2021 · 被引用 28 次
相关 Paper
- Text-Adaptive Multiple Visual Prototype Matching for Video-Text RetrievalChengzhi Lin, Ancong Wu, Junwei Liang, Jun Zhang 等NeurIPS 2022 · 被引用 52 次
- Dual Alignment Unsupervised Domain Adaptation for Video-Text RetrievalXiaoshuai Hao, Wanqian Zhang, Dayan Wu, Fei Zhu 等CVPR 2023
- Multimodal Aligned Semantic Knowledge for Unpaired Image-text MatchingLaiguo Yin, Yixin Zhang, YUQING SUN, Lizhen CuiICLR 2026
- Unsupervised Domain Adaptation for Referring Semantic SegmentationHaonan Shi, Wenwen Pan, Zhou Zhao, Mingmin Zhang 等ACM MM 2023 · 被引用 5 次
- Retrieval Across Any Domains via Large-scale Pre-trained ModelJiexi Yan, Zhihui Yin, Chenghao Xu, Cheng Deng 等ICML 2024
