Unsupervised Cross-Domain Image Retrieval with Semantic-Attended Mixture-of-Experts
Kai Wang, Jiayang Liu, Xing Xu, Jingkuan Song, Xin Liu, Heng Tao Shen
Abstract
Unsupervised cross-domain image retrieval is designed to facilitate the retrieval between images in different domains in an unsupervised way. Without the guidance of labels, both intra-domain semantic learning and inter-domain semantic alignment pose significant challenges to the model's learning process. The resolution of these challenges relies on the accurate capture of domain-invariant semantic features by the model. Based on this consideration, we propose our Semantic-Attended Mixture of Experts (SA-MoE) model. Leveraging the proficiency of MoE network in capturing visual features, we enhance the model's focus on semantically relevant features through a series of strategies. We first utilize the self-attention mechanism of Vision Transformer to adaptively collect information with different weights on instances from different domains. In addition, we introduce contextual semantic association metrics to more accurately measure the semantic relatedness between instances. By utilizing the association metrics, secondary clustering is performed in the feature space to reinforce semantic relationships. Finally, we employ the metrics for information selection on the fused data to remove the semantic noise. We conduct extensive experiments on three widely used datasets. The consistent comparison results with existing methods indicate that our model possesses the state-of-the-art performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- Text-Phase Synergy Network with Dual Priors for Unsupervised Cross-Domain Image RetrievalJing Yang, Hui Xue, Shipeng Zhu, Pengfei FangCVPR 2026
- Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context LearningTianci Luo, Haohao Pan, Jinpeng Wang, Niu Lian et al.CVPR 2026
Related papers
- Local Precise Refinement: A Dual-Gated Mixture-of-Experts for Enhancing Foundation Model Generalization against Spectral ShiftsXi Chen, Maojun Zhang, Yu Liu, Shen YanCVPR 2026 · 3 citations
- DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object DetectionChanghan Liu, Xunzhi Xiang, Zixuan Duan, Wenbin Li et al.NeurIPS 2025 · 8 citations
- SEEN-DA: SEmantic ENtropy guided Domain-aware Attention for Domain Adaptive Object DetectionHaochen Li, Rui Zhang, Hantao Yao, Xin Zhang et al.CVPR 2025
- Structure-Aware Semantic-Aligned Network for Universal Cross-Domain RetrievalJialin Tian, Xing Xu, Kai Wang, Zuo Cao et al.SIGIR 2022 · 7 citations
- Generalizable Person Re-Identification With Relevance-Aware Mixture of ExpertsYongxing Dai, Xiaotong Li, Jun Liu, Zekun Tong et al.CVPR 2021
