Retrieval Across Any Domains via Large-scale Pre-trained Model
Jiexi Yan, Zhihui Yin, Chenghao Xu, Cheng Deng, Heng Huang
Abstract
In order to enhance the generalization ability towards unseen domains, universal cross-domain image retrieval methods require a training dataset encompassing diverse domains, which is costly to assemble. Given this constraint, we introduce a novel problem of data-free adaptive crossdomain retrieval, eliminating the need for real images during training. Towards this goal, we propose a novel Text-driven Knowledge Integration (TKI) method, which exclusively utilizes a pre-trained vision-language model to implement an "aggregation after expansion" training strategy. Specifically, we extract diverse implicit domain-specific information through a set of learnable domain word vectors. Subsequently, a domain-agnostic universal projection, equipped with a non-Euclidean multi-layer perceptron, can be optimized using these assorted text descriptions through the text-proxied domain aggregation. Leveraging the cross-modal transferability phenomenon of the shared latent space, we can integrate the trained domain-agnostic universal projection with the pre-trained visual encoder to extract the features of the input image for the following retrieval during testing. Extensive experimental results on several benchmark datasets demonstrate the superiority of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b788a7ce-a392-44ee-a0d9-346196809cf0Cited by top-tier papers3
- Hystar: Hypernetwork-driven Style-adaptive Retrieval via Dynamic SVD ModulationYujia Cai, Boxuan Li, Chenghao Xu, Jiexi YanICLR 2026
- Towards Robust Edge Model Adaptation via Elastic Architecture SearchXianhang Chu, Xu Yang, Kun Wei, Xi WangAAAI 2026
- Channel-masked Asymmetric Distribution Matching for Cross-Domain Generalized Dataset DistillationQi Liu, Chenghao Xu, Jiexi Yan, Guangtao Lyu et al.AAAI 2026
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- DePro: Domain Ensemble using Decoupled Prompts for Universal Cross-Domain RetrievalKaixiang Chen, Pengfei Fang, Hui XueSIGIR 2025 · 2 citations
- Domain Adaptive Hashing Retrieval via VLM Assisted Pseudo-Labeling and Dual Space AdaptationJingyao Li, Zhanshan Li, Shuai LüNeurIPS 2025 · 1 citation
- Text-Phase Synergy Network with Dual Priors for Unsupervised Cross-Domain Image RetrievalJing Yang, Hui Xue, Shipeng Zhu, Pengfei FangCVPR 2026
- Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal RetrievalRui Cai, Zhiyu Dong, Jianfeng Dong, Xun WangAAAI 2025 · 1 citation
- IFSeg: Image-free Semantic Segmentation via Vision-Language ModelSukmin Yun, Seong Hyeon Park, Paul Hongsuck Seo, Jinwoo ShinCVPR 2023
