Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation
Xinyao Li, Yuke Li, Zhekai Du, Fengling Li, Ke Lu, Jingjing Li
Abstract
Large vision-language models (VLMs) like CLIP have demonstrated good zero-shot learning performance in the unsupervised domain adaptation task. Yet, most transfer approaches for VLMs focus on either the language or visual branches, overlooking the nuanced interplay between both modalities. In this work, we introduce a Unified Modality Separation (UniMoS) framework for unsupervised domain adaptation. Leveraging insights from modality gap studies, we craft a nimble modality separation network that distinctly disentangles CLIP's features into languageassociated and vision-associated components. Our proposed Modality-Ensemble Training (MET) method fosters the exchange of modality-agnostic information while maintaining modality-specific nuances. We align features across domains using a modality discriminator. Comprehensive evaluations on three benchmarks reveal our approach sets a new state-of-the-art with minimal computational costs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a994271-15cc-430a-9d99-605b2b9f7111Cited by top-tier papers10
- CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain AdaptationMainak Singha, Sarthak Mehrotra, Paolo Casari, Subhasis Chaudhuri et al.CVPR 2026 · 2 citations
- Domain Adaptive Hashing Retrieval via VLM Assisted Pseudo-Labeling and Dual Space AdaptationJingyao Li, Zhanshan Li, Shuai LüNeurIPS 2025 · 1 citation
- Cross-Domain Attribute Alignment with CLIP: A Rehearsal-Free Approach for Class-Incremental Unsupervised Domain AdaptationKerun Mi, Guoliang Kang, Guangyu Li, Lin Zhao et al.ACM MM 2025 · 1 citation
- Progressive Distribution Bridging: Unsupervised Adaptation for Large-Scale Pre-Trained Models via Adaptive Auxiliary DataWeinan He, Yixin Zhang, Zilei WangICCV 2025 · 1 citation
- End-to-End Knowledge Distillation for Unsupervised Domain Adaptation with Large Vision-language ModelsYangtao Wang, Xingwei Deng, Yanzhao Xie, Weilong Peng et al.AAAI 2026 · 1 citation
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang et al.ICML 2024 · 7 citations
- Understanding Transferable Representation Learning and Zero-shot Transfer in CLIPZixiang Chen, Yihe Deng, Yuanzhi Li, Quanquan GuICLR 2024 · 21 citations
- UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language ModelsJiachen Liang, Ruibing Hou, Minyang Hu, Hong Chang et al.NeurIPS 2024 · 4 citations
- DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only TrainingWei Li, Linchao Zhu, Longyin Wen, Yi YangICLR 2023 · 24 citations
- Intra-Modal Proxy Learning for Zero-Shot Visual Categorization with CLIPQi Qian, Yuanhong Xu, Juhua HuNeurIPS 2023 · 34 citations
