OSAN: A One-Stage Alignment Network to Unify Multimodal Alignment and Unsupervised Domain Adaptation
Ye Liu, Lingfeng Qiao, Changchong Lu, Di Yin, Chen Lin, Haoyuan Peng, Bo Ren
Abstract
Extending from unimodal to multimodal is a critical challenge for unsupervised domain adaptation (UDA). Two major problems emerge in unsupervised multimodal domain adaptation: domain adaptation and modality alignment. An intuitive way to handle these two problems is to fulfill these tasks in two separate stages: aligning modalities followed by domain adaptation, or vice versa. However, domains and modalities are not associated in most existing two-stage studies, and the relationship between them is not leveraged which can provide complementary information to each other. In this paper, we unify these two stages into one to align domains and modalities simultaneously. In our model, a tensor-based alignment module (TAL) is presented to explore the relationship between domains and modalities. By this means, domains and modalities can interact sufficiently and guide them to utilize complementary information for better results. Furthermore, to establish a bridge between domains, a dynamic domain generator (DDG) module is proposed to build transitional samples by mixing the shared information of two domains in a self-supervised manner, which helps our model learn a domain-invariant common representation space. Extensive experiments prove that our method can achieve superior performance in two real-world applications. The code will be publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08e6431c-dc17-4b64-8f44-08ecbd8400e1Cited by top-tier papers5
- DGMamba: Domain Generalization via Generalized State Space ModelShaocong Long, Qianyu Zhou, Xiangtai Li, Xuequan Lu et al.ACM MM 2024 · 15 citations
- Attention Bootstrapping for Multi-Modal Test-Time AdaptationYusheng Zhao, Junyu Luo, Xiao Luo, Jinsheng Huang et al.AAAI 2025 · 5 citations
- Distinguish Then Exploit: Source-free Open Set Domain Adaptation via Weight Barcode Estimation and Sparse Label AssignmentWeiming Liu, Jun Dan, Fan Wang, Xinting Liao et al.CVPR 2025
- Democratizing Clinical Risk Prediction with Cross-Cohort Cross-Modal Knowledge TransferQiannan Zhang, Manqi Zhou, Zilong Bai, Chang Su et al.NeurIPS 2025
- Active Multi-source Domain Adaptation for Multimodal Fake News DetectionYanping Chen, Weijie Shi, Mengze Li, Yue Cui et al.AAAI 2026
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 737 citations
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh et al.ACL 2020 · 584 citations
- Adversarial Domain Adaptation with Domain MixupMinghao Xu, Jian Zhang, Bingbing Ni, Teng Li et al.AAAI 2020 · 499 citations
Related papers
- Self-supervised Exclusive Learning for 3D Segmentation with Cross-Modal Unsupervised Domain AdaptationYachao Zhang, Miaoyu Li, Yuan Xie, Cuihua Li et al.ACM MM 2022 · 22 citations
- Towards Unsupervised Domain Bridging via Image Degradation in Semantic SegmentationWangkai Li, Rui Sun, Huayu Mai, Tianzhu ZhangNeurIPS 2025 · 8 citations
- Mix-DANN and Dynamic-Modal-Distillation for Video Domain AdaptationYuehao Yin, Bin Zhu, Jingjing Chen, Lechao Cheng et al.ACM MM 2022 · 7 citations
- Bridging Domain Generalization to Multimodal Domain Generalization via Unified RepresentationsHai Huang, Yan Xia, Sashuai Zhou, Hanting Wang et al.ICCV 2025 · 2 citations
- Dual Alignment Unsupervised Domain Adaptation for Video-Text RetrievalXiaoshuai Hao, Wanqian Zhang, Dayan Wu, Fei Zhu et al.CVPR 2023
