SimMMDG: A Simple and Effective Framework for Multi-modal Domain Generalization
Hao Dong, Ismail Nejjar, Han Sun, Eleni N. Chatzi, Olga Fink
Abstract
In real-world scenarios, achieving domain generalization (DG) presents significant challenges as models are required to generalize to unknown target distributions. Generalizing to unseen multi-modal distributions poses even greater difficulties due to the distinct properties exhibited by different modalities. To overcome the challenges of achieving domain generalization in multi-modal scenarios, we propose SimMMDG, a simple yet effective multi-modal DG framework. We argue that mapping features from different modalities into the same embedding space impedes model generalization. To address this, we propose splitting the features within each modality into modality-specific and modality-shared components. We employ supervised contrastive learning on the modality-shared features to ensure they possess joint properties and impose distance constraints on modality-specific features to promote diversity. In addition, we introduce a cross-modal translation module to regularize the learned features, which can also be used for missing-modality generalization. We demonstrate that our framework is theoretically well-supported and achieves strong performance in multi-modal DG on the EPIC-Kitchens dataset and the novel Human-Animal-Cartoon (HAC) dataset introduced in this paper. Our source code and HAC dataset are available at https://github.com/donghao51/SimMMDG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23fcebc8-5d6c-4b0a-b236-ceff8c644e20Cited by top-tier papers24
- MultiOOD: Scaling Out-of-Distribution Detection for Multiple ModalitiesHao Dong, Yue Zhao, Eleni N. Chatzi, Olga FinkNeurIPS 2024 · 43 citations
- Secure On-Device Video OOD Detection without BackpropagationShawn Li, Peilin Cai, Yuxiao Zhou, Zhiyu Ni et al.ICCV 2025 · 28 citations
- Cross-modal Representation Flattening for Multi-modal Domain GeneralizationYunfeng Fan, Wenchao Xu, Haozhao Wang, Song GuoNeurIPS 2024 · 21 citations
- Extremely Simple Multimodal Outlier Synthesis for Out-of-Distribution Detection and SegmentationMoru Liu, Hao Dong, Jessica Kelly, Olga Fink et al.NeurIPS 2025 · 12 citations
- Charts Are Not Images: On the Challenges of Scientific Chart EditingShawn Li, Ryan Rossi, Sungchul Kim, Sunav Choudhary et al.ICLR 2026 · 11 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- Bridging Domain Generalization to Multimodal Domain Generalization via Unified RepresentationsHai Huang, Yan Xia, Sashuai Zhou, Hanting Wang et al.ICCV 2025 · 2 citations
- Towards Multimodal Domain Generalization with Few LabelsHongzhao Li, Hao Dong, Hualei Wan, Shupan Li et al.CVPR 2026 · 2 citations
- Towards Robust Multimodal Domain Generalization via Modality-Domain Joint Adversarial TrainingHongzhao Li, Hualei Wan, Liangzhi Zhang, Mingyuan Jiu et al.ACM MM 2025 · 1 citation
- MER-DG: Modality-Entropy Regularization for Multimodal Domain GeneralizationYavuz Yarici, Ghassan AlRegibICML 2026
- BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic SegmentationMiaoyu Li, Yachao Zhang, Xu Ma, Yanyun Qu et al.ICCV 2023 · 22 citations
