SimMMDG: A Simple and Effective Framework for Multi-modal Domain Generalization
Hao Dong, Ismail Nejjar, Han Sun, Eleni N. Chatzi, Olga Fink
摘要
In real-world scenarios, achieving domain generalization (DG) presents significant challenges as models are required to generalize to unknown target distributions. Generalizing to unseen multi-modal distributions poses even greater difficulties due to the distinct properties exhibited by different modalities. To overcome the challenges of achieving domain generalization in multi-modal scenarios, we propose SimMMDG, a simple yet effective multi-modal DG framework. We argue that mapping features from different modalities into the same embedding space impedes model generalization. To address this, we propose splitting the features within each modality into modality-specific and modality-shared components. We employ supervised contrastive learning on the modality-shared features to ensure they possess joint properties and impose distance constraints on modality-specific features to promote diversity. In addition, we introduce a cross-modal translation module to regularize the learned features, which can also be used for missing-modality generalization. We demonstrate that our framework is theoretically well-supported and achieves strong performance in multi-modal DG on the EPIC-Kitchens dataset and the novel Human-Animal-Cartoon (HAC) dataset introduced in this paper. Our source code and HAC dataset are available at https://github.com/donghao51/SimMMDG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- MultiOOD: Scaling Out-of-Distribution Detection for Multiple ModalitiesHao Dong, Yue Zhao, Eleni N. Chatzi, Olga FinkNeurIPS 2024 · 被引用 43 次
- Secure On-Device Video OOD Detection without BackpropagationShawn Li, Peilin Cai, Yuxiao Zhou, Zhiyu Ni 等ICCV 2025 · 被引用 28 次
- Cross-modal Representation Flattening for Multi-modal Domain GeneralizationYunfeng Fan, Wenchao Xu, Haozhao Wang, Song GuoNeurIPS 2024 · 被引用 21 次
- Extremely Simple Multimodal Outlier Synthesis for Out-of-Distribution Detection and SegmentationMoru Liu, Hao Dong, Jessica Kelly, Olga Fink 等NeurIPS 2025 · 被引用 12 次
- Charts Are Not Images: On the Challenges of Scientific Chart EditingShawn Li, Ryan Rossi, Sungchul Kim, Sunav Choudhary 等ICLR 2026 · 被引用 11 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- Bridging Domain Generalization to Multimodal Domain Generalization via Unified RepresentationsHai Huang, Yan Xia, Sashuai Zhou, Hanting Wang 等ICCV 2025 · 被引用 2 次
- Towards Multimodal Domain Generalization with Few LabelsHongzhao Li, Hao Dong, Hualei Wan, Shupan Li 等CVPR 2026 · 被引用 2 次
- Towards Robust Multimodal Domain Generalization via Modality-Domain Joint Adversarial TrainingHongzhao Li, Hualei Wan, Liangzhi Zhang, Mingyuan Jiu 等ACM MM 2025 · 被引用 1 次
- MER-DG: Modality-Entropy Regularization for Multimodal Domain GeneralizationYavuz Yarici, Ghassan AlRegibICML 2026
- BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic SegmentationMiaoyu Li, Yachao Zhang, Xu Ma, Yanyun Qu 等ICCV 2023 · 被引用 22 次
