Towards Multimodal Domain Generalization with Few Labels
Hongzhao Li, Hao Dong, Hualei Wan, Shupan Li, Mingliang Xu, Muhammad Haris Khan
Abstract
Multimodal models ideally should generalize to unseen domains while remaining data-efficient to reduce annotation costs. To this end, we introduce and study a new problem, Semi-Supervised Multimodal Domain Generalization (SSMDG), which aims to learn robust multimodal models from multi-source data with few labeled samples. We observe that existing approaches fail to address this setting effectively: multimodal domain generalization methods cannot exploit unlabeled data, semi-supervised multimodal learning methods ignore domain shifts, and semi-supervised domain generalization methods are confined to single-modality inputs. To overcome these limitations, we propose a unified framework featuring three key components: Consensus-Driven Consistency Regularization, which obtains reliable pseudo-labels through confident fused-unimodal consensus; Disagreement-Aware Regularization, which effectively utilizes ambiguous non-consensus samples; and Cross-Modal Prototype Alignment, which enforces domain-and modality-invariant representations while promoting robustness under missing modalities via cross-modal translation. We further establish the first SS-MDG benchmarks, on which our method consistently outperforms strong baselines in both standard and missingmodality scenarios. Our benchmarks and code are available at https : / / github . com / lihongzhao99 / SSMDG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27ba8bfe-3d5c-409e-b41a-ef320b4b8859Cited by top-tier papers1
Ask how each one uses itBuilds on18
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Domain Generalization with MixStyleKaiyang Zhou, Yongxin Yang, Yu Qiao, Tao XiangICLR 2021 · 986 citations
Related papers
- SimMMDG: A Simple and Effective Framework for Multi-modal Domain GeneralizationHao Dong, Ismail Nejjar, Han Sun, Eleni N. Chatzi et al.NeurIPS 2023 · 80 citations
- Bridging Domain Generalization to Multimodal Domain Generalization via Unified RepresentationsHai Huang, Yan Xia, Sashuai Zhou, Hanting Wang et al.ICCV 2025 · 2 citations
- Steady Progress Beats Stagnation: Mutual Aid of Foundation and Conventional Models in Mixed Domain Semi-Supervised Medical Image SegmentationQinghe Ma, Jian Zhang, Zekun Li, Lei Qi et al.CVPR 2025
- STiL: Semi-supervised Tabular-Image Learning for Comprehensive Task-Relevant Information Exploration in Multimodal ClassificationSiyi Du, Xinzhe Luo, Declan P. O'Regan, Chen QinCVPR 2025
- Towards Generalizing to Unseen Domains with Few LabelsChamuditha Jayanga Galappaththige, Sanoojan Baliah, Malitha Gunawardhana, Muhammad Haris KhanCVPR 2024
