Semi-Supervised Multimodal Classification Through Learning from Modal and Strategic Complementarities
Junchi Chen, Richong Zhang, Junfan Chen
Abstract
Supervised multimodal classification has been proven to outperform unimodal classification in the image-text domain. However, this task is highly dependent on abundant labeled data. To perform multimodal classification in data-insufficient scenarios, in this study, we explore semi-supervised multimodal classification (SSMC) that only requires a small amount of labeled data and plenty of unlabeled data. Specifically, we first design baseline SSMC models by combining known semi supervised pseudo-labeling methods with the two most commonly used modal fusion strategies, i.e. feature-level fusion and label-level aggregation. Based on our investigation and empirical study of the baselines, we discover two complementarities that may benefit SSMC if properly exploited: the predictions from different modalities (modal complementarity) and modal fusion strategies for pseudo-labeling (strategic complementarity). Therefore, we propose a Modal and Strategic Complementarity (MSC) framework for SSMC. Concretely, to exploit modal complementarity, we propose to learn reliability weights for the predictions from different modalities and refine the fusion scores. To learn from strategic complementarity, we introduce a dual KL divergence loss to guide the balance of quantity and quality of pseudo-labeled data selection. Extensive empirical studies demonstrate the effectiveness of the proposed framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- FreeMatch: Self-adaptive Thresholding for Semi-supervised LearningYidong Wang, Hao Chen, Qiang Heng, Wenxin Hou et al.ICLR 2023 · 139 citations
- Expanding Large Pre-trained Unimodal Models with Multimodal Information Injection for Image-Text Multimodal ClassificationTao Liang, Guosheng Lin, Mingyang Wan, Tianrui Li et al.CVPR 2022 · 39 citations
Related papers
- A Semantic Knowledge Complementarity based Decoupling Framework for Semi-supervised Class-imbalanced Medical Image SegmentationZheng Zhang, Guanchun Yin, Bo Zhang, Wu Liu et al.CVPR 2025
- Multi-faceted Complementary Learning for Incomplete Multi-view Multi-label ClassificationXinyu Xiao, Peixi Peng, Qiang Wang, Chao Xing et al.ACM MM 2025
- On Uni-Modal Feature Learning in Supervised Multi-Modal LearningChenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu et al.ICML 2023 · 79 citations
- Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment RecognitionWuyou Xia, Guoli Jia, Sicheng Zhao, Jufeng YangCVPR 2025
- Enhancing Semi-Supervised Learning with Cross-Modal KnowledgeHui Zhu, Yongchun Lü, Hongbin Wang, Xunyi Zhou et al.ACM MM 2022 · 4 citations
