Semi-Supervised Multimodal Classification Through Learning from Modal and Strategic Complementarities
Junchi Chen, Richong Zhang, Junfan Chen
摘要
Supervised multimodal classification has been proven to outperform unimodal classification in the image-text domain. However, this task is highly dependent on abundant labeled data. To perform multimodal classification in data-insufficient scenarios, in this study, we explore semi-supervised multimodal classification (SSMC) that only requires a small amount of labeled data and plenty of unlabeled data. Specifically, we first design baseline SSMC models by combining known semi supervised pseudo-labeling methods with the two most commonly used modal fusion strategies, i.e. feature-level fusion and label-level aggregation. Based on our investigation and empirical study of the baselines, we discover two complementarities that may benefit SSMC if properly exploited: the predictions from different modalities (modal complementarity) and modal fusion strategies for pseudo-labeling (strategic complementarity). Therefore, we propose a Modal and Strategic Complementarity (MSC) framework for SSMC. Concretely, to exploit modal complementarity, we propose to learn reliability weights for the predictions from different modalities and refine the fusion scores. To learn from strategic complementarity, we introduce a dual KL divergence loss to guide the balance of quantity and quality of pseudo-labeled data selection. Extensive empirical studies demonstrate the effectiveness of the proposed framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- FreeMatch: Self-adaptive Thresholding for Semi-supervised LearningYidong Wang, Hao Chen, Qiang Heng, Wenxin Hou 等ICLR 2023 · 被引用 139 次
- Expanding Large Pre-trained Unimodal Models with Multimodal Information Injection for Image-Text Multimodal ClassificationTao Liang, Guosheng Lin, Mingyang Wan, Tianrui Li 等CVPR 2022 · 被引用 39 次
相关 Paper
- A Semantic Knowledge Complementarity based Decoupling Framework for Semi-supervised Class-imbalanced Medical Image SegmentationZheng Zhang, Guanchun Yin, Bo Zhang, Wu Liu 等CVPR 2025
- Multi-faceted Complementary Learning for Incomplete Multi-view Multi-label ClassificationXinyu Xiao, Peixi Peng, Qiang Wang, Chao Xing 等ACM MM 2025
- On Uni-Modal Feature Learning in Supervised Multi-Modal LearningChenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu 等ICML 2023 · 被引用 79 次
- Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment RecognitionWuyou Xia, Guoli Jia, Sicheng Zhao, Jufeng YangCVPR 2025
- Enhancing Semi-Supervised Learning with Cross-Modal KnowledgeHui Zhu, Yongchun Lü, Hongbin Wang, Xunyi Zhou 等ACM MM 2022 · 被引用 4 次
