Multimodal Knowledge Expansion
Zihui Xue, Sucheng Ren, Zhengqi Gao, Hang Zhao
Abstract
The popularity of multimodal sensors and the accessibility of the Internet have brought us a massive amount of unlabeled multimodal data. Since existing datasets and well-trained models are primarily unimodal, the modality gap between a unimodal network and unlabeled multi-modal data poses an interesting problem: how to transfer a pre-trained unimodal network to perform the same task with extra unlabeled multimodal data? In this work, we propose multimodal knowledge expansion (MKE), a knowledge distillation-based framework to effectively utilize multimodal data without requiring labels. Opposite to traditional knowledge distillation, where the student is designed to be lightweight and inferior to the teacher, we observe that a multimodal student model consistently rectifies pseudo labels and generalizes better than its teacher. Extensive experiments on four tasks and different modalities verify this finding. Furthermore, we connect the mechanism of MKE to semi-supervised learning and offer both empirical and theoretical explanations to understand the expansion capability of a multimodal student.1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2f27f3c-42ba-4bed-a522-79f652c89c2eCited by top-tier papers14
- Modality-aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence DetectionJiashuo Yu, Jinyu Liu, Ying Cheng, Rui Feng et al.ACM MM 2022 · 64 citations
- Multimodal Distillation for Egocentric Action RecognitionGorjan Radevski, Dusan Grujicic, Matthew B. Blaschko, Marie-Francine Moens et al.ICCV 2023 · 40 citations
- CrossMatch: Source-Free Domain Adaptive Semantic Segmentation via Cross-Modal Consistency TrainingYifang Yin, Wenmiao Hu, Zhenguang Liu, Guanfeng Wang et al.ICCV 2023 · 21 citations
- The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge DistillationZihui Xue, Zhengqi Gao, Sucheng Ren, Hang ZhaoICLR 2023 · 12 citations
- MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge DistillationHui Li, Pengfei Yang, Juanyang Chen, Le Dong et al.ACM MM 2025 · 4 citations
Builds on10
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Confidence Regularized Self-TrainingYang Zou, Zhiding Yu, Xiaofeng Liu, B. V. K. Vijaya Kumar et al.ICCV 2019 · 901 citations
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen et al.NeurIPS 2021 · 404 citations
Related papers
- C2KD: Bridging the Modality Gap for Cross-Modal Knowledge DistillationFushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang et al.CVPR 2024
- Multimodal Learning with Incomplete Modalities by Knowledge DistillationQi Wang, Liang Zhan, Paul M. Thompson, Jiayu ZhouKDD 2020 · 80 citations
- Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and AlgorithmT. K Tran, Duc Chu Anh, Quang Hung Pham, Phi Le Nguyen et al.ICML 2026
- Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior KnowledgeLong Zhao, Xi Peng, Yuxiao Chen, Mubbasir Kapadia et al.CVPR 2020
- MMANet: Margin-Aware Distillation and Modality-Aware Regularization for Incomplete Multimodal LearningShicai Wei, Chunbo Luo, Yang LuoCVPR 2023
