Modality-Balanced Learning for Multimedia Recommendation
Jinghao Zhang, Guofan Liu, Qiang Liu, Shu Wu, Liang Wang
摘要
Many recommender models have been proposed to investigate how to incorporate multimodal content information into traditional collaborative filtering framework effectively. The use of multimodal information is expected to provide more comprehensive information and lead to superior performance. However, the integration of multiple modalities often encounters the modal imbalance problem: since the information in different modalities is unbalanced, optimizing the same objective across all modalities leads to the under-optimization problem of the weak modalities with a slower convergence rate or lower performance. Even worse, we find that in multimodal recommendation models, all modalities suffer from the problem of insufficient optimization. To address these issues, we propose a Counterfactual Knowledge Distillation method that could solve the imbalance problem and make the best use of all modalities. Through modality-specific knowledge distillation, it could guide the multimodal model to learn modality-specific knowledge from uni-modal teachers. We also design a novel generic-and-specific distillation loss to guide the multimodal student to learn wider-and-deeper knowledge from teachers. Additionally, to adaptively recalibrate the focus of the multimodal model towards weaker modalities during training, we estimate the causal effect of each modality on the training objective using counterfactual inference techniques, through which we could determine the weak modalities, quantify the imbalance degree and re-weight the distillation loss accordingly. Our method could serve as a plug-and-play module for both late-fusion and early-fusion backbones. Extensive experiments on six backbones show that our proposed method can improve the performance by a large margin. The source code will be released at https://github.com/CRIPAC-DIG/Balanced-Multimodal-Rec
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product UnderstandingZhanheng Nie, Chenghan Fu, Daoze Zhang, Junxian Wu 等CVPR 2026 · 被引用 9 次
- Tackling Device Data Distribution Real-time Shift via Prototype-based Parameter EditingZheqi Lv, Wenqiao Zhang, Kairui Fu, Qi Tian 等ACM MM 2025
- Evaluating and Steering Modality Preferences in Multi-modal LLMsYu Zhang, Jinlong Ma, Yongshuai Hou, Xuefeng Bai 等ICML 2026
它引用的顶会 Paper28
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- Graph-Refined Convolutional Network for Multimedia Recommendation with Implicit FeedbackYinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He 等ACM MM 2020 · 被引用 374 次
- Mining Latent Structures for Multimedia RecommendationJinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu 等ACM MM 2021 · 被引用 350 次
- Bootstrap Latent Representations for Multi-modal RecommendationXin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng 等WWW 2023 · 被引用 326 次
- Contrastive Learning for Cold-Start RecommendationYinwei Wei, Xiang Wang, Qi Li, Liqiang Nie 等ACM MM 2021 · 被引用 321 次
相关 Paper
- Enhancing Adversarial Robustness of Multi-modal Recommendation via Modality BalancingYu Shang, Chen Gao, Jiansheng Chen, Depeng Jin 等ACM MM 2023 · 被引用 9 次
- G2D: Boosting Multimodal Learning with Gradient-Guided DistillationMohammed Rakib, Arunkumar BagavathiICCV 2025 · 被引用 1 次
- Semantic-Guided Feature Distillation for Multimodal RecommendationFan Liu, Huilin Chen, Zhiyong Cheng, Liqiang Nie 等ACM MM 2023 · 被引用 24 次
- Multimodal Learning with Incomplete Modalities by Knowledge DistillationQi Wang, Liang Zhan, Paul M. Thompson, Jiayu ZhouKDD 2020 · 被引用 80 次
- Bidirectional Counterfactual Distillation for Review-Based RecommendationSheng Sang, Shujie Li, Shuaiyang Li, Kang Liu 等AAAI 2026
