PromptMM: Multi-Modal Knowledge Distillation for Recommendation with Prompt-Tuning
Wei Wei, Jiabin Tang, Lianghao Xia, Yangqin Jiang, Chao Huang
Abstract
Multimedia online platforms (e.g., Amazon, TikTok) have greatly benefited from the incorporation of multimedia (e.g., visual, textual, and acoustic) content into their personal recommender systems. These modalities provide intuitive semantics that facilitate modality-aware user preference modeling. However, two key challenges in multi-modal recommenders remain unresolved: i) The introduction of multi-modal encoders with a large number of additional parameters causes overfitting, given high-dimensional multimodal features provided by extractors (e.g., ViT, BERT). ii) Side information inevitably introduces inaccuracies and redundancies, which skew the modality-interaction dependency from reflecting true user preference. To tackle these problems, we propose to simplify and empower recommenders through Multi-modal Knowledge Distillation (PromptMM) with the prompt-tuning that enables adaptive quality distillation. Specifically, PromptMM conducts model compression through distilling u-i edge relationship and multimodal node content from cumbersome teachers to relieve students from the additional feature reduction parameters. To bridge the semantic gap between multi-modal context and collaborative signals for empowering the overfitting teacher, soft prompt-tuning is introduced to perform student task-adaptive. Additionally, to adjust the impact of inaccuracies in multimedia data, a disentangled multi-modal list-wise distillation is developed with modalityaware re-weighting mechanism. Experiments on real-world data demonstrate PromptMM's superiority over existing techniques. Ablation tests confirm the effectiveness of key components. Additional tests show the efficiency and effectiveness. To facilitate the result reproducibility, our implementation is released via link: https://github.com/HKUDS/PromptMM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c91d45e-93f3-4aaf-bdff-19b1453e2aeaCited by top-tier papers14
- DiffMM: Multi-Modal Diffusion Model for RecommendationYangqin Jiang, Lianghao Xia, Wei Wei, Da Luo et al.ACM MM 2024 · 92 citations
- MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender SystemsYibiao Wei, Jie Zou, Weikang Guo, Guoqing Wang et al.SIGIR 2025 · 11 citations
- Structured Spectral Reasoning for Frequency-Adaptive Multimodal RecommendationWei Yang, Rui Zhong, Yiqun Chen, Chi Lu et al.NeurIPS 2025 · 10 citations
- The Best is Yet to Come: Graph Convolution in the Testing Phase for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.ACM MM 2025 · 8 citations
- FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation LearningWei Yang, Rui Zhong, Yiqun Chen, Shixuan Li et al.ACM MM 2025 · 6 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- GraphMAE: Self-Supervised Masked Graph AutoencodersZhenyu Hou, Xiao Liu, Yukuo Cen, Yuxiao Dong et al.KDD 2022 · 533 citations
- Representation Learning with Large Language Models for RecommendationXubin Ren, Wei Wei, Lianghao Xia, Lixin Su et al.WWW 2024 · 385 citations
Related papers
- Modality-Balanced Learning for Multimedia RecommendationJinghao Zhang, Guofan Liu, Qiang Liu, Shu Wu et al.ACM MM 2024 · 21 citations
- PLIKD: Prompt Learning with Instance-aware Knowledge Distillation for Web-scale Semantic Image ClassificationJianye Xie, Chunhua Hu, Lianyong Qi, Fan Wang et al.WWW 2026
- Semantic-Guided Feature Distillation for Multimodal RecommendationFan Liu, Huilin Chen, Zhiyong Cheng, Liqiang Nie et al.ACM MM 2023 · 24 citations
- Bridging the Gap: Teacher-Assisted Wasserstein Knowledge Distillation for Efficient Multi-Modal RecommendationZiyi Zhuang, Hanwen Du, Hui Han, Youhua Li et al.WWW 2025
- UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationZhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang et al.AAAI 2022 · 61 citations
