MENTOR: Multi-level Self-supervised Learning for Multimodal Recommendation
Jinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li, Hewei Wang, Edith C. H. Ngai
Abstract
As multimedia information proliferates, multimodal recommendation systems have garnered significant attention. These systems leverage multimodal information to alleviate the data sparsity issue inherent in recommendation systems, thereby enhancing the accuracy of recommendations. Due to the natural semantic disparities among multimodal features, recent research has primarily focused on cross-modal alignment using self-supervised learning to bridge these gaps. However, aligning different modal features might result in the loss of valuable interaction information, distancing them from ID embeddings. It is crucial to recognize that the primary goal of multimodal recommendation is to predict user preferences, not merely to understand multimodal content. To this end, we propose a new Multi-level sElf-supervised learNing for mulTimOdal Recommendation (MENTOR) method, which effectively reduces the gap among modalities while retaining interaction information. Specifically, MENTOR begins by extracting representations from each modality using both heterogeneous user-item and homogeneous item-item graphs. It then employs a multilevel cross-modal alignment task, guided by ID embeddings, to align modalities across multiple levels while retaining historical interaction information. To balance effectiveness and efficiency, we further propose an optional general feature enhancement task that bolsters the general features from both structure and feature perspectives, thus enhancing the robustness of our model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45a7c815-d3ed-442c-b34d-ecf5327aec13Cited by top-tier papers16
- COHESION: Composite Graph Convolutional Network with Dual-Stage Fusion for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Wei Wang, Xiping Hu et al.SIGIR 2025 · 20 citations
- Structured Spectral Reasoning for Frequency-Adaptive Multimodal RecommendationWei Yang, Rui Zhong, Yiqun Chen, Chi Lu et al.NeurIPS 2025 · 10 citations
- The Best is Yet to Come: Graph Convolution in the Testing Phase for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.ACM MM 2025 · 8 citations
- MDVT: Enhancing Multimodal Recommendation with Model-Agnostic Multimodal-Driven Virtual TripletsJinfeng Xu, Zheyu Chen, Jinze Li, Shuo Yang et al.KDD 2025 · 4 citations
- VI-MMRec: Similarity-Aware Training Cost-free Virtual User-Item Interactions for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.KDD 2026 · 3 citations
Builds on12
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Self-supervised Graph Learning for RecommendationJiancan Wu, Xiang Wang, Fuli Feng, Xiangnan He et al.SIGIR 2021 · 1,476 citations
- Are Graph Augmentations Necessary?: Simple Graph Contrastive Learning for RecommendationJunliang Yu, Hongzhi Yin, Xin Xia, Tong Chen et al.SIGIR 2022 · 658 citations
- Graph-Refined Convolutional Network for Multimedia Recommendation with Implicit FeedbackYinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He et al.ACM MM 2020 · 374 citations
Related papers
- Multi-Modal Self-Supervised Learning for RecommendationWei Wei, Chao Huang, Lianghao Xia, Chuxu ZhangWWW 2023 · 256 citations
- DiffMM: Multi-Modal Diffusion Model for RecommendationYangqin Jiang, Lianghao Xia, Wei Wei, Da Luo et al.ACM MM 2024 · 92 citations
- Semantic-Guided Feature Distillation for Multimodal RecommendationFan Liu, Huilin Chen, Zhiyong Cheng, Liqiang Nie et al.ACM MM 2023 · 24 citations
- Multi-view Semantic Contrastive Alignment for Multimodal RecommendationJiuqiang Li, Hongjun WangWWW 2026
- Multimodal-aware Multi-intention Learning for RecommendationWei Yang, Qingchen YangACM MM 2024 · 4 citations
