MDVT: Enhancing Multimodal Recommendation with Model-Agnostic Multimodal-Driven Virtual Triplets
Jinfeng Xu, Zheyu Chen, Jinze Li, Shuo Yang, Hewei Wang, Yijie Li, Mengran Li, Puzhen Wu, Edith C. H. Ngai
Abstract
The data sparsity problem significantly hinders the performance of recommender systems, as traditional models rely on limited historical interactions to learn user preferences and item properties. While incorporating multimodal information can explicitly represent these preferences and properties, existing works often use it only as side information, failing to fully leverage its potential. In this paper, we propose MDVT, a model-agnostic approach that constructs multimodal-driven virtual triplets to provide valuable supervision signals, effectively mitigating the data sparsity problem in multimodal recommendation systems. To ensure high-quality virtual triplets, we introduce three tailored warm-up threshold strategies: static, dynamic, and hybrid. The static warm-up threshold strategy exhaustively searches for the optimal number of warm-up epochs but is time-consuming and computationally intensive. The dynamic warm-up threshold strategy adjusts the warm-up period based on loss trends, improving efficiency but potentially missing optimal performance. The hybrid strategy combines both, using the dynamic strategy to find the approximate optimal number of warm-up epochs and then refining it with the static strategy in a narrow hyper-parameter space. Once the warm-up threshold is satisfied, the virtual triplets are used for joint model optimization by our enhanced pair-wise loss function without causing significant gradient skew. Extensive experiments on multiple real-world datasets demonstrate that integrating MDVT into advanced multimodal recommendation models effectively alleviates the data sparsity problem and improves recommendation performance, particularly in sparse data scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbf604a7-eae2-4a52-8abf-a0c7707c6240Cited by top-tier papers4
- The Best is Yet to Come: Graph Convolution in the Testing Phase for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.ACM MM 2025 · 8 citations
- VI-MMRec: Similarity-Aware Training Cost-free Virtual User-Item Interactions for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.KDD 2026 · 3 citations
- Personalized Parameter-Efficient Fine-Tuning of Foundation Models for Multimodal RecommendationSunwoo Kim, Hyunjin Hwang, Kijung ShinWWW 2026 · 1 citation
- Well Begun is Half Done: Training-Free and Model-Agnostic Semantically Guaranteed User Representation Initialization for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.SIGIR 2026
Builds on14
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Graph-Refined Convolutional Network for Multimedia Recommendation with Implicit FeedbackYinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He et al.ACM MM 2020 · 374 citations
- Mining Latent Structures for Multimedia RecommendationJinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu et al.ACM MM 2021 · 350 citations
- Bootstrap Latent Representations for Multi-modal RecommendationXin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng et al.WWW 2023 · 326 citations
Related papers
- Online Item Cold-Start Recommendation with Popularity-Aware Meta-LearningYunze Luo, Yuezihan Jiang, Yinjie Jiang, Gaode Chen et al.KDD 2025 · 8 citations
- Semantic-Guided Feature Distillation for Multimodal RecommendationFan Liu, Huilin Chen, Zhiyong Cheng, Liqiang Nie et al.ACM MM 2023 · 24 citations
- M²VAE: Multi-Modal Multi-View Variational Autoencoder for Cold-start Item RecommendationChuan He, Yongchao Liu, Qiang Li, Chuntao Hong et al.AAAI 2026 · 1 citation
- MENTOR: Multi-level Self-supervised Learning for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.AAAI 2025 · 21 citations
- MoToRec: Sparse-Regularized Multimodal Tokenization for Cold-Start RecommenderJialin Liu, Zhaorui Zhang, Ray C. C. CheungAAAI 2026
