SGP4SR: Separated-Modality Guided User Preference Learning for Multimodal Sequential Recommendation
Changhong Li, Zhiqiang Guo, Guohui Li, Zhong Yang, Chuhang Hong
Abstract
With the booming development of multimodal data (e.g., image, text) on internet platforms, multimodal sequential recommendation methods continue to emerge. Most existing methods incorporate item modal features as auxiliary information, typically concatenating them to learn unified user representations. However, these methods directly use modal features for representation learning, neglecting the impact of inherent modal noise. We argue that internal-modal noise and cross-modal noise hinder the acquisition of more accurate user representations. To address this problem, we propose SGP4SR - Separated-modality Guided user Preference learning for multimodal Sequential Recommendation. Globally, the user preference modeling is carried out from a separated-modality perspective to alleviate cross-modal noise. Locally, for each individual modality, we use item relationship graphs and user interest centers, aggregated with ID embeddings, to replace direct modal features, thereby mitigating internal-modal noise. Finally, user representations from both separated-modality and multimodal perspectives participate in prediction independently. In experiments conducted on four real-world datasets, our method outperforms state-of-the-art approaches, achieving an average performance improvement of up to 8.84% over the best baseline. The comprehensive experiments further validate the superior noise tolerance and robustness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2abe0505-5d53-45b5-a643-21cb8bb46eb8Builds on13
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh et al.ACL 2020 · 584 citations
- Sequential Recommendation with Graph Neural NetworksJianxin Chang, Chen Gao, Yu Zheng, Yiqun Hui et al.SIGIR 2021 · 435 citations
- Mining Latent Structures for Multimedia RecommendationJinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu et al.ACM MM 2021 · 350 citations
- Towards Universal Sequence Representation Learning for Recommender SystemsYupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li et al.KDD 2022 · 245 citations
- A Tale of Two Graphs: Freezing and Denoising Graph Structures for Multimodal RecommendationXin Zhou, Zhiqi ShenACM MM 2023 · 234 citations
Related papers
- Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential RecommendationShengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang et al.WWW 2025 · 29 citations
- DMMD4SR: Diffusion Model-based Multi-level Multimodal Denoising for Sequential RecommendationWeihai Lu, Li YinACM MM 2025 · 11 citations
- MISSRec: Pre-training and Transferring Multi-modal Interest-aware Sequence Representation for RecommendationJinpeng Wang, Ziyun Zeng, Yunxiao Wang, Yuting Wang et al.ACM MM 2023 · 62 citations
- Sign-Aware Multimodal Graph RecommendationYahong Lian, Haotian Tian, Chunyao Song, Tingjian GeAAAI 2026
- MENTOR: Multi-level Self-supervised Learning for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.AAAI 2025 · 21 citations
