Harnessing Multimodal Large Language Models for Multimodal Sequential Recommendation
Yuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang, Hengruo Zhang, Peijun Zhu, Runlong Yu, Kai Zhang, Hui Xiong
Abstract
Recent advances in Large Language Models (LLMs) have demonstrated significant potential in the field of Recommendation Systems (RSs). Most existing studies have focused on converting user behavior logs into textual prompts and leveraging techniques such as prompt tuning to enable LLMs for recommendation tasks. Meanwhile, research interest has recently grown in multimodal recommendation systems that integrate data from images, text, and other sources using modality fusion techniques. This introduces new challenges to the existing LLM-based recommendation paradigm which relies solely on text modality information. Moreover, although Multimodal Large Language Models (MLLMs) capable of processing multi-modal inputs have emerged, how to equip MLLMs with multi-modal recommendation capabilities remains largely unexplored. To this end, in this paper, we propose the Multimodal Large Language Model-enhanced Multimodal Sequential Recommendation (MLLM-MSR) model. To capture the dynamic user preference, we design a twostage user preference summarization method. Specifically, we first utilize an MLLM-based item-summarizer to extract image feature given an item and convert the image into text. Then, we employ a recurrent user preference summarization generation paradigm to capture the dynamic changes in user preferences based on an LLM-based user-summarizer. Finally, to enable the MLLM for multi-modal recommendation task, we propose to fine-tune a MLLM-based recommender using Supervised Fine-Tuning (SFT) techniques. Extensive evaluations across various datasets validate the effectiveness of MLLM-MSR, showcasing its superior ability to capture and adapt to the evolving dynamics of user preferences. The code of our work is publicly available at https://github.com/YuyangYe/MLLM-MSR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bcd02e5-7add-4e2a-925d-7f97279191ccCited by top-tier papers24
- Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential RecommendationShengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang et al.WWW 2025 · 29 citations
- When Large Vision Language Models Meet Multimodal Sequential Recommendation: An Empirical StudyPeilin Zhou, Chao Liu, Jing Ren, Xinfeng Zhou et al.WWW 2025 · 21 citations
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential RecommendationYu Wang, Yonghui Yang, Le Wu, Yi Zhang et al.SIGIR 2026 · 9 citations
- FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation LearningWei Yang, Rui Zhong, Yiqun Chen, Shixuan Li et al.ACM MM 2025 · 6 citations
- When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video RecommendationSiran Chen, Boyu Chen, Chenyun Yu, Yi Ouyang et al.AAAI 2026 · 5 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Generating Images with Multimodal Language ModelsJing Yu Koh, Daniel Fried, Russ SalakhutdinovNeurIPS 2023 · 403 citations
- Representation Learning with Large Language Models for RecommendationXubin Ren, Wei Wei, Lianghao Xia, Lixin Su et al.WWW 2024 · 385 citations
- Learning Vector-Quantized Item Representation for Transferable Sequential RecommendersYupeng Hou, Zhankui He, Julian J. McAuley, Wayne Xin ZhaoWWW 2023 · 256 citations
Related papers
- Harnessing Large Language Models for Text-Rich Sequential RecommendationZhi Zheng, Wenshuo Chao, Zhaopeng Qiu, Hengshu Zhu et al.WWW 2024 · 114 citations
- Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?Sein Kim, Hongseok Kang, Kibum Kim, Jiwan Kim et al.KDD 2025 · 3 citations
- MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal RecommendationYuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan et al.SIGIR 2026 · 1 citation
- MSR-Rec: Multi-Step Reasoning-Enhanced LLM for Sequential RecommendationTuo Wang, Meng Jian, Ge Shi, Lifang Wu et al.AAAI 2026
- LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential RecommendationYingzhi He, Xiaohao Liu, An Zhang, Yunshan Ma et al.KDD 2025 · 2 citations
