Harnessing Multimodal Large Language Models for Multimodal Sequential Recommendation
Yuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang, Hengruo Zhang, Peijun Zhu, Runlong Yu, Kai Zhang, Hui Xiong
摘要
Recent advances in Large Language Models (LLMs) have demonstrated significant potential in the field of Recommendation Systems (RSs). Most existing studies have focused on converting user behavior logs into textual prompts and leveraging techniques such as prompt tuning to enable LLMs for recommendation tasks. Meanwhile, research interest has recently grown in multimodal recommendation systems that integrate data from images, text, and other sources using modality fusion techniques. This introduces new challenges to the existing LLM-based recommendation paradigm which relies solely on text modality information. Moreover, although Multimodal Large Language Models (MLLMs) capable of processing multi-modal inputs have emerged, how to equip MLLMs with multi-modal recommendation capabilities remains largely unexplored. To this end, in this paper, we propose the Multimodal Large Language Model-enhanced Multimodal Sequential Recommendation (MLLM-MSR) model. To capture the dynamic user preference, we design a twostage user preference summarization method. Specifically, we first utilize an MLLM-based item-summarizer to extract image feature given an item and convert the image into text. Then, we employ a recurrent user preference summarization generation paradigm to capture the dynamic changes in user preferences based on an LLM-based user-summarizer. Finally, to enable the MLLM for multi-modal recommendation task, we propose to fine-tune a MLLM-based recommender using Supervised Fine-Tuning (SFT) techniques. Extensive evaluations across various datasets validate the effectiveness of MLLM-MSR, showcasing its superior ability to capture and adapt to the evolving dynamics of user preferences. The code of our work is publicly available at https://github.com/YuyangYe/MLLM-MSR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential RecommendationShengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang 等WWW 2025 · 被引用 29 次
- When Large Vision Language Models Meet Multimodal Sequential Recommendation: An Empirical StudyPeilin Zhou, Chao Liu, Jing Ren, Xinfeng Zhou 等WWW 2025 · 被引用 21 次
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential RecommendationYu Wang, Yonghui Yang, Le Wu, Yi Zhang 等SIGIR 2026 · 被引用 9 次
- FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation LearningWei Yang, Rui Zhong, Yiqun Chen, Shixuan Li 等ACM MM 2025 · 被引用 6 次
- When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video RecommendationSiran Chen, Boyu Chen, Chenyun Yu, Yi Ouyang 等AAAI 2026 · 被引用 5 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Generating Images with Multimodal Language ModelsJing Yu Koh, Daniel Fried, Russ SalakhutdinovNeurIPS 2023 · 被引用 403 次
- Representation Learning with Large Language Models for RecommendationXubin Ren, Wei Wei, Lianghao Xia, Lixin Su 等WWW 2024 · 被引用 385 次
- Learning Vector-Quantized Item Representation for Transferable Sequential RecommendersYupeng Hou, Zhankui He, Julian J. McAuley, Wayne Xin ZhaoWWW 2023 · 被引用 256 次
相关 Paper
- Harnessing Large Language Models for Text-Rich Sequential RecommendationZhi Zheng, Wenshuo Chao, Zhaopeng Qiu, Hengshu Zhu 等WWW 2024 · 被引用 114 次
- Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?Sein Kim, Hongseok Kang, Kibum Kim, Jiwan Kim 等KDD 2025 · 被引用 3 次
- MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal RecommendationYuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan 等SIGIR 2026 · 被引用 1 次
- MSR-Rec: Multi-Step Reasoning-Enhanced LLM for Sequential RecommendationTuo Wang, Meng Jian, Ge Shi, Lifang Wu 等AAAI 2026
- LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential RecommendationYingzhi He, Xiaohao Liu, An Zhang, Yunshan Ma 等KDD 2025 · 被引用 2 次
