Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation
Yu Wang, Yonghui Yang, Le Wu, Yi Zhang, Fei Liu, Richang Hong
Abstract
Recent advances in Large Language Models (LLMs) have opened new avenues for sequential recommendation by enabling natural language reasoning over user behavior sequences. A common approach formulates recommendation as a language modeling task, where interaction histories are transformed into prompts and user preferences are learned via supervised fine-tuning. However, these methods operate solely in the textual modality and often miss users' fine-grained interests, especially when shaped by rich visual signals such as product images or movie posters. Multimodal Large Language Models (MLLMs) offer a promising alternative by aligning text and vision in a shared semantic space. A prevalent training paradigm applies Supervised Fine-Tuning (SFT) followed by Direct Preference Optimization (DPO) to model user preferences. Yet, two core challenges remain: 1) Imbalanced sample hardness, where random negative sampling causes overfitting on easy examples and undertraining on hard ones; 2) Cross-modal semantic bias, where the fixed reference model in DPO prevents the policy model from correcting modality misalignments-especially over long sequences. To address these issues, we propose a Multimodal LLM framework that integrates Hardness-aware and Noise-regularized preference optimization for Recommendation (HaNoRec) . Specifically, HaNoRec dynamically adjusts optimization weights based on both the estimated hardness of each training sample and the policy model's real-time responsiveness, prioritizing harder examples. It further introduces Gaussian-perturbed distribution optimization on output logits to enhance cross-modal semantic consistency and reduce modality bias inherited from the reference model. Experiments on three benchmarks demonstrate its superiority over state-of-the-art methods. The code is available at http://github.com/wangyu0627/HaNoRec.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 635e74b7-5d1c-4f55-bbc9-30d6fcc2eeb0Cited by top-tier papers2
- DIAURec: Dual-Intent Space Representation Optimization for RecommendationYu Zhang, Yiwen Zhang, Yi Zhang, Lei SangSIGIR 2026
- ProMax: Exploring the Potential of LLM-derived Profiles with Distribution Shaping for Recommender SystemsYi Zhang, Yiwen Zhang, Kai Zheng, Tong Chen et al.SIGIR 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
Related papers
- Harnessing Multimodal Large Language Models for Multimodal Sequential RecommendationYuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang et al.AAAI 2025 · 68 citations
- MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal RecommendationYuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan et al.SIGIR 2026 · 1 citation
- LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential RecommendationYingzhi He, Xiaohao Liu, An Zhang, Yunshan Ma et al.KDD 2025 · 2 citations
- Boosting Guided Diffusion with Large Language Models for Multimodal Sequential RecommendationTe Song, Lianyong Qi, Weiming Liu, Fan Wang et al.ACM MM 2025 · 1 citation
- Factorized Latent Reasoning for LLM-based RecommendationTianqi Gao, Chengkai Huang, Zihan Wang, Cao Liu et al.SIGIR 2026
