MMMLP: Multi-modal Multilayer Perceptron for Sequential Recommendations
Jiahao Liang, Xiangyu Zhao, Muyang Li, Zijian Zhang, Wanyu Wang, Haochen Liu, Zitao Liu
Abstract
Sequential recommendation aims to offer potentially interesting products to users by capturing their historical sequence of interacted items. Although it has facilitated extensive physical scenarios, sequential recommendation for multi-modal sequences has long been neglected. Multi-modal data that depicts a user’s historical interactions exists ubiquitously, such as product pictures, textual descriptions, and interacted item sequences, providing semantic information from multiple perspectives that comprehensively describe a user’s preferences. However, existing sequential recommendation methods either fail to directly handle multi-modality or suffer from high computational complexity. To address this, we propose a novel Multi-Modal Multi-Layer Perceptron (MMMLP) for maintaining multi-modal sequences for sequential recommendation. MMMLP is a purely MLP-based architecture that consists of three modules - the Feature Mixer Layer, Fusion Mixer Layer, and Prediction Layer - and has an edge on both efficacy and efficiency. Extensive experiments show that MMMLP achieves state-of-the-art performance with linear complexity. We also conduct ablating analysis to verify the contribution of each component. Furthermore, compatible experiments are devised, and the results show that the multi-modal representation learned by our proposed model generally benefits other recommendation models, emphasizing our model’s ability to handle multi-modal information. We have made our code available online to ease reproducibility1.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c84fa48c-72de-4289-9f7d-873a8c2940e2Cited by top-tier papers23
- LLM-ESR: Large Language Models Enhancement for Long-tailed Sequential RecommendationQidong Liu, Xian Wu, Yejing Wang, Zijian Zhang et al.NeurIPS 2024 · 154 citations
- MISSRec: Pre-training and Transferring Multi-modal Interest-aware Sequence Representation for RecommendationJinpeng Wang, Ziyun Zeng, Yunxiao Wang, Yuting Wang et al.ACM MM 2023 · 62 citations
- LLM4Rerank: LLM-based Auto-Reranking Framework for RecommendationsJingtong Gao, Bo Chen, Xiangyu Zhao, Weiwen Liu et al.WWW 2025 · 50 citations
- Disentangling ID and Modality Effects for Session-based RecommendationXiaokun Zhang, Bo Xu, Zhaochun Ren, Xiaochen Wang et al.SIGIR 2024 · 32 citations
- SIGMA: Selective Gated Mamba for Sequential RecommendationZiwei Liu, Qidong Liu, Yejing Wang, Wanyu Wang et al.AAAI 2025 · 31 citations
Related papers
- CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.ICDE 2026 · 1 citation
- Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential RecommendationShengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang et al.WWW 2025 · 29 citations
- CMCLRec: Cross-modal Contrastive Learning for User Cold-start Sequential RecommendationXiaolong Xu, Hongsheng Dong, Lianyong Qi, Xuyun Zhang et al.SIGIR 2024 · 56 citations
- Online Distillation-enhanced Multi-modal Transformer for Sequential RecommendationWei Ji, Xiangyan Liu, An Zhang, Yinwei Wei et al.ACM MM 2023 · 32 citations
- Multi-Modal Multi-Behavior Sequential Recommendation with Conditional Diffusion-Based Feature DenoisingXiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li et al.SIGIR 2025 · 21 citations
