Multimodal-aware Multi-intention Learning for Recommendation
Wei Yang, Qingchen Yang
Abstract
Whether it is an e-commerce platform or a short video platform, the effective use of multi-modal data plays an important role in the recommendation system. More and more researchers are exploring how to effectively use multimodal signals to entice more users to buy goods or watch short videos. Some studies have added multimodal features as side information to the model and achieved certain results. In practice, the purchase behavior of users mainly depends on some personalized intentions. However, it is difficult for neural networks to process noise information and extract high-level intention information effectively. To investigate the benefits of latent intentions and leverage them effectively for recommendation, we propose a Multimodal-aware Multi-intention Learning method for recommendation (MMIL). Specifically, we first construct a multi-intention recommendation framework based on the probability distribution relationship among predicted objectives and intentions, while avoiding intention overfitting. We then design an intention representation learning module to learn accurate multiple intention representations based on intention prototypes. Further, we propose a multi-modal intention perceptron module to learn multi-modal intention representations. In addition, we design self-supervised intention-based and representation-based contrastive objectives to achieve cross-modal representation alignment. On three real-world data sets, the proposed MMIL method outperforms other advanced techniques. The effectiveness of intention modeling and intention alignment is verified by comprehensive experiments. The source code is available at: https://github.com/ml-mindset/MMIL.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5f916f74-7ea3-4dd0-b6a7-4d9fefad24dbCited by top-tier papers5
- Structured Spectral Reasoning for Frequency-Adaptive Multimodal RecommendationWei Yang, Rui Zhong, Yiqun Chen, Chi Lu et al.NeurIPS 2025 · 10 citations
- FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation LearningWei Yang, Rui Zhong, Yiqun Chen, Shixuan Li et al.ACM MM 2025 · 6 citations
- Leveraging Multimodal Data and Side Users for Diffusion Cross-Domain RecommendationFan Zhang, Jinpeng Chen, Huan Li, Senzhang Wang et al.ACM MM 2025 · 5 citations
- DiMA: Distinguishing Resident and Tourist Preferences via Multi-Modal LLM Alignment for Out-of-Town Cross-Domain RecommendationFan Zhang, Jinpeng Chen, Tao Wang, Huan Li et al.AAAI 2026
- TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal RecommendationWei Yang, Rui Zhong, Zihan Lin, Xiaodan Wang et al.SIGIR 2026
Related papers
- Multi-Modal Self-Supervised Learning for RecommendationWei Wei, Chao Huang, Lianghao Xia, Chuxu ZhangWWW 2023 · 256 citations
- Modal-aware Bias Constrained Contrastive Learning for Multimodal RecommendationWei Yang, Zhengru Fang, Tianle Zhang, Shiguang Wu et al.ACM MM 2023 · 15 citations
- MENTOR: Multi-level Self-supervised Learning for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.AAAI 2025 · 21 citations
- Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential RecommendationShengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang et al.WWW 2025 · 29 citations
- Disentangling Long and Short-Term Interests for RecommendationYu Zheng, Chen Gao, Jianxin Chang, Yanan Niu et al.WWW 2022 · 128 citations
