DMGIN: How Multimodal LLMs Enhance Large Recommendation Models for Lifelong User Post-click Behaviors
Zhuoxing Wei, Qingchen Xie, Qi Liu, Jingsong Yu
Abstract
Modeling user interest based on lifelong user behavior sequences is crucial for enhancing Click-Through Rate (CTR) prediction. However, long post-click behavior sequences themselves pose severe performance issues: the sheer volume of data leads to high computational costs and inefficiencies in model training and inference. Traditional methods address this by introducing two-stage approaches, but this compromises model effectiveness due to incomplete utilization of the full sequence context. More importantly, integrating multimodal embeddings into existing large recommendation models (LRM) presents significant challenges: These embeddings often exacerbate computational burdens and mismatch with LRM architectures. To address these issues and enhance the model's efficiency and accuracy, we introduce Deep Multimodal Group Interest Network (DMGIN). Given the observation that user post-click behavior sequences contain a large number of repeated items with varying behaviors and timestamps, DMGIN employs Multimodal LLMs(MLLM) for grouping to reorganize complete lifelong post-click behavior sequences more effectively, with almost no additional computational overhead, as opposed to directly introducing multimodal embeddings. To mitigate the potential information loss from grouping, we have implemented two key strategies. First, we analyze behaviors within each group using both interest statistics and intra-group transformers to capture group traits. Second, apply inter-group transformers to temporally ordered groups to capture the evolution of user group interests. Our extensive experiments on both industrial and public datasets confirm the effectiveness and efficiency of DMGIN. The A/B test in our LBS advertising system shows that DMGIN improves CTR by 4.7% and Revenue per Mile by 2.3%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- Diffusion-based Multi-modal Synergy Interest Network for Click-through Rate PredictionXiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li et al.SIGIR 2025 · 15 citations
- From Abstract to Details: A Generative Multimodal Fusion Framework for RecommendationFangxiong Xiao, Lixi Deng, Jingjing Chen, Houye Ji et al.ACM MM 2022 · 15 citations
- Dual Graph enhanced Embedding Neural Network for CTR PredictionWei Guo, Rong Su, Renhao Tan, Huifeng Guo et al.KDD 2021 · 72 citations
- HyMiRec: A Hybrid Multi-interest Learning Framework for LLM-based Sequential RecommendationJingyi Zhou, Cheng Chen, Kai Zuo, Manjie Xu et al.WWW 2026 · 2 citations
- Neighbour Interaction based Click-Through Rate Prediction via Graph-masked TransformerErxue Min, Yu Rong, Tingyang Xu, Yatao Bian et al.SIGIR 2022 · 46 citations
