From Abstract to Details: A Generative Multimodal Fusion Framework for Recommendation
Fangxiong Xiao, Lixi Deng, Jingjing Chen, Houye Ji, Xiaorui Yang, Zhuoye Ding, Bo Long
Abstract
In E-commerce recommendation, Click-Through Rate (CTR) prediction has been extensively studied in both academia and industry to enhance user experience and platform benefits. At present, most popular CTR prediction methods are concatenation-based models that represent items by simply merging multiple heterogeneous features including ID, visual, and text features into a large vector. As these heterogeneous modalities have moderately different properties, directly concatenating them without mining the correlation and reducing the redundancy are unlikely to achieve the optimal fusion results. Besides, these concatenation-based models treat all modalities equally for each user and overlook the fact that users tend to pay unequal attention to information of various modalities when browsing items in the real scenario. To address the above issues, this paper proposes a generative multimodal fusion framework (GMMF) for CTR prediction task. To eliminate the redundancy and strength the complementary of multimodal features, GMMF generates the new visual and text representations by a Difference-Set network (DSN). These representations are non-overlapping with the information conveyed by ID embedding. Specifically, DSN maps ID embedding into visual and text modalities and depicts the difference between multiple modalities based on their properties. Besides, GMMF learns unequal weights to multiple modalities with a Modal-Interest network (MIN) modeling users' preference on heterogeneous modalities. These weights reflect the usual habits and hobbies of users. Finally, We conduct extensive experiments on both public and collected industrial datasets, and the results show that GMMF greatly improves performance and achieves state-of-the-art performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5796065d-9e03-431d-ad50-3c7366b8036bCited by top-tier papers2
- Multi-Modal Multi-Behavior Sequential Recommendation with Conditional Diffusion-Based Feature DenoisingXiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li et al.SIGIR 2025 · 21 citations
- Diffusion-based Multi-modal Synergy Interest Network for Click-through Rate PredictionXiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li et al.SIGIR 2025 · 15 citations
Related papers
- Adversarial Multimodal Representation Learning for Click-Through Rate PredictionXiang Li, Chao Wang, Jiwei Tan, Xiaoyi Zeng et al.WWW 2020 · 61 citations
- Deep Match to Rank Model for Personalized Click-Through Rate PredictionZequn Lyu, Yu Dong, Chengfu Huo, Weijun RenAAAI 2020 · 73 citations
- Dual Graph enhanced Embedding Neural Network for CTR PredictionWei Guo, Rong Su, Renhao Tan, Huifeng Guo et al.KDD 2021 · 72 citations
- Neighbour Interaction based Click-Through Rate Prediction via Graph-masked TransformerErxue Min, Yu Rong, Tingyang Xu, Yatao Bian et al.SIGIR 2022 · 46 citations
- MaskFusion: Feature Augmentation for Click-Through Rate Prediction via Input-adaptive Mask FusionChao Liao, Jianchao Tan, Jiyuan Jia, Yi Guo et al.ICLR 2023
