From Abstract to Details: A Generative Multimodal Fusion Framework for Recommendation
Fangxiong Xiao, Lixi Deng, Jingjing Chen, Houye Ji, Xiaorui Yang, Zhuoye Ding, Bo Long
摘要
In E-commerce recommendation, Click-Through Rate (CTR) prediction has been extensively studied in both academia and industry to enhance user experience and platform benefits. At present, most popular CTR prediction methods are concatenation-based models that represent items by simply merging multiple heterogeneous features including ID, visual, and text features into a large vector. As these heterogeneous modalities have moderately different properties, directly concatenating them without mining the correlation and reducing the redundancy are unlikely to achieve the optimal fusion results. Besides, these concatenation-based models treat all modalities equally for each user and overlook the fact that users tend to pay unequal attention to information of various modalities when browsing items in the real scenario. To address the above issues, this paper proposes a generative multimodal fusion framework (GMMF) for CTR prediction task. To eliminate the redundancy and strength the complementary of multimodal features, GMMF generates the new visual and text representations by a Difference-Set network (DSN). These representations are non-overlapping with the information conveyed by ID embedding. Specifically, DSN maps ID embedding into visual and text modalities and depicts the difference between multiple modalities based on their properties. Besides, GMMF learns unequal weights to multiple modalities with a Modal-Interest network (MIN) modeling users' preference on heterogeneous modalities. These weights reflect the usual habits and hobbies of users. Finally, We conduct extensive experiments on both public and collected industrial datasets, and the results show that GMMF greatly improves performance and achieves state-of-the-art performance.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Multi-Modal Multi-Behavior Sequential Recommendation with Conditional Diffusion-Based Feature DenoisingXiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li 等SIGIR 2025 · 被引用 21 次
- Diffusion-based Multi-modal Synergy Interest Network for Click-through Rate PredictionXiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li 等SIGIR 2025 · 被引用 15 次
相关 Paper
- Adversarial Multimodal Representation Learning for Click-Through Rate PredictionXiang Li, Chao Wang, Jiwei Tan, Xiaoyi Zeng 等WWW 2020 · 被引用 61 次
- Deep Match to Rank Model for Personalized Click-Through Rate PredictionZequn Lyu, Yu Dong, Chengfu Huo, Weijun RenAAAI 2020 · 被引用 73 次
- Dual Graph enhanced Embedding Neural Network for CTR PredictionWei Guo, Rong Su, Renhao Tan, Huifeng Guo 等KDD 2021 · 被引用 72 次
- Neighbour Interaction based Click-Through Rate Prediction via Graph-masked TransformerErxue Min, Yu Rong, Tingyang Xu, Yatao Bian 等SIGIR 2022 · 被引用 46 次
- MaskFusion: Feature Augmentation for Click-Through Rate Prediction via Input-adaptive Mask FusionChao Liao, Jianchao Tan, Jiyuan Jia, Yi Guo 等ICLR 2023
