Adversarial Multimodal Representation Learning for Click-Through Rate Prediction
Xiang Li, Chao Wang, Jiwei Tan, Xiaoyi Zeng, Dan Ou, Bo Zheng
Abstract
For better user experience and business effectiveness, Click-Through Rate (CTR) prediction has been one of the most important tasks in E-commerce. Although extensive CTR prediction models have been proposed, learning good representation of items from multimodal features is still less investigated, considering an item in E-commerce usually contains multiple heterogeneous modalities. Previous works either concatenate the multiple modality features, that is equivalent to giving a fixed importance weight to each modality; or learn dynamic weights of different modalities for different items through technique like attention mechanism. However, a problem is that there usually exists common redundant information across multiple modalities. The dynamic weights of different modalities computed by using the redundant information may not correctly reflect the different importance of each modality. To address this, we explore the complementarity and redundancy of modalities by considering modality-specific and modality-invariant features differently. We propose a novel Multimodal Adversarial Representation Network (MARN) for the CTR prediction task. A multimodal attention network first calculates the weights of multiple modalities for each item according to its modality-specific features. Then a multimodal adversarial network learns modality-invariant representations where a double-discriminators strategy is introduced. Finally, we achieve the multimodal item representations by combining both modality-specific and modality-invariant representations. We conduct extensive experiments on both public and industrial datasets, and the proposed method consistently achieves remarkable improvements to the state-of-the-art methods. Moreover, the approach has been deployed in an operational E-commerce system and online A/B testing further demonstrates the effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4df6aace-69c5-4e15-b537-7fbb42e9c4adCited by top-tier papers6
- Learning Modality-Specific and -Agnostic Representations for Asynchronous Multimodal Language SequencesDingkang Yang, Haopeng Kuang, Shuai Huang, Lihua ZhangACM MM 2022 · 64 citations
- Diffusion-based Multi-modal Synergy Interest Network for Click-through Rate PredictionXiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li et al.SIGIR 2025 · 15 citations
- MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product UnderstandingZhanheng Nie, Chenghan Fu, Daoze Zhang, Junxian Wu et al.CVPR 2026 · 9 citations
- Micro-video Tagging via Jointly Modeling Social Influence and Tag RelationXiao Wang, Tian Gan, Yinwei Wei, Jianlong Wu et al.ACM MM 2022 · 8 citations
- Cross-Contrastive Clustering for Multimodal Attributed Graphs with Dual Graph FilteringHaoran Zheng, Renchi Yang, Hongtao Wang, Jianliang XuKDD 2026 · 7 citations
Related papers
- From Abstract to Details: A Generative Multimodal Fusion Framework for RecommendationFangxiong Xiao, Lixi Deng, Jingjing Chen, Houye Ji et al.ACM MM 2022 · 15 citations
- Multi-Domain Deep Learning from a Multi-View Perspective for Cross-Border E-commerce SearchYiqian Zhang, Yinfu Feng, Wen-Ji Zhou, Yunan Ye et al.AAAI 2024 · 9 citations
- HIEN: Hierarchical Intention Embedding Network for Click-Through Rate PredictionZuowu Zheng, Changwang Zhang, Xiaofeng Gao, Guihai ChenSIGIR 2022 · 16 citations
- Enhancing CTR Prediction with Context-Aware Feature Representation LearningFangye Wang, Yingxu Wang, Dongsheng Li, Hansu Gu et al.SIGIR 2022 · 43 citations
- Dual Adversarial Graph Neural Networks for Multi-label Cross-modal RetrievalShengsheng Qian, Dizhan Xue, Huaiwen Zhang, Quan Fang et al.AAAI 2021 · 72 citations
