Multimodal Counterfactual Learning Network for Multimedia-based Recommendation
Shuaiyang Li, Dan Guo, Kang Liu, Richang Hong, Feng Xue
Abstract
Multimedia-based recommendation (MMRec) utilizes multimodal content (images, textual descriptions, etc.) as auxiliary information on historical interactions to determine user preferences. Most MMRec approaches predict user interests by exploiting a large amount of multimodal contents of user-interacted items, ignoring the potential effect of multimodal content of user-uninteracted items. As a matter of fact, there is a small portion of user preference-irrelevant features in the multimodal content of user-interacted items, which may be a kind of spurious correlation with user preferences, thereby degrading the recommendation performance. In this work, we argue that the multimodal content of user-uninteracted items can be further exploited to identify and eliminate the user preference-irrelevant portion inside user-interacted multimodal content, for example by counterfactual inference of causal theory. Going beyond multimodal user preference modeling only using interacted items, we propose a novel model called Multimodal Counterfactual Learning Network (MCLN), in which user-uninteracted items' multimodal content is additionally exploited to further purify the representation of user preference-relevant multimodal content that better matches the user's interests, yielding state-of-the-art performance. Extensive experiments are conducted to validate the effectiveness and rationality of MCLN. We release the complete codes of MCLN at https://github.com/hfutmars/MCLN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 855f6a3e-a8a8-49fe-a762-a673898105e2Cited by top-tier papers2
- Generating with Fairness: A Modality-Diffused Counterfactual Framework for Incomplete Multimodal RecommendationsJin Li, Shoujin Wang, Qi Zhang, Shui Yu et al.WWW 2025 · 26 citations
- Bidirectional Counterfactual Distillation for Review-Based RecommendationSheng Sang, Shujie Li, Shuaiyang Li, Kang Liu et al.AAAI 2026
Builds on13
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Revisiting Graph Based Collaborative Filtering: A Linear Residual Graph Convolutional Network ApproachLei Chen, Le Wu, Richang Hong, Kun Zhang et al.AAAI 2020 · 634 citations
- Causal Intervention for Leveraging Popularity Bias in RecommendationYang Zhang, Fuli Feng, Xiangnan He, Tianxin Wei et al.SIGIR 2021 · 431 citations
- Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identificationYongming Rao, Guangyi Chen, Jiwen Lu, Jie ZhouICCV 2021 · 330 citations
- Model-Agnostic Counterfactual Reasoning for Eliminating Popularity Bias in Recommender SystemTianxin Wei, Fuli Feng, Jiawei Chen, Ziwei Wu et al.KDD 2021 · 246 citations
Related papers
- LightGT: A Light Graph Transformer for Multimedia RecommendationYinwei Wei, Wenqi Liu, Fan Liu, Xiang Wang et al.SIGIR 2023 · 71 citations
- Learning Hybrid Behavior Patterns for Multimedia RecommendationZongshen Mu, Yueting Zhuang, Jie Tan, Jun Xiao et al.ACM MM 2022 · 47 citations
- Invariant Representation Learning for Multimedia RecommendationXiaoyu Du, Zike Wu, Fuli Feng, Xiangnan He et al.ACM MM 2022 · 45 citations
- STARLINE: Contrastive Learning with Modality-Aware Graph Refinement for Effective Multimedia RecommendationTaeri Kim, Sohee Ban, Hyunjoon Kim, Sang-Wook KimKDD 2025 · 1 citation
- MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal RecommendationYuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan et al.SIGIR 2026 · 1 citation
