Bootstrap Latent Representations for Multi-modal Recommendation
Xin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng, Chunyan Miao, Pengwei Wang, Yuan You, Feijun Jiang
摘要
This paper studies the multi-modal recommendation problem, where the item multi-modality information (e.g., images and textual descriptions) is exploited to improve the recommendation accuracy. Besides the user-item interaction graph, existing state-of-the-art methods usually use auxiliary graphs (e.g., user-user or item-item relation graph) to augment the learned representations of users and/or items. These representations are often propagated and aggregated on auxiliary graphs using graph convolutional networks, which can be prohibitively expensive in computation and memory, especially for large graphs. Moreover, existing multi-modal recommendation methods usually leverage randomly sampled negative examples in Bayesian Personalized Ranking (BPR) loss to guide the learning of user/item representations, which increases the computational cost on large graphs and may also bring noisy supervision signals into the training process. To tackle the above issues, we propose a novel self-supervised multi-modal recommendation model, dubbed BM3, which requires neither augmentations from auxiliary graphs nor negative samples. Specifically, BM3 first bootstraps latent contrastive views from the representations of users and items with a simple dropout augmentation. It then jointly optimizes three multimodal objectives to learn the representations of users and items by reconstructing the user-item interaction graph and aligning modality features under both inter-and intra-modality perspectives. BM3 alleviates both the need for contrasting with negative examples and the complex graph augmentation from an additional target network for contrastive view generation. We show BM3 outperforms prior recommendation models on three datasets with Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper55
- A Tale of Two Graphs: Freezing and Denoising Graph Structures for Multimodal RecommendationXin Zhou, Zhiqi ShenACM MM 2023 · 被引用 234 次
- Multi-View Graph Convolutional Network for Multimedia RecommendationPenghang Yu, Zhiyi Tan, Guanming Lu, Bing-Kun BaoACM MM 2023 · 被引用 181 次
- LGMRec: Local and Global Graph Learning for Multimodal RecommendationZhiqiang Guo, Jianjun Li, Guohui Li, Chaoyang Wang 等AAAI 2024 · 被引用 164 次
- DiffMM: Multi-Modal Diffusion Model for RecommendationYangqin Jiang, Lianghao Xia, Wei Wei, Da Luo 等ACM MM 2024 · 被引用 92 次
- Layer-refined Graph Convolutional Networks for RecommendationXin Zhou, Donghui Lin, Yong Liu, Chunyan MiaoICDE 2023 · 被引用 81 次
它引用的顶会 Paper17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding 等ICML 2020 · 被引用 1,910 次
相关 Paper
- Modal-aware Bias Constrained Contrastive Learning for Multimodal RecommendationWei Yang, Zhengru Fang, Tianle Zhang, Shiguang Wu 等ACM MM 2023 · 被引用 15 次
- Multi-Modal Self-Supervised Learning for RecommendationWei Wei, Chao Huang, Lianghao Xia, Chuxu ZhangWWW 2023 · 被引用 256 次
- MENTOR: Multi-level Self-supervised Learning for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li 等AAAI 2025 · 被引用 21 次
- Refining Contrastive Learning and Homography Relations for Multi-Modal RecommendationShouxing Ma, Yawen Zeng, Shiqing Wu, Guandong XuACM MM 2025 · 被引用 3 次
- Multi-view Semantic Contrastive Alignment for Multimodal RecommendationJiuqiang Li, Hongjun WangWWW 2026
