Mining Latent Structures for Multimedia Recommendation
Jinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu, Shuhui Wang, Liang Wang
Abstract
Multimedia content is of predominance in the modern Web era. Investigating how users interact with multimodal items is a continuing concern within the rapid development of recommender systems. The majority of previous work focuses on modeling user-item interactions with multimodal features included as side information. However, this scheme is not well-designed for multimedia recommendation. Specifically, only collaborative item-item relationships are implicitly modeled through high-order item-user-item relations. Considering that items are associated with rich contents in multiple modalities, we argue that the latent semantic item-item structures underlying these multimodal contents could be beneficial for learning better item representations and further boosting recommendation. To this end, we propose a LATent sTructure mining method for multImodal reCommEndation, which we term LATTICE for brevity. To be specific, in the proposed LATTICE model, we devise a novel modality-aware structure learning layer, which learns item-item structures for each modality and aggregates multiple modalities to obtain latent item graphs. Based on the learned latent graphs, we perform graph convolutions to explicitly inject high-order item affinities into item representations. These enriched item representations can then be plugged into existing collaborative filtering methods to make more accurate recommendations. Extensive experiments on three real-world datasets demonstrate the superiority of our method over state-of-the-art multimedia recommendation methods and validate the efficacy of mining latent item-item relationships from multimodal features.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers80
- Bootstrap Latent Representations for Multi-modal RecommendationXin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng et al.WWW 2023 · 326 citations
- Multi-Modal Self-Supervised Learning for RecommendationWei Wei, Chao Huang, Lianghao Xia, Chuxu ZhangWWW 2023 · 256 citations
- A Tale of Two Graphs: Freezing and Denoising Graph Structures for Multimodal RecommendationXin Zhou, Zhiqi ShenACM MM 2023 · 234 citations
- Multi-level Cross-view Contrastive Learning for Knowledge-aware Recommender SystemDing Zou, Wei Wei, Xian-Ling Mao, Ziyang Wang et al.SIGIR 2022 · 226 citations
- Multi-View Graph Convolutional Network for Multimedia RecommendationPenghang Yu, Zhiyi Tan, Guanming Lu, Bing-Kun BaoACM MM 2023 · 181 citations
Builds on10
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Graph Contrastive Learning with Adaptive AugmentationYanqiao Zhu, Yichen Xu, Feng Yu, Qiang Liu et al.WWW 2021 · 1,415 citations
- Graph Structure Learning for Robust Graph Neural NetworksWei Jin, Yao Ma, Xiaorui Liu, Xianfeng Tang et al.KDD 2020 · 604 citations
- Iterative Deep Graph Learning for Graph Neural Networks: Better and Robust Node EmbeddingsYu Chen, Lingfei Wu, Mohammed J. ZakiNeurIPS 2020 · 559 citations
- AM-GCN: Adaptive Multi-channel Graph Convolutional NetworksXiao Wang, Meiqi Zhu, Deyu Bo, Peng Cui et al.KDD 2020 · 464 citations
Related papers
- Learning Hybrid Behavior Patterns for Multimedia RecommendationZongshen Mu, Yueting Zhuang, Jie Tan, Jun Xiao et al.ACM MM 2022 · 47 citations
- MELON: Learning Multi-Aspect Modality Preferences for Accurate Multimedia RecommendationDongho Jeong, Taeri Kim, Donghyeon Cho, Sang-Wook KimSIGIR 2025 · 2 citations
- STARLINE: Contrastive Learning with Modality-Aware Graph Refinement for Effective Multimedia RecommendationTaeri Kim, Sohee Ban, Hyunjoon Kim, Sang-Wook KimKDD 2025 · 1 citation
- Subspace-Aware Graph Construction and Contrastive Alignment for Multimodal Recommendation with Large Language ModelsHaodong Li, Lianyong Qi, Weiming Liu, Fan Wang et al.AAAI 2026
- LightGT: A Light Graph Transformer for Multimedia RecommendationYinwei Wei, Wenqi Liu, Fan Liu, Xiang Wang et al.SIGIR 2023 · 71 citations
