Subspace-Aware Graph Construction and Contrastive Alignment for Multimodal Recommendation with Large Language Models
Haodong Li, Lianyong Qi, Weiming Liu, Fan Wang, Chong Li, Shengye Pang, Wenwen Gong, Yanwei Xu, Xiaoxiao Chi, Yang Zhang, Xiaokang Zhou
摘要
Multimedia content offers additional context for recommender systems to better understand user interests. Existing studies on multimodal recommendation primarily focus on constructing item-item semantic graphs. However, most of these methods capture only shallow semantic structures based on feature similarity and struggle to model more complex or cross-entity semantic relationships (e.g., user-item). Moreover, in these methods, collaborative signals often dominate and suppress semantic knowledge, which limits its role in representation learning. To address these issues, we propose SCALE, a novel framework that combines subspace-aware graph construction and contrastive alignment for multimodal recommendation with large language models. Specifically, we first use large language models and encoders to extract user and item features. Following the subspace clustering assumption, we apply the Orthogonal Matching Pursuit algorithm to mine complex semantic structures within the item-item, user-user, and user-item spaces, and integrate them into a unified semantic graph. We then perform graph convolution on both the semantic and interaction graphs, and aggregate the results for recommendation. Furthermore, contrastive losses are employed to enhance semantic fusion and alignment. Extensive experiments on five real-world datasets demonstrate that SCALE significantly outperforms state-of-the-art multimodal recommendation models, highlighting its effectiveness in modeling complex relationships and integrating semantic knowledge with collaborative signals.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- Representation Learning with Large Language Models for RecommendationXubin Ren, Wei Wei, Lianghao Xia, Lixin Su 等WWW 2024 · 被引用 385 次
- Mining Latent Structures for Multimedia RecommendationJinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu 等ACM MM 2021 · 被引用 350 次
- Bootstrap Latent Representations for Multi-modal RecommendationXin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng 等WWW 2023 · 被引用 326 次
- A Tale of Two Graphs: Freezing and Denoising Graph Structures for Multimodal RecommendationXin Zhou, Zhiqi ShenACM MM 2023 · 被引用 234 次
相关 Paper
- CCLRec: Consensus-driven Contrastive Learning for LLM-enhanced Graph RecommendationTing Guo, Dongyu Pei, Litiao Qiu, Xiaoying Liao 等ICML 2026
- MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal RecommendationYuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan 等SIGIR 2026 · 被引用 1 次
- MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender SystemsYibiao Wei, Jie Zou, Weikang Guo, Guoqing Wang 等SIGIR 2025 · 被引用 11 次
- STARLINE: Contrastive Learning with Modality-Aware Graph Refinement for Effective Multimedia RecommendationTaeri Kim, Sohee Ban, Hyunjoon Kim, Sang-Wook KimKDD 2025 · 被引用 1 次
- DaRec: A Disentangled Alignment Framework for Large Language Model and Recommender SystemXihong Yang, Heming Jing, Zixing Zhang, Jindong Wang 等ICDE 2025 · 被引用 2 次
