MusicRec: Multi-modal Semantic-Enhanced Identifier with Collaborative Signals for Generative Recommendation
Yuqiu Zhao, Lei Shi, Yan Zhong, Feifei Kou, Pengfei Zhang, Jiwei Zhang, Mingying Xu, Yanchao Liu
Abstract
Generative recommendation as a new paradigm is influencing the current development of recommender systems. It aims to assign identifiers that capture richer semantic and collaborative information to items, and subsequently predict item identifiers via autoregressive generation using Large Language Models (LLMs). Existing approaches primarily tokenize item text into codebooks with preserved semantic IDs through RQ-VAE, or separately tokenize different modality features of items. However, existing tokenization methods face two major challenges: (1) Learning decoupled multi-modal features limits the quality of the semantic representation. (2) Ignoring collaborative signals from interaction history limits the comprehensiveness of identifiers. To address these limitations, we propose a multi-modal semantic-enhanced identifier with collaborative signals for generative recommendation, named MusicRec. In MusicRec, we propose a tokenization approach based on shared-specific modal fusion, enabling the generated identifiers to preserve semantic information more comprehensively from all modalities. In addition, we incorporate collaborative signals from user interactions to guide identifier generation, preserving collaborative patterns in the semantic representation space. Extensive experiments on three public datasets demonstrate that MusicRec achieves state-of-the-art performance compared to existing baseline methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9911787-69a6-429d-a79a-8dbe25dc74e3Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan et al.NeurIPS 2023 · 474 citations
- Learning Vector-Quantized Item Representation for Transferable Sequential RecommendersYupeng Hou, Zhankui He, Julian J. McAuley, Wayne Xin ZhaoWWW 2023 · 256 citations
- Adapting Large Language Models by Integrating Collaborative Semantics for RecommendationBowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen et al.ICDE 2024 · 132 citations
Related papers
- Order-agnostic Identifier for Large Language Model-based Generative RecommendationXinyu Lin, Haihan Shi, Wenjie Wang, Fuli Feng et al.SIGIR 2025 · 15 citations
- Understanding Generative Recommendation with Semantic IDs from a Model-scaling ViewJingzhe Liu, Liam Collins, Jiliang Tang, Tong Zhao et al.KDD 2026 · 17 citations
- Universal Item Tokenization for Transferable Generative RecommendationBowen Zheng, Hongyu Lu, Yu Chen, Wayne Xin Zhao et al.SIGIR 2026
- Multi-Aspect Cross-modal Quantization for Generative RecommendationFuwei Zhang, Xiaoyu Liu, Dongbo Xi, Jishen Yin et al.AAAI 2026 · 2 citations
- Learning Decomposed Contextual Token Representations from Pretrained and Collaborative Signals for Generative RecommendationYifan Liu, Yaokun Liu, Zelin Li, Zhenrui Yue et al.SIGIR 2026
