Multi-Aspect Cross-modal Quantization for Generative Recommendation
Fuwei Zhang, Xiaoyu Liu, Dongbo Xi, Jishen Yin, Huan Chen, Peng Yan, Fuzhen Zhuang, Zhao Zhang
Abstract
Generative Recommendation (GR) has emerged as a new paradigm in recommender systems. This approach relies on quantized representations to discretize item features, modeling users’ historical interactions as sequences of discrete tokens. Based on these tokenized sequences, GR predicts the next item by employing next-token prediction methods. The challenges of GR lie in constructing high-quality semantic identifiers (IDs) that are hierarchically organized, minimally conflicting, and conducive to effective generative model training. However, current approaches remain limited in their ability to harness multimodal information and to capture the deep and intricate interactions among diverse modalities, both of which are essential for learning high-quality semantic IDs and for effectively training GR models. To address this, we propose Multi-Aspect Cross-modal quantization for generative Recommendation (MACRec), which introduces multimodal information and incorporates it into both semantic ID learning and generative model training from different aspects. Specifically, we first introduce cross-modal quantization during the ID learning process, which effectively reduces conflict rates and thus improves codebook usability through the complementary integration of multimodal information. In addition, to further enhance the generative ability of our GR model, we incorporate multi-aspect cross-modal alignments, including the implicit and explicit alignments. Finally, we conduct extensive experiments on three well-known recommendation datasets to demonstrate the effectiveness of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37928366-88c3-4e60-8cea-09d6b7f30d9bCited by top-tier papers2
- SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative RecommendationWei Chen, Xingyu Guo, Shuang Li, Fuwei Zhang et al.ICML 2026 · 2 citations
- CARD: Non-Uniform Quantization of Visual Semantic Unit for Generative RecommendationYibiao Wei, Jie Zou, Pengfei Zhang, Xiao Ao et al.SIGIR 2026 · 1 citation
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan et al.NeurIPS 2023 · 474 citations
- Graph-Refined Convolutional Network for Multimedia Recommendation with Implicit FeedbackYinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He et al.ACM MM 2020 · 374 citations
- Autoregressive Image Generation using Residual QuantizationDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho et al.CVPR 2022 · 184 citations
- Adapting Large Language Models by Integrating Collaborative Semantics for RecommendationBowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen et al.ICDE 2024 · 132 citations
Related papers
- MusicRec: Multi-modal Semantic-Enhanced Identifier with Collaborative Signals for Generative RecommendationYuqiu Zhao, Lei Shi, Yan Zhong, Feifei Kou et al.AAAI 2026
- Universal Item Tokenization for Transferable Generative RecommendationBowen Zheng, Hongyu Lu, Yu Chen, Wayne Xin Zhao et al.SIGIR 2026
- Understanding Generative Recommendation with Semantic IDs from a Model-scaling ViewJingzhe Liu, Liam Collins, Jiliang Tang, Tong Zhao et al.KDD 2026 · 17 citations
- Multimodal Quantitative Language for Generative RecommendationJianyang Zhai, Zi-Feng Mai, Chang-Dong Wang, Feidiao Yang et al.ICLR 2025
- UniGCRec: Unified User-Item Quantization for Generative Cross-Domain RecommendationChaoyue Ding, Jiahao Liu, Dongsheng Li, Shengkang Gu et al.KDD 2026
