Improving LLMs for Recommendation with Out-Of-Vocabulary Tokens
Ting-Ji Huang, Jia-Qi Yang, Chunxu Shen, Kai-Qi Liu, De-Chuan Zhan, Han-Jia Ye
摘要
Characterizing users and items through vector representations is crucial for various tasks in recommender systems. Recent approaches attempt to apply Large Language Models (LLMs) in recommendation through a question&answer format, where real users and items (e.g., Item No.2024) are represented with in-vocabulary tokens (e.g., "item", "20", "24"). However, since LLMs are typically pretrained on natural language tasks, these in-vocabulary tokens lack the expressive power for distinctive users and items, thereby weakening the recommendation ability even after fine-tuning on recommendation tasks. In this paper, we explore how to effectively tokenize users and items in LLM-based recommender systems. We emphasize the role of out-of-vocabulary (OOV) tokens in addition to the in-vocabulary ones and claim the memorization of OOV tokens that capture correlations of users/items as well as diversity of OOV tokens. By clustering the learned representations from historical user-item interactions, we make the representations of user/item combinations share the same OOV tokens if they have similar properties. Furthermore, integrating these OOV tokens into the LLM's vocabulary allows for better distinction between users and items and enhanced capture of user-item relationships during fine-tuning on downstream tasks. Our proposed framework outperforms existing state-of-the-art methods across various downstream recommendation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- GFlowGR: Fine-tuning Generative Recommendation Frameworks with Generative Flow NetworksYejing Wang, Shengyu Zhou, Jinyu Lu, Qidong Liu 等SIGIR 2026 · 被引用 1 次
- AgentVocab: Structure-Aware Vocabulary Adaptation for Efficient LLM AgentsKai Bian, Haosi Mo, Xuebo Liu, Shuangyong Song 等ICML 2026
它引用的顶会 Paper8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- Contrastive Learning for Sequential RecommendationXu Xie, Fei Sun, Zhaoyang Liu, Shiwen Wu 等ICDE 2022 · 被引用 674 次
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan 等NeurIPS 2023 · 被引用 474 次
相关 Paper
- Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic TokenizationGuanghan Li, Xun Zhang, Yufei Zhang, Yifan Yin 等AAAI 2025 · 被引用 18 次
- Adapting Large Language Models by Integrating Collaborative Semantics for RecommendationBowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen 等ICDE 2024 · 被引用 132 次
- From Token to Item: Enhancing Large Language Models for Recommendation via Item-aware Attention MechanismXiaokun Zhang, Bowei He, Jiamin Chen, Ziqiang Cui 等WWW 2026
- IDGenRec: LLM-RecSys Alignment with Textual ID LearningJuntao Tan, Shuyuan Xu, Wenyue Hua, Yingqiang Ge 等SIGIR 2024 · 被引用 45 次
- MSL: Not All Tokens Are What You Need for Tuning LLM as a RecommenderBohao Wang, Feng Liu, Jiawei Chen, Xingyu Lou 等SIGIR 2025 · 被引用 5 次
