Improving LLMs for Recommendation with Out-Of-Vocabulary Tokens
Ting-Ji Huang, Jia-Qi Yang, Chunxu Shen, Kai-Qi Liu, De-Chuan Zhan, Han-Jia Ye
Abstract
Characterizing users and items through vector representations is crucial for various tasks in recommender systems. Recent approaches attempt to apply Large Language Models (LLMs) in recommendation through a question&answer format, where real users and items (e.g., Item No.2024) are represented with in-vocabulary tokens (e.g., "item", "20", "24"). However, since LLMs are typically pretrained on natural language tasks, these in-vocabulary tokens lack the expressive power for distinctive users and items, thereby weakening the recommendation ability even after fine-tuning on recommendation tasks. In this paper, we explore how to effectively tokenize users and items in LLM-based recommender systems. We emphasize the role of out-of-vocabulary (OOV) tokens in addition to the in-vocabulary ones and claim the memorization of OOV tokens that capture correlations of users/items as well as diversity of OOV tokens. By clustering the learned representations from historical user-item interactions, we make the representations of user/item combinations share the same OOV tokens if they have similar properties. Furthermore, integrating these OOV tokens into the LLM's vocabulary allows for better distinction between users and items and enhanced capture of user-item relationships during fine-tuning on downstream tasks. Our proposed framework outperforms existing state-of-the-art methods across various downstream recommendation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8e1e261-1bb7-4700-a0b0-97a8d4714c77Cited by top-tier papers2
- GFlowGR: Fine-tuning Generative Recommendation Frameworks with Generative Flow NetworksYejing Wang, Shengyu Zhou, Jinyu Lu, Qidong Liu et al.SIGIR 2026 · 1 citation
- AgentVocab: Structure-Aware Vocabulary Adaptation for Efficient LLM AgentsKai Bian, Haosi Mo, Xuebo Liu, Shuangyong Song et al.ICML 2026
Builds on8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Contrastive Learning for Sequential RecommendationXu Xie, Fei Sun, Zhaoyang Liu, Shiwen Wu et al.ICDE 2022 · 674 citations
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan et al.NeurIPS 2023 · 474 citations
Related papers
- Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic TokenizationGuanghan Li, Xun Zhang, Yufei Zhang, Yifan Yin et al.AAAI 2025 · 18 citations
- Adapting Large Language Models by Integrating Collaborative Semantics for RecommendationBowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen et al.ICDE 2024 · 132 citations
- From Token to Item: Enhancing Large Language Models for Recommendation via Item-aware Attention MechanismXiaokun Zhang, Bowei He, Jiamin Chen, Ziqiang Cui et al.WWW 2026
- IDGenRec: LLM-RecSys Alignment with Textual ID LearningJuntao Tan, Shuyuan Xu, Wenyue Hua, Yingqiang Ge et al.SIGIR 2024 · 45 citations
- MSL: Not All Tokens Are What You Need for Tuning LLM as a RecommenderBohao Wang, Feng Liu, Jiawei Chen, Xingyu Lou et al.SIGIR 2025 · 5 citations
