Catalog-Native LLM: Speaking Item-ID dialect with Less Entanglement for Recommendation
Reza Shirkavand, Xiaokai Wei, Chen Wang, Zheng Hui, Heng Huang, Michelle Gong
Abstract
While collaborative filtering delivers predictive accuracy and efficiency, and Large Language Models (LLMs) enable expressive and generalizable reasoning, modern recommendation systems must bring these strengths together. Growing user expectations, such as natural-language queries and transparent explanations, further highlight the need for a unified approach. However, doing so is nontrivial. Collaborative signals are often token-efficient but semantically opaque, while LLMs are semantically rich but struggle to model implicit user preferences when trained only on textual inputs. This paper introduces Item-ID + Natural-language Mixtureof-Experts Language Model (IDIOMoE), which treats item interaction histories as a native dialect within the language space, enabling collaborative signals to be understood in the same way as natural language. By splitting the Feed Forward Network of each block of a pretrained LLM into a separate text expert and an item expert with token-type gating, our method avoids destructive interference between text and catalog modalities. IDIOMoE demonstrates strong recommendation performance across both public and proprietary datasets, while preserving the text understanding of the pretrained model. INTRODUCTION Recommendation systems shape what people read, watch, buy, learn, and play. As AI shifts from static predictors to reasoning agents capable of following instructions, recommendation is also evolving from ranking fixed lists to assisting users in exploring, planning, and deciding. This trend is visible in practice: Amazon's Rufus provides LLM-powered conversational shopping (Amazon, 2024); Meta's Llama-3 assistant is embedded in WhatsApp, Instagram, and Facebook for task planning (Meta, 2024); and Netflix is adopting foundation-model approaches for personalization and LLM-based conversational retrieval (Netflix, 2025; Zhu et al., 2025) . These examples motivate bringing LLM knowledge and instruction-following into recommenders while preserving the collaborative patterns that make them accurate at scale. Conventional recommenders like collaborative filtering (CF) (Koren et al., 2009) , content-based (CB) (Lops et al., 2011) , and sequential models (Kang & McAuley, 2018; Sun et al., 2019; Zhai et al., 2024) perform well within their scope when data are abundant, but they depend heavily on the quality of logs and item attributes. They remain vulnerable to popularity bias (Abdollahpouri et al., 2019) , struggle to integrate heterogeneous signals (text, behavior, and context), and cannot support natural language queries. Pre-trained LLMs offer complementary strengths: they bring broad world knowledge, can follow natural-language instructions, and can reason about multi-objective trade-offs. Yet a fundamental gap remains. LLM pretraining centers on semantic understanding, whereas recommendation requires modeling collaborative preference patterns. The key challenge is leveraging LLMs for preference understanding without disrupting their semantic competence. Recent work has tried to bridge this gap by extending LLM vocabularies with item IDs (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b14677dd-7946-4891-93a7-c23c94355768Builds on34
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
Related papers
- Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic TokenizationGuanghan Li, Xun Zhang, Yufei Zhang, Yifan Yin et al.AAAI 2025 · 18 citations
- Collaborative Large Language Model for Recommender SystemsYaochen Zhu, Liang Wu, Qi Guo, Liangjie Hong et al.WWW 2024 · 150 citations
- Large Language Models meet Collaborative Filtering: An Efficient All-round LLM-based Recommender SystemSein Kim, Hongseok Kang, Seungyoon Choi, Donghyun Kim et al.KDD 2024 · 107 citations
- Token-level Collaborative Alignment for LLM-based Generative RecommendationFake Lin, Binbin Hu, Zhi Zheng, Xi Zhu et al.WWW 2026 · 1 citation
- Language Representations Can be What Recommenders Need: Findings and PotentialsLeheng Sheng, An Zhang, Yi Zhang, Yuxin Chen et al.ICLR 2025
