SLMRec: Distilling Large Language Models into Small for Sequential Recommendation
Wujiang Xu, Qitian Wu, Zujie Liang, Jiaojiao Han, Xuying Ning, Yunxiao Shi, Wenfang Lin, Yongfeng Zhang
摘要
Sequential Recommendation (SR) task involves predicting the next item a user is likely to interact with, given their past interactions. The SR models examine the sequence of a user's actions to discern more complex behavioral patterns and temporal dynamics. Recent research demonstrates the great impact of LLMs on sequential recommendation systems, either viewing sequential recommendation as language modeling or serving as the backbone for user representation. Although these methods deliver outstanding performance, there is scant evidence of the necessity of a large language model and how large the language model is needed, especially in the sequential recommendation scene. Meanwhile, due to the huge size of LLMs, it is inefficient and impractical to apply a LLM-based model in realworld platforms that often need to process billions of traffic logs daily. In this paper, we explore the influence of LLMs' depth by conducting extensive experiments on large-scale industry datasets. Surprisingly, our motivational experiments reveal that most intermediate layers of LLMs are redundant, indicating that pruning the remaining layers can still maintain strong performance. Motivated by this insight, we empower small language models for SR, namely SLMREC, which adopt a simple yet effective knowledge distillation method. Moreover, SLMREC is orthogonal to other post-training efficiency techniques, such as quantization and pruning, so that they can be leveraged in combination. Comprehensive experimental results illustrate that the proposed SLMREC model attains the best performance using only 13% of the parameters found in LLM-based recommendation models, while simultaneously achieving up to 6.6x and 8.0x speedups in training and inference time costs, respectively. Besides, we provide a theoretical justification for why small language models can perform comparably to large language models in SR. The source code and datasets are available at the URL 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive ActivationsChaofan Gan, Yuanpeng Tu, Xi Chen, Tieyuan Chen 等NeurIPS 2025 · 被引用 22 次
- Interactive Recommendation Agent with Active User CommandsJiakai Tang, Wen Chen, Yujie Luo, Xunke Xi 等KDD 2026 · 被引用 14 次
- Mixture of Sequence: Theme-Aware Mixture-of-Experts for Long-Sequence RecommendationXiao Lin, Zhicheng Tang, Weilin Cong, Mengyue Hang 等WWW 2026 · 被引用 3 次
- FeDecider: An LLM-Based Framework for Federated Cross-Domain RecommendationXinrui He, Ting-Wei Li, Tianxin Wei, Xuying Ning 等WWW 2026 · 被引用 2 次
- Bat: Efficient Generative Recommender Serving with Bipartite AttentionJie Sun, Shaohang Wang, Zimo Zhang, Zhengyu Liu 等ASPLOS 2026 · 被引用 1 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
相关 Paper
- DELRec: Distilling Sequential Pattern to Enhance LLMs-Based Sequential RecommendationHaoyi Zhang, Guohao Sun, Jinhu Lu, Guanfeng Liu 等ICDE 2025 · 被引用 1 次
- Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?Sein Kim, Hongseok Kang, Kibum Kim, Jiwan Kim 等KDD 2025 · 被引用 3 次
- LLM4RSR: Large Language Models as Data Correctors for Robust Sequential RecommendationYatong Sun, Xiaochun Yang, Zhu Sun, Yan Wang 等AAAI 2025 · 被引用 2 次
- Breaking the Bottleneck: User-Specific Optimization and Real-Time Inference Integration for Sequential RecommendationWenjia Xie, Hao Wang, Minghao Fang, Ruize Yu 等KDD 2025
- Can Small Language Models be Good Reasoners for Sequential Recommendation?Yuling Wang, Changxin Tian, Binbin Hu, Yanhua Yu 等WWW 2024 · 被引用 69 次
