Efficiency Unleashed: Inference Acceleration for LLM-based Recommender Systems with Speculative Decoding
Yunjia Xi, Hangyu Wang, Bo Chen, Jianghao Lin, Menghui Zhu, Weiwen Liu, Ruiming Tang, Zhewei Wei, Weinan Zhang, Yong Yu
Abstract
The past few years have witnessed a growing interest in LLM-based recommender systems (RSs), although their industrial deployment remains in a preliminary stage. Most existing deployments leverage LLMs offline as feature enhancers, generating augmented knowledge for downstream tasks. However, in recommendation scenarios with numerous users and items, even offline knowledge generation with LLMs demands significant time and computational resources. This inefficiency arises from the autoregressive nature of LLMs. A promising solution is speculative decoding, a Draft-Then-Verify approach that increases the number of tokens generated per decoding step. In this work, we first identify recommendation knowledge generation as a highly fitting use case for retrieval-based speculative decoding. Then, we discern its two characteristics: (1) the vast number of items and users in RSs leads to retrieval inefficiency, and (2) RSs exhibit high diversity tolerance for LLM-generated text. Building on these insights, we introduce Lossless Acceleration via Speculative Decoding for LLM-based Recommender Systems (LASER), which features a Customized Retrieval Pool to enhance retrieval efficiency and Relaxed Verification to improve the acceptance rate of draft tokens. LASER achieves a 3-5x speedup on public datasets and saves about 67% of computational resources during the online A/B test on a large-scale advertising scenario with lossless downstream recommendation performance. Our code is available at https://github.com/YunjiaXi/LASER
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- A Survey of Large Language Model-Based Search AgentsYunjia Xi, Jianghao Lin, Yongzhao Xiao, Zheli Zhou et al.ACL 2026 · 1,216 citations
- EARN: Efficient Inference Acceleration for LLM-based Generative Recommendation by Register TokensChaoqun Yang, Xinyu Lin, Wenjie Wang, Yongqi Li et al.KDD 2025 · 1 citation
- Modular Representation Compression: Adapting LLM Representations for Efficient and Effective RecommendationYunjia Xi, Menghui Zhu, Jianghao Lin, Bo Chen et al.SIGIR 2026
Builds on20
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng et al.ICML 2024 · 669 citations
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 424 citations
Related papers
- Efficient Inference for Large Language Model-based Generative RecommendationXinyu Lin, Chaoqun Yang, Wenjie Wang, Yongqi Li et al.ICLR 2025 · 1 citation
- RAPID: Long-Context Inference with Retrieval-Augmented Speculative DecodingGuanzheng Chen, Qilong Feng, Jinjie Ni, Xin Li et al.ICML 2025
- SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative DecodingLinye Wei, Shuzhang Zhong, Songqiang Xu, Runsheng Wang et al.DAC 2025 · 7 citations
- SAM Decoding: Speculative Decoding via Suffix AutomatonYuxuan Hu, Ke Wang, Xiaokang Zhang, Fanjin Zhang et al.ACL 2025
- PEARL: Parallel Speculative Decoding with Adaptive Draft LengthTianyu Liu, Yun Li, Qitan Lv, Kai Liu et al.ICLR 2025
