BEAR: Towards Beam-Search-Aware Optimization for Recommendation with Large Language Models
Weiqin Yang, Bohao Wang, Zhenxiang Xu, Jiawei Chen, Shengjia Zhang, Jingbang Chen, Canghong Jin, Can Wang
Abstract
Recent years have seen a rapid surge in research leveraging Large Language Models (LLMs) for recommendation. These methods typically employ supervised fine-tuning (SFT) to adapt LLMs to recommendation scenarios, and utilize beam search during inference to efficiently retrieve 𝐵 top-ranked recommended items. However, we identify a critical training-inference inconsistency: while SFT optimizes the overall probability of positive items, it does not guarantee that such items will be retrieved by beam search even if they possess high overall probabilities. Due to the greedy pruning mechanism, beam search can prematurely discard a positive item once its prefix probability is insufficient.
To address this inconsistency, we propose BEAR (Beam-SEarch-Aware Regularization), a novel fine-tuning objective that explicitly accounts for beam search behavior during training. Rather than directly simulating beam search for each instance during training, which is computationally prohibitive, BEAR enforces a relaxed necessary condition: each token in a positive item must rank within the top-𝐵 candidate tokens at each decoding step. This objective effectively mitigates the risk of incorrect pruning while incurring negligible computational overhead compared to standard SFT. Extensive experiments across four real-world datasets demonstrate that BEAR significantly outperforms strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7aaffd95-8006-400b-91b1-3744fe246ccdCited by top-tier papers4
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo et al.ACL 2026 · 42 citations
- Spend Search Where It Pays: Value-Guided Structured Sampling and Optimization for Generative RecommendationJie Jiang, Yangru Huang, Zeyu Wang, Changping Wang et al.KDD 2026 · 3 citations
- The Pitfall of Scaling Up: Uncovering and Mitigating Popularity Bias Amplification in Scaling Transformer-based RecommendersWeiqin Yang, Yue Pan, Chongming Gao, Sheng Zhou et al.KDD 2026
- SpecTran: Spectral-Aware Transformer-based Adapter for LLM-Enhanced Sequential RecommendationYu Cui, Feng Liu, Zhaoxiang Wang, Changwang Zhang et al.SIGIR 2026
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
Related papers
- Does LLM Focus on the Right Words? Mitigating Context Bias in LLM-based RecommendersBohao Wang, Jiawei Chen, Feng Liu, Changwang Zhang et al.WWW 2026 · 1 citation
- APAO: Bridging the Training-Inference Gap in Generative Recommendation via Adaptive Prefix-Aware OptimizationYuanqing Yu, Yifan Wang, Weizhi Ma, Zhiqiang Guo et al.KDD 2026 · 4 citations
- Process-Supervised LLM Recommenders via Flow-guided TuningChongming Gao, Mengyao Gao, Chenxiao Fan, Shuai Yuan et al.SIGIR 2025 · 6 citations
- MSL: Not All Tokens Are What You Need for Tuning LLM as a RecommenderBohao Wang, Feng Liu, Jiawei Chen, Xingyu Lou et al.SIGIR 2025 · 5 citations
- Unifying Search and Recommendation in LLMs via Gradient Multi-Subspace TuningJujia Zhao, Zihan Wang, Shuaiqun Pan, Suzan Verberne et al.SIGIR 2026 · 1 citation
