Beyond Static Best-of-N: Bayesian List-wise Alignment for LLM-based Recommendation
Ruijun Chen, Chongming Gao, Jiawei Chen, Weiqin Yang, Xiangnan He
Abstract
Large Language Models have revolutionized recommender systems (LLM4Rec) by leveraging their generative capabilities to model complex user preferences. However, existing LLM4Rec methods primarily rely on token-level objectives, making it difficult to optimize list-level and non-differentiable metrics (e.g., NDCG, fairness) that define actual recommendation quality. While Best-of-N (BoN) directly optimizes these metrics during inference, its high computational cost hinders real-world deployment. To address this, BoN Alignment aims to distill the search capability into the model itself, yet current approaches suffer from two critical limitations: (1) Indiscriminate Supervision, where the static reference fails to distinguish the relative quality of candidates exceeding its empirical range, leading to a loss of ranking guidance; and (2) Gradient Decay, where the effective supervision signal rapidly diminishes as the evolving policy improves, resulting in inefficient optimization.
To overcome these challenges, we propose BLADE (Bayesian List-wise Alignment via Dynamic Estimation). Unlike static approaches, BLADE introduces a Bayesian framework that continuously updates the target distribution by fusing historical priors with dynamic evidence from the model's current rollouts. This mechanism constructs a self-evolving target that adapts to the model's growing capabilities, ensuring the training signal remains informative throughout the learning process. Extensive experiments on three real-world datasets demonstrate that BLADE significantly outperforms state-of-the-art baselines. Crucially, it breaks the static performance upper bound, achieving sustained gains in both ranking accuracy (Recall, NDCG) and complex list-wise metrics (Fairness, Diversity). The code is available via https://github. com/RegionCh/BLADE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cce31812-ffc9-4d1d-a61b-adcf13c95b19Cited by top-tier papers2
- Uncertainty-aware Generative RecommendationChenxiao Fan, Chongming Gao, Yaxin Gong, Haoyan Liu et al.KDD 2026 · 2 citations
- The Pitfall of Scaling Up: Uncovering and Mitigating Popularity Bias Amplification in Scaling Transformer-based RecommendersWeiqin Yang, Yue Pan, Chongming Gao, Sheng Zhou et al.KDD 2026
Builds on14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Representation Learning with Large Language Models for RecommendationXubin Ren, Wei Wei, Lianghao Xia, Lixin Su et al.WWW 2024 · 385 citations
- AgentCF: Collaborative Learning with Autonomous Language Agents for Recommender SystemsJunjie Zhang, Yupeng Hou, Ruobing Xie, Wenqi Sun et al.WWW 2024 · 164 citations
- BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n SamplingLin Gui, Cristina Garbacea, Victor VeitchNeurIPS 2024 · 138 citations
Related papers
- Align³GR: Unified Multi-Level Alignment for LLM-based Generative RecommendationWencai Ye, Mingjie Sun, Shuhang Chen, Wenjin Wu et al.AAAI 2026 · 2 citations
- Token-level Collaborative Alignment for LLM-based Generative RecommendationFake Lin, Binbin Hu, Zhi Zheng, Xi Zhu et al.WWW 2026 · 1 citation
- Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic TokenizationGuanghan Li, Xun Zhang, Yufei Zhang, Yifan Yin et al.AAAI 2025 · 18 citations
- Efficient Inference for Large Language Model-based Generative RecommendationXinyu Lin, Chaoqun Yang, Wenjie Wang, Yongqi Li et al.ICLR 2025 · 1 citation
- Adaptive Mix Preference Optimization for Generative RecommendationJunbo Qi, Yanyan Zou, Xuanhua Yang, Sulong Xu et al.SIGIR 2026
