Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
Gholamali Aminian, Idan Shenfeld, Amir R. Asadi, Ahmad Beirami, Youssef Mroueh
Abstract
A simple yet effective method for inference-time alignment of generative models is Best-of-N (BoN), where N outcomes are sampled from a reference policy, evaluated using a proxy reward model, and the highest-scoring one is selected. While prior work argues that BoN is almost optimal in reward vs KL tradeoffs, the effectiveness of BoN depends critically on the quality of the proxy reward model used for selection. For this purpose, we study BoN through a smooth version known as Soft Best-of-N (SBoN) and develop a theoretical framework to address this gap. We analyze the scaling behaviour of BoN by providing bounds on the KL divergence between the SBoN policy and the reference policy, offering insights into how performance varies with the number of samples. We also study the regret gap, i.e., the gap between the expected true reward under the optimal (tilted) policy and the SBoN policy. Our theoretical and empirical findings show that smoothing helps SBoN mitigate reward overoptimization, especially when the quality of the proxy reward is low.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1565cfdf-bc6d-40a8-b6f9-303d39ec8465Cited by top-tier papers4
- Taming Imperfect Process Verifiers: A Sampling Perspective on BacktrackingDhruv Rohatgi, Abhishek Shetty, Donya Saless, Yuchen Li et al.ICLR 2026 · 15 citations
- Guided Speculative Inference for Efficient Test-Time Alignment of LLMsJonathan Geuter, Youssef Mroueh, David Alvarez-MelisICLR 2026 · 12 citations
- KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample ComplexityGholamali Aminian, Amir Reza Asadi, Idan Shenfeld, Youssef MrouehNeurIPS 2025 · 11 citations
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference ScalingQiwei Di, Kaixuan Ji, Xuheng Li, Heyang Zhao et al.ICLR 2026 · 7 citations
Builds on18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
Related papers
- Theoretical guarantees on the best-of-n alignment policyAhmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander Nicholas D'Amour et al.ICML 2025
- Variational Best-of-N AlignmentAfra Amini, Tim Vieira, Elliott Ash, Ryan CotterellICLR 2025 · 1 citation
- BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n SamplingLin Gui, Cristina Garbacea, Victor VeitchNeurIPS 2024 · 138 citations
- InfAlign: Inference-aware language model alignmentAnanth Balashankar, Ziteng Sun, Jonathan Berant, Jacob Eisenstein et al.ICML 2025
- Inference-Time Reward Hacking in Large Language ModelsHadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju et al.NeurIPS 2025 · 38 citations
