Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
Audrey Huang, Adam Block, Qinghua Liu, Nan Jiang, Akshay Krishnamurthy, Dylan J. Foster
摘要
Inference-time computation offers a powerful axis for scaling the performance of language models. However, naively increasing computation in techniques like Best-of-N sampling can lead to performance degradation due to reward hacking. Toward a theoretical understanding of how to best leverage additional computation, we focus on inference-time alignment, which we formalize as the problem of improving the quality of responses drawn from a pre-trained policy, given a prompt of interest and access to an imperfect reward model. We analyze the performance of inference-time alignment algorithms in terms of (i) response quality, and (ii) compute, and provide new results that highlight the importance of the pre-trained policy's coverage over high-quality responses for performance and compute scaling: 1. We show that Best-of-N alignment with an ideal choice for N can achieve optimal performance under stringent notions of coverage, but provably suffers from reward hacking when N is large, and fails to achieve tight guarantees under more realistic coverage conditions. 2. We introduce InferenceTimePessimism, a new algorithm which mitigates reward hacking through more sophisticated use of inference-time compute, implementing the principle of pessimism in the face of uncertainty via rejection sampling; we prove that its performance is optimal and does not degrade with N , a property that we call scaling-monotonic. We complement our theoretical results with an experimental evaluation that demonstrate the benefits of InferenceTimePessimism across a variety of tasks and models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-TrainingJunkai Zhang, Zihao Wang, Lin Gui, Swarnashree Mysore Sathyendra 等ICLR 2026 · 被引用 48 次
- Energy-Based Transformers are Scalable Learners and ThinkersAlexi Gladstone, Ganesh Nanduru, Md Mofijul Islam, Peixuan Han 等ICLR 2026 · 被引用 38 次
- Inference-Time Reward Hacking in Large Language ModelsHadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju 等NeurIPS 2025 · 被引用 38 次
- The Coverage Principle: How Pre-Training Enables Post-TrainingFan Chen, Audrey Huang, Noah Golowich, Sadhika Malladi 等ICLR 2026 · 被引用 28 次
- Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution SharpeningXiaotong Ji, Rasul Tutunov, Matthieu Zimmer, Haitham Bou AmmarICML 2026 · 被引用 17 次
它引用的顶会 Paper40
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
相关 Paper
- From Curiosity to Caution: Mitigating Reward Hacking for Best-of- with PessimismZhuohao Yu, Steven Z. Wu, Adam BlockICLR 2026 · 被引用 3 次
- Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo MethodsIsha Puri, Shivchander Sudalairaj, Guangxuan Xu, Abhishek Bhandwaldar 等NeurIPS 2025 · 被引用 8 次
- InfAlign: Inference-aware language model alignmentAnanth Balashankar, Ziteng Sun, Jonathan Berant, Jacob Eisenstein 等ICML 2025
- Inference-Aware Prompt Optimization for Aligning Black-Box Large Language ModelsSaaduddin Mahmud, Mason Nakamura, Kyle Hollins Wray, Shlomo ZilbersteinAAAI 2026
- OptScale: Probabilistic Optimality for Inference-time ScalingYoukang Wang, Jian Wang, Rubing Chen, Xiao-Yong WeiAAAI 2026 · 被引用 2 次
