Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
Qiwei Di, Kaixuan Ji, Xuheng Li, Heyang Zhao, Quanquan Gu
Abstract
LLM inference often generates a batch of candidates for a prompt and selects one via strategies like majority voting or Best-of- N (BoN). For difficult tasks, this single-shot selection often underperforms. Consequently, evaluations commonly report Pass@: the agent may submit up to responses, and only the best of them is used when computing regret. Motivated by this, we study inference scaling in the more general Pass@ inference setting, and prove that neither majority voting nor BoN exhibits the desirable scaling with and the sampling budget . Combining the advantages of majority voting and BoN, we propose a new inference strategy called Best-of-Majority (BoM), with a pivotal step that restricts the candidates to the responses with high frequency in the samples before selecting the top- rewards. We prove that when the sampling budget is , the regret of BoM is , where is the coverage coefficient, is the estimation error of the reward model, and is the estimation error of reward at the optimal response. We further establish a matching lower bound, certifying that our algorithm is minimax optimal. Beyond optimality, BoM has a key advantage: unlike majority voting and BoN, its performance does not degrade when increasing . Experimental results of inference on math problems show BoM outperforming both majority voting and BoN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 608880f8-45de-4a26-9ca3-5cf39b995a8aCited by top-tier papers1
Ask how each one uses itBuilds on32
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Inference-Aware Prompt Optimization for Aligning Black-Box Large Language ModelsSaaduddin Mahmud, Mason Nakamura, Kyle Hollins Wray, Shlomo ZilbersteinAAAI 2026
- Majority of the Bests: Improving Best-of-N via BootstrappingAmin Rakhsha, Kanika Madan, Tianyu Zhang, Amir-massoud Farahmand et al.NeurIPS 2025 · 6 citations
- Best-of-Infinity: Asymptotic Performance of Test-Time LLM EnsemblingJunpei Komiyama, Daisuke Oba, Masafumi OyamadaICLR 2026 · 2 citations
- Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability TheoryYexiang Liu, Zekun Li, Zhi Fang, Nan Xu et al.ACL 2025 · 12 citations
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-SolvingYangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck et al.ICLR 2025
