Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
Yiming Wang, Pei Zhang, Siyuan Huang, Baosong Yang, Zhuosheng Zhang, Fei Huang, Rui Wang
Abstract
Test-time scaling enhances large language model performance by allocating additional compute resources during inference. Best-of-N (BoN) sampling serves as a common sampling-based scaling technique, broadening the search space in parallel to find better solutions from the model distribution. However, its cost-performance trade-off is still underexplored. Two main challenges limit the efficiency of BoN sampling: (1) Generating N full samples consumes substantial GPU memory, reducing inference capacity under limited resources. (2) Reward models add extra memory and latency overhead, and training strong reward models introduces potential training data costs. Although some studies have explored efficiency improvements, none have addressed both challenges at once. To address this gap, we propose Self-Truncation Best-of-N (ST-BoN), a decoding method that avoids fully generating all N samples and eliminates the need for reward models. It leverages early sampling consistency in the model's internal states to identify the most promising path and truncate suboptimal ones. In terms of cost, ST-BoN reduces dynamic GPU memory usage by over 80% and inference latency by 50%. In terms of cost-performance trade-off, ST-BoN achieves the same performance as Full-BoN while saving computational cost by 70%-80%, and under the same cost, it can improve accuracy by 3-4 points.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e74a18a5-7d10-42c1-8a64-7ca187eeefcfCited by top-tier papers11
- TrimR: Verifier-based Training-Free Thinking Trimming for Efficient Test-Time ScalingWeizhe Lin, Xing Li 023, Zhiyuan Yang, Xiaojin Fu et al.ICLR 2026 · 14 citations
- VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic ModelWenhao Li, Xiu Su, Yichao Cao, Hongyan Xu et al.ICML 2026 · 13 citations
- MUR: Momentum Uncertainty guided Reasoning for Large Language ModelsHang Yan, Fangzhi Xu, Rongman Xu, Yifei Li et al.ACL 2026 · 12 citations
- Revisiting Model Interpolation for Efficient ReasoningTaiqiang Wu, Runming Yang, Tao Liu, Jiahao Wang et al.ACL 2026 · 8 citations
- Efficient Reasoning with Balanced ThinkingYulin Li, Tengyao Tu, Li Ding, Junjie Wang et al.ICLR 2026 · 7 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time ScalingShengyin Sun, Yiming Li, Xing Li, Yingzhao Lian et al.ICLR 2026 · 6 citations
- ROC-n-reroll: How verifier imperfection affects test-time scalingFlorian E. Dorner, Yatong Chen, André F Cruz, Fanny YangICLR 2026 · 13 citations
- A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM ReasoningZhi Zhou, Tan Yuhao, Zenan Li, Yuan Yao et al.NeurIPS 2025 · 17 citations
- Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic ConfidenceAmirhosein Ghasemabadi, Keith G. Mills, Baochun Li, Di NiuACL 2026 · 9 citations
- Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability TheoryYexiang Liu, Zekun Li, Zhi Fang, Nan Xu et al.ACL 2025 · 12 citations
