Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step Reasoning
Yiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Xinglin Wang, Bin Sun, Heda Wang, Kan Li
摘要
Self-consistency (SC) has been a widely used decoding strategy for chain-ofthought reasoning. Despite bringing significant performance improvements across a variety of multi-step reasoning tasks, it is a high-cost method that requires multiple sampling with the preset size. In this paper, we propose a simple and scalable sampling process, Early-Stopping Self-Consistency (ESC), to greatly reduce the cost of SC without sacrificing performance. On this basis, one control scheme for ESC is further derivated to dynamically choose the performance-cost balance for different tasks and models. To demonstrate ESC's effectiveness, we conducted extensive experiments on three popular categories of reasoning tasks: arithmetic, commonsense and symbolic reasoning over language models with varying scales. The empirical results show that ESC reduces the average number of sampling of chain-of-thought reasoning by a significant margin on six benchmarks, including MATH (-33.8%), GSM8K (-80.1%), StrategyQA (-76.8%), CommonsenseQA (-78.5%), Coin Flip (-84.2%) and Last Letters (-67.4%), while attaining comparable performances * . INTRODUCTION Large language models (LLMs) have exhibited strong reasoning capabilities (Bubeck et al., 2023) , especially with chain-of-thought (CoT) prompting (Wei et al., 2022) . Based on this, Wang et al. ( 2023 ) introduce a simple decoding strategy called self-consistency (SC) to further improve reasoning performances, which takes advantage of the fact that complex reasoning tasks typically allow for more than one reasoning paths leading to the correct answer. In contrast to the standard chainof-thought prompting which only generates the greedy one, this method samples multiple reasoning paths according to the predetermined sample size, and then derives the final answer through votingbased scheme. However, despite generally leading to improvements, the SC strategy incurs a significant overhead proportional to the number of sampled outputs, under the assumption that the sampled outputs are of equal length. Taking the most popular arithmetic reasoning benchmark for LLMs as an example, testing the entire MATH dataset with SC (sampling size is 64 as Lewkowycz et al. ( 2022 )) costs about 2000$ through GPT-4 API, which is a significant burden for many researchers and organizations. Therefore, it is essential to minimize the cost of SC while maintaining performance. The process of generating multiple samples in SC can be viewed as approximating the true answer distribution predicted by the language model under a specific sampling temperature. Then the most † Equal contributions. ‡ Corresponding author. * Our code and data have been released on https://github.com/Yiwei98/ESC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu 等ICLR 2026 · 被引用 250 次
- Deep Think with ConfidenceYichao Fu, Xuewei Wang, Hao Zhang, Yuandong Tian 等ICLR 2026 · 被引用 171 次
- S-GRPO: Early Exit via Reinforcement Learning in Reasoning ModelsMuzhi Dai, Chenxu Yang, Qingyi SiNeurIPS 2025 · 被引用 100 次
- Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early DecodingYiming Wang, Pei Zhang, Siyuan Huang, Baosong Yang 等NeurIPS 2025 · 被引用 66 次
- Efficiently Scaling LLM Reasoning Programs with CertaindexYichao Fu, Junda Chen, Siqi Zhu, Zheyu Fu 等NeurIPS 2025 · 被引用 42 次
它引用的顶会 Paper8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMsPranjal Aggarwal, Aman Madaan, Yiming Yang, MausamEMNLP 2023 · 被引用 5 次
- Optimal Self-Consistency for Efficient Reasoning with Large Language ModelsAustin Feng, Marius Alonso, Ambroise Odonnat, Vasilii Feofanov 等ICML 2026 · 被引用 6 次
- Think Faster Than Words: Efficient LLM Chain-of-Thought Reasoning via Dynamic Shortcut DecodingFan Liu, Yanhao Wang, Min Zhang, Zhikang Chen 等ACL 2026
- THE PATH OF LEAST RESISTANCE: GUIDING LLM REASONING TRAJECTORIES WITH PREFIX CONSENSUSIshan Jindal, Sai Prashanth Akuthota, Jayant Taneja, Sachin Dev SharmaICLR 2026 · 被引用 1 次
- Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer ReasoningJungsuk Oh, Jay-Yoon LeeAAAI 2026 · 被引用 2 次
