Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
Renliang Sun, Wei Cheng, Dawei Li, Haifeng Chen, Wei Wang
Abstract
Chain-of-Thought (CoT) reasoning has driven recent gains of large language models (LLMs) on reasoning-intensive tasks by externalizing intermediate steps. However, excessive or redundant reasoning -- so-called overthinking -- can increase inference costs and lead LLMs toward incorrect conclusions. In this paper, we present REFRAIN (lective-edundancy for daptive ference), a training-free framework that adaptively determines when to stop reasoning to mitigate overthinking. REFRAIN integrates a two-stage stop discriminator to identify reflective yet redundant reasoning and a sliding-window Upper Confidence Bound (SW-UCB) multi-armed bandit controller to dynamically adjust stopping thresholds according to problem difficulty without supervision or fine-tuning. Across four representative benchmarks and two model families, REFRAIN reduces token usage by 20-55% while maintaining or improving accuracy compared to standard CoT prompting. Extensive ablation and robustness analyses demonstrate its stability across models, scorers, and prompt variations. In summary, our findings highlight when-to-stop as a new and practical axis of test-time scaling -- enabling models to reason not just more, but just enough.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 838064e7-c30f-4cdd-9f81-0c0e24888ac0Cited by top-tier papers6
- The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics AnalysisZihao Wei, Liang Pang, Jiahao Liu, Wenjie Shi et al.ACL 2026 · 14 citations
- Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and FailuresYi Hu, Jiaqi Gu, Ruxin Wang, Zijun Yao et al.ACL 2026 · 5 citations
- When to Think and When to Look: Uncertainty-Guided LookbackJing Bi, Filippos Bellos, JunJia Guo, Yayuan Li et al.CVPR 2026 · 4 citations
- Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought ReasoningRenos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille et al.ICML 2026 · 4 citations
- Think Better, Not Longer: Token-Level Marginal Utility for Efficient Reasoning in Large Reasoning ModelsJiawei Li, Yang Gao, Huashan Sun, Chong FengACL 2026
Builds on6
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu et al.ICLR 2026 · 250 citations
- Deep Think with ConfidenceYichao Fu, Xuewei Wang, Hao Zhang, Yuandong Tian et al.ICLR 2026 · 171 citations
- Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step ReasoningYiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan et al.ICLR 2024 · 101 citations
- Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection SuppressionJiameng Huang, Baijiong Lin, Guhao Feng, Jierun Chen et al.AAAI 2026 · 21 citations
- ConCISE: Confidence-guided Compression in Step-by-step Efficient ReasoningZiqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang et al.EMNLP 2025 · 1 citation
Related papers
- Stop When Further Reasoning Won’t Help: Attention-State Adaptive Generation in Reasoning ModelsJiakai Li, KE QIN, Rongzheng Wang, Yizhuo Ma et al.ICML 2026
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du et al.ICLR 2026 · 225 citations
- Think Faster Than Words: Efficient LLM Chain-of-Thought Reasoning via Dynamic Shortcut DecodingFan Liu, Yanhao Wang, Min Zhang, Zhikang Chen et al.ACL 2026
- Re2: Unlocking LLM Reasoning via Reinforcement Learning with Re-solvingPinzheng Wang, Shuli Xu, Juntao Li, Yu Luo et al.ICLR 2026 · 7 citations
- Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired SketchingSimon A. Aytes, Jinheon Baek, Sung Ju HwangEMNLP 2025 · 3 citations
