Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
Renliang Sun, Wei Cheng, Dawei Li, Haifeng Chen, Wei Wang
摘要
Chain-of-Thought (CoT) reasoning has driven recent gains of large language models (LLMs) on reasoning-intensive tasks by externalizing intermediate steps. However, excessive or redundant reasoning -- so-called overthinking -- can increase inference costs and lead LLMs toward incorrect conclusions. In this paper, we present REFRAIN (lective-edundancy for daptive ference), a training-free framework that adaptively determines when to stop reasoning to mitigate overthinking. REFRAIN integrates a two-stage stop discriminator to identify reflective yet redundant reasoning and a sliding-window Upper Confidence Bound (SW-UCB) multi-armed bandit controller to dynamically adjust stopping thresholds according to problem difficulty without supervision or fine-tuning. Across four representative benchmarks and two model families, REFRAIN reduces token usage by 20-55% while maintaining or improving accuracy compared to standard CoT prompting. Extensive ablation and robustness analyses demonstrate its stability across models, scorers, and prompt variations. In summary, our findings highlight when-to-stop as a new and practical axis of test-time scaling -- enabling models to reason not just more, but just enough.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics AnalysisZihao Wei, Liang Pang, Jiahao Liu, Wenjie Shi 等ACL 2026 · 被引用 14 次
- Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and FailuresYi Hu, Jiaqi Gu, Ruxin Wang, Zijun Yao 等ACL 2026 · 被引用 5 次
- When to Think and When to Look: Uncertainty-Guided LookbackJing Bi, Filippos Bellos, JunJia Guo, Yayuan Li 等CVPR 2026 · 被引用 4 次
- Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought ReasoningRenos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille 等ICML 2026 · 被引用 4 次
- Think Better, Not Longer: Token-Level Marginal Utility for Efficient Reasoning in Large Reasoning ModelsJiawei Li, Yang Gao, Huashan Sun, Chong FengACL 2026
它引用的顶会 Paper6
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu 等ICLR 2026 · 被引用 250 次
- Deep Think with ConfidenceYichao Fu, Xuewei Wang, Hao Zhang, Yuandong Tian 等ICLR 2026 · 被引用 171 次
- Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step ReasoningYiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan 等ICLR 2024 · 被引用 101 次
- Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection SuppressionJiameng Huang, Baijiong Lin, Guhao Feng, Jierun Chen 等AAAI 2026 · 被引用 21 次
- ConCISE: Confidence-guided Compression in Step-by-step Efficient ReasoningZiqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang 等EMNLP 2025 · 被引用 1 次
相关 Paper
- Stop When Further Reasoning Won’t Help: Attention-State Adaptive Generation in Reasoning ModelsJiakai Li, KE QIN, Rongzheng Wang, Yizhuo Ma 等ICML 2026
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du 等ICLR 2026 · 被引用 225 次
- Think Faster Than Words: Efficient LLM Chain-of-Thought Reasoning via Dynamic Shortcut DecodingFan Liu, Yanhao Wang, Min Zhang, Zhikang Chen 等ACL 2026
- Re2: Unlocking LLM Reasoning via Reinforcement Learning with Re-solvingPinzheng Wang, Shuli Xu, Juntao Li, Yu Luo 等ICLR 2026 · 被引用 7 次
- Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired SketchingSimon A. Aytes, Jinheon Baek, Sung Ju HwangEMNLP 2025 · 被引用 3 次
