Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
Jiameng Huang, Baijiong Lin, Guhao Feng, Jierun Chen, Di He, Lu Hou
摘要
Recent Large Reasoning Language Models (LRLMs) employ long chain-of-thought reasoning with complex reflection behaviors, typically signaled by specific trigger words (e.g., "Wait" and "Alternatively") to enhance performance. However, these reflection behaviors can lead to the overthinking problem where the generation of redundant reasoning steps that unnecessarily increase token usage, raise inference costs, and reduce practical utility. In this paper, we propose Certainty-Guided Reflection Suppression (CGRS), a novel method that mitigates overthinking in LRLMs while maintaining reasoning accuracy. CGRS operates by dynamically suppressing the model's generation of reflection triggers when it exhibits high confidence in its current response, thereby preventing redundant reflection cycles without compromising output quality. Our approach is model-agnostic, requires no retraining or architectural modifications, and can be integrated seamlessly with existing autoregressive generation pipelines. Extensive experiments across four reasoning benchmarks (i.e., AIME24, AMC23, MATH500, and GPQA-D) demonstrate CGRS's effectiveness: it reduces token usage by an average of 18.5% to 41.9% while preserving accuracy and also achieves the optimal balance between length reduction and performance compared to state-of-the-art baselines. These results hold consistently across model architectures (e.g., DeepSeek-R1-Distill series, QwQ-32B, and Qwen3 family) and scales (4B to 32B parameters), highlighting CGRS's practical value for efficient reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMsIvo Petrov, Jasper Dekoninck, Martin VechevICML 2026 · 被引用 25 次
- The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics AnalysisZihao Wei, Liang Pang, Jiahao Liu, Wenjie Shi 等ACL 2026 · 被引用 14 次
- Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought ReasoningRenliang Sun, Wei Cheng, Dawei Li, Haifeng Chen 等ACL 2026 · 被引用 11 次
- When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient ReasoningYang Xiang, Yixin Ji, Ruotao Xu, Dan Qiao 等ACL 2026 · 被引用 3 次
- SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive ThinkingWeiyang Huang, Xuefeng Bai, Kehai Chen, Xinyang Chen 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper7
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 被引用 270 次
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu 等ICLR 2026 · 被引用 250 次
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMsJingyao Wang, Wenwen Qiang, Zeen Song, Changwen Zheng 等NeurIPS 2025 · 被引用 13 次
相关 Paper
- Think Better, Not Longer: Token-Level Marginal Utility for Efficient Reasoning in Large Reasoning ModelsJiawei Li, Yang Gao, Huashan Sun, Chong FengACL 2026
- Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated PenaltyZewei Yu, Lirong Gao, Yuke Zhu, Bo Zheng 等ICLR 2026 · 被引用 2 次
- VeriThinker: Learning to Verify Makes Reasoning Model EfficientZigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu 等NeurIPS 2025 · 被引用 30 次
- Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RLSongjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian 等NeurIPS 2025 · 被引用 69 次
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking TokensWei-Lin Chen, Liqian Peng, Tian Tan, Chao Zhao 等ICML 2026 · 被引用 20 次
