Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models
Sohyun An, Ruochen Wang, Tianyi Zhou, Cho-Jui Hsieh
摘要
While recent success of large reasoning models (LRMs) significantly advanced LLMs' reasoning capability by optimizing the final answer accuracy using reinforcement learning, they may also drastically increase the output length due to overthinking-characterized by unnecessarily complex reasoning paths that waste computation and potentially degrade the performance. We hypothesize that such inefficiencies stem from LRMs' limited capability to dynamically select the proper modular reasoning strategies, termed thinking patterns at the right position. To investigate this hypothesis, we propose a dynamic optimization framework that segments model-generated reasoning paths into distinct thinking patterns, systematically identifying and promoting beneficial patterns that improve the answer while removing detrimental ones. Empirical analysis confirms that our optimized thinking paths yield more concise yet sufficiently informative trajectories, enhancing reasoning efficiency by reducing attention FLOPs by up to 47% while maintaining accuracy for originally correct responses. Moreover, a non-trivial portion of originally incorrect responses are transformed into correct ones, achieving a 15.6% accuracy improvement with reduced length. Motivated by the improvement brought by the optimized thinking paths, we apply a preference optimization technique supported by a pairwise dataset contrasting suboptimal and optimal reasoning paths. Experimental evaluations across multiple mathematical reasoning benchmarks reveal that our method notably reduces computational overhead while simultaneously improving reasoning accuracy, achieving up to a 12% accuracy improvement and reducing token usage from approximately 5,000 to 3,000 tokens. Q: If a and b are integers such that ab ≡ 17 mod 20, then what is the remainder when (a+10)(b+10) is divided by 20? (Answer: 17) Wait , 10a + 10b can be factored as 10 x (a + b). So, 10(a + b) mod 20. Hmm, so that term is 10 times something, and modulo 20. I remember that if you have a multiple of 10 modulo 20, it's either 0 or 10. Because 10 * even number is 0 mod 20, and 10 * odd number is 10 mod 20. So, 10(a + b) is either 0 or 10 mod 20, depending on whether (a + b) is even or odd. So, we have 17 + 10(a + b) mod 20. So, that would be: If a + b is even, then 10(a + b) ≡ 0 mod 20, so the total is 17 + 0 = 17 mod 20. If a + b is odd, then 10(a + b) ≡ 10 mod 20, so the total is 17 + 10 = 27 mod 20, which is 7. Keep Pattern 2 Ptn 2 Wait , so depending on whether a + b is even or odd, the result is either 17 or 7. … it's possible that a + b can be either even or odd, so we might have two possible remainders. Remove Pattern 3 Ptn 3 Wait, maybe I can express (a + 10)(b + 10) as ab + 10a + 10b + 100, which we already did. ... So, 10(a + b) mod 20 is either 0 or 10, depending on a + b. … Remove Pattern 11 Ptn 11 Wait , but maybe we can find the parity of a + b. Let's think about that. Given that ab ≡ 17 mod 20. So, ab is congruent to 17 mod 20. Let's think about the possible values of a and b modulo 20. Since 20 is the modulus, a and b can each be from 0 to 19. ... both a and b are odd, so their sum a + b is even. Because odd + odd = even. So, a + b is even. Therefore, 10*(a + b) mod 20 is 0, since 10even = multiple of 20. Therefore, 10(a + b) ≡ 0 mod 20. So, going back, (a + 10)(b + 10) ≡ 17 + 0 = 17 mod 20. Keep Pattern 12 Ptn 12 But hold on , earlier I thought it could be 17 or 7, but now since both a and b are odd, their sum is even, so 10(a + b) is 0 mod 20, so 17 + 0 is 17. So, the remainder is 17. Keep Pattern 13 Ptn 13 … … Wait , this is getting a bit tedious, but I think the key point is that a and b must both be odd because 17 is odd and 20 is even, so their product has to be odd, so both a and b are odd. … Ptn 37
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- PEAR: Phase Entropy Aware Reward for Efficient ReasoningChen Huang, Wei Lu, Wenxuan ZhangICLR 2026 · 被引用 7 次
- Optimizing Test-Time Compute via Meta Reinforcement FinetuningYuxiao Qu, Matthew Y. R. Yang, Amrith Setlur, Lewis Tunstall 等ICML 2025
- Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Significant Gains in Reasoning Efficiency in Large Language ModelsQiguang Chen, Dengyun Peng, Jinhao Liu, Huikang Su 等AAAI 2026
- Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning ModelRong Bao, Bo Wang, Xiao Wang, Hongyu Li 等AAAI 2026
- Quantifying and Understanding Uncertainty in Large Reasoning ModelsYangyi Li, Chenxu Zhao, Mengdi HuaiACL 2026
它引用的顶会 Paper8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 被引用 270 次
相关 Paper
- Incentivizing Dual Process Thinking for Efficient Large Language Model ReasoningXiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang 等NeurIPS 2025 · 被引用 25 次
- SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive ThinkingWeiyang Huang, Xuefeng Bai, Kehai Chen, Xinyang Chen 等ACL 2026 · 被引用 3 次
- Think Better, Not Longer: Token-Level Marginal Utility for Efficient Reasoning in Large Reasoning ModelsJiawei Li, Yang Gao, Huashan Sun, Chong FengACL 2026
- MuTIS: Enhancing Reasoning Efficiency through Multi Turn Intervention Sampling in Reinforcement LearningWenshuo Zhao, Haoxing Zhai, Xinyu Qiu, Zhenting Qi 等EMNLP 2025 · 被引用 3 次
- Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RLSongjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian 等NeurIPS 2025 · 被引用 69 次
