Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models
Sohyun An, Ruochen Wang, Tianyi Zhou, Cho-Jui Hsieh
Abstract
While recent success of large reasoning models (LRMs) significantly advanced LLMs' reasoning capability by optimizing the final answer accuracy using reinforcement learning, they may also drastically increase the output length due to overthinking-characterized by unnecessarily complex reasoning paths that waste computation and potentially degrade the performance. We hypothesize that such inefficiencies stem from LRMs' limited capability to dynamically select the proper modular reasoning strategies, termed thinking patterns at the right position. To investigate this hypothesis, we propose a dynamic optimization framework that segments model-generated reasoning paths into distinct thinking patterns, systematically identifying and promoting beneficial patterns that improve the answer while removing detrimental ones. Empirical analysis confirms that our optimized thinking paths yield more concise yet sufficiently informative trajectories, enhancing reasoning efficiency by reducing attention FLOPs by up to 47% while maintaining accuracy for originally correct responses. Moreover, a non-trivial portion of originally incorrect responses are transformed into correct ones, achieving a 15.6% accuracy improvement with reduced length. Motivated by the improvement brought by the optimized thinking paths, we apply a preference optimization technique supported by a pairwise dataset contrasting suboptimal and optimal reasoning paths. Experimental evaluations across multiple mathematical reasoning benchmarks reveal that our method notably reduces computational overhead while simultaneously improving reasoning accuracy, achieving up to a 12% accuracy improvement and reducing token usage from approximately 5,000 to 3,000 tokens. Q: If a and b are integers such that ab ≡ 17 mod 20, then what is the remainder when (a+10)(b+10) is divided by 20? (Answer: 17) Wait , 10a + 10b can be factored as 10 x (a + b). So, 10(a + b) mod 20. Hmm, so that term is 10 times something, and modulo 20. I remember that if you have a multiple of 10 modulo 20, it's either 0 or 10. Because 10 * even number is 0 mod 20, and 10 * odd number is 10 mod 20. So, 10(a + b) is either 0 or 10 mod 20, depending on whether (a + b) is even or odd. So, we have 17 + 10(a + b) mod 20. So, that would be: If a + b is even, then 10(a + b) ≡ 0 mod 20, so the total is 17 + 0 = 17 mod 20. If a + b is odd, then 10(a + b) ≡ 10 mod 20, so the total is 17 + 10 = 27 mod 20, which is 7. Keep Pattern 2 Ptn 2 Wait , so depending on whether a + b is even or odd, the result is either 17 or 7. … it's possible that a + b can be either even or odd, so we might have two possible remainders. Remove Pattern 3 Ptn 3 Wait, maybe I can express (a + 10)(b + 10) as ab + 10a + 10b + 100, which we already did. ... So, 10(a + b) mod 20 is either 0 or 10, depending on a + b. … Remove Pattern 11 Ptn 11 Wait , but maybe we can find the parity of a + b. Let's think about that. Given that ab ≡ 17 mod 20. So, ab is congruent to 17 mod 20. Let's think about the possible values of a and b modulo 20. Since 20 is the modulus, a and b can each be from 0 to 19. ... both a and b are odd, so their sum a + b is even. Because odd + odd = even. So, a + b is even. Therefore, 10*(a + b) mod 20 is 0, since 10even = multiple of 20. Therefore, 10(a + b) ≡ 0 mod 20. So, going back, (a + 10)(b + 10) ≡ 17 + 0 = 17 mod 20. Keep Pattern 12 Ptn 12 But hold on , earlier I thought it could be 17 or 7, but now since both a and b are odd, their sum is even, so 10(a + b) is 0 mod 20, so 17 + 0 is 17. So, the remainder is 17. Keep Pattern 13 Ptn 13 … … Wait , this is getting a bit tedious, but I think the key point is that a and b must both be odd because 17 is odd and 20 is even, so their product has to be odd, so both a and b are odd. … Ptn 37
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29ab07ca-56af-491a-b998-c747c570f89aCited by top-tier papers7
- PEAR: Phase Entropy Aware Reward for Efficient ReasoningChen Huang, Wei Lu, Wenxuan ZhangICLR 2026 · 7 citations
- Optimizing Test-Time Compute via Meta Reinforcement FinetuningYuxiao Qu, Matthew Y. R. Yang, Amrith Setlur, Lewis Tunstall et al.ICML 2025
- Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Significant Gains in Reasoning Efficiency in Large Language ModelsQiguang Chen, Dengyun Peng, Jinhao Liu, Huikang Su et al.AAAI 2026
- Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning ModelRong Bao, Bo Wang, Xiao Wang, Hongyu Li et al.AAAI 2026
- Quantifying and Understanding Uncertainty in Large Reasoning ModelsYangyi Li, Chenxu Zhao, Mengdi HuaiACL 2026
Builds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 270 citations
Related papers
- Incentivizing Dual Process Thinking for Efficient Large Language Model ReasoningXiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang et al.NeurIPS 2025 · 25 citations
- SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive ThinkingWeiyang Huang, Xuefeng Bai, Kehai Chen, Xinyang Chen et al.ACL 2026 · 3 citations
- Think Better, Not Longer: Token-Level Marginal Utility for Efficient Reasoning in Large Reasoning ModelsJiawei Li, Yang Gao, Huashan Sun, Chong FengACL 2026
- MuTIS: Enhancing Reasoning Efficiency through Multi Turn Intervention Sampling in Reinforcement LearningWenshuo Zhao, Haoxing Zhai, Xinyu Qiu, Zhenting Qi et al.EMNLP 2025 · 3 citations
- Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RLSongjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian et al.NeurIPS 2025 · 69 citations
