The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE Training
Weize Chen, Jiarui Yuan, Tailin Jin, Ning Ding, Huimin Chen, Zhiyuan Liu, Maosong Sun
Abstract
Recent large language models (LLMs) exhibit impressive reasoning but often overthink, generating excessively long responses that hinder efficiency. We introduce DIET (DIfficulty-AwarE Training), a framework that systematically cuts these "token calories" by integrating on-the-fly problem difficulty into the reinforcement learning (RL) process. DIET dynamically adapts token compression strategies by modulating token penalty strength and conditioning target lengths on estimated task difficulty, to optimize the performance-efficiency trade-off. We also theoretically analyze the pitfalls of naive reward weighting in group-normalized RL algorithms like GRPO, and propose Advantage Weighting technique, which enables stable and effective implementation of these difficulty-aware objectives. Experimental results demonstrate that DIET significantly reduces token counts while simultaneously improving reasoning performance. Beyond raw token reduction, we show two crucial benefits largely overlooked by prior work: (1) DIET leads to superior inference scaling. By maintaining high per-sample quality with fewer tokens, it enables better scaling performance via majority voting with more samples under fixed computational budgets, an area where other methods falter. (2) DIET enhances the natural positive correlation between response length and problem difficulty, ensuring verbosity is appropriately allocated, unlike many existing compression methods that disrupt this relationship. Our analyses provide a principled and effective framework for developing more efficient, practical, and high-performing LLMs. Our code is available at https://github.com/thunlp/DIET.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cd37afc-52ac-43fe-8748-349dee6ffe28Cited by top-tier papers5
- ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy ShapingShuang Chen, Hangyu Guo, Yimeng Ye, Shijue Huang et al.ICLR 2026 · 23 citations
- Self-Aligned Reward: Towards Effective and Efficient ReasonersPeixuan Han, ADIT KRISHNAN, Gerald Friedland, Jiaxuan You et al.ICLR 2026 · 10 citations
- Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty QuantificationShaohao Rui, Kaitao Chen, Weijie Ma, Xiaosong WangICML 2026
- Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning ModelRong Bao, Bo Wang, Xiao Wang, Hongyu Li et al.AAAI 2026
- SafeAdapt: Safety Alignment with Adaptive Thinking Allocation for Large Reasoning ModelsJiazheng Song, Junxu Liu, Jian Lou, Jinfei LiuUSENIX Security 2026
Builds on10
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 270 citations
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du et al.ICLR 2026 · 225 citations
- C3oT: Generating Shorter Chain-of-Thought Without Compromising EffectivenessYu Kang, Xianghui Sun, Liangyu Chen, Wei ZouAAAI 2025 · 162 citations
- Can Language Models Learn to Skip Steps?Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Cheng Jiayang et al.NeurIPS 2024 · 92 citations
Related papers
- Sample More to Think Less: Group Filtered Policy Optimization for Concise ReasoningVaishnavi Shrivastava, Ahmed Hassan Awadallah, Vidhisha Balachandran, Shivam Garg et al.ICLR 2026 · 85 citations
- Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement LearningHanbing Liu, Lang Cao, Yuanyi Ren, Mengyu Zhou et al.ACL 2026 · 5 citations
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy EntropyHongze Tan, Zihan Wang, Jianfei Pan, Jinghao Lin et al.ICML 2026 · 53 citations
- Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step EntropyZeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen et al.ICLR 2026 · 35 citations
- LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning ModelYanhao Li, Lu Ma, Jiaran Zhang, Lexiang Tang et al.ACL 2026 · 8 citations
