Demystifying Entropy Control in LLM RL Training: Theoretical Analysis and Dynamic Scheduling
Jingchu Gai, Guanning Zeng, Huaqing ZHANG, Han Zhong, Yige Hong, Andrej Risteski, Aditi Raghunathan
Abstract
We investigate a pivotal yet debated component of reinforcement learning (RL) for training large language models (LLMs): controlling entropy (increasing or decreasing it) during RL fine-tuning. The existing literature presents a dichotomy: some studies posit that increasing entropy facilitates exploration, whereas others argue that decreasing entropy enhances performance. Crucially, we observe that the impact of entropy regularization exhibits significant heterogeneity across different tasks. In this paper, we resolve this conflict by identifying the governing factor of optimal entropy control. We define Entropy Discrepancy (Definition 1) and demonstrate that this metric dictates the appropriate direction of regularization. Guided by this insight, we derive a principled dynamic scheduling method that adaptively modulates the entropy coefficient, seamlessly switching between maximization and minimization as training evolves. Extensive experiments confirm the correlation between Entropy Discrepancy and the efficacy of entropy control. Furthermore, our adaptive method yields substantial improvements over vanilla RL, achieving a 6.7 percentage-point Pass@K gain on AIME24 and a 17.52 percentage-point gain on KNK5 at Pass@1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dedcaa2f-c7d1-44b9-ac2f-ab08851736ffBuilds on21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
Related papers
- On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language ModelsShumin Wang, Yuexiang Xie, Wenhao Zhang, Yuchang Sun et al.ICML 2026 · 7 citations
- Differential Smoothing Mitigates Sharpening and Improves LLM ReasoningJingchu Gai, Guanning Zeng, Huaqing ZHANG, Aditi RaghunathanICML 2026 · 13 citations
- Reasoning with Exploration: An Entropy PerspectiveDaixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai et al.AAAI 2026 · 216 citations
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo et al.ACL 2026 · 42 citations
- Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement LearningTongxi Wang, Zhuoyang Xia, Xinran Chen, Shan LiuICML 2026
