Demystifying Entropy Control in LLM RL Training: Theoretical Analysis and Dynamic Scheduling
Jingchu Gai, Guanning Zeng, Huaqing ZHANG, Han Zhong, Yige Hong, Andrej Risteski, Aditi Raghunathan
摘要
We investigate a pivotal yet debated component of reinforcement learning (RL) for training large language models (LLMs): controlling entropy (increasing or decreasing it) during RL fine-tuning. The existing literature presents a dichotomy: some studies posit that increasing entropy facilitates exploration, whereas others argue that decreasing entropy enhances performance. Crucially, we observe that the impact of entropy regularization exhibits significant heterogeneity across different tasks. In this paper, we resolve this conflict by identifying the governing factor of optimal entropy control. We define Entropy Discrepancy (Definition 1) and demonstrate that this metric dictates the appropriate direction of regularization. Guided by this insight, we derive a principled dynamic scheduling method that adaptively modulates the entropy coefficient, seamlessly switching between maximization and minimization as training evolves. Extensive experiments confirm the correlation between Entropy Discrepancy and the efficacy of entropy control. Furthermore, our adaptive method yields substantial improvements over vanilla RL, achieving a 6.7 percentage-point Pass@K gain on AIME24 and a 17.52 percentage-point gain on KNK5 at Pass@1.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
相关 Paper
- On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language ModelsShumin Wang, Yuexiang Xie, Wenhao Zhang, Yuchang Sun 等ICML 2026 · 被引用 7 次
- Differential Smoothing Mitigates Sharpening and Improves LLM ReasoningJingchu Gai, Guanning Zeng, Huaqing ZHANG, Aditi RaghunathanICML 2026 · 被引用 13 次
- Reasoning with Exploration: An Entropy PerspectiveDaixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai 等AAAI 2026 · 被引用 216 次
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo 等ACL 2026 · 被引用 42 次
- Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement LearningTongxi Wang, Zhuoyang Xia, Xinran Chen, Shan LiuICML 2026
