Lune

ICML2026Top-tier venue

Demystifying Entropy Control in LLM RL Training: Theoretical Analysis and Dynamic Scheduling

Jingchu Gai, Guanning Zeng, Huaqing ZHANG, Han Zhong, Yige Hong, Andrej Risteski, Aditi Raghunathan

2026Year

Abstract

We investigate a pivotal yet debated component of reinforcement learning (RL) for training large language models (LLMs): controlling entropy (increasing or decreasing it) during RL fine-tuning. The existing literature presents a dichotomy: some studies posit that increasing entropy facilitates exploration, whereas others argue that decreasing entropy enhances performance. Crucially, we observe that the impact of entropy regularization exhibits significant heterogeneity across different tasks. In this paper, we resolve this conflict by identifying the governing factor of optimal entropy control. We define Entropy Discrepancy (Definition 1) and demonstrate that this metric dictates the appropriate direction of regularization. Guided by this insight, we derive a principled dynamic scheduling method that adaptively modulates the entropy coefficient, seamlessly switching between maximization and minimization as training evolves. Extensive experiments confirm the correlation between Entropy Discrepancy and the efficacy of entropy control. Furthermore, our adaptive method yields substantial improvements over vanilla RL, achieving a 6.7 percentage-point Pass@K gain on AIME24 and a 17.52 percentage-point gain on KNK5 at Pass@1.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext dedcaa2f-c7d1-44b9-ac2f-ab08851736ff

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines