Decoupling regularization from the action space
Sobhan Mohammadpour, Emma Frejinger, Pierre-Luc Bacon
Abstract
Regularized reinforcement learning (RL), particularly the entropy-regularized kind, has gained traction in optimal control and inverse RL. While standard unregularized RL methods remain unaffected by changes in the number of actions, we show that it can severely impact their regularized counterparts. This paper demonstrates the importance of decoupling the regularizer from the action space: that is, to maintain a consistent level of regularization regardless of how many actions are involved to avoid over-regularization. Whereas the problem can be avoided by introducing a task-specific temperature parameter, it is often undesirable and cannot solve the problem when action spaces are state-dependent. In the state-dependent action context, different states with varying action spaces are regularized inconsistently. We introduce two solutions: a static temperature selection approach and a dynamic counterpart, universally applicable where this problem arises. Implementing these changes improves performance on the DeepMind control suite in static and dynamic temperature regimes and a biological sequence design task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f05cf6a-d1d6-4f67-ae76-8d1811e50ce6Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup et al.NeurIPS 2021 · 565 citations
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun et al.NeurIPS 2022 · 316 citations
- Twice regularized MDPs and the equivalence between robustness and regularizationEsther Derman, Matthieu Geist, Shie MannorNeurIPS 2021 · 68 citations
- Extreme Q-Learning: MaxEnt RL without EntropyDivyansh Garg, Joey Hejna, Matthieu Geist, Stefano ErmonICLR 2023 · 5 citations
Related papers
- REValueD: Regularised Ensemble Value-Decomposition for Factorisable Markov Decision ProcessesDavid Ireland, Giovanni MontanaICLR 2024 · 6 citations
- Regularization Matters in Policy Optimization - An Empirical Study on Continuous ControlZhuang Liu, Xuanlin Li, Bingyi Kang, Trevor DarrellICLR 2021 · 8 citations
- Convergence Theorems for Entropy-Regularized and Distributional Reinforcement LearningYash Jhaveri, Harley Wiltzer, Patrick Shafto, Marc G. Bellemare et al.NeurIPS 2025 · 3 citations
- Negatively Correlated Ensemble Reinforcement Learning for Online Diverse Game Level GenerationZiqi Wang, Chengpeng Hu, Jialin Liu, Xin YaoICLR 2024 · 8 citations
- Demystifying Entropy Control in LLM RL Training: Theoretical Analysis and Dynamic SchedulingJingchu Gai, Guanning Zeng, Huaqing ZHANG, Han Zhong et al.ICML 2026
