Decoupling regularization from the action space
Sobhan Mohammadpour, Emma Frejinger, Pierre-Luc Bacon
摘要
Regularized reinforcement learning (RL), particularly the entropy-regularized kind, has gained traction in optimal control and inverse RL. While standard unregularized RL methods remain unaffected by changes in the number of actions, we show that it can severely impact their regularized counterparts. This paper demonstrates the importance of decoupling the regularizer from the action space: that is, to maintain a consistent level of regularization regardless of how many actions are involved to avoid over-regularization. Whereas the problem can be avoided by introducing a task-specific temperature parameter, it is often undesirable and cannot solve the problem when action spaces are state-dependent. In the state-dependent action context, different states with varying action spaces are regularized inconsistently. We introduce two solutions: a static temperature selection approach and a dynamic counterpart, universally applicable where this problem arises. Implementing these changes improves performance on the DeepMind control suite in static and dynamic temperature regimes and a biological sequence design task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup 等NeurIPS 2021 · 被引用 565 次
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun 等NeurIPS 2022 · 被引用 316 次
- Twice regularized MDPs and the equivalence between robustness and regularizationEsther Derman, Matthieu Geist, Shie MannorNeurIPS 2021 · 被引用 68 次
- Extreme Q-Learning: MaxEnt RL without EntropyDivyansh Garg, Joey Hejna, Matthieu Geist, Stefano ErmonICLR 2023 · 被引用 5 次
相关 Paper
- REValueD: Regularised Ensemble Value-Decomposition for Factorisable Markov Decision ProcessesDavid Ireland, Giovanni MontanaICLR 2024 · 被引用 6 次
- Regularization Matters in Policy Optimization - An Empirical Study on Continuous ControlZhuang Liu, Xuanlin Li, Bingyi Kang, Trevor DarrellICLR 2021 · 被引用 8 次
- Convergence Theorems for Entropy-Regularized and Distributional Reinforcement LearningYash Jhaveri, Harley Wiltzer, Patrick Shafto, Marc G. Bellemare 等NeurIPS 2025 · 被引用 3 次
- Negatively Correlated Ensemble Reinforcement Learning for Online Diverse Game Level GenerationZiqi Wang, Chengpeng Hu, Jialin Liu, Xin YaoICLR 2024 · 被引用 8 次
- Demystifying Entropy Control in LLM RL Training: Theoretical Analysis and Dynamic SchedulingJingchu Gai, Guanning Zeng, Huaqing ZHANG, Han Zhong 等ICML 2026
