Flat Reward in Policy Parameter Space Implies Robust Reinforcement Learning
Hyun-Kyu Lee, Sung Whan Yoon
Abstract
Investigating flat minima on loss surfaces in parameter space is well-documented in the supervised learning context, highlighting its advantages for model generalization. However, limited attention has been paid to the reinforcement learning (RL) context, where the impact of flatter reward landscapes in policy parameter space remains largely unexplored. Beyond merely extrapolating from supervised learning, which suggests a link between flat reward landscapes and enhanced generalization, we aim to formally connect the flatness of the reward surface to the robustness of RL models. In policy models where a deep neural network determines actions, flatter reward landscapes in response to parameter perturbations lead to consistent rewards even when actions are perturbed. Moreover, robustness to action perturbations further enhances robustness against other variations, such as changes in state transition probabilities and reward functions. We extensively simulate various RL environments, confirming the consistent benefits of flatter reward landscapes in enhancing the robustness of RL under diverse conditions, including action selection, transition dynamics, and reward functions. The code for these experiments is available at https://github.com/HK-05/flatreward-RRL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- SHAPO: Sharpness-Aware Policy Optimization for Safe ExplorationKaustubh Mani, Yann Pequignot, Vincent Mai, Liam PaullICLR 2026 · 3 citations
- Improving Model-Based Reinforcement Learning by Converging to Flatter MinimaShrinivas Ramasubramanian, Benjamin Freed, Alexandre Capone, Jeff G. SchneiderNeurIPS 2025 · 3 citations
- Reward Sharpness-Aware Fine-Tuning for Diffusion ModelsKwanyoung Kim, Byeongsu SimCVPR 2026 · 1 citation
- Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference OptimizationHaocheng Luo, Zehang Deng, Thanh-Toan Do, Mehrtash Harandi et al.ICLR 2026 · 1 citation
- -SPPO: Semantic-Calibrated Self-Play Preference OptimizationXiwen Chen, Wenhui Zhu, Jingjing Wang, Peijie Qiu et al.ICML 2026
Builds on16
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- SWAD: Domain Generalization by Seeking Flat MinimaJunbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho et al.NeurIPS 2021 · 630 citations
- Overcoming Catastrophic Forgetting in Incremental Few-Shot Learning by Finding Flat MinimaGuangyuan Shi, Jiaxin Chen, Wenlong Zhang, Li-Ming Zhan et al.NeurIPS 2021 · 229 citations
- Generalized Federated Learning via Sharpness Aware MinimizationZhe Qu, Xingyu Li, Rui Duan, Yao Liu et al.ICML 2022 · 219 citations
- Robust Reinforcement Learning for Continuous Control with Model MisspecificationDaniel J. Mankowitz, Nir Levine, Rae Jeong, Abbas Abdolmaleki et al.ICLR 2020 · 138 citations
Related papers
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires et al.ICML 2023 · 162 citations
- Connected Superlevel Set in (Deep) Reinforcement Learning and its Application to Minimax TheoremsSihan Zeng, Thinh T. Doan, Justin RombergNeurIPS 2023 · 4 citations
- Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous ControlNate Rahn, Pierluca D'Oro, Harley Wiltzer, Pierre-Luc Bacon et al.NeurIPS 2023 · 11 citations
- Deep Reinforcement Learning Policies Learn Shared Adversarial Features across MDPsEzgi KorkmazAAAI 2022 · 33 citations
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 244 citations
