Analyzing Generalization in Policy Networks: A Case Study with the Double-Integrator System
Ruining Zhang, Haoran Han, Maolong Lv, Qisong Yang, Jian Cheng
摘要
Extensive utilization of deep reinforcement learning (DRL) policy networks in diverse continuous control tasks has raised questions regarding performance degradation in expansive state spaces where the input state norm is larger than that in the training environment. This paper aims to uncover the underlying factors contributing to such performance deterioration when dealing with expanded state spaces, using a novel analysis technique known as state division. In contrast to prior approaches that employ state division merely as a post-hoc explanatory tool, our methodology delves into the intrinsic characteristics of DRL policy networks. Specifically, we demonstrate that the expansion of state space induces the activation function to exhibit saturability, resulting in the transformation of the state division boundary from nonlinear to linear. Our analysis centers on the paradigm of the double-integrator system, revealing that this gradual shift towards linearity imparts a control behavior reminiscent of bang-bang control. However, the inherent linearity of the division boundary prevents the attainment of an ideal bang-bang control, thereby introducing unavoidable overshooting. Our experimental investigations, employing diverse RL algorithms, establish that this performance phenomenon stems from inherent attributes of the DRL policy network, remaining consistent across various optimization algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto 等NeurIPS 2020 · 被引用 833 次
- Network Randomization: A Simple Technique for Generalization in Deep Reinforcement LearningKimin Lee, Kibok Lee, Jinwoo Shin, Honglak LeeICLR 2020 · 被引用 191 次
- WCSAC: Worst-Case Soft Actor Critic for Safety-Constrained Reinforcement LearningQisong Yang, Thiago D. Simão, Simon H. Tindemans, Matthijs T. J. SpaanAAAI 2021 · 被引用 168 次
- Improving Generalization in Reinforcement Learning with Mixture RegularizationKaixin Wang, Bingyi Kang, Jie Shao, Jiashi FengNeurIPS 2020 · 被引用 143 次
相关 Paper
- Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli PoliciesTim Seyde, Igor Gilitschenski, Wilko Schwarting, Bartolomeo Stellato 等NeurIPS 2021 · 被引用 59 次
- LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal ControlXujie Song, Jingliang Duan, Wenxuan Wang, Shengbo Eben Li 等ICML 2023 · 被引用 19 次
- Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous ControlNate Rahn, Pierluca D'Oro, Harley Wiltzer, Pierre-Luc Bacon 等NeurIPS 2023 · 被引用 11 次
- Understanding the Evolution of Linear Regions in Deep Reinforcement LearningSetareh Cohan, Nam Hee Kim, David Rolnick, Michiel van de PanneNeurIPS 2022 · 被引用 10 次
- Is High Variance Unavoidable in RL? A Case Study in Continuous ControlJohan Bjorck, Carla P. Gomes, Kilian Q. WeinbergerICLR 2022 · 被引用 36 次
