Is High Variance Unavoidable in RL? A Case Study in Continuous Control
Johan Bjorck, Carla P. Gomes, Kilian Q. Weinberger
Abstract
Reinforcement learning (RL) experiments have notoriously high variance, and minor details can have disproportionately large effects on measured outcomes. This is problematic for creating reproducible research and also serves as an obstacle when applying RL to sensitive real-world applications. In this paper, we investigate causes for this perceived instability. To allow for an in-depth analysis, we focus on a specifically popular setup with high variance -continuous control from pixels with an actor-critic agent. In this setting, we demonstrate that poor outlier runs which completely fail to learn are an important source of variance, but that weight initialization and initial exploration are not at fault. We show that one cause for these outliers is unstable network parametrization which leads to saturating nonlinearities. We investigate several fixes to this issue and find that simply normalizing penultimate features is surprisingly effective. For sparse tasks, we also find that partially disabling clipped double Q-learning decreases variance. By combining fixes we significantly decrease variances, lowering the average standard deviation across 21 tasks by a factor > 3 for a state-of-the-art agent. This demonstrates that the perceived variance is not necessarily inherent to RL. Instead, it may be addressed via simple modifications and we argue that developing low-variance agents is an important goal for the RL community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 144a5aee-81bb-4949-b800-61d8b46f68ccCited by top-tier papers11
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon et al.ICML 2022 · 269 citations
- Value Gradient weighted Model-Based Reinforcement LearningClaas Voelcker, Victor Liao, Animesh Garg, Amir-massoud FarahmandICLR 2022 · 37 citations
- Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay BuffersGautham Vasan, Mohamed Elsayed, Seyed Alireza Azimi, Jiamin He et al.NeurIPS 2024 · 27 citations
- The Impact of Task Underspecification in Evaluating Deep Reinforcement LearningVindula Jayawardana, Catherine Tang, Sirui Li, Dajiang Suo et al.NeurIPS 2022 · 16 citations
- A Long N-step Surrogate Stage Reward for Deep Reinforcement LearningJunmin Zhong, Ruofan Wu, Jennie SiNeurIPS 2023 · 9 citations
Builds on18
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
Related papers
- Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous ControlNate Rahn, Pierluca D'Oro, Harley Wiltzer, Pierre-Luc Bacon et al.NeurIPS 2023 · 11 citations
- LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal ControlXujie Song, Jingliang Duan, Wenxuan Wang, Shengbo Eben Li et al.ICML 2023 · 19 citations
- Towards Deeper Deep Reinforcement Learning with Spectral NormalizationJohan Bjorck, Carla P. Gomes, Kilian Q. WeinbergerNeurIPS 2021 · 26 citations
- Efficient Continuous Control with Double Actors and Regularized CriticsJiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu LiAAAI 2022 · 69 citations
- Learning Multimodal Behaviors from Scratch with Diffusion Policy GradientSteven Li, Rickmer Krohn, Tao Chen, Anurag Ajay et al.NeurIPS 2024 · 61 citations
