A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning
Jacob Adkins, Michael Bowling, Adam White
Abstract
The performance of modern reinforcement learning algorithms critically relies on tuning ever-increasing numbers of hyperparameters. Often, small changes in a hyperparameter can lead to drastic changes in performance, and different environments require very different hyperparameter settings to achieve state-of-the-art performance reported in the literature. We currently lack a scalable and widely accepted approach to characterizing these complex interactions. This work proposes a new empirical methodology for studying, comparing, and quantifying the sensitivity of an algorithm's performance to hyperparameter tuning for a given set of environments. We then demonstrate the utility of this methodology by assessing the hyperparameter sensitivity of several commonly used normalization variants of PPO. The results suggest that several algorithmic performance improvements may, in fact, be a result of an increased reliance on hyperparameter tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e0766ba-96c2-46d5-a6af-ddab6877aa77Cited by top-tier papers4
- Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular DesignXingyu Su, Xiner Li, Masatoshi Uehara, Sunwoo Kim et al.ICLR 2026 · 10 citations
- SPO: Sequential Monte Carlo Policy OptimisationMatthew Macfarlane, Edan Toledo, Donal Byrne, Paul Duckworth et al.NeurIPS 2024 · 8 citations
- SoK: The Pitfalls of Deep Reinforcement Learning for CybersecurityShae McFadden, Myles Foley, Elizabeth Bates, Ilias Tsingenopoulos et al.USENIX Security 2026 · 7 citations
- Improving Value Estimation Critically Enhances Vanilla Policy GradientTao Wang, Ruipeng Zhang, Sicun GaoICML 2025
Builds on4
- Discovered Policy OptimisationChris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz et al.NeurIPS 2022 · 134 citations
- Hyperparameters in Reinforcement Learning and How To Tune ThemTheresa Eimer, Marius Lindauer, Roberta RaileanuICML 2023 · 96 citations
- Evaluating the Performance of Reinforcement Learning AlgorithmsScott M. Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang et al.ICML 2020 · 59 citations
- Sample-Efficient Automated Deep Reinforcement LearningJörg K. H. Franke, Gregor Köhler, André Biedenkapp, Frank HutterICLR 2021 · 49 citations
Related papers
- The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning NetworksWalter Mayor, Johan S. Obando-Ceron, Aaron C. Courville, Pablo Samuel CastroICML 2025
- Hyperparameter Selection for Imitation LearningLéonard Hussenot, Marcin Andrychowicz, Damien Vincent, Robert Dadashi et al.ICML 2021 · 20 citations
- Gray-Box Gaussian Processes for Automated Reinforcement LearningGresa Shala, André Biedenkapp, Frank Hutter, Josif GrabockaICLR 2023
- On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning ImplementationsRajdeep Singh Hundal, Yan Xiao, Xiaochun Cao, Jin Song Dong et al.ICSE 2025
- Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 TricksRyan Sullivan, Akarsh Kumar, Shengyi Huang, John P. Dickerson et al.NeurIPS 2023 · 11 citations
