Hyperparameters in Reinforcement Learning and How To Tune Them
Theresa Eimer, Marius Lindauer, Roberta Raileanu
Abstract
In order to improve reproducibility, deep reinforcement learning (RL) has been adopting better scientific practices such as standardized evaluation metrics and reporting. However, the process of hyperparameter optimization still varies widely across papers, which makes it challenging to compare RL algorithms fairly. In this paper, we show that hyperparameter choices in RL can significantly affect the agent's final performance and sample efficiency, and that the hyperparameter landscape can strongly depend on the tuning seed which may lead to overfitting. We therefore propose adopting established best practices from Au-toML, such as the separation of tuning and testing seeds, as well as principled hyperparameter optimization (HPO) across a broad search space. We support this by comparing multiple state-of-theart HPO tools on a range of RL algorithms and environments to their hand-tuned counterparts, demonstrating that HPO approaches often have higher performance and lower compute overhead. As a result of our findings, we recommend a set of best practices for the RL community, which should result in stronger empirical results with fewer computational costs, better reproducibility, and thus faster progress. In order to encourage the adoption of these practices, we provide plug-andplay implementations of the tuning algorithms used in this paper at https://github.com/ facebookresearch/how-to-autorl .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccb17885-ef0e-434b-aceb-1d345027f3ccCited by top-tier papers18
- A Method for Evaluating Hyperparameter Sensitivity in Reinforcement LearningJacob Adkins, Michael Bowling, Adam WhiteNeurIPS 2024 · 36 citations
- ReLU to the Rescue: Improve Your On-Policy Actor-Critic with Positive AdvantagesAndrew Jesson, Chris Lu, Gunshi Gupta, Nicolas Beltran-Velez et al.ICML 2024 · 11 citations
- RARE: Retrieval-Augmented Reasoning ModelingZhengren Wang, Jiayang Yu, Dongsheng Ma, Zhe Chen et al.KDD 2026 · 9 citations
- Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL FinetuningDesai Xie, Jiahao Li, Hao Tan, Xin Sun et al.CVPR 2024 · 9 citations
- Enhancing Tactile-based Reinforcement Learning for Robotic ControlElle Miller, Trevor McInroe, David Abel, Oisin Mac Aodha et al.NeurIPS 2025 · 9 citations
Builds on17
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann et al.ICML 2020 · 584 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 198 citations
Related papers
- Sample-Efficient Automated Deep Reinforcement LearningJörg K. H. Franke, Gregor Köhler, André Biedenkapp, Frank HutterICLR 2021 · 49 citations
- Adaptive Q-Network: On-the-fly Target Selection for Deep Reinforcement LearningThéo Vincent, Fabian Wahren, Jan Peters, Boris Belousov et al.ICLR 2025
- Gray-Box Gaussian Processes for Automated Reinforcement LearningGresa Shala, André Biedenkapp, Frank Hutter, Josif GrabockaICLR 2023
- ULTHO: Ultra-Lightweight Yet Efficient Hyperparameter Optimization in Deep Reinforcement LearningMingqi Yuan, Bo Li, Xin Jin, Wenjun ZengICCV 2025
- Hyperparameter Optimization Is Deceiving Us, and How to Stop ItA. Feder Cooper, Yucheng Lu, Jessica Zosa Forde, Christopher De SaNeurIPS 2021 · 40 citations
