Hyperparameters in Reinforcement Learning and How To Tune Them
Theresa Eimer, Marius Lindauer, Roberta Raileanu
摘要
In order to improve reproducibility, deep reinforcement learning (RL) has been adopting better scientific practices such as standardized evaluation metrics and reporting. However, the process of hyperparameter optimization still varies widely across papers, which makes it challenging to compare RL algorithms fairly. In this paper, we show that hyperparameter choices in RL can significantly affect the agent's final performance and sample efficiency, and that the hyperparameter landscape can strongly depend on the tuning seed which may lead to overfitting. We therefore propose adopting established best practices from Au-toML, such as the separation of tuning and testing seeds, as well as principled hyperparameter optimization (HPO) across a broad search space. We support this by comparing multiple state-of-theart HPO tools on a range of RL algorithms and environments to their hand-tuned counterparts, demonstrating that HPO approaches often have higher performance and lower compute overhead. As a result of our findings, we recommend a set of best practices for the RL community, which should result in stronger empirical results with fewer computational costs, better reproducibility, and thus faster progress. In order to encourage the adoption of these practices, we provide plug-andplay implementations of the tuning algorithms used in this paper at https://github.com/ facebookresearch/how-to-autorl .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- A Method for Evaluating Hyperparameter Sensitivity in Reinforcement LearningJacob Adkins, Michael Bowling, Adam WhiteNeurIPS 2024 · 被引用 36 次
- ReLU to the Rescue: Improve Your On-Policy Actor-Critic with Positive AdvantagesAndrew Jesson, Chris Lu, Gunshi Gupta, Nicolas Beltran-Velez 等ICML 2024 · 被引用 11 次
- RARE: Retrieval-Augmented Reasoning ModelingZhengren Wang, Jiayang Yu, Dongsheng Ma, Zhe Chen 等KDD 2026 · 被引用 9 次
- Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL FinetuningDesai Xie, Jiahao Li, Hao Tan, Xin Sun 等CVPR 2024 · 被引用 9 次
- Enhancing Tactile-based Reinforcement Learning for Robotic ControlElle Miller, Trevor McInroe, David Abel, Oisin Mac Aodha 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper17
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann 等ICML 2020 · 被引用 584 次
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 305 次
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 被引用 198 次
相关 Paper
- Sample-Efficient Automated Deep Reinforcement LearningJörg K. H. Franke, Gregor Köhler, André Biedenkapp, Frank HutterICLR 2021 · 被引用 49 次
- Adaptive Q-Network: On-the-fly Target Selection for Deep Reinforcement LearningThéo Vincent, Fabian Wahren, Jan Peters, Boris Belousov 等ICLR 2025
- Gray-Box Gaussian Processes for Automated Reinforcement LearningGresa Shala, André Biedenkapp, Frank Hutter, Josif GrabockaICLR 2023
- ULTHO: Ultra-Lightweight Yet Efficient Hyperparameter Optimization in Deep Reinforcement LearningMingqi Yuan, Bo Li, Xin Jin, Wenjun ZengICCV 2025
- Hyperparameter Optimization Is Deceiving Us, and How to Stop ItA. Feder Cooper, Yucheng Lu, Jessica Zosa Forde, Christopher De SaNeurIPS 2021 · 被引用 40 次
