Hyperparameter Selection for Imitation Learning
Léonard Hussenot, Marcin Andrychowicz, Damien Vincent, Robert Dadashi, Anton Raichuk, Sabela Ramos, Nikola Momchev, Sertan Girgin, Raphaël Marinier, Lukasz Stafiniak, Manu Orsini, Olivier Bachem
摘要
We address the issue of tuning hyperparameters (HPs) for imitation learning algorithms in the context of continuous-control, when the underlying reward function of the demonstrating expert cannot be observed at any time. The vast literature in imitation learning mostly considers this reward function to be available for HP selection, but this is not a realistic setting. Indeed, would this reward function be available, it could then directly be used for policy training and imitation would not be necessary. To tackle this mostly ignored problem, we propose a number of possible proxies to the external reward. We evaluate them in an extensive empirical study (more than 10'000 agents across 9 environments) and make practical recommendations for selecting HPs. Our results show that while imitation learning algorithms are sensitive to HP choices, it is often possible to select good enough HPs through a proxy to the reward function.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- What Matters for Adversarial Imitation Learning?Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent 等NeurIPS 2021 · 被引用 106 次
- Continuous Control with Action Quantization from DemonstrationsRobert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin 等ICML 2022 · 被引用 32 次
- Uni[MASK]: Unified Inference in Sequential Decision ProblemsMicah Carroll, Orr Paradise, Jessy Lin, Raluca Georgescu 等NeurIPS 2022 · 被引用 29 次
- Showing Your Offline Reinforcement Learning Work: Online Evaluation Budget MattersVladislav Kurenkov, Sergey KolesnikovICML 2022 · 被引用 25 次
- Imitating Human Behaviour with Diffusion ModelsTim Pearce, Tabish Rashid, Anssi Kanervisto, David Bignell 等ICLR 2023 · 被引用 23 次
它引用的顶会 Paper4
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Decoupling Representation Learning from Reinforcement LearningAdam Stooke, Kimin Lee, Pieter Abbeel, Michael LaskinICML 2021 · 被引用 389 次
- Predictive Information Accelerates Learning in RLKuang-Huei Lee, Ian Fischer, Anthony Z. Liu, Yijie Guo 等NeurIPS 2020 · 被引用 82 次
- Primal Wasserstein Imitation LearningRobert Dadashi, Léonard Hussenot, Matthieu Geist, Olivier PietquinICLR 2021 · 被引用 41 次
相关 Paper
- Imitation Learning by Reinforcement LearningKamil CiosekICLR 2022 · 被引用 22 次
- Imitation by Predicting ObservationsAndrew Jaegle, Yury Sulsky, Arun Ahuja, Jake Bruce 等ICML 2021 · 被引用 16 次
- On Covariate Shift of Latent Confounders in Imitation and Reinforcement LearningGuy Tennenholtz, Assaf Hallak, Gal Dalal, Shie Mannor 等ICLR 2022 · 被引用 16 次
- A Method for Evaluating Hyperparameter Sensitivity in Reinforcement LearningJacob Adkins, Michael Bowling, Adam WhiteNeurIPS 2024 · 被引用 36 次
- Inverse Reinforcement Learning in a Continuous State Space with Formal GuaranteesGregory Dexter, Kevin Bello, Jean HonorioNeurIPS 2021 · 被引用 9 次
