Goal-Conditioned Agents that Learn Everything All at Once
Michael Matthews, Matthew Jackson, Michael Beukman, Thomas Foster, Alistair Letcher, Scott Fujimoto, Cédric Colas, Jakob Foerster
摘要
A goal-conditioned reinforcement learning agent exploring an environment will see a wealth of information throughout a trajectory, most of which is discarded when only performing on-policy updates with respect to the commanded goal. Allgoals learning, where each transition is used for learning off-policy with respect to every goal, allows agents to extract maximal information, however it is usually computationally infeasible when done via naïve relabelling. This can be overcome by jointly outputting values and actions for every goal at once, allowing for efficient, parallel all-goals updates with a single pass through the network, in a process we call Learning Everything all at Once (LEO). We show that this approach significantly outperforms other methods on goalconditioned Craftax and is competitive with existing baselines on continuous control environments, while achieving a > 250× speed-up compared to all-goals relabelling. We then go on to show that this approach can be made even more powerful by using LEO as a teacher network, rather than a direct actor. We hope that, by unlocking allgoals learning at scale, LEO can serve as a useful tool for RL practitioners in complex environments. We open source our code 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 331 次
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair 等ICML 2020 · 被引用 303 次
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu 等NeurIPS 2020 · 被引用 251 次
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu 等ICLR 2021 · 被引用 222 次
- Benchmarking the Spectrum of Agent CapabilitiesDanijar HafnerICLR 2022 · 被引用 193 次
相关 Paper
- Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural NetworksFabio Pardo, Vitaly Levdik, Petar KormushevAAAI 2020 · 被引用 4 次
- Goal-Conditioned Q-learning as Knowledge DistillationAlexander Levine, Soheil FeiziAAAI 2023 · 被引用 4 次
- Knowledge Transfer in Multi-Task Deep Reinforcement Learning for Continuous ControlZhiyuan Xu, Kun Wu, Zhengping Che, Jian Tang 等NeurIPS 2020 · 被引用 58 次
- Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement LearningMichael T. Matthews, Michael Beukman, Benjamin Ellis, Mikayel Samvelyan 等ICML 2024 · 被引用 71 次
- Learning to Reach Goals via DiffusionVineet Jain, Siamak RavanbakhshICML 2024 · 被引用 11 次
