Discovery of Options via Meta-Learned Subgoals
Vivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu, Junhyuk Oh, Iurii Kemaev, Hado van Hasselt, David Silver, Satinder Singh
摘要
Temporal abstractions in the form of options have been shown to help reinforcement learning (RL) agents learn faster. However, despite prior work on this topic, the problem of discovering options through interaction with an environment remains a challenge. In this paper, we introduce a novel meta-gradient approach for discovering useful options in multi-task RL environments. Our approach is based on a manager-worker decomposition of the RL agent, in which a manager maximises rewards from the environment by learning a task-dependent policy over both a set of task-independent discovered-options and primitive actions. The option-reward and termination functions that define a subgoal for each option are parameterised as neural networks and trained via meta-gradients to maximise their usefulness. Empirical analysis on gridworld and DeepMind Lab tasks show that: (1) our approach can discover meaningful and diverse temporally-extended options in multi-task RL domains, (2) the discovered options are frequently used by the agent while learning to solve the training tasks, and (3) that the discovered options help a randomly initialised manager learn faster in completely new tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Deep Hierarchical Planning from PixelsDanijar Hafner, Kuang-Huei Lee, Ian Fischer, Pieter AbbeelNeurIPS 2022 · 被引用 153 次
- CivRealm: A Learning and Reasoning Odyssey in Civilization for Decision-Making AgentsSiyuan Qi, Shuo Chen, Yexin Li, Xiangyu Kong 等ICLR 2024 · 被引用 35 次
- Pre-Training Goal-based Models for Sample-Efficient Reinforcement LearningHaoqi Yuan, Zhancun Mu, Feiyang Xie, Zongqing LuICLR 2024 · 被引用 26 次
- Discovering General Reinforcement Learning Algorithms with Adversarial Environment DesignMatthew Thomas Jackson, Minqi Jiang, Jack Parker-Holder, Risto Vuorio 等NeurIPS 2023 · 被引用 23 次
- Goal-Conditioned On-Policy Reinforcement LearningXudong Gong, Dawei Feng, Kele Xu, Bo Ding 等NeurIPS 2024 · 被引用 16 次
它引用的顶会 Paper1
相关 Paper
- Unveiling Options with Neural Network DecompositionMahdi Alikhasi, Levi LelisICLR 2024 · 被引用 2 次
- Options of Interest: Temporal Abstraction with Interest FunctionsKhimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon 等AAAI 2020 · 被引用 51 次
- Flexible Option LearningMartin Klissarov, Doina PrecupNeurIPS 2021 · 被引用 38 次
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh 等NeurIPS 2020 · 被引用 90 次
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu 等NeurIPS 2020 · 被引用 154 次
