Unveiling Options with Neural Network Decomposition
Mahdi Alikhasi, Levi Lelis
摘要
In reinforcement learning, agents often learn policies for specific tasks without the ability to generalize this knowledge to related tasks. This paper introduces an algorithm that attempts to address this limitation by decomposing neural networks encoding policies for Markov Decision Processes into reusable sub-policies, which are used to synthesize temporally extended actions, or options. We consider neural networks with piecewise linear activation functions, so that they can be mapped to an equivalent tree that is similar to oblique decision trees. Since each node in such a tree serves as a function of the input of the tree, each sub-tree is a sub-policy of the main policy. We turn each of these sub-policies into options by wrapping it with while-loops of varied number of iterations. Given the large number of options, we propose a selection mechanism based on minimizing the Levin loss for a uniform policy on these options. Empirical results in two grid-world domains where exploration can be difficult confirm that our method can identify useful options, thereby accelerating the learning process on similar but different tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi 等NeurIPS 2020 · 被引用 364 次
- Exploration in Reinforcement Learning with Deep Covering OptionsYuu Jinnai, Jee Won Park, Marlos C. Machado, George Dimitri KonidarisICLR 2020 · 被引用 64 次
- Same State, Different Task: Continual Reinforcement Learning without InterferenceSamuel Kessler, Jack Parker-Holder, Philip J. Ball, Stefan Zohren 等AAAI 2022 · 被引用 57 次
- Reinforcement Learning with Competitive Ensembles of Information-Constrained PrimitivesAnirudh Goyal, Shagun Sodhani, Jonathan Binas, Xue Bin Peng 等ICLR 2020 · 被引用 55 次
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 被引用 51 次
相关 Paper
- Discovery of Options via Meta-Learned SubgoalsVivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu 等NeurIPS 2021 · 被引用 38 次
- On the Role of Weight Sharing During Deep Option LearningMatthew Riemer, Ignacio Cases, Clemens Rosenbaum, Miao Liu 等AAAI 2020 · 被引用 22 次
- Data-efficient Hindsight Off-policy Option LearningMarkus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe 等ICML 2021 · 被引用 48 次
- Hierarchies of Reward MachinesDaniel Furelos-Blanco, Mark Law, Anders Jonsson, Krysia Broda 等ICML 2023 · 被引用 15 次
- Policy Caches with Successor FeaturesMark W. Nemecek, Ron ParrICML 2021 · 被引用 19 次
