Unveiling Options with Neural Network Decomposition
Mahdi Alikhasi, Levi Lelis
Abstract
In reinforcement learning, agents often learn policies for specific tasks without the ability to generalize this knowledge to related tasks. This paper introduces an algorithm that attempts to address this limitation by decomposing neural networks encoding policies for Markov Decision Processes into reusable sub-policies, which are used to synthesize temporally extended actions, or options. We consider neural networks with piecewise linear activation functions, so that they can be mapped to an equivalent tree that is similar to oblique decision trees. Since each node in such a tree serves as a function of the input of the tree, each sub-tree is a sub-policy of the main policy. We turn each of these sub-policies into options by wrapping it with while-loops of varied number of iterations. Given the large number of options, we propose a selection mechanism based on minimizing the Levin loss for a uniform policy on these options. Empirical results in two grid-world domains where exploration can be difficult confirm that our method can identify useful options, thereby accelerating the learning process on similar but different tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b19f8178-2a4b-4777-970d-4671e0b48c7dCited by top-tier papers1
Ask how each one uses itBuilds on9
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- Exploration in Reinforcement Learning with Deep Covering OptionsYuu Jinnai, Jee Won Park, Marlos C. Machado, George Dimitri KonidarisICLR 2020 · 64 citations
- Same State, Different Task: Continual Reinforcement Learning without InterferenceSamuel Kessler, Jack Parker-Holder, Philip J. Ball, Stefan Zohren et al.AAAI 2022 · 57 citations
- Reinforcement Learning with Competitive Ensembles of Information-Constrained PrimitivesAnirudh Goyal, Shagun Sodhani, Jonathan Binas, Xue Bin Peng et al.ICLR 2020 · 55 citations
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 51 citations
Related papers
- Discovery of Options via Meta-Learned SubgoalsVivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu et al.NeurIPS 2021 · 38 citations
- On the Role of Weight Sharing During Deep Option LearningMatthew Riemer, Ignacio Cases, Clemens Rosenbaum, Miao Liu et al.AAAI 2020 · 22 citations
- Data-efficient Hindsight Off-policy Option LearningMarkus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe et al.ICML 2021 · 48 citations
- Hierarchies of Reward MachinesDaniel Furelos-Blanco, Mark Law, Anders Jonsson, Krysia Broda et al.ICML 2023 · 15 citations
- Policy Caches with Successor FeaturesMark W. Nemecek, Ron ParrICML 2021 · 19 citations
