Context-Specific Representation Abstraction for Deep Option Learning
Marwa Abdulhai, Dong-Ki Kim, Matthew Riemer, Miao Liu, Gerald Tesauro, Jonathan P. How
Abstract
Hierarchical reinforcement learning has focused on discovering temporally extended actions, such as options, that can provide benefits in problems requiring extensive exploration. One promising approach that learns these options end-to-end is the option-critic (OC) framework. We examine and show in this paper that OC does not decompose a problem into simpler sub-problems, but instead increases the size of the search over policy space with each option considering the entire state space during learning. This issue can result in practical limitations of this method, including sample inefficient learning. To address this problem, we introduce Context-Specific Representation Abstraction for Deep Option Learning (CRADOL), a new framework that considers both temporal abstraction and context-specific representation abstraction to effectively reduce the size of the search over policy space. Specifically, our method learns a factored belief state representation that enables each option to learn a policy over only a subsection of the state space. We test our method against hierarchical, non-hierarchical, and modular recurrent neural network baselines, demonstrating significant sample efficiency improvements in challenging partially observable environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 072f73f2-4304-422c-8a8b-6799175df066Cited by top-tier papers3
- Continual Learning In Environments With Polynomial Mixing TimesMatthew Riemer, Sharath Chandra Raparthy, Ignacio Cases, Gopeshh Subbaraj et al.NeurIPS 2022 · 18 citations
- Balancing Context Length and Mixing Times for Reinforcement Learning at ScaleMatthew Riemer, Khimya Khetarpal, Janarthanan Rajendran, Sarath ChandarNeurIPS 2024 · 7 citations
- Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous InferenceMatthew Riemer, Gopeshh Subbaraj, Glen Berseth, Irina RishICLR 2025
Builds on3
- Recurrent Independent MechanismsAnirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani et al.ICLR 2021 · 357 citations
- Options of Interest: Temporal Abstraction with Interest FunctionsKhimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon et al.AAAI 2020 · 51 citations
- On the Role of Weight Sharing During Deep Option LearningMatthew Riemer, Ignacio Cases, Clemens Rosenbaum, Miao Liu et al.AAAI 2020 · 22 citations
Related papers
- Flexible Option LearningMartin Klissarov, Doina PrecupNeurIPS 2021 · 38 citations
- DHRL: A Graph-Based Approach for Long-Horizon and Sparse Hierarchical Reinforcement LearningSeungjae Lee, Jigang Kim, Inkyu Jang, H. Jin KimNeurIPS 2022 · 33 citations
- Autonomous Option Invention for Continual Hierarchical Reinforcement Learning and PlanningRashmeet Kaur Nayyar, Siddharth SrivastavaAAAI 2025 · 7 citations
- Possibility Before Utility: Learning And Using Hierarchical AffordancesRobby Costales, Shariq Iqbal, Fei ShaICLR 2022 · 5 citations
- Active Hierarchical Exploration with Stable Subgoal Representation LearningSiyuan Li, Jin Zhang, Jianhao Wang, Yang Yu et al.ICLR 2022 · 28 citations
