Flexible Option Learning
Martin Klissarov, Doina Precup
Abstract
Temporal abstraction in reinforcement learning (RL), offers the promise of improving generalization and knowledge transfer in complex environments, by propagating information more efficiently over time. Although option learning was initially formulated in a way that allows updating many options simultaneously, using off-policy, intra-option learning (Sutton, Precup & Singh, 1999) , many of the recent hierarchical reinforcement learning approaches only update a single option at a time: the option currently executing. We revisit and extend intra-option learning in the context of deep reinforcement learning, in order to enable updating all options consistent with current primitive action choices, without introducing any additional estimates. Our method can therefore be naturally adopted in most hierarchical RL frameworks. When we combine our approach with the option-critic algorithm for option discovery, we obtain significant improvements in performance and data-efficiency across a wide variety of domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf98402e-90a1-4dde-8806-71deb2fa1859Cited by top-tier papers9
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu et al.ICLR 2024 · 97 citations
- Deep Laplacian-based Options for Temporally-Extended ExplorationMartin Klissarov, Marlos C. MachadoICML 2023 · 31 citations
- Code as Reward: Empowering Reinforcement Learning with VLMsDavid Venuto, Mohammad Sami Nur Islam, Martin Klissarov, Doina Precup et al.ICML 2024 · 29 citations
- Autonomous Option Invention for Continual Hierarchical Reinforcement Learning and PlanningRashmeet Kaur Nayyar, Siddharth SrivastavaAAAI 2025 · 7 citations
- Skill Disentanglement for Imitation Learning from Suboptimal DemonstrationsTianxiang Zhao, Wenchao Yu, Suhang Wang, Lu Wang et al.KDD 2023 · 5 citations
Builds on7
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 184 citations
- Option Discovery using Deep Skill ChainingAkhil Bagaria, George KonidarisICLR 2020 · 126 citations
- Options of Interest: Temporal Abstraction with Interest FunctionsKhimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon et al.AAAI 2020 · 51 citations
- Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance SamplingYao Liu, Pierre-Luc Bacon, Emma BrunskillICML 2020 · 49 citations
- Data-efficient Hindsight Off-policy Option LearningMarkus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe et al.ICML 2021 · 48 citations
Related papers
- Context-Specific Representation Abstraction for Deep Option LearningMarwa Abdulhai, Dong-Ki Kim, Matthew Riemer, Miao Liu et al.AAAI 2022 · 14 citations
- Average-Reward Learning and Planning with OptionsYi Wan, Abhishek Naik, Richard S. SuttonNeurIPS 2021 · 12 citations
- On the Role of Weight Sharing During Deep Option LearningMatthew Riemer, Ignacio Cases, Clemens Rosenbaum, Miao Liu et al.AAAI 2020 · 22 citations
- An Efficient Transfer Learning Framework for Multiagent Reinforcement LearningTianpei Yang, Weixun Wang, Hongyao Tang, Jianye Hao et al.NeurIPS 2021 · 34 citations
- Discovery of Options via Meta-Learned SubgoalsVivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu et al.NeurIPS 2021 · 38 citations
