Options of Interest: Temporal Abstraction with Interest Functions
Khimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon, Doina Precup
摘要
Temporal abstraction refers to the ability of an agent to use behaviours of controllers which act for a limited, variable amount of time. The options framework describes such behaviours as consisting of a subset of states in which they can initiate, an internal policy and a stochastic termination condition. However, much of the subsequent work on option discovery has ignored the initiation set, because of difficulty in learning it from data. We provide a generalization of initiation sets suitable for general function approximation, by defining an interest function associated with an option. We derive a gradient-based learning algorithm for interest functions, leading to a new interest-option-critic architecture. We investigate how interest functions can be leveraged to learn interpretable and reusable temporal abstractions. We demonstrate the efficacy of the proposed approach through quantitative and qualitative results, in both discrete and continuous environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu 等ICLR 2024 · 被引用 97 次
- What can I do here? A Theory of Affordances in Reinforcement LearningKhimya Khetarpal, Zafarali Ahmed, Gheorghe Comanici, David Abel 等ICML 2020 · 被引用 60 次
- Reset-Free Lifelong Learning with Skill-Space PlanningKevin Lu, Aditya Grover, Pieter Abbeel, Igor MordatchICLR 2021 · 被引用 42 次
- Flexible Option LearningMartin Klissarov, Doina PrecupNeurIPS 2021 · 被引用 38 次
- Deep Laplacian-based Options for Temporally-Extended ExplorationMartin Klissarov, Marlos C. MachadoICML 2023 · 被引用 31 次
相关 Paper
- Discovery of Options via Meta-Learned SubgoalsVivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu 等NeurIPS 2021 · 被引用 38 次
- Average-Reward Learning and Planning with OptionsYi Wan, Abhishek Naik, Richard S. SuttonNeurIPS 2021 · 被引用 12 次
- On the Role of Weight Sharing During Deep Option LearningMatthew Riemer, Ignacio Cases, Clemens Rosenbaum, Miao Liu 等AAAI 2020 · 被引用 22 次
- Context-Specific Representation Abstraction for Deep Option LearningMarwa Abdulhai, Dong-Ki Kim, Matthew Riemer, Miao Liu 等AAAI 2022 · 被引用 14 次
- Temporally Abstract Partial ModelsKhimya Khetarpal, Zafarali Ahmed, Gheorghe Comanici, Doina PrecupNeurIPS 2021 · 被引用 17 次
