Continual Auxiliary Task Learning
Matthew McLeod, Chunlok Lo, Matthew Schlegel, Andrew Jacobsen, Raksha Kumaraswamy, Martha White, Adam White
Abstract
Learning auxiliary tasks, such as multiple predictions about the world, can provide many benefits to reinforcement learning systems. A variety of off-policy learning algorithms have been developed to learn such predictions, but as yet there is little work on how to adapt the behavior to gather useful data for those off-policy predictions. In this work, we investigate a reinforcement learning system designed to learn a collection of auxiliary tasks, with a behavior policy learning to take actions to improve those auxiliary predictions. We highlight the inherent non-stationarity in this continual auxiliary task learning problem, for both prediction learners and the behavior learner. We develop an algorithm based on successor features that facilitates tracking under non-stationary rewards, and prove the separation into learning successor features and rewards provides convergence rate improvements. We conduct an in-depth study into the resulting multi-prediction learning system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Learning Successor Features the Simple WayRaymond Chua, Arna Ghosh, Christos Kaplanis, Blake A. Richards et al.NeurIPS 2024 · 14 citations
- Exploring through Random Curiosity with General Value FunctionsAditya A. Ramesh, Louis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberNeurIPS 2022 · 13 citations
- Adaptive Interest for Emphatic Reinforcement LearningMartin Klissarov, Rasool Fakoor, Jonas W. Mueller, Kavosh Asadi et al.NeurIPS 2022 · 3 citations
- Adaptive Exploration for Data-Efficient General Value Function EvaluationsArushi Jain, Josiah Hanna, Doina PrecupNeurIPS 2024 · 3 citations
- Discerning Temporal Difference LearningJianfei MaAAAI 2024 · 2 citations
Builds on8
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo et al.ICLR 2020 · 349 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement LearningSilviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie et al.ICML 2020 · 145 citations
- Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) OptimismWang Chi Cheung, David Simchi-Levi, Ruihao ZhuICML 2020 · 114 citations
- Generating Adjacency-Constrained Subgoals in Hierarchical Reinforcement LearningTianren Zhang, Shangqi Guo, Tian Tan, Xiaolin Hu et al.NeurIPS 2020 · 112 citations
Related papers
- The Value-Improvement Path: Towards Better Representations for Reinforcement LearningWill Dabney, André Barreto, Mark Rowland, Robert Dadashi et al.AAAI 2021 · 76 citations
- Optimistic Linear Support and Successor Features as a Basis for Optimal Policy TransferLucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da SilvaICML 2022 · 36 citations
- Chaining Value Functions for Off-Policy LearningSimon Schmitt, John Shawe-Taylor, Hado van HasseltAAAI 2022 · 5 citations
- Distributional Successor Features Enable Zero-Shot Policy OptimizationChuning Zhu, Xinqi Wang, Tyler Han, Simon S. Du et al.NeurIPS 2024 · 11 citations
- Deep Reinforcement Learning amidst Continual Structured Non-StationarityAnnie Xie, James Harrison, Chelsea FinnICML 2021 · 43 citations
