Discovering Temporally-Aware Reinforcement Learning Algorithms
Matthew Thomas Jackson, Chris Lu, Louis Kirsch, Robert Tjarko Lange, Shimon Whiteson, Jakob Nicolaus Foerster
Abstract
Recent advancements in meta-learning have enabled the automatic discovery of novel reinforcement learning algorithms parameterized by surrogate objective functions. To improve upon manually designed algorithms, the parameterization of this learned objective function must be expressive enough to represent novel principles of learning (instead of merely recovering already established ones) while still generalizing to a wide range of settings outside of its meta-training distribution. However, existing methods focus on discovering objective functions that, like many widely used objective functions in reinforcement learning, do not take into account the total number of steps allowed for training, or "training horizon". In contrast, humans use a plethora of different learning objectives across the course of acquiring a new ability. For instance, students may alter their studying techniques based on the proximity to exam deadlines and their self-assessed capabilities. This paper contends that ignoring the optimization time horizon significantly restricts the expressive potential of discovered learning algorithms. We propose a simple augmentation to two existing objective discovery approaches that allows the discovered algorithm to dynamically update its objective function throughout the agent's training procedure, resulting in expressive schedules and increased generalization across different training horizons. In the process, we find that commonly used meta-gradient approaches fail to discover such adaptive objective functions while evolution strategies discover highly dynamic learning rules. We demonstrate the effectiveness of our approach on a wide range of tasks and analyze the resulting learned algorithms, which we find effectively balance exploration and exploitation by modifying the structure of their learning rules throughout the agent's lifetime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Discovering Preference Optimization Algorithms with and for Large Language ModelsChris Lu, Samuel Holt, Claudio Fanconi, Alex J. Chan et al.NeurIPS 2024 · 41 citations
- Can Learned Optimization Make Reinforcement Learning Less Difficult?Alexander David Goldie, Chris Lu, Matthew Thomas Jackson, Shimon Whiteson et al.NeurIPS 2024 · 18 citations
- Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam TimestepsBenjamin Ellis, Matthew Thomas Jackson, Andrei Lupu, Alexander David Goldie et al.NeurIPS 2024 · 14 citations
- EvIL: Evolution Strategies for Generalisable Imitation LearningSilvia Sapora, Gokul Swamy, Chris Lu, Yee Whye Teh et al.ICML 2024 · 10 citations
- Behaviour DistillationAndrei Lupu, Chris Lu, Jarek Liesen, Robert Tjarko Lange et al.ICLR 2024 · 8 citations
Builds on14
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon et al.ICML 2022 · 269 citations
- Structured State Space Models for In-Context Reinforcement LearningChris Lu, Yannick Schroecker, Albert Gu, Emilio Parisotto et al.NeurIPS 2023 · 164 citations
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu et al.NeurIPS 2020 · 154 citations
- Understanding and Preventing Capacity Loss in Reinforcement LearningClare Lyle, Mark Rowland, Will DabneyICLR 2022 · 151 citations
- Discovered Policy OptimisationChris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz et al.NeurIPS 2022 · 134 citations
Related papers
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh et al.NeurIPS 2020 · 90 citations
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 132 citations
- Discovering Evolution Strategies via Meta-Black-Box OptimizationRobert Tjarko Lange, Tom Schaul, Yutian Chen, Tom Zahavy et al.ICLR 2023 · 21 citations
- Discovering General Reinforcement Learning Algorithms with Adversarial Environment DesignMatthew Thomas Jackson, Minqi Jiang, Jack Parker-Holder, Risto Vuorio et al.NeurIPS 2023 · 23 citations
- Adaptive Interest for Emphatic Reinforcement LearningMartin Klissarov, Rasool Fakoor, Jonas W. Mueller, Kavosh Asadi et al.NeurIPS 2022 · 3 citations
